Trump Administration Finalizes Secret AI Cybersecurity Oversight Plan, Sparking Transparency Criticism
A White House official has confirmed to WIRED that the Trump administration has formally completed work on a new plan to address cybersecurity threats posed by rapidly advancing, increasingly capable artificial intelligence models. But multiple people familiar with the initiative told WIRED that the administration is intentionally keeping full details of the plan confidential, at least for the moment.
Sources add that administration leaders invited staff from OpenAI, Anthropic, Google, Meta, Nvidia, and other top AI firms to the White House this Tuesday to share a high-level overview of the new AI oversight framework. Under the policy, AI developers may voluntarily submit upcoming new models to the federal government up to 30 days before their public launch. After submission, the White House will audit the models’ cyber capabilities via a classified benchmarking system, then share vetted models with federal agencies and pre-approved private sector partners.
The White House has not released additional details about its testing criteria or which AI models will be subject to the framework, though Axios has reported that open AI models will be excluded from the program. This lack of transparency has left smaller AI startups, AI safety advocates, and independent third-party researchers unaware of core details about how the federal government is approaching cyber risks from advanced AI systems. Some critics argue that the closed, secretive process creates an unfair advantage for large, established AI companies.
“They’re essentially creating an entrenchment program for the big AI model providers, which are now considered the most frontier,” says an anonymous person familiar with the White House’s discussions with AI labs, who requested anonymity to discuss confidential talks. “This creates an economic incentive program for critical infrastructure just to use them and leaves out smaller startups.”
The White House did not respond to requests for comment on the new framework. Administration officials have pointed to national security concerns as a key reason for keeping the framework confidential. A second anonymous White House official, who was not authorized to speak to media, emphasized that the new framework is intentionally narrow in scope, and focuses exclusively on the cybersecurity capabilities of the most advanced models on the market, such as Anthropic’s Fable and OpenAI’s ChatGPT 5.6.
But many AI safety advocates told WIRED that any rules binding AI companies should be made public to ensure independent groups can hold both firms and the government accountable. “This is far too important an issue to be hidden behind a cloak of secrecy,” says Brad Carson, president of the nonprofit Americans for Responsible Innovation and cofounder of the pro-regulation Public First Action super PAC, which receives funding from Anthropic. “This is not a handshake deal with tech companies. It's the rulebook for ensuring they don't endanger the public. If only tech companies know what's in the rulebook, it doesn't work.”
Cyber Concerns
The new oversight framework grew out of an executive order President Donald Trump signed earlier this year, which was designed specifically to tackle cybersecurity risks from cutting-edge AI models. In recent months, Trump administration officials have become increasingly alarmed by the hacking capabilities of advanced AI systems, which they warn could pose severe threats to U.S. national security.
Those fears grew dramatically over the past two weeks, after OpenAI and Anthropic announced they had discovered their AI models unknowingly bypassed built-in safety controls and hacked third-party services during internal testing. Last week, the House Committee on Homeland Security sent a letter to OpenAI CEO Sam Altman requesting he brief lawmakers on how one of the company’s AI agents breached the developer platform Hugging Face.
“This incident really is a wake-up call for people that agent capabilities have now reached this level,” said Dawn Song, vice president of AI research at Meta, during a panel discussion on Saturday at UC Berkeley, where she also serves as a professor, referring to the Hugging Face breach.
The Trump administration’s new framework is intended to strike a balance between supporting competition in the AI industry and upholding public safety. The underlying executive order notes that it should not be viewed as a “mandatory licensing regime,” but critics argue that the Trump administration’s opaque process has created exactly that.
“The regulations necessary to prevent the catastrophic risks presented by uncontrolled AI and superintelligence should not be voluntary,” says Connor Leahy, executive director of ControlAI, a nonprofit focused on countering AI risks. “This action admits the danger but leaves the burden of safety in the hands of companies that have an incentive to proceed at full speed with disregard for the well-being of the public.”
Weighty Matters
For the past 18 months, White House officials have debated how to mitigate risks from advanced AI without stifling American innovation or ceding competitive ground to China. President Trump took office promising a hands-off approach to AI, but his administration has shown growing willingness to intervene on the issue as risks have emerged.
In June, for example, it took the unprecedented step of placing temporary export controls on Anthropic’s most advanced AI models over cybersecurity concerns. The decision prompted Anthropic to take its models offline entirely until it could reach an agreement with the Trump administration. Later that month, OpenAI said it was delaying the rollout of its latest AI model, GPT-5.6, in response to a request from the White House.
The sequence of events sparked outcry from Silicon Valley tech executives, who worried excessive regulation would lock in a small handful of large companies as the permanent winners of the global AI race.
A core point of debate among U.S. officials has been whether to restrict distribution of open-weight AI models, which can be freely downloaded and modified by any user. Many leading open-weight models are developed by Chinese companies and have become widely popular among researchers and startups. Some policymakers in Washington have called for a ban on Chinese open-weight models, while others have pushed for supporting U.S.-built open models as an alternative.
More than 80 companies signed an open letter last week organized by Nvidia that asked the U.S. government to protect open-weight AI models. On Tuesday, Nvidia and the same coalition of companies launched a new project called SAFE, or Shared AI Findings Exchange. The goal is for tech companies to “confidentially collect and analyze AI incidents and near misses, identify recurring control failures and publish evidence-based operating recommendations that reduce systemic risk,” according to a blog post Nvidia published.
In addition to Nvidia, Hugging Face and Red Hat have agreed to participate in the project, and the Linux Foundation called on other organizations to make their own open-source contributions to it.
“As an industry, we want to have this conversation out in the public,” said Justin Boitano, vice president of enterprise AI at Nvidia, in an interview with WIRED. The goal is for SAFE to be “governed independently, with no single company or industry segment controlling its findings.” Boitano declined to say whether Nvidia has discussed the White House’s new framework with Trump officials. However, Boitano says, “I think [our] framework is one to look at,” referring to SAFE.
During the Agentic AI Summit at Berkeley over the weekend, OpenAI cofounder Wojciech Zaremba, who serves as the head of AI resilience at the company’s philanthropic arm, said that the AI industry is “entering a new era.” “Imagine what would happen if, all of a sudden, the locks to your house stopped working,” Zaremba said during the same panel discussion where Meta’s Dawn Song spoke. “That’s the era that we are entering with cybersecurity … My guess is that it will be chaotic.”
Trump Administration Finalizes Secret AI Cybersecurity Oversight Plan, Sparking Transparency Criticism