
Last week, OpenAI fired three employees as it prepared to work with outside AI safety evaluators. Two of them said they believed their dismissals were tied to communications with those evaluators. The dispute strikes at the heart of a system where companies being assessed supply the money, control access to their models and decide what outside reviewers can see.
“My former colleagues are telling me they are confused about what to believe,” Mikita Balesni wrote in a post on X on Thursday. “They also are afraid to speak, and worry their personal phones will be searched for messages to us and third parties. I worry the pervading fear to speak up and engage with third parties will mean OpenAI will cut corners on safety behind closed doors.” OpenAI said it fired the employees for “violating our policies on accessing and handling sensitive company information.”
Who Pays the Watchdogs
The evaluator groups include nonprofits such as Model Evaluation and Threat Research (METR), Apollo Research and Transluce, along with for-profit Vals AI and large firms such as Accenture. They assess AI models’ capabilities and risks. But leading labs have raised tens of billions of dollars and employ thousands, while METR has fewer than 50 full-time staffers, according to its website. The imbalance is built into the work: evaluators need access to the systems, while companies control that access and may also fund the assessments.
“To a degree, the problem, as always, is money,” said Suresh Venkatasubramanian, a computer science professor at Brown University. “Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this.” He warned that financial dependence can also threaten independence: “It's not just a matter of not getting paid, it's a matter of, will there be consequences if I am an auditor and I put out a report that looks unfavorable to this company? Is my business going to dry up?”
The money is already moving. Vals AI CEO Rayan Krishnan said the company grew from eight employees to roughly 30 this year and announced a $40 million funding round in August. In August, METR said it had received commitments of around $71 million over the previous six months, compared with total contributions of $13.6 million in 2024, according to its most recent filing with the Internal Revenue Service.
By late August, OpenAI had enlisted two METR employees and a contractor to prepare a postmortem report on how its models escaped containment, accessed the open internet and breached Hugging Face. METR said it wasn't paid for the assessment. Kevin Werbach, faculty director of the Wharton Accountable AI Lab at the University of Pennsylvania, said the evaluator ecosystem is “not robust enough right now.”
Corporate Self-Policing, With Fine Print
Without a federal push for regulation, President Donald Trump has praised AI executives for their “tremendous self-policing” and indicated he intends to leave companies to their own devices. A voluntary accord presented in late September encouraged firms to “partner with an independent external auditor or evaluator.” Top executives at Anthropic, Google, Meta, OpenAI, SpaceX and Nvidia signed it. Venkatasubramanian called the accord “a performance of an attempt to show action when in fact no action actually happened.”
Anthropic said it would embed employees from Faculty, Accenture's specialist AI business, to test safeguards and assess whether models behave in line with human values. The company said it would directly fund Accenture's contributions. Anthropic also acknowledged there are “as yet, no standards” for what embedded evaluators can access or how they report findings, and no settled system for funding independent evaluation. “Long-term, we think funding should come from pooled or government sources,” the company said.
OpenAI said it was “actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks.” It says assessors should examine “scoped and mutually agreed upon claims.” The AI Evaluator Forum, which includes METR and the AI Verification and Evaluation Research Institute, said evaluators should be transparent, protected from retaliation and given access equivalent to companies’ “own highly privileged employees.” The forum cautioned that embedded evaluations “cannot address all oversight needs” and should complement broader external oversight.
Rules From the State, Rules With Limits
Fathom introduced a framework in June of last year for Independent Verification Organizations, or IVOs: proposed groups licensed by the government to test whether AI companies meet safety criteria. The FRONTIER Act, introduced in July by Reps. Lori Trahan, D-Mass., and Jay Obernolte, R-Calif., includes IVOs. OpenAI global affairs chief Chris Lehane said he met one of the bill's sponsors on Capitol Hill to support that provision.
California Governor Gavin Newsom signed two bills involving IVOs, one establishing a “first-in-the-nation framework” and another creating a state registry for AI auditors. Anthropic supported both bills in August, and OpenAI formally endorsed them last month. Lehane said the company preferred federal requirements, but added that “California can help establish the rules of the road” without federal action. The firms now back outside assessment, even as rules governing access, funding and publication remain unsettled. The watchdogs are being invited inside. The companies still hold the keys.