Technology

AI's third-party safety evaluators step into the spotlight as funding and access questions loom

Third-party AI safety evaluators assessing model testing at an independent research organization
Small nonprofits like METR and Apollo Research are being asked to police the AI industry, but questions over funding, access and reporting remain unresolved.

A handful of small third-party evaluators, long a quiet corner of the multitrillion-dollar artificial intelligence industry, are being asked to help safeguard it as a dispute unfolds between Anthropic and OpenAI over whether the labs can protect their advanced models while growing their businesses.

The labs are seeking support from groups including Model Evaluation and Threat Research (METR), Apollo Research and Transluce. The evaluators, most of them nonprofits, have mainly assessed model capabilities and risks and flagged cases where the technology misbehaves. They are still establishing themselves in an industry drawing record capital and releasing new models faster than ever.

With no federal regulatory push, the evaluators have taken on outsized importance. Anthropic CEO Dario Amodei pledged last month to embed independent evaluators at his company, a move OpenAI CEO Sam Altman quickly endorsed. President Donald Trump backed the idea, as did most of the largest U.S. tech companies. But how the third parties will be funded, how much access they will receive and what their reporting structure will look like remain unanswered.

“To a degree, the problem, as always, is money,” Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC in an interview. “Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this.”

Anthropic, OpenAI and the infrastructure partners profiting from the AI boom are currently setting the rules. Critics compare the arrangement to asking the largest banks to guard against a financial crisis, or letting pharmaceutical companies sell drugs without regulatory clearance.

Voluntary oversight

Trump recently praised AI executives for their “tremendous self-policing” and signaled he intends to leave companies to their own devices rather than impede an industry driving the economy and the stock market. A voluntary accord he presented in late September encourages AI companies to “partner with an independent external auditor or evaluator.”

The conversation began with Amodei’s viral essay last month, in which he called for a “slower pace” in advanced model development after researchers left Anthropic and raised concerns about existential threats from the technology.

Friction has already emerged as the labs move to install evaluators.

OpenAI last week fired three employees for “violating our policies on accessing and handling sensitive company information,” according to a spokesperson. Two of them, Mikita Balesni and Tomek Korbak, said they believe they were dismissed because of how they communicated with third-party evaluators.

“My former colleagues are telling me they are confused about what to believe,” Balesni wrote in a post on X on Thursday. “They also are afraid to speak, and worry their personal phones will be searched for messages to us and third parties. I worry the pervading fear to speak up and engage with third parties will mean OpenAI will cut corners on safety behind closed doors.”

OpenAI disputed that characterization. In a post on Friday, the company said it is “actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks.”

“We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work,” OpenAI wrote.

An OpenAI spokesperson said in an emailed statement that the upcoming work “builds on existing collaboration with independent safety organizations,” including METR and Redwood Research.

A maturing field

The evaluator ecosystem consists largely of small organizations such as METR and Apollo Research, alongside larger accounting and auditing firms like Accenture.

AI labs have worked with evaluators in limited capacities, but Andrew Freedman, CEO of the policy nonprofit Fathom, said the field is maturing quickly.

“I’ve worked in politics and policy for the last 20 years of my life, and I’ve never seen an issue move so fast on so many different political spectrums,” Freedman told CNBC in an interview. He said he expects an “influx of capital” into the ecosystem.

The money is already arriving. Rayan Krishnan, CEO of the independent evaluator Vals AI, said his for-profit startup, which builds benchmarks measuring how models perform on industry-specific tasks, grew from eight employees to roughly 30 this year and announced a $40 million funding round in August.

METR, a nonprofit, said in August that it had raised commitments of around $71 million over the prior six months, up from total 2024 contributions of $13.6 million, according to its most recent filing with the Internal Revenue Service.

METR’s profile rose further late that month, when OpenAI enlisted two of its employees and a contractor to produce a postmortem report on how the company’s models escaped containment, accessed the open internet and breached the open-source developer platform Hugging Face. METR said it did not accept payment from OpenAI for the assessment.

Kevin Werbach, faculty director of the Wharton Accountable AI Lab at the University of Pennsylvania, said the ecosystem is “not robust enough right now.” METR employs fewer than 50 full-time staffers, according to its website, while the leading labs have raised tens of billions of dollars and employ thousands of people.

“If you want true third-party evaluation, you need true independence financially and otherwise,” Venkatasubramanian said. “It’s not just a matter of not getting paid, it’s a matter of, will there be consequences if I am an auditor and I put out a report that looks unfavorable to this company? Is my business going to dry up?”

Anthropic embeds reviewers

Anthropic acknowledged the complexity in a blog post last month announcing it will embed employees from Faculty, Accenture’s specialist AI business, to test safeguards and assess whether models behave in line with human values. Anthropic said that “given the importance and urgency of this work,” it will fund Accenture’s contributions directly.

“There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation,” Anthropic said. “Long-term, we think funding should come from pooled or government sources.”

The company said it is in discussions with METR and other nonprofit evaluators planning to use their own funding to pilot “elements” of embedded evaluation.

In June of last year, Fathom introduced a marketplace framework for Independent Verification Organizations, or IVOs — groups that would be licensed by the government and authorized to test whether AI companies meet safety criteria.

Freedman said government oversight is essential because, without it, third-party evaluators can become dependent on the large AI labs for revenue, incentivizing them to “start rubber stamping stuff” to stay in favor.

IVOs are a key provision of the “Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting” (FRONTIER) Act, introduced in July by Reps. Lori Trahan, D-Mass., and Jay Obernolte, R-Calif. Fathom helped draft language for the bill and provided technical expertise, Freedman said.

OpenAI global affairs chief Chris Lehane told reporters in September that he met one of the bill’s sponsors on Capitol Hill to express support for the IVO provision.

“It was important for them to hear that and hear it from us, and we wanted to be really clear about that,” Lehane said, according to reports.

States act first

Lawmakers in California, Connecticut and Virginia have moved to implement IVOs, and states including Massachusetts are weighing independent safety evaluations more broadly.

California Governor Gavin Newsom recently signed two IVO-related bills, one establishing a “first-in-the-nation framework” and the other creating a state registry for AI auditors. Anthropic backed both bills in August, and OpenAI formally endorsed them last month, the same day Newsom signed them into law.

Writing in a blog post at the time, Lehane said OpenAI “prefer[s] independent technical assessments to be required at the federal level,” but that in the absence of federal action, “California can help establish the rules of the road.”

Freedman said it will be “really difficult” for companies such as OpenAI and Anthropic to work out how to engage independent evaluators on their own. With the government’s role unclear, “it’s a muscle worth developing in the interim,” he said.

A ‘morally binding’ accord

For now, the industry’s closest thing to standards is the one-page accord Trump called “morally binding” at a luncheon he hosted for tech leaders at the White House late last month.

The document states that “every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public,” and encourages signees to work with an “independent external auditor or evaluator to carry out independent assessments.”

Top executives at Anthropic, Google, Meta, OpenAI, SpaceX and Nvidia signed the accord, a rare show of unity among leaders who have held conflicting views on managing AI’s risks. The executives still must chart their own paths forward.

“It was a performance of an attempt to show action when in fact no action actually happened,” Venkatasubramanian said. “The things that they promise to do are things they should have been doing already, and, in fact, have claimed that they were doing in the past.”

In his September essay, Amodei said Anthropic will give evaluators desks, access badges, company laptops and permissions “mostly comparable” with those of internal risk assessment teams. Contracts will grant evaluators “the right to publish key findings,” with Anthropic reserving “the narrow ability” to redact certain security-sensitive or confidential information.

“This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers,” Amodei wrote.

OpenAI published its own proposal days later, saying evaluators should work on “scoped and mutually agreed upon claims for assessment,” explain their methodology and standards clearly, demonstrate relevant technical expertise and disclose conflicts of interest.

The AI Evaluator Forum, a group that includes METR and the AI Verification and Evaluation Research Institute (AVERI), published a public letter last month titled “Minimum Conditions for Embedding Evaluators.” The letter said evaluators should be transparent, shielded from retaliation and granted access equivalent to AI companies’ “own highly privileged employees.”

“Embedded evaluations cannot address all oversight needs and should be treated as a complement to, rather than a replacement for, broader efforts by frontier AI companies to expand external oversight,” the letter said.

Freedman said he has seen a shift in posturing from OpenAI and Anthropic in recent months, largely because they have realized they cannot roll out advanced systems without public trust.

“I don’t think you need to trust that they’ve suddenly turned altruistic or that there’s anything but corporations acting like corporations,” Freedman said.

Werbach said that points to the central problem: OpenAI and Anthropic are, first and foremost, competing with each other as they move toward the public markets and seek trillion-dollar-plus valuations.

“There is a tremendous amount of personal distrust between those two companies,” Werbach said. “Even though there’s also tremendous agreement about the need for this kind of evaluation to happen.”

COMMENTS

Leave a comment

MORE NEWS

Flash floods sweep away houses and bridges in central Nepal

A flash flood along the Budhi Gandaki River in Nepal's Gorkha district has destroyed four bridges and nearly two dozen houses and damaged a hydropower project, with five people still missing after 49 were rescued, according to the army and the district's top official.