Home Artificial intelligence AI’s quiet safety gatekeepers are stepping into the spotlight
Artificial intelligence

AI’s quiet safety gatekeepers are stepping into the spotlight

Share


Co-founder and CEO of Anthropic Dario Amodei looks on as US President Donald Trump speaks to the press after a meeting with technology executives about artificial intelligence at the White House in Washington, DC, on Sept. 29, 2026.

Kent Nishimura | AFP | Getty Images

Two months ago, independent evaluators occupied a relatively sleepy corner of the multitrillion-dollar artificial intelligence industry. Now they’re being asked to come to its rescue. 

While Anthropic and OpenAI are the heart of a fierce debate over whether they can safeguard their advanced models and grow their businesses simultaneously, the companies are seeking support from a handful of small third-party groups like Model Evaluation and Threat Research (METR), Apollo Research and Transluce.

The evaluators, which mostly operate as nonprofits, are still finding their footing in an industry where capital is flowing at historic levels and new models are rolling out faster than ever. Their primary role has been to assess AI model capabilities and risks, and to call attention to instances where the technology behaves badly.

In the absence of a federal push for regulations, evaluators have taken on outsized importance. Anthropic CEO Dario Amodei pledged to embed independent evaluators in his company last month – a move that OpenAI CEO Sam Altman quickly endorsed. President Donald Trump supported the idea, as did most of the largest U.S. tech companies. But left unanswered are questions about how those third parties should be funded, what level of access they will have and what the reporting structure will ultimately look like.

“To a degree, the problem, as always, is money,” Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC in an interview. “Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this.”

Right now, Anthropic, OpenAI and the infrastructure partners that are profiting from the AI boom are writing the rules. Critics say that’s like asking the biggest banks to protect us from a financial crisis or allowing pharmaceutical companies to put drugs on the market without regulatory clearance.

President Trump recently lauded AI executives for their “tremendous self-policing,” and signaled that he intends to leave companies to their own devices, unwilling to impede the growth of the industry that’s driving the economy and stock market. Trump encouraged AI companies to “partner with an independent external auditor or evaluator” as part of a voluntary accord he presented in late September.

Dissecting Trump's 'morally binding' AI order

It’s a conversation that Amodei kicked off In his viral essay last month, when he called for a “slower pace” in advanced model development after researchers left his company and voiced their concerns about the existential threats the technology poses.

As the AI labs move to put evaluators in place, friction is already starting to emerge.

OpenAI fired three employees last week for “violating our policies on accessing and handling sensitive company information,” according to a spokesperson. Two of those employees, Mikita Balesni and Tomek Korbak, said they believe they were dismissed because of how they communicated with third-party evaluators. 

“My former colleagues are telling me they are confused about what to believe,” Balesni wrote in a post on X on Thursday. “They also are afraid to speak, and worry their personal phones will be searched for messages to us and third parties. I worry the pervading fear to speak up and engage with third parties will mean OpenAI will cut corners on safety behind closed doors.”

OpenAI disputed that characterization and said in a post on Friday that it’s “actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks.”

“We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work,” OpenAI wrote. 

An OpenAI spokesperson said in an emailed statement that its upcoming work with evaluators “builds on existing collaboration with independent safety organizations,” including METR and Redwood Research.

Anthropic didn’t respond to CNBC’s request for comment.

‘I’ve never seen an issue move so fast’

The AI evaluator ecosystem consists mostly of small organizations, including METR and Apollo Research, and larger accounting and auditing firms like Accenture. 

AI labs have been working with evaluators in limited capacities, but Andrew Freedman, CEO of policy nonprofit Fathom, said the field is quickly maturing. 

“I’ve worked in politics and policy for the last 20 years of my life, and I’ve never seen an issue move so fast on so many different political spectrums,” Freedman told CNBC in an interview. He said he expects an “influx of capital” to flow into the ecosystem.

Rayan Krishnan, CEO of independent evaluator Vals AI, said his for-profit startup, which builds benchmarks to measure how AI models perform on industry-specific tasks, has grown from eight employees to roughly 30 this year, and in August announced a $40 million funding round.

METR, a nonprofit, announced in August that it had raised commitments of around $71 million over the last six months. That’s up from total 2024 contributions of $13.6 million, according to the group’s most recent filing with the Internal Revenue Service.  

By late that month, METR’s profile had risen further. OpenAI enlisted two of its employees and a contractor to put together a postmortem report detailing how the company’s models escaped containment, accessed the open internet and breached open-source developer platform Hugging Face. METR said it did not accept payment from OpenAI for the assessment. 

Andrey Rudakov | Bloomberg | Getty Images

Kevin Werbach, faculty director of the Wharton Accountable AI Lab at the University of Pennsylvania, said the ecosystem is “not robust enough right now.” METR, for example, employs fewer than 50 full-time staffers, according to its website. 

The power imbalance between the small evaluators and the leading labs that have raised tens of billions of dollars and employ thousands of people raises questions surrounding potential conflicts.

“If you want true third-party evaluation, you need true independence financially and otherwise,” said Venkatasubramanian. “It’s not just a matter of not getting paid, it’s a matter of, will there be consequences if I am an auditor and I put out a report that looks unfavorable to this company? Is my business going to dry up?” 

Anthropic acknowledged the complexity in a blog post last month, as it announced it will embed employees from Faculty, Accenture’s specialist AI business, to test safeguards and assess whether models will behave in line with human values. Anthropic said that “given the importance and urgency of this work,” it will fund Accenture’s contributions directly.

“There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find. There is also no settled system for funding independent evaluation,” Anthropic said. “Long-term, we think funding should come from pooled or government sources.”

Anthropic said it’s in discussions with METR and other nonprofit evaluators that are planning to use their own funding to pilot “elements” of embedded evaluation. 

Will the government step in?

In June of last year, Fathom introduced a marketplace framework for Independent Verification Organizations, or IVOs. These groups would be licensed by the government and authorized to test whether AI companies are meeting various safety criteria. 

Freedman, the group’s CEO, said government oversight is key because otherwise third-party evaluators can become beholden to the large AI labs for revenue, incentivizing them to “start rubber stamping stuff” to maintain favor. 

Some lawmakers are on board.

IVOs are a key provision of the ”Frontier Risk Oversight, National Transparency, Independent Evaluation, and Reporting” (FRONTIER) Act, which Reps. Lori Trahan, D-Mass., and Jay Obernolte, R-Calif., introduced in July. Fathom helped draft language and provided technical expertise for the bill, Freedman said.  

OpenAI global affairs chief Chris Lehane told reporters in September that he sat down with one of the bill’s sponsors on Capitol Hill to express support for the IVO provision.

“It was important for them to hear that and hear it from us, and we wanted to be really clear about that,” Lehane said, according to reports. 

Meanwhile, lawmakers in California, Connecticut and Virginia have taken steps to implement IVOs, and states like Massachusetts are weighing independent safety evaluations more broadly. 

Gavin Newsom, Governor of California, speaks during a press briefing at the United Nations in New York City on Sept. 22, 2026.

Michael Kappeler | Picture Alliance | Getty Images

California Governor Gavin Newsom recently signed two bills involving IVOs, one establishing a “first-in-the-nation framework,” and the other creating a state registry for AI auditors. Anthropic threw its support behind both bills in August, and OpenAI formally endorsed them last month, the same day Newsom signed them into law. 

Lehane wrote in a blog post at the time that “we prefer independent technical assessments to be required at the federal level,” but in the absence of federal action, “California can help establish the rules of the road.” 

Freedman said he thinks it will be “really difficult” for companies like OpenAI and Anthropic to work out how to engage with independent evaluators on their own. However, with the government’s role unclear, “it’s a muscle worth developing in the interim,” he said.

For now, the closest thing the industry has to a set of standards is what Trump called a “morally binding” agreement at a luncheon he hosted for tech leaders at the White House late last month.

The one-page accord says that “every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public.” It also encourages signees to work with an “independent external auditor or evaluator to carry out independent assessments.”

The document was signed by top execs at Anthropic, Google, Meta, OpenAI, SpaceX and Nvidia, a rare show of solidarity between leaders who have shared conflicting views on addressing AI’s risks. The executives still have to chart their own paths forward.

“It was a performance of an attempt to show action when in fact no action actually happened,” Venkatasubramanian said. “The things that they promise to do are things they should have been doing already, and, in fact, have claimed that they were doing in the past.”

Tech leaders sign AI agreement

Amodei, in his September essay, said Anthropic will equip evaluators with desks, access badges, company laptops, and permissions that are “mostly comparable” with internal risk assessment teams. Additionally, evaluators will be supported with contracts that give them “the right to publish key findings,” with Anthropic reserving “the narrow ability” to redact certain security-sensitive or confidential information.

“This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers,” Amodei wrote.

OpenAI published its own proposal days later, and said evaluators should work on “scoped and mutually agreed upon claims for assessment,” clearly explain their methodology and standards, demonstrate relevant technical expertise and disclose conflicts of interest.

The AI Evaluator Forum, which includes METR, the AI Verification and Evaluation Research Institute (AVERI), and other groups, published a public letter last month titled, “Minimum Conditions for Embedding Evaluators.”

The letter said evaluators should be transparent, shielded from retaliation and granted access equivalent to AI companies’ “own highly privileged employees.”

“Embedded evaluations cannot address all oversight needs and should be treated as a complement to, rather than a replacement for, broader efforts by frontier AI companies to expand external oversight,” the letter said.

Freedman said he’s seen a shift in posturing out of OpenAI and Anthropic in recent months, largely because they’ve realized they won’t be able to roll out their advanced systems without the public’s trust. 

“I don’t think you need to trust that they’ve suddenly turned altruistic or that there’s anything but corporations acting like corporations,” Freedman said.

That underscores perhaps the central problem, Werbach said. OpenAI and Anthropic are, first and foremost, competing with each other as they march toward the public markets and seek trillion-dollar-plus valuations.

“There is a tremendous amount of personal distrust between those two companies,” Werbach said. “Even though there’s also tremendous agreement about the need for this kind of evaluation to happen.” 

WATCH: Bradley Tusk on Anthropic IPO: Why add public market pressure if safety is your top priority?

Bradley Tusk on Anthropic IPO: Why add public market pressure if safety is your top priority?



Source link

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Articles
Artificial intelligence

Questions mount over what an AI ‘slowdown’ would look like

We also can't deny the fact that underpinning this entire tech revolution...

Artificial intelligence

Novelist accused of using AI to write book removed from French prize list | Prix Goncourt

The organisers of France’s most prestigious literary prize said they had removed...

Artificial intelligence

Microsoft CEO Satya Nadella Calls For ‘Emergency Brake’ On Advanced AI

Microsoft Corp. Chief Executive Officer Satya Nadella said companies should treat powerful...