Glossary term

Evaluator Access

Permanent, employee-level access for independent outside organisations to a lab’s frontier models and internal systems, so that safety and capability claims can be checked before and after release. Proposed as a commitment by Anthropic and OpenAI in September 2026; the evaluators have not yet been named.

AI-generated — produced automatically by Closelook’s systems under this site’s editorial policy.

What it means

Today most outside testing of frontier models happens through limited API access for a few weeks before release. Evaluator access means something stronger: named organisations with standing access to models, training data descriptions, internal evaluations and, in the strongest form, the systems the lab uses to monitor its own models — the same access an employee would have. The evaluator can then say publicly whether a model does what the lab claims and whether it shows capabilities the lab has not disclosed.

The commitment is meaningful only when the evaluators are named, their access is verified and their findings can be published without the lab’s approval. As of mid-September 2026 none of the three had been done.

Why it matters for the AI trade

Evaluator access is the mechanism that would turn the pacing call from a promise into a constraint: an evaluator that can see a lab’s internal capability results can also see whether the lab is pacing. For investors in the labs’ suppliers and customers it is the difference between a managed frontier that can be verified and one that is asserted. It is also the first place a real drift or reward-hacking finding would surface publicly.

How Closelook uses it

The recursive self-improvement read keeps a dated paragraph on what has been named and what has not; the names of the evaluator organisations are one of the two facts we are waiting for, with the next Gemini release notes the other.

Common questions

Who would the evaluators be?
Not announced as of 15 September 2026. Candidates discussed in public include government AI safety institutes and independent research organisations that already run pre-release tests under narrower terms.
Is this the same as regulation?
No. It is a voluntary commitment by the labs. Regulation would make access and disclosure a legal requirement; the pacing essay asked for common standards but did not propose a law.
Why does it matter for investors?
It is how a slower frontier would be verified rather than asserted. Suppliers’ valuations depend on the cadence of frontier releases; an evaluator with real access is the only outsider who would know the cadence had changed before the orders did.