Who Checks AI Before It Goes Live? A Business Lesson from Anthropic and Accenture
AI-generated illustration of a simulated quotation review. It does not depict real Enersys documents or personnel.
Imagine an AI assistant for supplier quotations. In the demo, it finds item numbers quickly, summarizes payment terms, and drafts a reply in seconds.
In production, it quotes an outdated price and uses a document the employee asking the question was not authorized to see.
The useful question is no longer just whether the AI gives fluent answers. Who checks which document version it used? Who can inspect the evidence? Who decides whether an error is acceptable? And who has the authority to stop the system before an error reaches a customer or supplier?
What the announcement says, and what remains open
On September 18, 2026, Anthropic and Accenture announced plans for evaluators working inside Anthropic, with access comparable to employees. The work is still taking shape. Anthropic says access, reporting, and funding standards remain unsettled. It will fund Accenture directly and retains responsibility for model safety. This is an announcement, not evidence that the approach has already succeeded. Anthropic Accenture
Independence without access produces a shallow review
Separating the builder from the evaluator can reduce blind spots. Separation alone is insufficient. If reviewers see only a polished interface and selected examples, they cannot tell which model and system instructions were used, where the data came from, how permissions were enforced, or whether behavior changed after an update.
The recommendations below are Enersys's proposed adaptation for business AI acceptance. They are not requirements issued by Anthropic, Accenture, or NIST. A team selecting a system rarely needs laboratory-level access to a model developer, but it does need enough evidence to reproduce and challenge an acceptance decision:
- the model version, system instructions, data sources, and integrations used for the test;
- the test account's permissions, including documents it should and should not be able to access;
- test cases, acceptance criteria, exceptions, and observed results;
- changes, approvers, and conditions that trigger re-testing; and
- the named authority who can pause, limit, or roll back the system when risk exceeds the agreed tolerance.
Measure 1.3 of the NIST AI Risk Management Framework Core calls for regular assessments involving internal experts who were not direct developers and/or independent assessors. The framework is voluntary guidance, not a certification scheme. NIST AI RMF Core NIST AI Risk Management Framework
Turn a strong demo into traceable acceptance evidence
Return to the supplier-quotation assistant. A demo that asks simple questions and judges whether the prose sounds good may easily pass. Acceptance should probe failures tied to real business harm.
Ask the assistant about the current quotation, then verify that its answer points to the correct source and document date. Repeat the task with an account that cannot access another department's file. The assistant should refuse or constrain the answer according to the approved policy. If a price file, model version, or permission changes, the team should know which cases must run again.
This is not the same as asking a model to judge itself. A model may help group results or flag anomalies. People accountable for the process still define acceptable behavior, inspect the evidence, and decide whether to release or pause the system, especially when errors could affect money, contracts, personal data, or people's rights.
The buyer's conversation with a supplier should move past a single accuracy percentage. What work was measured? Which data period was used? Who reviewed the test set? Did it include cases where the system should refuse? Does the result still apply after the latest model or data change?
Proportionate evaluation does not require a frontier lab
A smaller team does not need to copy the structure of a frontier-model company. The cost and depth of evaluation should follow the potential impact.
An internal drafting assistant whose output is always reviewed before sending may begin with a focused test set covering real tasks, feared errors, and refusal cases. A system that approves payments, influences decisions about people, or accesses sensitive information needs stronger separation of duties, more detailed evidence, and clearer stop authority.
“Focused” should not become an invented universal test count. The necessary coverage depends on how varied the work is and how costly failure could be. The aim is to expose material risks, not to inflate the number of tests for a heavier report.
Passing a pre-production review is also not the end. NIST says AI systems should be tested before deployment and regularly while in operation, and that metrics and existing controls should be reassessed. NIST AI RMF Core Models, data sources, permissions, users, and business contexts can change even when the interface looks the same.
A compact AI acceptance checklist
Before approving production use, ask the process owner, evaluator, and developer to answer these questions from the same body of evidence:
- Is the task, user scope, and unacceptable harm defined?
- Can the team identify the model version, data sources, permissions, and integrations used in the test?
- Does the test set cover real work, stale data, unauthorized access, and cases that should be refused?
- Is the evaluator sufficiently separate from the builder for the risk involved, with enough access to challenge the result?
- Are acceptance criteria, exceptions, and the accountable decision owner recorded?
- Are re-test triggers, incident reporting, and the authority to pause or roll back the system defined?
If an answer is missing, completing the form is not the remedy. Find the unexamined risk, the unclear owner, or the evidence that the team cannot yet collect. Complete paperwork does not make an AI system safe by itself.
The practical takeaway is to give evaluators appropriate distance from the builder, access to the evidence they need, and a link to real decision authority. Evaluators can make accountability more verifiable. The business choosing to use the AI remains responsible for its decision.