When staff need to search several systems to help a customer, an AI assistant may save time spent opening documents one by one. But before an organization expands that assistant beyond its first team, “How many people use it?” is only one question. It should also ask which parts of the work changed, who handles mistakes, and whether the improvement justifies the cost.
On October 1, 2026, Anthropic announced that Barclays would expand its use of Claude across operations, software development, and legacy modernization. The bank expects Claude Code adoption to reach half of its developer population by the end of 2026 and a majority of software engineers by 2027. Anthropic’s announcement also says that Barclays’ Colleague Knowledge Assistant has been live since 2025, has been adopted by more than 16,000 colleagues, and has handled over one million searches. In Global Markets, teams use Claude to classify, enrich, and route about 120,000 incoming emails per day.
These figures were published by Anthropic in a partnership announcement with Barclays. They are not an independent evaluation, and they do not, by themselves, show how much money was saved or whether customers became more satisfied. Read them as an example of AI connected to real work, not as a target another organization should match.
Measure the workflow, not just the number of accounts
The announcement describes two different kinds of work. In one, employees search for knowledge to answer customer questions. In another, a system helps classify and route emails in Global Markets. Neither example means that AI makes every decision or completes every transaction for an employee.
For an organization testing AI, define the unit of measurement as a specific workflow: “search an approved knowledge base, then send the answer to a staff member for review” or “classify an incoming email before routing it to the responsible team.” Record where the task starts and ends, who uses it, what data the system can see, which decisions remain with people, and which cases should be handed back for review.
If the only measure is how many employees have access to an AI tool, the team still does not know whether the original work became faster, whether answer quality changed, or whether difficult cases were simply shifted to other employees.
Set a baseline before comparing results
Before introducing AI to a team, capture the existing workflow over a period that represents ordinary workload. Useful starting measures may include incoming volume, time from receipt to completion, the share of items corrected or handed off, and reasons work gets delayed. After the pilot begins, keep collecting the same measures and add ones relevant to AI, such as answers staff can use without editing, emails routed to the wrong team, items requiring a second review, system cost, and time spent resolving exceptions.
Choose measures based on the business outcome the team needs. It does not need to track everything at once. If the goal is faster knowledge search, search time is only one part: the answer must also be accurate and usable by the employee. If the goal is email routing, track misroutes and the time spent handling cases the system cannot confidently classify.
These measures are suggestions for an organization’s own pilot, not results disclosed by Barclays. Do not compare numbers from different workflows or periods with different workload conditions as if they were equivalent.
Decide when the system hands work back to a person
A live workflow will encounter cases where an AI system should not continue: conflicting information, missing documents, a customer question outside the knowledge base, or an email with financial consequences or implications for customer rights. Teams should define in advance when the AI must stop, who takes the case, what information that person needs, and where the reason and review outcome will be recorded.
Start with reading, searching, or classification tasks that can be reversed. Changes to records, external messages, approvals, or consequential status updates need review and approval appropriate to their risk. Limit system permissions to what the workflow needs. Requiring staff to check every low-risk case can make the cost outweigh the benefit; letting every case run automatically can hide consequential errors. The handoff design should reflect impact and reversibility.
Use risk management as a review cycle
NIST identifies the AI Risk Management Framework Core page as an excerpt from AI RMF 1.0 (2023) and says a revised version is in progress. It organizes risk work under Govern, Map, Measure, and Manage, which NIST describes as iterative and context-dependent rather than a checklist or required sequence. NIST says the framework and Playbook are voluntary, so organizations can select the parts that fit their resources and risk level (NIST AI RMF Core).
For a small AI pilot, that may mean assigning a process owner and someone with authority to stop the work; understanding users, data, and impacts before selecting an AI system; testing quality in conditions close to real use; and reviewing failures after launch. If results miss a threshold set by the team or the workflow changes, the team should be able to adjust the system, narrow its scope, or stop it.
Before expanding beyond the first team
- Where does this task start and end, and who owns the final result?
- What information is required, and can access be limited to that information?
- How will the team compare the pilot with its existing baseline without using prompt counts as a proxy for work quality?
- Which errors require the AI to stop, who receives the handoff, and who can approve the next step?
- If quality declines, costs rise, or the process changes, how will the team narrow or stop the deployment?
Barclays’ adoption and email figures show the scale of work it is addressing. For a Thai team, a practical starting point is to define one workflow, measure quality and human handoffs, then decide whether to expand. If your team could pilot one process this quarter, which would you choose, and who would own its measures and approvals?