The difficult question is not whether an AI agent can act. It is what it may decide, what evidence it must keep and when it must stop.
The difficult question about an AI agent is not whether it can act. It is what it is allowed to decide.
Most demonstrations avoid that question. The agent reads an email, looks something up, drafts a response and completes a task. It feels autonomous because the path is smooth. In a real business, the path is not smooth. A supplier misses a deadline. Two systems disagree. A threshold is breached. The evidence is incomplete. Someone needs to decide whether to continue, escalate or stop.
That is where agency becomes an operating-model question rather than a technology feature.
The architecture is less new than the language
An interesting procurement case revisited in 2026 describes software agents, distributed fulfilment decisions, monitoring and progressive escalation in a patent filed in 2000. The valuable lesson is not that somebody predicted the current market. It is that the underlying loop already worked without a large language model.
The system could find an appropriate fulfilment source, initiate the next step, watch for the expected response and escalate when it did not arrive. Modern AI can read messier information and handle broader context, but it did not invent the need for clear states, thresholds, ownership and feedback.
Capability was not the whole constraint then. It is not the whole constraint now.
Four levels of decision right
I find it useful to separate an agent's authority into four levels.
- Read. The agent may retrieve information, classify it and assemble context. It cannot change the underlying process.
- Recommend. The agent may propose a decision and show its evidence. A named person remains the approver.
- Act. The agent may execute within explicit limits: value, confidence, data source, customer type, time window or risk class.
- Escalate. The agent must stop and route the case when evidence is missing, rules conflict, a threshold is breached or the expected result does not arrive.
“Human in the loop” is too vague on its own. Which human, at which point, reviewing what evidence, within what time? Without those answers, it is reassurance rather than control.
Our earlier piece on being in the loop or on the loop covers the same distinction from an analytics governance perspective.
Put a deterministic spine underneath the agent
Language models are useful where the world is messy: interpreting an email, matching a service note to a case, summarising a contract or recognising that two descriptions probably refer to the same item. Financial controls, approval thresholds and posting rules are different. They need precision.
The practical design is often probabilistic at the edges and deterministic through the spine. Let AI interpret the unstructured signal. Pass the result into governed business rules. Log the evidence, the rule applied, the action taken and the next expected state.
This is close to the strongest point in Workday's CFO buyer guide: context, governance and explainability have to travel with the action. I would make the point vendor-neutral. An agent should inherit least-privilege access, operate on the same governed definitions as the process, and leave evidence an auditor or process owner can follow.
Design for absence, not just activity
Many systems can react when something happens. The more interesting control is noticing when something should have happened and did not.
A shipment confirmation fails to arrive. An invoice remains unmatched. A forecast owner has not responded. A data pipeline completes but the expected volume is missing. Those are absences. They require the system to know the expected state, the time allowed and the escalation route.
This is why an agent architecture begins with the operating loop, not the chat interface:
- 1What event starts the work?
- 2What evidence is required?
- 3What decision may the agent make?
- 4What state should exist afterwards?
- 5How long should that take?
- 6What happens if the state never appears?
If those answers are unclear, the agent is not ready for autonomy.
The audit trail is part of the product
For every material action, retain the input, the relevant business context, the rule or policy used, the decision, the confidence where applicable, the actor identity and the result. Do not treat this as a compliance export to be added later. It is how the business diagnoses failure, improves performance and earns permission to widen the agent's authority.
The same discipline underpins AI built on trusted data. An agent acting on a disputed measure or an unexplained rule is not autonomous. It is unsupervised.
Start narrower than the demo suggests
Begin with one bounded decision where the evidence is available, the rules are understood and the cost of an exception is manageable. Give the agent read and recommendation rights first. Observe the error modes. Introduce action rights only for cases that meet a clear threshold, and make escalation the default for everything else.
Autonomy should be earned with evidence, not declared in a roadmap.
The most useful agent will not be the one that appears most intelligent in a demonstration. It will be the one that knows what it may do, what it must prove and when it must stop.
If you are moving an agent from prototype into a process that matters, talk to us. We can help define the data, controls and decision rights before the interface gets ahead of the operation.
Related reading
- Context is the bottleneck, not the model
- AI and analytics for the mid-market
- Applied AI and machine learning
Simon Devine
Founder, Hopton Analytics
Part of the Hopton Analytics team, delivering governed analytics programmes for UK mid-market organisations.
