When an answer becomes an action
Ask an AI tool to draft an email and a person can read it before anything happens. Ask an AI agent to send the email, update the customer record, schedule a follow-up, and alert another team, and the accountability question changes.
The issue is no longer limited to whether the output was accurate. Someone also needs to understand whether the action was allowed, whether it made sense in the circumstances, and who owns the result.
Organizations are already testing that boundary. McKinsey’s 2025 State of AI survey found that 62 percent of respondents said their organizations were at least experimenting with AI agents. In any single business function, however, no more than 10 percent reported scaling them. Plenty of companies are curious. Far fewer have learned what it feels like when an agent becomes part of daily work.
No agent acts alone
“Autonomous” can sound as if the software has stepped outside human control. It has not. People choose the systems an agent can reach, the data it can read, the credentials it can use, and the goals it is asked to pursue. Vendors make another set of choices about how the product behaves.
NIST underscored this point when it launched the AI Agent Standards Initiative. The agency describes agents as capable of autonomous action while also emphasizing that they operate on behalf of users.
That phrase—“on behalf of”—deserves attention. An employee might start the task, a system owner might grant access, and an executive might sponsor the use case. The resulting action still goes out under the organization’s name.
A permission setting tells software what it can reach. It does not settle what the software has the business authority to decide.
One action can have several owners
Imagine that an agent sends the wrong notice to a customer. The error could begin with an ambiguous instruction, an unsuitable model response, excessive system access, a missing policy exception, or a product change made by a vendor. Several causes can overlap.
Inside the organization, responsibility is split among the people who selected, configured, approved, and used the system. The customer sees something simpler: the company acted.
A technically precise account of every component can still leave leaders without a clear answer about who was expected to prevent the event, who could have stopped it, and who now has to explain it.
Access can add up
Identity and access systems were designed to connect actions to people and software. Agents make that job harder because they can decide how to pursue a goal after access has been granted.
The NCCoE’s 2026 concept paper examines how identity standards and best practices could apply to software and AI agents. It calls out identification, authorization, auditing, and non-repudiation as areas that need attention.
An agent might begin with an employee’s delegated access, call another service under a shared account, and then write to a third system with a different set of rules. Each permission can look ordinary on its own. Together, they can give the agent practical authority that no single application was built to evaluate.
A login can identify the connection. It says nothing about whose judgment the agent was carrying out.
Accuracy is only part of judgment
An agent can follow an instruction exactly and still make a poor business decision.
An experienced employee might recognize that a customer has already been promised an exception, that the timing of a message is unusually sensitive, or that a policy written for routine cases does not fit this one. Those details often live in conversation, history, and judgment rather than in a field the agent can read.
Organizations have always tried to turn judgment into repeatable processes. AI agents bring the gaps into sharper focus because software can now move through steps that once forced a person to pause and interpret the situation.
Completing the task correctly and handling the situation well are not always the same achievement.
Human review can be real or ceremonial
“Human in the loop” sounds reassuring, but it can describe very different working arrangements.
One person might approve every action in advance. Another might review a weekly sample. Someone else might receive an alert only when the agent recognizes an exception. In each case a human is involved, but the level of supervision is plainly different.
Volume changes the picture too. A reviewer who can study ten actions might only glance at a queue of five hundred. On the other hand, forcing approval for every routine, reversible step can remove much of the benefit the agent was meant to provide.
Meaningful oversight depends on what the person can see, how much time they have, whether they understand the stakes, and whether they can actually stop the action. Merely placing a name next to the process does not create control.
The customer does not see the vendor stack
An enterprise agent can combine a model, an orchestration tool, a business application, an identity provider, internal data, and outside implementation work. Contracts divide responsibility across those providers. The person affected by an action rarely sees those boundaries.
If a customer receives an improper message or an employee record changes incorrectly, pointing to one supplier’s terms does not explain why the organization allowed the action. The vendor relationships matter, but they do not replace the organization’s own accountability.
Leaders need a coherent account of what happened even when the technology was assembled from many parts. A collection of technically correct, vendor-specific explanations may still fall short of that.
Logs show activity, not always reasoning
After an incident, a log might confirm that an application was called, a record changed, or a message was sent. It might not show which information shaped the agent’s plan, how it interpreted the instruction, which policy applied at the time, or when a person had a chance to intervene.
NIST’s May 2026 analysis of public responses on AI agent security found broad agreement among commenters that agents introduce new security threats and that established cybersecurity practices will need to adapt.
The report’s focus is practical: existing controls need to adapt when model output can trigger real actions. A conventional audit trail might be too thin. Leaders may need to explain the sequence of technical events and why the organization permitted that sequence in the first place.
Accountability follows the use, not the label
The same agent technology can summarize public research in one department and alter sensitive customer records in another. Calling both uses “agentic AI” says little about the responsibility attached to each one.
The stakes change with the data involved, the action’s reversibility, the time available for intervention, and the harm an error could cause. Formal duties will dominate in some settings. In others, the harder issue will be limited staff knowledge, heavy vendor dependence, or ownership split across several teams.
That is why a single accountability model is unlikely to fit every agent or every organization. The label is broad; the authority granted in a specific use is what matters.
Before the agent acts on the organization’s behalf
AI agents do not make human responsibility disappear. They spread pieces of authority across users, system owners, vendors, policies, and software, then carry the result into the real world.
Those relationships are easier to overlook during a controlled experiment. They become harder to ignore once an agent touches more systems, more data, and more people.
Before asking whether an agent can be trusted to act, leadership has a more basic question to answer: which decisions has the organization allowed it to make, and who will stand behind those decisions when the outcome affects someone?