Institute for Operational Assurance

Applied methodology

Does your AI agent operate as intended?

An AI agent can perform a task efficiently while still performing the wrong task, using the wrong information, exceeding its authority or failing to escalate when human judgement is required.

An AI Agent Operational Assurance Audit examines whether an agent’s actual behaviour remains aligned with its intended purpose, authority, requirements and operational boundaries—and whether effective Assurance Points exist to identify and manage material divergence.

Agent Operational Intention

Execution Gap

Agent Operational Reality

The agent’s Execution Gap is the difference between its Operational Intention and Operational Reality.

The operational risk

AI increases both the speed of execution and the speed of divergence.

AI agents are moving beyond generating information. They can retrieve data, use tools, make recommendations, initiate actions, coordinate tasks and participate in consequential workflows.

Traditional AI governance may define policies, approve technology and manage technical or regulatory risk. That does not by itself establish whether an agent behaves as intended during real work.

Operational Assurance addresses the execution question:

Is this agent doing what the organisation intends, within the authority and boundaries it has been given, under the conditions in which it actually operates?
  1. 01

    Purpose drift

    The agent’s behaviour moves away from its intended role or outcome.

  2. 02

    Authority drift

    The agent acts beyond its permissions, autonomy or delegated decision rights.

  3. 03

    Context failure

    The agent uses incomplete, outdated, conflicting or inappropriate information.

  4. 04

    Escalation failure

    The agent continues autonomously when human judgement, approval or intervention is required.

The audit

Compare intended behaviour with actual behaviour.

An AI Agent Operational Assurance Audit is a structured, evidence-led assessment of whether an AI agent’s actual behaviour remains aligned with its Operational Intention and whether adequate Assurance Points exist to identify and manage material divergence.

Agent Operational Intention

  • Purpose
  • Authority
  • Information
  • Permissions
  • Boundaries
  • Escalation

Compare

Execution Gap

Agent Operational Reality

  • Decisions
  • Actions
  • Tool use
  • Handoffs
  • Exceptions
  • Outcomes
Agent Operational Intention is compared with Agent Operational Reality to identify the Execution Gap.

For an AI agent, Operational Intention includes

  • its intended purpose and outcomes
  • assigned role and responsibilities
  • permitted and prohibited actions
  • decision authority and autonomy
  • approved information sources
  • tool and system access
  • required human approvals
  • exception and escalation requirements
  • evidence and accountability expectations

Operational Reality is what the agent actually does across real tasks, decisions, tool use, handoffs, exceptions and interactions with people or other systems.

The audit compares the two to identify the agent’s Execution Gap.

Audit scope

The conditions governing operational behaviour.

  1. 01

    Purpose and intended outcome

    Is the agent’s role clearly defined? Is success expressed as an operational outcome rather than only a technical task?

  2. 02

    Instructions and context

    Does the agent receive clear, current and relevant instructions? Can conflicting requirements be identified and resolved?

  3. 03

    Information and sources

    Which information may the agent use? Are sources approved, current, accessible and appropriate to the task?

  4. 04

    Authority and permissions

    What may the agent decide, recommend, change, approve or initiate? Are prohibited actions and autonomy boundaries explicit?

  5. 05

    Human oversight and escalation

    When must the agent pause, request approval or escalate? Is the responsible human role clear and available?

  6. 06

    Behaviour and outcomes

    What does evidence from real or representative tasks reveal about the agent’s decisions, actions, exceptions, handoffs and results?

  7. 07

    Evidence and accountability

    Can the organisation reconstruct what the agent did, why it acted, which information it used and where human involvement occurred?

The method

Understand. Observe. Compare. Assure. Adapt.

  1. 01UnderstandOperational Intention
  2. 02ObserveOperational Reality
  3. 03CompareThe Execution Gap
  4. 04AssureMoments that matter
  5. 05AdaptLearn and improve
Adapt feeds evidence and learning back into Understand
Continuous Operational Assurance moves through Understand, Observe, Compare, Assure and Adapt. Adapt then feeds learning back into Understand.
  1. 01

    Understand

    Establish the agent’s Operational Intention.

    Define its purpose, intended outcome, authority, permitted sources, tool access, boundaries, human approvals, escalation requirements and expected evidence.

  2. 02

    Observe

    Establish the agent’s Operational Reality.

    Examine permitted evidence from actual or representative work, including instructions, inputs, outputs, decisions, actions, tool use, handoffs, exceptions and human interventions.

  3. 03

    Compare

    Identify the agent’s Execution Gap.

    Determine where actual behaviour materially diverges from intended purpose, requirements, authority or operational boundaries.

  4. 04

    Assure

    Evaluate and design Assurance Points.

    Determine where the agent’s behaviour should be guided, verified, constrained, approved, interrupted or escalated. Recommend the Minimum Effective Intervention justified by the evidence.

  5. 05

    Adapt

    Improve and reassess.

    Update the agent, its instructions, permissions, sources, surrounding workflow or original Operational Intention. Reassess after material changes and learn from subsequent outcomes.

Evidence-led assessment

Assess the agent as it operates—not only as it was designed.

Design documentation and governance approvals establish intended conditions. They do not prove how the agent behaves during real execution.

Subject to permissions, scope and availability, the audit may examine:

  • agent purpose and role documentation
  • system prompts and governing instructions
  • knowledge sources and retrieval configuration
  • permissions and tool access
  • workflow and integration design
  • approval and escalation rules
  • representative tasks and test scenarios
  • decision and action logs
  • exception and failure cases
  • human-agent handoffs
  • observed outcomes
  • previous incidents or remediation activity

Evidence requirements should remain proportionate to the agent’s operational authority, risk and impact.

Where assurance happens

Place assurance around consequential agent behaviour.

The audit identifies the material moments where an AI agent’s activity can be evaluated against Operational Intention and, where necessary, influenced.

Potential Assurance Points may exist

  • before an agent accepts a task
  • when it selects information or a tool
  • before a consequential decision or action
  • when confidence or available evidence is insufficient
  • when a policy, authority or permission boundary is reached
  • during a human-agent handoff
  • when an exception occurs
  • before an external commitment is made

An Assurance Point does not always require human approval. Depending on the evidence and risk, it may provide context, verify a condition, restrict an action, record evidence or trigger escalation.

  1. 1Task received
  2. APContext gathered
  3. APDecision
  4. APAction
  5. 5Outcome
Illustrative—not a universal control design.

Proportionate response

Strengthen assurance without defaulting to maximum restriction.

Where the audit identifies a material Execution Gap, the response should be proportionate.

A Minimum Effective Intervention may include:

  • clarifying the agent’s purpose or instructions
  • improving access to current, approved information
  • removing inappropriate information or tool access
  • narrowing permissions or autonomy
  • introducing verification before a consequential action
  • improving confidence thresholds
  • requiring human approval in selected circumstances
  • strengthening exception detection or escalation
  • improving evidence and logging
  • redesigning the surrounding workflow where lesser intervention is insufficient

The objective is not to remove useful autonomy. It is to ensure that autonomy remains bounded, observable and aligned with Operational Intention.

Explore the Principles of Operational Assurance

Potential outputs

A practical basis for assurance and remediation.

  1. 01

    Agent Operational Intention Profile

    The agent’s purpose, intended outcomes, authority, permissions, boundaries, sources, approvals and escalation requirements.

  2. 02

    Agent Operational Reality Map

    An evidence-led account of how the agent behaves across relevant tasks, decisions, actions, exceptions and handoffs.

  3. 03

    Execution Gap Register

    A prioritised record of material divergence between intended and observed behaviour.

  4. 04

    Assurance Point Register

    The moments where the agent’s behaviour should be guided, verified, constrained, approved, evidenced or escalated.

  5. 05

    Minimum Effective Intervention Roadmap

    Prioritised recommendations for improving alignment without introducing unnecessary friction or control.

Optional outputs may include

  • human-agent responsibility map
  • evidence and observability requirements
  • remediation priorities
  • readiness conditions for deployment or expanded autonomy
  • continuous-assurance requirements
  • recommended reassessment triggers

Outputs depend on the agreed scope and available evidence. The audit does not produce an invented audit score, certification badge or pass/fail rating.

Audit triggers

Assess before deployment—and when the operating conditions change.

  1. 01

    Before deployment

    Assess whether the agent’s intended role, boundaries and Assurance Points are sufficiently defined before operational use.

  2. 02

    Before expanding autonomy

    Evaluate whether evidence supports giving the agent additional tools, authority or decision rights.

  3. 03

    After a material change

    Reassess when instructions, models, data sources, integrations, permissions or workflows change materially.

  4. 04

    After an incident or unexpected outcome

    Reconstruct what happened, identify the Execution Gap and strengthen the relevant Assurance Points.

The appropriate reassessment cadence depends on the agent’s authority, risk, rate of change and operational impact.

Complementary assurance

Operational behaviour is one part of responsible AI.

An AI Agent Operational Assurance Audit complements—not replaces—other forms of AI assurance.

  1. 01

    AI governance

    Establishes policies, accountabilities and organisational requirements.

  2. 02

    Technical and model evaluation

    Examines performance, robustness, security and technical behaviour.

  3. 03

    Privacy, legal and regulatory assessment

    Examines applicable obligations and legal risk.

  4. 04

    Operational Assurance

    Examines whether the agent behaves as intended while participating in real work.

The distinctive question is not only whether the AI system is technically capable or formally approved. It is whether the agent operates as intended within the workflow in which it has been placed.

The audit does not claim authority over legal, regulatory, cybersecurity or Internal Audit conclusions.

What happens next

A point-in-time audit can establish the foundation. Continuous assurance maintains it.

The audit identifies the agent’s Operational Intention, Operational Reality, material Execution Gaps and required Assurance Points.

Those findings can then support ongoing Operational Assurance as the agent performs work, operating conditions change and evidence accumulates.

  1. 01

    Audit

    Establish intention, reality and gaps.

  2. 02

    Assurance design

    Define Assurance Points and Minimum Effective Interventions.

  3. 03

    Continuous Operational Assurance

    Observe, compare, assure and adapt during execution.

A bounded operational assessment

What the audit does—and does not claim.

The audit is

  • an evidence-led assessment of operational alignment
  • focused on a defined AI agent or agent-enabled workflow
  • based on explicit criteria and permitted evidence
  • concerned with both agent behaviour and its operating environment
  • intended to support practical assurance and remediation

The audit is not

  • a certification
  • a guarantee that an agent cannot fail
  • a formal Internal Audit opinion
  • a legal or regulatory determination
  • a cybersecurity penetration test
  • a complete model-safety evaluation
  • continuous surveillance of every action

Begin with one agent

Can you demonstrate that the agent operates as intended?

Start with one AI agent performing a consequential, recurring and bounded role. Establish its Operational Intention, examine its Operational Reality and identify the Assurance Points required to keep the two aligned.