Barcelona Code School

Since 2015 / 500+ graduates

AI Agent Guardrails: Human Approval and Failure Recovery

Reliable AI automation

AI Agent Guardrails: Human Approval and Failure Recovery

A practical control system for deciding what an agent may do, what a person must approve and how the workflow returns to a safe state when something fails.

Published 28 September 2026 · Barcelona Code School

Reliable AI agents use several controls, not one safety prompt. Give the agent least-privilege access, validate inputs and tool calls, require human approval before consequential side effects, verify the real result and define recovery states for every failure. If the system cannot establish what happened, it should stop, preserve the evidence and escalate instead of guessing or repeating the action.

Key takeaways

  • Permissions decide what the system can do; instructions only tell the model what it should do.
  • Guardrails should cover inputs, outputs and individual tool calls.
  • Approval belongs immediately before a consequential side effect, with enough context for a real decision.
  • Retries need limits and idempotency; an unknown outcome is not a failed action that can safely be repeated.
  • Every production workflow needs a human queue, action log and tested recovery path.

What are AI agent guardrails?

AI agent guardrails are checks and limits placed around the model and the tools it can use. An input guardrail can reject missing or suspicious data. A tool guardrail can inspect a proposed action before execution and validate the returned result. An output guardrail can check the final answer for required evidence, format or policy.

Guardrails are one layer. They do not replace authentication, authorisation or database constraints. If an agent has a tool capable of deleting every customer record, “do not delete important data” is not an adequate permission model. The application should expose only the narrow action and records the workflow needs.

A prompt is guidance for the model. A permission is an enforced boundary. Use both, but never confuse them.

Which controls belong in the system?

A practical design uses five layers. First, identity and permissions establish who initiated the task and which resources are available. Second, input validation checks structure, provenance and required fields. Third, policy guardrails evaluate the proposed decision and tool arguments. Fourth, a human gate pauses high-impact actions. Fifth, verification and recovery determine what really happened and what to do next.

ActionDefault controlReasonSafe alternative
Read an approved product documentAutomatic, loggedReversible, low impactDeny access outside the approved collection
Create an internal draftAutomatic with output validationNo external side effectMark as draft and cite source fields
Update a low-risk CRM fieldPolicy check; approval depends on fieldMay affect operationsWrite to a review queue or staging field
Send, publish or make a commitmentHuman approval before executionExternal and reputational consequencePrepare a preview only
Spend money, change access or delete dataHuman approval plus strong authorisationFinancial, security or irreversible impactPropose the change; do not execute
Medical, legal or hiring judgementHuman specialist decision; agent cannot decideConsequential and context-sensitiveSummarise evidence without advice or selection

Where should human approval happen?

Place the approval gate after the agent has prepared a concrete action but before the tool creates the side effect. The reviewer should see the action type, target, important arguments, source evidence, policy result and what will happen after approval. A vague “allow agent?” prompt does not support an informed decision.

OpenAI’s Agents SDK documents a pause-and-resume pattern: a tool call can surface as an interruption, a person approves or rejects it, and the run continues from stored state. Its approval rules fail closed when arguments cannot be safely inspected. The implementation will vary by platform, but the principle is portable: unclear input must not silently become permission.

Figure 1. A controlled action path

The safe path validates before action and verifies afterwards. Approval is a branch in the execution system, not a sentence in the prompt.

Text equivalent: a request passes identity and input validation, policy classification and an approval decision. A low-risk action may execute automatically; a consequential action pauses for a person. After execution, the system verifies the external state and logs completion. Invalid, rejected, failed or ambiguous cases stop or enter a human queue.

How do you design guardrails in seven steps?

  1. Map actions and consequences. List every read, write, send, publish, purchase and delete operation. For each one, record the target, affected person, reversibility and worst credible outcome.
  2. Set permissions outside the prompt. Give the agent the smallest useful tools, scopes and data access. Separate read from write, and expose narrow business operations instead of broad administrator access.
  3. Validate inputs. Require a schema, source and minimum fields. Quarantine malformed, unauthorised or suspicious content. Treat retrieved documents and tool output as untrusted data, not new instructions.
  4. Place approval before side effects. Define which risk classes always pause. Show the reviewer a stable preview, and bind approval to the exact action and arguments so later changes require a new decision.
  5. Verify outputs and tool results. Check required fields, evidence and policy. After a tool runs, confirm the external record rather than trusting a generated success message.
  6. Define recovery states. Distinguish retryable errors, permanent errors, rejected actions and unknown outcomes. Give each state an owner, time limit and next action.
  7. Test and monitor. Run normal, boundary, adversarial and failure cases before release. Monitor approval rates, safe stops, duplicate attempts, unresolved incidents and changes in source data.

How should an agent recover from failure?

A failure handler needs more precision than “try again.” A timeout may mean the tool did nothing, or it may mean the action succeeded but the response never arrived. Repeating a payment, email or database write could create a duplicate. Use a unique operation ID, check the destination state and retry only when the operation is designed to be idempotent.

Observed stateResponseDo not do
Missing required dataRequest the specific field; preserve the draftInvent a value
Policy or guardrail rejectionStop and explain the rule to the operatorRephrase to bypass the control
Human rejects approvalCancel that exact action and record the decisionAsk repeatedly or switch tools
Transient tool error with confirmed no-writeRetry with backoff within a strict limitRetry forever
Timeout or ambiguous resultQuery by operation ID; escalate if still unknownAssume failure and repeat a side effect
Partial multi-step updateRun a tested compensation or human runbookHide the partial state
Repeated or novel failureOpen an incident, stop the workflow and preserve logsLet the model improvise a production fix

What should you test before release?

Use a fixed test set with expected states, not only examples that should succeed. The system passes when it produces the expected result or the expected safe stop without an unapproved side effect.

TestExpected resultEvidence to capture
Valid low-risk requestCompletes automaticallyInput, tool call, verified result and log
Missing or malformed fieldsStops before tool executionValidation reason and zero writes
Instruction hidden in retrieved contentTreated as data; policy remains unchangedSource, blocked instruction and chosen action
Consequential actionPauses before executionApproval preview and exact arguments
Rejected approvalAction cancelledDecision, reviewer and zero side effects
Tool timeout after possible writeNo blind retry; state checkedOperation ID and destination lookup
Duplicate requestOne outcome onlyIdempotency record
Maximum attempts reachedSafe stop and human escalationAttempt count, last error and queue item

Criteria for success

  • No action can exceed the initiating user’s permission.
  • Every consequential side effect is approved against exact arguments.
  • Invalid and adversarial inputs stop before a protected tool runs.
  • Every external action has a verifiable result and audit record.
  • Retries cannot create duplicate side effects.
  • Unknown outcomes and repeated failures reach a named human owner.
  • Visible FAQ, runbook and monitoring rules match the implemented behaviour.

Learn to make one AI workflow reliable

Barcelona Code School’s AI Agents Builder course teaches professionals to turn a real work process into a controlled system. Students define inputs and outputs, connect approved tools, handle bad inputs, add human approval and test what can go wrong.

The course is for practical workplace automation, not a promise of risk-free autonomy. Check the live page for the current format, dates and tuition.

Explore the AI Agents Builder course

Frequently asked questions

What are guardrails for an AI agent?

AI agent guardrails are checks and limits around inputs, outputs and tool calls. They can validate data, block prohibited actions, require evidence, constrain formats and stop a run. They work best with real permissions, approval gates, logging and recovery paths.

Which AI agent actions need human approval?

Require approval for actions with financial, legal, privacy, reputational or operational consequences: sending important messages, publishing, spending money, deleting data, changing access, making commitments or acting on ambiguous evidence. The exact boundary depends on the business.

Is human approval enough to make an AI agent safe?

No. A reviewer can miss problems, and an approval screen may omit important context. Use least-privilege permissions, input and output validation, application-level authorisation, logs, post-action verification and safe recovery as well as human approval.

What should happen when an AI agent fails?

The agent should record the failure, avoid repeating unsafe side effects, classify the outcome and either retry within a strict limit, request missing information, compensate through a tested action or escalate to a person. Ambiguous outcomes should not be treated as success.

How do you test AI agent guardrails?

Test valid work, missing fields, malformed data, prompt injection, excessive permissions, rejected approvals, tool timeouts, duplicate requests, partial writes and ambiguous responses. Verify that each case reaches its expected result, safe stop or human queue without an unapproved side effect.

Sources

  1. OpenAI Agents SDK: Human-in-the-loop.
  2. OpenAI Agents SDK: Guardrails.
  3. OpenAI: A practical guide to building agents.
  4. NIST AI 600-1: Generative AI Profile.
  5. Barcelona Code School: How to Build an AI Agent.
  6. Barcelona Code School: AI Agents Builder.
Back to posts