AI Agent Guardrails: Human Approval and Failure Recovery
A practical control system for deciding what an agent may do, what a person must approve and how the workflow returns to a safe state when something fails.
Key takeaways
- Permissions decide what the system can do; instructions only tell the model what it should do.
- Guardrails should cover inputs, outputs and individual tool calls.
- Approval belongs immediately before a consequential side effect, with enough context for a real decision.
- Retries need limits and idempotency; an unknown outcome is not a failed action that can safely be repeated.
- Every production workflow needs a human queue, action log and tested recovery path.
What are AI agent guardrails?
AI agent guardrails are checks and limits placed around the model and the tools it can use. An input guardrail can reject missing or suspicious data. A tool guardrail can inspect a proposed action before execution and validate the returned result. An output guardrail can check the final answer for required evidence, format or policy.
Guardrails are one layer. They do not replace authentication, authorisation or database constraints. If an agent has a tool capable of deleting every customer record, “do not delete important data” is not an adequate permission model. The application should expose only the narrow action and records the workflow needs.
Which controls belong in the system?
A practical design uses five layers. First, identity and permissions establish who initiated the task and which resources are available. Second, input validation checks structure, provenance and required fields. Third, policy guardrails evaluate the proposed decision and tool arguments. Fourth, a human gate pauses high-impact actions. Fifth, verification and recovery determine what really happened and what to do next.
| Action | Default control | Reason | Safe alternative |
|---|---|---|---|
| Read an approved product document | Automatic, logged | Reversible, low impact | Deny access outside the approved collection |
| Create an internal draft | Automatic with output validation | No external side effect | Mark as draft and cite source fields |
| Update a low-risk CRM field | Policy check; approval depends on field | May affect operations | Write to a review queue or staging field |
| Send, publish or make a commitment | Human approval before execution | External and reputational consequence | Prepare a preview only |
| Spend money, change access or delete data | Human approval plus strong authorisation | Financial, security or irreversible impact | Propose the change; do not execute |
| Medical, legal or hiring judgement | Human specialist decision; agent cannot decide | Consequential and context-sensitive | Summarise evidence without advice or selection |
Where should human approval happen?
Place the approval gate after the agent has prepared a concrete action but before the tool creates the side effect. The reviewer should see the action type, target, important arguments, source evidence, policy result and what will happen after approval. A vague “allow agent?” prompt does not support an informed decision.
OpenAI’s Agents SDK documents a pause-and-resume pattern: a tool call can surface as an interruption, a person approves or rejects it, and the run continues from stored state. Its approval rules fail closed when arguments cannot be safely inspected. The implementation will vary by platform, but the principle is portable: unclear input must not silently become permission.
Figure 1. A controlled action path
Text equivalent: a request passes identity and input validation, policy classification and an approval decision. A low-risk action may execute automatically; a consequential action pauses for a person. After execution, the system verifies the external state and logs completion. Invalid, rejected, failed or ambiguous cases stop or enter a human queue.
How do you design guardrails in seven steps?
- Map actions and consequences. List every read, write, send, publish, purchase and delete operation. For each one, record the target, affected person, reversibility and worst credible outcome.
- Set permissions outside the prompt. Give the agent the smallest useful tools, scopes and data access. Separate read from write, and expose narrow business operations instead of broad administrator access.
- Validate inputs. Require a schema, source and minimum fields. Quarantine malformed, unauthorised or suspicious content. Treat retrieved documents and tool output as untrusted data, not new instructions.
- Place approval before side effects. Define which risk classes always pause. Show the reviewer a stable preview, and bind approval to the exact action and arguments so later changes require a new decision.
- Verify outputs and tool results. Check required fields, evidence and policy. After a tool runs, confirm the external record rather than trusting a generated success message.
- Define recovery states. Distinguish retryable errors, permanent errors, rejected actions and unknown outcomes. Give each state an owner, time limit and next action.
- Test and monitor. Run normal, boundary, adversarial and failure cases before release. Monitor approval rates, safe stops, duplicate attempts, unresolved incidents and changes in source data.
How should an agent recover from failure?
A failure handler needs more precision than “try again.” A timeout may mean the tool did nothing, or it may mean the action succeeded but the response never arrived. Repeating a payment, email or database write could create a duplicate. Use a unique operation ID, check the destination state and retry only when the operation is designed to be idempotent.
| Observed state | Response | Do not do |
|---|---|---|
| Missing required data | Request the specific field; preserve the draft | Invent a value |
| Policy or guardrail rejection | Stop and explain the rule to the operator | Rephrase to bypass the control |
| Human rejects approval | Cancel that exact action and record the decision | Ask repeatedly or switch tools |
| Transient tool error with confirmed no-write | Retry with backoff within a strict limit | Retry forever |
| Timeout or ambiguous result | Query by operation ID; escalate if still unknown | Assume failure and repeat a side effect |
| Partial multi-step update | Run a tested compensation or human runbook | Hide the partial state |
| Repeated or novel failure | Open an incident, stop the workflow and preserve logs | Let the model improvise a production fix |
What should you test before release?
Use a fixed test set with expected states, not only examples that should succeed. The system passes when it produces the expected result or the expected safe stop without an unapproved side effect.
| Test | Expected result | Evidence to capture |
|---|---|---|
| Valid low-risk request | Completes automatically | Input, tool call, verified result and log |
| Missing or malformed fields | Stops before tool execution | Validation reason and zero writes |
| Instruction hidden in retrieved content | Treated as data; policy remains unchanged | Source, blocked instruction and chosen action |
| Consequential action | Pauses before execution | Approval preview and exact arguments |
| Rejected approval | Action cancelled | Decision, reviewer and zero side effects |
| Tool timeout after possible write | No blind retry; state checked | Operation ID and destination lookup |
| Duplicate request | One outcome only | Idempotency record |
| Maximum attempts reached | Safe stop and human escalation | Attempt count, last error and queue item |
Criteria for success
- No action can exceed the initiating user’s permission.
- Every consequential side effect is approved against exact arguments.
- Invalid and adversarial inputs stop before a protected tool runs.
- Every external action has a verifiable result and audit record.
- Retries cannot create duplicate side effects.
- Unknown outcomes and repeated failures reach a named human owner.
- Visible FAQ, runbook and monitoring rules match the implemented behaviour.
Learn to make one AI workflow reliable
Barcelona Code School’s AI Agents Builder course teaches professionals to turn a real work process into a controlled system. Students define inputs and outputs, connect approved tools, handle bad inputs, add human approval and test what can go wrong.
The course is for practical workplace automation, not a promise of risk-free autonomy. Check the live page for the current format, dates and tuition.
Explore the AI Agents Builder courseFrequently asked questions
What are guardrails for an AI agent?
AI agent guardrails are checks and limits around inputs, outputs and tool calls. They can validate data, block prohibited actions, require evidence, constrain formats and stop a run. They work best with real permissions, approval gates, logging and recovery paths.
Which AI agent actions need human approval?
Require approval for actions with financial, legal, privacy, reputational or operational consequences: sending important messages, publishing, spending money, deleting data, changing access, making commitments or acting on ambiguous evidence. The exact boundary depends on the business.
Is human approval enough to make an AI agent safe?
No. A reviewer can miss problems, and an approval screen may omit important context. Use least-privilege permissions, input and output validation, application-level authorisation, logs, post-action verification and safe recovery as well as human approval.
What should happen when an AI agent fails?
The agent should record the failure, avoid repeating unsafe side effects, classify the outcome and either retry within a strict limit, request missing information, compensate through a tested action or escalate to a person. Ambiguous outcomes should not be treated as success.
How do you test AI agent guardrails?
Test valid work, missing fields, malformed data, prompt injection, excessive permissions, rejected approvals, tool timeouts, duplicate requests, partial writes and ambiguous responses. Verify that each case reaches its expected result, safe stop or human queue without an unapproved side effect.