Barcelona Code School

Since 2015 / 500+ graduates

How to Test AI Agents Before Production

AI agent evaluation guide

How to Test AI Agents Before Production

A good answer is only one test. Production readiness also requires correct tools, protected permissions, safe failures, repeatable evaluation and observable outcomes.

Published 5 October 2026 · Barcelona Code School

Test an AI agent as a complete system, not only as a model response. Build a versioned dataset of normal, edge, adversarial and failure cases; assert deterministic rules first; score language quality separately; inspect tool choice and arguments; verify approval and retry behaviour; then release gradually with monitoring and rollback. Every production incident should become a regression test.

Key takeaways

  • Define the job, risk boundary and pass criteria before collecting test cases.
  • Test output quality, tool behaviour, permissions and recovery as separate layers.
  • Use deterministic assertions whenever the correct result can be expressed exactly.
  • Keep a frozen regression set and a changing discovery set.
  • A staged release, monitoring and rollback plan are part of evaluation.

What exactly are you testing?

An agent is more than a prompt. It includes input handling, context retrieval, model reasoning, tool selection, tool arguments, permissions, approvals, external writes, retries, logging and human handoff. A beautiful answer can hide a wrong record update; a correct classification can still be unsafe if the agent used an unauthorised tool.

Start with a one-sentence contract: “Given these inputs and approved sources, the system returns this schema, may use these read tools, and must ask for approval before these writes.” Tests should map back to that contract.

LayerQuestionExample metric or assertion
InputIs data valid and authorised?Missing consent stops before model call
OutputIs the result useful and grounded?Schema pass, category accuracy, source coverage
ToolsDid it call the right tool correctly?Allowed tool, exact ID, valid arguments
PermissionsCould it exceed its role?No write without bound approval
RecoveryDoes failure return to a safe state?No duplicate action after timeout
OperationsCan owners observe and stop it?Trace completeness, alert time, rollback test

A practical evaluation pipeline

Figure 1. From contract to monitored release

Evaluation combines exact software checks with model-quality measurement and operational readiness.

Text equivalent: define the system contract and risks, run a representative dataset through exact assertions and quality graders, inspect every tool and approval path, simulate failures, enforce release thresholds and deploy gradually. Monitor real outcomes and turn failures into tests.

Test the AI agent step by step

  1. Write the contract and risk boundary. Name required inputs, output schema, approved sources, allowed read tools, approval-required writes, forbidden actions and the human owner for every exception.
  2. Build a representative dataset. Include common requests, rare valid cases, incomplete inputs, ambiguous language, conflicting records, multiple languages, adversarial instructions and historical failures. Remove or protect personal data.
  3. Add deterministic assertions. Test schema, enums, required IDs, arithmetic, policy gates, tool allowlists, approval state and duplicate prevention with exact checks. Do not ask another model to grade what code can prove.
  4. Score model quality. Measure task-specific dimensions such as category correctness, evidence coverage, retrieval relevance, groundedness, usefulness and appropriate refusal. Define the rubric and examples before comparing variants.
  5. Test tools and permissions. Inspect tool choice, arguments, record IDs, scopes and order. Confirm that text in a request or retrieved document cannot add a tool, change credentials or bypass approval.
  6. Inject failures and adversarial inputs. Simulate timeouts before and after writes, malformed responses, stale sources, conflicting data, rate limits, prompt injection, duplicate events and denied or expired approval.
  7. Run regression and staged release. Freeze a core set that every change must pass. Keep a separate discovery set for new edge cases. Release to shadow mode or a small traffic slice with a clear rollback trigger.
  8. Monitor and learn. Track business outcomes, unsafe-action attempts, handoffs, latency, cost, overrides and silent failures. Add every incident and surprising near miss to the regression set.

Build the dataset around decisions, not polished examples

A test set that contains only well-written happy paths will exaggerate quality. Group cases by the decision the system must make and the failure it must resist.

  • Normal: representative requests in expected formats.
  • Boundary: values immediately above and below thresholds.
  • Missing: absent IDs, consent, evidence or required fields.
  • Ambiguous: several plausible categories or records.
  • Adversarial: attempts to override instructions, exfiltrate data or gain tools.
  • Operational: API outages, partial writes, duplicated events and rate limits.
  • Policy: actions that must refuse, escalate or wait for approval.

Store the dataset version, prompt, model, workflow revision, tool versions and scorecard version with every run. Without configuration lineage, a better or worse score is difficult to explain.

Choose metrics that expose trade-offs

MetricWhat it revealsWatch out for
Task successWhether the requested job completedCan hide unsafe paths
Category precision and recallFalse accepts versus false rejectsOne accuracy number hides imbalance
GroundednessWhether claims follow supplied sourcesGood wording can still cite the wrong source
Tool correctnessRight tool, argument and recordFinal answer may look correct despite wrong call
Escalation qualityWhether a person receives usable contextToo many safe-looking handoffs can make system useless
Latency and costOperational sustainabilityOptimising them can reduce quality or safety

n8n's current evaluation documentation distinguishes small, visually reviewed evaluations during development from metric-based evaluation over larger datasets. Its metric workflow can score dimensions and compare runs; availability varies by plan, so confirm the current product documentation.

Release gates and rollback

Decide thresholds before looking at the final candidate. A release gate can require zero unauthorised writes, 100% approval compliance, complete trace IDs, minimum classification scores for critical routes and no regression beyond an agreed tolerance.

Do not average away a catastrophic failure. One unauthorised payment, privacy leak or account change can block release even when the overall score is high.

Start in shadow mode, draft-only mode or with a small internal group. Compare the agent with the existing process, sample outputs by risk level and prove that the system can be disabled without losing pending work.

Minimum pre-production test matrix

TestExpected resultRelease condition
Valid common requestCorrect result and traceMeets task threshold
Missing required identifierStops or asks specificallyNo guessed record
Prompt injectionIgnored as dataNo permission or tool change
Denied approvalNo actionDenial recorded
Timeout after write requestExternal-state verificationNo duplicate write
Conflicting sourcesConflict surfacedNo unsupported merge
Duplicate eventIdempotent resultOne external action
Monitoring disabledRelease blockedOwner, alerts and rollback ready

Success criteria

  • The dataset represents normal traffic, edge cases and the highest-impact risks.
  • Every exact requirement has a deterministic assertion.
  • Tool traces prove which records and permissions were used.
  • Approval, denial, timeout and retry paths are tested end to end.
  • Release thresholds and catastrophic blockers are written in advance.
  • Production monitoring has named owners and a tested rollback route.
  • Incidents and overrides continuously improve the regression set.

Build complete AI workflows, not isolated demos

A production-ready agent needs data contracts, tools, RAG where appropriate, permissions, human approvals, evaluation, failure recovery and monitoring. Barcelona Code School’s four-week AI Agent & Automation Bootcamp brings those pieces together in live, instructor-led projects.

If you only need one personal workflow, a free tutorial may be enough. The bootcamp is for people who want to design and explain connected automation systems for real work.

Explore the AI Agent & Automation Bootcamp

Frequently asked questions

How do you test an AI agent?

Define the complete system contract, build a representative dataset, add deterministic assertions, score task-specific model quality, inspect tools and permissions, inject failures, run regression tests and release gradually with monitoring and rollback.

What should be in an AI agent test dataset?

Include normal, boundary, missing-data, ambiguous, multilingual, adversarial, policy and operational-failure cases. Add historical incidents and surprising production examples after protecting personal data.

Should you use another AI model to grade agent outputs?

Use model-based grading for nuanced qualities when the rubric is clear, but prefer deterministic checks for schemas, enums, arithmetic, IDs, permissions and tool calls. Calibrate model graders against human-reviewed examples.

What is the difference between an eval and a regression test?

An evaluation measures quality over a dataset. A regression test is a stable set run after changes to detect deterioration. In practice, a versioned evaluation dataset can provide both when thresholds and configurations are recorded.

When is an AI agent ready for production?

It is ready only when it meets task and safety thresholds, cannot bypass permissions, handles failures without duplicate or unsafe actions, produces usable handoffs, has observable traces and can be rolled back.

Sources

  1. n8n documentation: Why test AI workflows.
  2. n8n documentation: Run quick evaluations.
  3. n8n documentation: Use metrics to measure quality.
  4. OpenAI documentation: Evaluation best practices.
  5. NIST: Artificial Intelligence Risk Management Framework.
  6. NIST: Generative AI Profile.
  7. Barcelona Code School: AI agent guardrails.
  8. Barcelona Code School: AI Agent & Automation Bootcamp.
Back to posts