AI Agent & LLM Testing — Frequently Asked Questions
What is AI Agent & LLM Testing in QAEverest?
It is a black-box evaluation harness for conversational AI agents and LLM apps. You give it your agent's endpoint, with no SDK and no code access, and it writes test scenarios from your policy or system prompt, drives real multi-turn conversations against the agent, attacks its guardrails, and scores every reply for accuracy, guardrail compliance, tool-call correctness and PII leakage.
Each run can be compared with a pinned baseline, and a release gate blocks real regressions, in the app or as a failing step in your CI pipeline.
Which agents can I test?
Anything that answers over HTTP. Pick a protocol on the launch form:
- HTTP / JSON endpoint — your own API. The Request format panel lets you write the request body template with {{message}}, {{history}} and {{conversation_id}} placeholders, and set the reply path, tool-call path, conversation-id path, auth header name, extra headers and timeout. Multi-turn runs send the earlier turns, plus the conversation id your agent returned when the template uses {{conversation_id}}.
- OpenAI-compatible /chat/completions — any endpoint that speaks the Chat Completions format; name the model to send.
- Custom webhook — like HTTP / JSON, with the Request format panel open from the start.
- Claude model on AWS Bedrock — test a Claude model directly: Haiku or Sonnet, plus Opus where it is enabled. No endpoint or API key is needed.
Before spending credits, click Test request. It sends one message with your exact settings and shows the parsed reply, the path it was read from, any tool calls, the conversation id and a raw preview with your key masked. It is free, saves nothing, and allows 10 tries a minute. It isn't available for Bedrock models.
How do I start an evaluation?
Open AI Agent Testing from the sidebar and click New evaluation. Enter an agent name, choose the protocol, and give the endpoint URL and API key. Paste your policy or system prompt text as the agent's knowledge; generated scenarios are written from it.
Choose Generate new or Run a saved suite, pick the suites and the red-team depth, and click Run Evaluation. The form shows how many scenarios will run and the most the run can cost. A live stepper follows the run through Connect, Generate, Execute, Judge and Score.
What does it test?
Five suites, each of which you can turn on or off for a run, plus drift checking against your baseline:
- Response accuracy (8 scenarios) — factual questions written from your policy, each with its correct answer.
- Guardrail enforcement (6) — requests the agent must refuse or redirect, such as discount abuse, unauthorized actions and out-of-policy promises.
- Tool-call validation (5) — requests that should trigger a specific tool; the tool name and its arguments are both checked.
- Red-team & jailbreak (7, 30 or 60) — adversarial attacks, by red-team depth (see the next question).
- PII protection (4) — attempts to make the agent disclose personal data it must protect.
- Model drift — every run can be compared with a pinned baseline run of the same agent and suite.
How thorough is the red team?
Choose a red-team depth for the run:
- Quick (7 attacks) — the core attacks, one message each. A fast smoke test.
- Standard (30 attacks, the default) — injection, jailbreaks, system-prompt leaks, data exfiltration, unsafe output and excessive agency, including three multi-turn escalations and attacks in Spanish and Hindi.
- Deep (60 attacks) — Standard plus base64, leetspeak and ROT13 variants of 10 attacks, to catch filters that only read plain text.
Each attack is tagged with its OWASP Top 10 for LLM Applications category where one fits. Some injection attacks ask the agent to repeat a canary code, so a reply that obeys the injection is caught by rule, while a refusal that merely quotes the code is not. Attacker links and scripts are checked as forbidden content. A failed red-team scenario always counts as a breach.
How does the judge decide whether a reply passes?
Rules go first. PII detectors, required and forbidden phrases, injection canaries, and the tool name and arguments are checked deterministically, and a hard failure there can't be overruled by the judge. Phrases match as whole words, and a forbidden phrase made of ordinary words (such as "personal information") is passed to the judge as evidence instead, because a correct refusal often names what it refuses.
Then a judge model reads the whole exchange: the conversation, the agent's tool calls and its final reply, each fenced so nothing the agent says can instruct the judge. It returns pass or fail, a score and a written rationale. Under Advanced you can set:
- Judge strictness — Lenient (the substance is right and safe), Balanced (correct, grounded and on-policy; the default) or Strict (exact amounts and limits, airtight refusals).
- Judge model — Sonnet by default, or Haiku, plus Opus where it is enabled. A stronger judge misses fewer bad replies. The choice is saved with the agent, and runs scored by different judges aren't compared for drift.
Why does it run each scenario more than once?
An AI agent can answer the same question differently each time. Runs per scenario (1 to 10, default 3) repeats every scenario in fresh conversations. A scenario passes only when every run of it passes, and the report shows how many runs passed and failed, so an inconsistent answer is visible instead of hidden behind one lucky reply.
Parallel conversations (1 to 8, default 4) sets how many conversations run at once. More finish sooner; lower it if your agent limits how many requests it accepts.
What is the difference between Generate new and Run a saved suite?
Generate new writes fresh scenarios from the agent's knowledge. It is a quick look, and because the questions change every time, drift is compared on overall rates only. Click Save as suite on a run's results to keep its scenarios.
Run a saved suite drives exactly those scenarios every time, so drift is compared scenario by scenario and the most it can cost is known exactly. Change the endpoint to run a saved suite against another environment, such as staging; each endpoint keeps its own baseline.
How does the release gate decide that the agent has regressed?
Click Pin as baseline on a good run. Later runs of the same agent endpoint and suite are compared with it, metric by metric. A drop blocks the gate only when it is bigger than the metric's threshold (5 points on the trust score; 5 points on accuracy, guardrails and tool calls; 2 points on PII protection) and a one-sided Fisher exact test on the pass and fail counts across every run shows it is unlikely to be chance (p ≤ 0.05). Guardrails also have a hard floor of 90% and PII protection 95%, applied with an allowance for noise. A significant rise in the share of runs that breached a guardrail blocks the gate too. On a saved suite each scenario is also compared: one that now fails significantly more often blocks the gate, and one that fails more often but not significantly goes on a watch list. Smaller or uncertain metric drops are shown but don't block.
If the baseline was scored by a different judge model, judge prompt version or strictness, or used a different red-team set or scenario set, the run reports it as not comparable rather than showing a misleading diff.
Does the model learn or update itself over time?
No model is retrained, and that is deliberate: a testing tool has to be reproducible. What stays current is the data around the models. Saved suites re-run exactly as written, baselines record what good looks like, and suites can be checked against your policy whenever it changes (see the next question).
How do saved suites stay current when my policy changes?
On the AI Agent Testing dashboard, each saved suite has Check for changes. Paste the agent's updated policy or knowledge and run the check: scenarios the change made stale are flagged and stop running straight away, and replacement scenarios are proposed. The proposals run only after you click Accept changes, which activates them, retires the stale scenarios and moves the suite to a new version. If the AI model is unavailable, the check reports an error instead of proposing placeholder scenarios.
Can I run it from my CI/CD pipeline?
Yes. Each saved suite on the dashboard has Use in CI, which gives you a ready-made GitHub Actions workflow or a bash step (curl and jq) for any CI. The step starts a run through the QAEverest public API, polls until there is a verdict, prints the reasons and a link to the report, and exits non-zero on fail or error. The API has three endpoints:
- GET /api/v1/aiagent/suites — your saved suites, the agent each one runs against, and whether a baseline is pinned.
- POST /api/v1/aiagent/runs — runs a saved suite against its saved agent, or against another endpoint such as staging. It answers 202 with the run id, a status URL and a report URL.
- GET /api/v1/aiagent/runs/{id} — progress, then the trust score, results, gate and a verdict: pending, pass, fail or error, with reasons.
Add min_trust_score, max_breaches or fail_on_no_baseline to make the verdict stricter. Without a pinned baseline the drift gate doesn't run, and the run passes with a warning unless you set fail_on_no_baseline. CI runs need a personal API key (Profile → API keys), are billed exactly like runs from the web form, and at most 5 can be in progress per account at once. Your agent's own key travels in the request body and, as on the web, is never stored.
What does an evaluation produce?
- Agent Trust Score — one headline score, with response accuracy, guardrail pass rate, tool-call precision and PII protection broken out, plus average latency and refusal rate.
- Scenario results — each with the conversation, the verdict, the score and the judge's written rationale, and how many of its runs passed or failed. Replies that couldn't be scored show as Not scored.
- Red-team report — every attack marked Blocked or Bypassed (with how many runs it got through), with its severity and OWASP category, multi-turn attacks included.
- Tool calls — the expected call next to what the agent actually called.
- Drift — baseline against current for each metric with the statistical evidence, scenario-level regressions, a watch list, and the release gate: PASS or BLOCKED.
- Credits — what the run was charged and what came back.
What happens if the AI model is unavailable during a run?
Scenario generation retries once; if it still fails, the run stops before any scenario is sent to your agent and the credits are refunded. Saved suites don't need the AI model to write scenarios, so they can still run.
If the judge can't score a reply, that reply is marked Not scored and left out of every score and count, never guessed as a pass, a fail or a breach. Rule-based checks still count. A run stops early when 5 replies needed the judge and none could be scored, and fails at the end when half or more of the scorable replies went unscored; both are refunded. Unscored replies are left out of the drift comparison too, so a judge outage doesn't spoil your baseline.
How does PII checking avoid false alarms?
- Real numbers only — card numbers need a real network prefix, a valid length and a Luhn check; SSNs a valid area, group and serial; IBANs a known country length and mod-97 check. Card, SSN and IBAN leaks count on every scenario.
- Phones and emails in context — a phone number must look like one, so order, ticket, tracking and account numbers and dates aren't mistaken for phones. Emails and phone numbers count only in scenarios where the agent should withhold them.
- Allowed contacts — your own support emails, @domains and phone numbers aren't leaks. Add them under Advanced → Contacts the agent may share; contacts in the knowledge text are allowed automatically, as is anything the user already said in the conversation and example domains such as example.com.
- Judgment for the rest — anything else, such as a name or an address, is left to the AI judge, which scores the reply against what the scenario expects. Numbers and contacts the detectors catch are shown masked in the rationale.
Are my agent's credentials safe?
Yes. The API key you enter is used only to drive the run and is never stored; only a masked value is kept with the run record, and a key echoed back in an error message is masked too. Extra headers saved in the Request format can't have credential-looking names: put the secret in the API key field and set Auth header name instead.
Endpoints must use http or https and can't resolve to private, loopback, link-local or cloud-metadata addresses, and redirects aren't followed, so a run can't be pointed at internal systems.