Skip to main content
AI & Quality5 min readQAEverest Team

How AI and your QA team work together: four hand-off points

AI can take on much of the legwork in testing while your QA team makes the calls. Four hand-off points — bugs, change maps, AI baselines and recommended test sets — and what each one should hand your team.

It's 9 a.m. on release day. Overnight, an exploratory run crawled the app, wrote and ran forty flows, and came back with nine bugs. A comparison with last week's crawl marked six pages as changed. The chatbot's regression run scored a little lower than the week before. And the pull request that's meant to go out today comes with fourteen recommended tests.

None of that is a decision yet. Every item is a question waiting for someone who knows the product. Are these bugs real? Were those changes intended? Is "a little lower" acceptable? Are fourteen tests enough for this change?

That isn't a gap in the tooling. It's the design. The most useful thing an AI testing tool can do is take on the legwork, then stop in the right place, holding the evidence a QA team needs to make the call.

Autonomy is the wrong scorecard

Demos tend to measure how far a tool goes on its own: it explored, it found, it filed, it fixed. That makes a good demo. In a real release process, though, every step a tool takes without a person is one of two things: work that's safe to automate, or a decision that has been made silently.

The silent decisions are the ones that hurt. A false bug lands in a developer's backlog. A baseline drifts without anyone noticing. A test set skips the one area the change actually touched. None of them looks like a failure at the time.

So a better question for any AI testing tool is this: where does it hand off, and what does it hand over?

A good hand-off has three properties:

  • It stops before anything other people will rely on. A ticket in someone's backlog, a reference that future runs are measured against, a release verdict.
  • It hands over evidence, not just a verdict. Enough for a person to agree or disagree in a minute, without re-running anything.
  • It says what it doesn't know. What it didn't reach, couldn't map or couldn't measure.

Here are the four places where that matters most.

Hand-off 1: a found bug is a candidate until someone reviews it

Autonomous exploration produces false positives. Some come from a check that misread the page. Many are the automation failing rather than the product: a step that timed out, an element the run couldn't reach. We wrote about this in "Your AI tester found five bugs. Four weren't real."

A bug filed straight into a tracker is a claim on someone's time. File a few that aren't real and people start doubting the ones that are.

What the hand-off should carry:

  • Steps to reproduce, with the failed step marked
  • Expected against actual, in plain words
  • The screenshot from the moment the step failed
  • Whether it looks like a product problem or an automation problem

The decision that stays with your QA team: is it real, how severe is it, and who owns it?

Hand-off 2: a map of what changed tells you where to look, not what's broken

After a release, a fresh crawl compared with the previous one can show which pages are new, which changed and which disappeared. It's tempting to read "changed" as "regressed", or to generate tests for everything that's new.

But change is the point of a release. The map's job is to direct attention, and it can only do that if it's honest about its own limits:

  • Not reached is not removed. If the earlier run never reached a page, its absence from the old crawl proves nothing.
  • Changing data is not a changed app. Twenty product pages built from one template are one page when you compare crawls.
  • Seen is not visited. Links the crawl saw but didn't follow should stay on the map, marked as such.

The decision that stays with your team: which changes were intended, which deserve a closer look today, and which flows should become regression tests.

Hand-off 3: a drift check needs a baseline someone chose to trust

Testing an AI agent or chatbot means accepting that its answers vary. "Worse" only means something relative to a reference, so the real question is who chooses the reference.

If a tool simply compares each run with the one before, a slow decline becomes invisible: every run is measured against one that was already a little worse. The reference has to be a run that somebody reviewed and is prepared to defend.

That makes the baseline a decision, not a setting:

  • Pin it deliberately, after reviewing the run's answers, not because it happened to be the latest.
  • Re-pin deliberately, when you accept a change such as a new model, a new prompt or a new policy, and record why.
  • Read drift as a signal, not a verdict. The judge that scores the answers is a model too, and it varies. A delta is a reason to look, not proof of a regression.

We covered how to keep that judge honest in "Your agent was jailbroken. Your eval said it passed."

The decision that stays with your team: which behaviour counts as the standard.

Picking tests from a pull request's diff is a prediction. It's a useful one, because running what a change touches gives feedback while the pull request is still open. But the mapping from changed files to requirements can be wrong, and some files won't map to anything at all.

So the recommendation has to explain itself:

  • Why each test is in the set: the requirement it covers and its business risk
  • How much of the mapped risk the set covers
  • Changed files it couldn't map to any requirement
  • Changed code that has no tests yet

The decision that stays with your team: add what the mapping missed, decide whether this change deserves a full regression, then run.

What doesn't need a hand-off

None of this means a person should approve every step. A tool that stops for approval at every click saves nobody any time.

A workable rule: automate the work that's reversible and visible, and hand off anything that writes into someone else's system, sets a reference others will rely on, or decides what ships.

The platform doesYour QA team decides
Crawls, writes flows, runs them, triages failuresWhich findings are real bugs, and files them
Maps pages and compares them with the last crawlWhich changes were intended, and where to look
Runs agent checks, scores them, measures driftWhich run is the baseline
Ranks tests by risk for a pull requestWhat runs before merge

The work an AI testing tool should do is the legwork: exploring, mapping, scoring and ranking. The call on what's real, what matters and what ships belongs to your QA team, and a good tool is built to hand it over cleanly.

See it on your own stories

1000+ free credits on signup, no card required — roughly 100 test-case generations before you pay for anything.

Start free