Testing AI Systems in 2026: Acceptance Criteria for Non-Deterministic Software

Knowledge Blog
Software quality team evaluating multiple AI outputs against acceptance thresholds

Traditional software testing is comfortable with a simple assertion: given input A, expect output B. Generative and other probabilistic AI systems can produce different outputs for the same input while still behaving acceptably.

That does not make them untestable. It changes what good acceptance evidence looks like.

ISTQB’s Certified Tester AI Testing syllabus version 2.0, released in 2026, reflects this evolving practice, including generative AI, non-determinism, statistical approaches and specialised testing techniques.

Replace single answers with acceptance properties

For a customer-support summariser, the exact wording may vary. The acceptance criteria can still require all material facts to be preserved, no unsupported account data to be added, prohibited sensitive content to be excluded and the summary to remain within a length range.

The Certified Software Quality Analyst provides a broader software quality foundation for professionals adapting assurance to AI-enabled systems.

Test distributions

Run representative cases multiple times where variability matters. Measure pass rate, error severity and confidence intervals rather than reporting one successful demo.

Define the sample and threshold before testing to reduce the temptation to move the target after seeing results.

Separate quality dimensions

An AI system can be accurate but unsafe, useful but biased, or robust but too slow. Define relevant dimensions such as correctness, relevance, robustness, security, fairness, privacy, performance and usability.

ISO/IEC 25010:2023 provides a software product quality model that can help teams think systematically about quality characteristics, even though AI-specific evaluation may require additional measures.

Include adversarial acceptance cases

Test prompt injection, malicious retrieval content, excessive inputs, ambiguous instructions and attempts to cross authority boundaries where relevant.

Security expertise such as the Certified AI Cyber Risk Assessor can help connect adversarial AI testing to business risk.

Define safe failure

Ask what the system should do when it lacks evidence, a tool is unavailable or a confidence threshold is missed. An explicit refusal or human escalation can be a successful test result.

Make regression testing version-aware

Record model version, prompt/configuration, retrieval sources, tool versions and evaluation set. When one changes materially, rerun the tests that protect important behaviour.

The Certified Agile Testing Foundation Professional can support teams integrating this evidence into iterative delivery rather than deferring it to the end.

Example acceptance statement

Instead of: “The assistant answers the refund question correctly.”

Use: “Across the approved refund-policy test set, at least the defined target proportion of responses must cite the correct policy condition; zero responses may invent a refund entitlement; prohibited personal data must never be disclosed; ambiguous cases must route to human support.”

The numbers should come from business risk and validation evidence, not from a generic benchmark.

Build a representative evaluation set

Include normal cases, boundary cases, rare but material scenarios and examples from groups or contexts where failure would matter. Keep a protected regression subset so developers do not optimise prompts only to the visible test data.

For generative systems, combine automated evaluators with calibrated human review where qualities such as usefulness or harmful ambiguity cannot be captured reliably by one metric. Validate the evaluator itself before trusting its score.

Define change triggers

A new foundation model, system prompt, retrieval corpus, tool permission or safety setting can invalidate previous evidence. Classify which changes require full, targeted or no revalidation.

This reduces two opposite risks: rerunning everything after trivial changes, or carrying old acceptance results into a materially different system.

Add operational acceptance

Production quality includes monitoring, incident response and rollback. The Certified Software Project Manager is relevant where test evidence must connect to release and lifecycle decisions.

Final takeaway

Non-deterministic software still needs deterministic accountability. Define acceptable properties, test distributions, include adversarial cases, design safe failure and rerun evidence after material change.

AI testing becomes credible when acceptance is tied to the risk and purpose of the system rather than to one impressive output.

Further reading

Tags :
acceptance criteria,AI testing,GenAI testing,non-deterministic software,Software Quality
Share This :

Responses

error:
The Case HQ Online
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.