# AI evaluation playbook A practical planning guide from Acadify AI. Adapt it to your system and business consequences; it is not a completed assessment or certification. ## 1. Define the decision Record the system version, workflow, users, permitted actions and accountable release owner. Decide whether the evaluation supports an initial launch, a model comparison or a regression review. Agree acceptance criteria before examining candidate results. Use separate criteria for critical failures and ordinary quality issues. ## 2. Assemble permitted evidence Inventory source documents, task examples, reference answers and execution traces. Record provenance, permission to use each input, version and review date. Redact unnecessary identifiers. Keep development examples separate from the held-out evaluation set; document duplicate detection and overlap checks. Include normal work, difficult cases, missing evidence and inputs that require clarification or refusal. ## 3. Choose checks for the workflow | Workflow | Checks | Evidence to retain | | --- | --- | --- | | RAG | Retrieval coverage, supported claims, correct citation and current document version | Question, retrieved passages, document IDs, answer and reviewer rationale | | Agents | Final state, permitted actions, approval gates, retries and recovery | Initial state, tool calls, responses, approvals and final state | | Model comparison | Same tasks, quality rubric, latency and cost boundaries | Candidate versions, configuration, paired task outcomes and run logs | | Datasets | Label agreement, provenance, duplicates, task coverage and protected splits | Schema, annotation instructions, disagreements and split manifest | | Training | Baseline comparison, held-out quality, regressions and leakage | Dataset versions, training configuration, checkpoint and held-out results | | Coding | Reproduction, executable success checks, regressions and environment isolation | Repository revision, patch, commands, exit codes and test output | | Multimodal | Field values, units, layout, table structure and source locations | Input version, expected fields, extracted result and evidence coordinates | ## 4. Freeze the run conditions Record model/provider version, prompt, tools, retrieval settings, dependencies and grader version. Define timeouts, retry policy and concurrency. State whether latency includes retrieval and tool execution. Record token pricing and date when estimating cost. Separate infrastructure failures from incorrect task outcomes. Repeat stochastic tasks enough to inspect variability; report trial counts rather than implying certainty from one run. ## 5. Calibrate grading Write observable pass/fail criteria and examples of borderline outcomes. Have domain reviewers independently grade a shared calibration subset. Resolve disagreements and record rubric changes before the main run. If using a model grader, compare its decisions with human review and inspect disagreements. Critical permission violations require direct evidence, not an aggregate score. ## 6. Report the useful breakdown Report task and trial counts, exclusions and missing results for each slice. Compare candidates on the same tasks. Show failures by consequence and workflow; retain reproducible examples. Document uncertainty, sample limitations and reviewer disagreement. Do not generalize a small convenience sample to every future user. ## 7. Investigate and rerun For each finding, separate observed evidence from a suspected cause. Assign an owner, proposed change and rerun condition. Rerun both affected cases and the broader baseline after a fix. Add newly observed production failures to a reviewed regression set without silently changing the historical comparison. ## 8. Hand off a decision Use the release-readiness checklist and report template. Record criteria met, unmet and untested, residual risks, review decisions, rollback conditions and artefact locations. Confirm who owns the suite and when it must be rerun. The accountable owner makes the release decision. ## First-use sequence 1. Complete the project brief with descriptions and redacted examples. 2. Select the relevant checks above and agree references with a domain reviewer. 3. Fill the report template during the run, preserving evidence for each finding. 4. Review the release checklist with engineering and the decision owner. Public policy examples use fictional records. They demonstrate exact evidence checks, not measured model performance.