# AI evaluation report — engagement template This is a report structure, not a completed evaluation. Fill the fields with evidence from the agreed engagement. Do not publish placeholders as measured results. ## 1. Decision and scope - Decision owner and intended release or business decision: - System, task and evaluated version: - In-scope workflows and excluded work: - Business consequences of failure: - Acceptance criteria agreed before the run: - Actions requiring human review or approval: ## 2. Inputs and references - Dataset version, provenance and permitted use: - Reference-answer owner and review process: - Ordinary, difficult and unanswerable task slices: - Language, region, document-format and role coverage: - Duplicates, protected splits and contamination checks: - Missing references and unresolved label disagreements: ## 3. Run conditions - Model/provider/version, prompt version and configuration: - Retrieval, tool, code and dependency versions: - Starting state, environment and access restrictions: - Number of tasks and repeated trials: - Grading logic, rubric version and human calibration: - Runtime failures distinguished from system failures: - Cost assumptions, timestamps and latency measurement boundaries: ## 4. Findings | Finding ID | Task / input ID | Observed output or state | Expected behaviour | Evidence | Severity | Grader / reviewer | | --- | --- | --- | --- | --- | --- | --- | Severity reflects the agreed business consequence, not a model-generated judgement alone. Include reproducing conditions for every material finding. ## 5. Baseline and comparison | Task slice | Sample size / trials | Baseline | Candidate | Difference | Uncertainty / limitations | | --- | --- | --- | --- | --- | --- | Separate measured values from estimates. Document variability and grader disagreement. A small sample or a single aggregate score may not support a release conclusion. ## 6. Recommendations and ownership | Finding ID | Proposed change | Evidence supporting the change | Owner | Rerun check | Priority | | --- | --- | --- | --- | --- | --- | Distinguish an observed failure from a suspected cause. Record alternative explanations and any engineering work outside the evaluation scope. ## 7. Decision record and handoff - Acceptance criteria met, unmet or untested: - Remaining risks, coverage gaps and uncertain results: - Domain and engineering reviewer decisions: - Deployment decision by the accountable owner: - Artefact locations, version history and rerun instructions: - Access removal, retention and deletion actions: - Conditions that require the evaluation to be repeated: Evaluation provides bounded evidence. This template is not a safety certification, legal assessment or guarantee of future model behaviour. Acadify AI · https://ai.acadifysolution.com/pages/methodology.html ## 8. Failure taxonomy and triage | Category | Observable evidence | Count / denominator | Consequence | Investigation owner | | --- | --- | --- | --- | --- | Suggested categories: missing source, unsupported claim, wrong reference version, extraction error, invalid tool parameters, unauthorized action, incomplete task, timeout and infrastructure error. Adapt categories to the workflow. Do not label a suspected root cause as proven without evidence. ## 9. Release and operations plan - Rollout boundaries and monitoring signals: - Fallback, escalation and human handoff paths: - Timeouts, retries and duplicate-action controls: - Rollback trigger and responsible owner: - Regression-suite owner and maintenance interval: - Rerun triggers after source, prompt, model or tool changes: ## 10. Reproduction manifest | Artefact | Version / hash | Location | Access owner | Purpose | | --- | --- | --- | --- | --- | Include inputs, references, environment configuration, grader, run logs and report. Record the exact rerun command, required dependencies and expected output. Keep secrets out of commands and public artefacts.