# Policy source evaluation — sample report **Demonstration only.** Fictional documents and hand-written candidate records. No model was queried. No client work, production dataset or measured model accuracy is represented. ## Objective Check three structured fields in a return-policy answer: the return window, cited policy and effective date. ## Reference For the fictional purchase dated 15 February 2026, policy-2026 is authoritative: 30 days, effective 1 January 2026. The retired policy-2025 specified 14 days. ## Method Each candidate is checked against the reference with exact equality. The browser and Python runner read the same JSON fixture. No API keys, network calls or external dependencies are required by the Python runner. ## Results | Candidate | Return window | Citation | Effective date | Checks passed | | --- | --- | --- | --- | --- | | current | Pass | Pass | Pass | 3/3 | | retired | Fail | Fail | Fail | 0/3 | | mixed | Pass | Fail | Fail | 1/3 | ## Finding A correct return window can accompany an incorrect citation. An answer-only check misses that distinction. Evidence checks help identify the mismatch. ## Next investigation in a real system Inspect which document was retrieved, whether version filters were applied, and how the answer referenced its source. Rerun date-sensitive questions after a change. This fixture does not identify the cause of any actual retrieval or model failure. ## Limitations Three invented records do not represent users. Exact structured checks do not evaluate natural-language claims, access controls, prompt injection, document retrieval or actual model behaviour. Counts are fixture results, not accuracy estimates. A real report needs coverage, run conditions, reviewer decisions and uncertainty. ## Reproduce Save `policy-evaluation-fixture.json` and `run_policy_checks.py` together. Run `python3 run_policy_checks.py`. Modify a local candidate and rerun to inspect the effect of each field. Prepared by Acadify AI, a practice of Acadify Solution. https://ai.acadifysolution.com/pages/research-examples.html ## Evidence and failure classification The authoritative structured reference is stored under `expected` in the JSON fixture. Candidate IDs `current`, `retired` and `mixed` map to the result rows above. `return_days`, `citation_id` and `effective_date` are separate checks: a passing answer value does not override a failing citation or date. The retired record fails all three fields. The mixed record preserves the 30-day window while citing the retired source and date. These findings show mismatched structured evidence; they do not prove how a real system selected a document. No retrieval trace or generated response is present. ## Suggested verification exercises 1. Run the fixture unchanged and compare the result counts with the table. 2. Change only the mixed record's citation to `policy-2026`: two fields should pass while the old date still fails. 3. Remove a candidate field: that check should fail instead of assuming the reference value. 4. Restore the fixture before comparing browser and Python results. ## Translating this into a real evaluation Define which policy applies to the transaction date, region and product. Preserve retrieved passages and policy versions. Check natural-language claims against those passages, and include questions with no valid evidence or conflicting sources. Have a domain reviewer resolve reference ambiguity. Test permissions separately. Document sample coverage and uncertainty before using outcomes to support a release decision.