Original PolicyOps research · Pilot edition

Can policy evidence be retrieved without hiding the route?

A transparent benchmark of official public-PDF ingestion, deterministic source retrieval, page-reference coverage and unsupported-answer controls.

Published 24 July 2026 · Methodology 1.0.0 · No AI used in scoring

8/8Real-document retrieval questions passed
30/30Generalisation scenarios passed
206/207Docling sections retained page references
5/5Unsupported controls remained bounded

What this study tests

Evidence engineering, not a compliance ranking.

The benchmark tests whether PolicyOps can recover expected source authority from controlled policy material while retaining an inspectable route to the evidence. It deliberately separates real official PDFs from synthetic source-inspired fixtures.

The real-document track exercises download, file validation, extraction, canonical sectioning, temporary persistence, governance review and deterministic retrieval. The generalisation track tests broader wording and negative controls without storing copied policy text.

Measured results

Both extraction routes found the expected authority.

The same source bytes were used for the current and Docling comparisons. Passing required the expected public document at rank one and reviewer-selected clause-signal coverage.

Current fallback

8/8 retrieval questions passed
Searchable documents
4/4
Page-referenced sections
0/543
Measured ingestion time
34.1s

Docling structured extraction

8/8 retrieval questions passed
Searchable documents
4/4
Page-referenced sections
206/207
Measured ingestion time
87.2s

Section counts are not a direct quality score: the extractors segment documents differently. Page-reference coverage is reported because it materially affects evidence inspection.

Generalisation and safety

The broader suite tests what should not be answered too.

The separate source-inspired suite contains 30 questions across gifts, conflicts, supplier conduct, security, confidentiality and anti-corruption. 25 are answerable source-selection tests and 5 are negative controls.

Overall scenarios30/30
Expected source at rank one25/25
Unsupported controls bounded5/5

Methodology

Designed to be inspected and rerun.

  1. 1
    Freeze the public source

    Record the final source URL, retrieval timestamp and SHA-256 digest. Stop the comparison if source bytes differ between extraction runs.

  2. 2
    Use temporary isolated storage

    Download and process source PDFs at runtime. Do not commit full PDFs or extracted document text.

  3. 3
    Run both extraction routes

    Measure searchable output, section structure, page-reference coverage and elapsed ingestion time.

  4. 4
    Apply reviewer-defined questions

    Require the expected source at rank one and sufficient evidence-pattern coverage for the real-document track.

  5. 5
    Test refusal boundaries separately

    Run source-inspired answerable scenarios and unsupported controls without treating synthetic content as real public-policy analysis.

Source register

Official documents, identified at run time.

The public PDFs are downloaded from their registered sources for each full research run. The published checksum allows later reviewers to identify the exact bytes used.

Limitations

A useful pilot, deliberately not overclaimed.

  • This is a pilot engineering benchmark, not a representative survey of the policy-management market.
  • The four real documents test ingestion and retrieval against official public PDFs; they do not support sector-wide prevalence claims.
  • The generalisation track uses synthetic mini-fixtures inspired by public policy structures and must not be described as analysis of the source organisations' actual wording.
  • Reviewer-selected questions and clause signals were not independently peer reviewed.
  • Results measure retrieval and evidence-location behaviour, not legal correctness, organisational compliance or the substantive quality of a policy.
  • Section counts differ between extractors because segmentation strategies differ; section totals are not a direct quality score.
  • Public URLs and source documents can change after the recorded retrieval timestamps and hashes.

Governance statement

Policy content remains evidence, not instruction.

No AI model produced or graded these results. The benchmark measures deterministic retrieval and evidence-location behaviour. Human review remains necessary before relying on policy interpretation or making an organisational decision.