Research
Design reports and capability notes from the Attestrum project. Every claim is byte-reproducible; the full source for each lives in the public repository.
Explainers
How Attestrum Works, End to End
A user walkthrough — from a folder of files to a proof anyone can check.
READ →
Deterministic by Construction
Seven disciplines that make every seal byte-identical across re-runs — and the precise claim they support.
READ →
Provenance Without Disclosure
Cryptographic training-data provenance that keeps the corpus private — a design-and-explanation report.
READ →
The Disclosure Floor and the Defensible Position
How verifiable training-data provenance backs the new transparency mandates — a positioning report.
READ →
Capabilities
Corpus-Version Diffing
What changed between two sealed corpus versions — added / removed / unchanged, composition shift, both Merkle roots — re-checkable byte-for-byte.
READ →
Benchmark Contamination Scanning
Did the eval set leak into the training corpus? A read-only scan flagging exact, near-duplicate, and embedded overlap.
READ →
Training-Content Summaries
What a sealed corpus is made of — modality, source-type, license, and language mix, weighted by count and bytes.
READ →
Intra-Corpus Near-Duplicate Rate
How much of a corpus is near-duplicated against itself — rate, cluster-size histogram, and example clusters.
READ →
Corpus Removal Evidence
Inclusion in the earlier version plus non-inclusion in the later one — proof a document was taken down, verifiable with stock cosign.
READ →