Radical transparency
How accurate is EMSist?
Most compliance tools ask you to trust them. Almost none publish how wrong they are, and fewer still say which of their numbers was measured on the data they tuned against. Ours are below with their confidence interval, the one we never tuned on first. You are going to take these findings to an auditor, and a figure you cannot interrogate is not much use to you.
Engine 2.6.0. The frozen-holdout figure is the one to hold us to; both are explained below, along with what neither yet proves.
The number to hold us to: 9.0 points, on documents held back from every scoring decision
First, what the engine is: 41 deterministic rules mapped to the clauses of the standard, derived from the standard itself and cross-validated against the BSI briefing and the published FDIS structure. It is not a model that learned what compliance looks like from examples. The corpus below does not teach it anything; it measures it. The same document always produces the same score. What the corpus does is tell us how far that score lands from a known answer. The figure that counts comes from 14 full-length documents held back from every scoring decision: mean absolute error 9.0 percentage points, 95% confidence interval 6.0 to 12.1, ten of the fourteen within ten points. For comparison, the error is 6.6 points on the 17 documents used to set the scoring constants (a length reference and the coverage thresholds). It is lower for a reason that has nothing to do with the engine being better, so we publish both and let you see the gap rather than making you ask about it.
What those documents are, and are not
Measuring error at all requires documents whose true readiness is known, and real manuals do not come with that. No corpus of real EMS documentation with independently verified readiness scores exists anywhere to test against. Every set we use is therefore expert-written to a specification, which is a real limit on what any of these numbers prove and we would rather say so. Across the three sets that is 72 labelled documents. The 14 held-back ones run from 3,000 to 12,700 words across eleven readiness bands and eleven sectors the engine had never seen; a further 41 shorter documents score 10.7 points and are useful for breadth but average around 2,800 words, so they under-represent the full-length manuals most organisations hold. At fourteen documents the headline interval is wide, and we publish the interval rather than a precise-looking number the evidence does not support.
Separately, we run it against real published manuals
Scoring well on tidy specimens proves little on its own, so the engine is also run against real published EMS manuals from operating organisations, with all the mess that implies: tables, scanned pages, inconsistent headings and years of accumulated appendices. Those manuals carry no verified readiness score, so they cannot tell us the engine's error; what they catch is extraction and matching failures a clean specimen never triggers. That testing is how we found, and fixed, a class of PDF whose invisible character encoding silently defeated multi-word matching, and it is why the engine's length reference is set from real manuals rather than from our own corpus.
Nothing changes without proof
Two separate checks run, and they answer different questions. The accuracy harness above asks whether the engine is right. A second corpus of 41 documents is re-scanned nightly and after every engine change, and compared against a saved baseline, to catch drift: an edit moving scores without anyone intending it. If it drifts beyond a small threshold the check fails and the change is blocked until a human reviews it. It has run daily since April 2026. A high average on the drift corpus is not a good score, it just means the engine got more generous, which is why we report the error figure and not that one.
Where the rules come from
The 41 tracked requirements are derived from the published clause structure, confirmed against the published prEN ISO 14001:2026 table of contents, and cross-validated against publicly available certification-body transition guidance and the standard working group's own briefings. No ISO copyrighted text is reproduced. Full-text validation against the purchased published standard is on our roadmap. We will update this page, and the claim, when it is complete. We would rather under-claim than over-claim.
What the engine cannot see
EMSist analyses documentation. It cannot verify that documented processes are actually followed on the ground, which is what your certification audit exists to do. Documents under roughly 200 words trigger an insufficient-evidence warning rather than a confident verdict. Scanned image-only PDFs are read with OCR, up to 40 pages per scan; where a document runs longer, the report names the pages it assessed rather than scoring the rest silently. The engine reports evidence it found and evidence it could not find; treat every finding as a pointer for professional judgement, not a substitute for it.
See it rather than trust it
The fastest way to judge the engine is to read a full report it produced, then run your own documentation through it for free and check the findings against what you know about your EMS.
Questions about our testing, or found a finding you disagree with? Email [email protected] — disputed findings get reviewed against the rule source, and corrections ship only after the nightly regression passes.