How the verification layer works

Every rail buyer's first question about language models is whether the output can be trusted. The honest answer is that a model's confidence is not evidence, so we built the checking into the system rather than leaving it to the reader. This page explains what the layer does, step by step.

Extracted claimfact · verbatim

Series retrofit of the onboard signalling equipment shall be completed for all 14 vehicles within 26 months of contract award.

ITT Vol. 2 · §4.3.1 · p. 87 · R-043

Adjacent text, re-read in the original

„Das erste nachgerüstete Fahrzeug ist spätestens im 14. Monat nach Auftragsvergabe wieder in Betrieb zu nehmen; die Serien-Nachrüstung erfolgt in Losen von jeweils zwei Fahrzeugen.“

Assessmentassessment · not verbatim

Two vehicles out of service at any time between month 14 and month 26. Read the availability clause in Annex C §2 against this.

One requirement from a tender, as the verification layer returns it: the claim quoted, its source cited, the surrounding text re-read in German, and the assessment kept apart from the fact.

One requirement from a tender, as the verification layer returns it.

What does the verification layer check?

Four things, in order.

  1. SourceITT Vol. 2 · §4.3.1 · p. 87
  2. Contextsentence before · sentence after · the clause
  3. Languagere-read in German
  4. Kindfact · verbatim

Source. Every extracted claim carries the document, section and page it came from. If a statement cannot be anchored, it is not presented as a fact.

Context. The text around the anchor is re-read to confirm the interpretation — the sentence before, the sentence after, the clause it sits in. Many misreadings come from a clause taken alone.

Language. Tenders in French, German, Spanish or Italian are checked in the original, not in a translation. Where a term matters ("shall guarantee" against "should aim for"), the original wording is shown beside the interpretation.

Kind. Each statement is classified: verbatim fact, paraphrase, or assessment. Assessments — "this obligation probably extends to obsolescence" — are labelled as such and never presented as findings.

Why does the wording matter as much as the facts?

Language models tend toward absolutes: "critical exposure", "must be resolved immediately". A bid team reading that phrasing either over-reacts or, after a few false alarms, stops reading. The layer normalises the tone to what an engineer would write — "worth a question to the authority rather than an assumption in the price" — so that severity is carried by the facts, not the adjectives.

model outputCRITICAL: the availability requirement is contradictory and creates severe penalty exposure. This must be resolved immediately.
after the verification layer

Availability is defined twice, differently (Annex C · §2; ITT Vol. 2 · §6.2). At the planned 98.3 %, the Annex C definition gives an exposure of about € 416 k, within a 4 % cap. Worth a question to the authority rather than an assumption in the price.

R-044 · Annex C · §2 · ITT Vol. 2 · §6.2

Fictional tender. The figure is the /expertise penalty calculator at its default plan.

Where does a person still come in?

Two places. Statements the layer cannot anchor, or classifies as assessment, are flagged for review rather than silently included. And scanned or poorly structured documents — annexes, drawings, tables — are routed to a human check, because that is where models fail most confidently.

Where it has been used

Across the Agentic Bid Stack on 2,500+ pages of tender documentation in six or more countries, and on a live bid managed inside a client's team, where it was the reason the team could act on the analysis instead of re-doing it.

Book a call — thirty minutes by video, if you would like to see it run on one of your own documents.