How VERA turns a news article into a calibrated, falsifiable forecast — and what is structurally forbidden from touching the number. This is the densest page on the site by design; the depth is the point.
Looking for the system architecture? The pipeline blueprint →
The number VERA publishes — the probability that a stated outcome occurs — is moved by only two inputs: a counted prior drawn from the historical record, and a bounded adjustment for the specifics of the case. Nothing else can reach it.
That matters because VERA produces a great deal besides the number: five schools of thought argue over the story, a counter-case is assembled against the forecast, market signals are read, a constraint network is mapped, actors are modelled. All of that is analysis a reader can weigh — and none of it is permitted to change the probability. The commentary informs the reader; it does not move the forecast.
This is not a stylistic promise. It is enforced in the code: every one of those surfaces is mechanically checked for zero write-access to the probability before it is allowed into the system, and the check runs on every change. The separation is the codebase's constitution, not its aspiration — which is why the product can afford to be this deep. The depth is quarantined from the number, so the number stays honest.
What that looks like in practice: when the market read diverges sharply from VERA's forecast, the divergence is measured, labelled, and shown to the reader as its own finding — and it moves the probability by nothing at all. When a school of thought argues forcefully that the house view is wrong, that dissent is published in full, next to the forecast it disputes — and the forecast does not shift to accommodate it. The commentary is there for the reader to weigh against the number; it is never allowed to become the number. Most systems blur that line because blurring it is easy and looks sophisticated. VERA draws it in the one place a line can actually hold: the code, checked automatically, every time.
When VERA estimates how likely an outcome is, it does not ask a language model to feel its way to a number. It starts from a base rate counted over history — how often this class of event has actually resolved the way in question, across every qualifying case in the peer-reviewed record.
Those counts are built over established academic datasets: the Uppsala Conflict Data Program for armed conflict, the Archigos record of national leaders and how their tenures end, the TIES record of economic sanctions episodes, and the International Crisis Behavior project for interstate crises. Each prior ships with a reproducible derivation script, an audit trail, and a Wilson 95% confidence interval — so the reader sees not just the rate but how much history stands behind it.
Where the record is rich enough, the count narrows to the specific adversary pair rather than a pooled average — adversary-pair-specific priors, with a typed fallback when a pair has too few historical cases to stand on its own. When VERA falls back, it says so, and says why. The language model proposes what evidence is present; the arithmetic that turns history into a prior is deterministic.
Concretely: if the question is whether a coercive diplomatic crisis of a given type resolves in a negotiated settlement, VERA does not estimate — it looks up how often crises in that reference class actually did, across every coded case, and reports the rate together with the sample size behind it. A base rate resting on four hundred cases and one resting on nine are both honest, but they are not equally load-bearing, and the confidence interval makes the difference legible rather than hiding it. This is the single most important reason a VERA number is not a language model's hunch: the starting point is a fact about history, retrievable and re-countable by anyone with the same datasets.
Naming the datasets is deliberate. A reader — or a reviewer — can go to Uppsala, to Archigos, to the crisis record, and check that the population VERA counted over is the population it claims. The prior is only as good as its reference class, so the reference class is stated, not assumed.
From the counted prior, VERA applies a single adjustment for the specifics of the case — the actors involved, the structural features the evidence surfaces — and arrives at the posterior probability. The whole chain is shown, not asserted: prior → adjustment → posterior, with each contribution labelled.
The adjustment is bounded by design. There is a hard cap on how far the case-specific factors can pull the prior. When the cap binds, the excess is intentionally clipped, not lost arithmetic — the reader can see that the model wanted to move further and the system refused to let it. Case detail can shade the base rate; it cannot overwhelm it.
Between prior and posterior sits the evidence — and it is drawn from far more than the article. The system reads the article against a much wider evidence layer: a quorum of independent sources reporting the same story, institutional feeds polled on a clock (war-risk circulars, UN and agency notices, defence and foreign-office releases), GDELT global event data, and market data. The language model proposes which factors that whole body of evidence supports and which way each one points; deterministic code then does the accounting. Factors that say the same thing twice are de-duplicated so a single fact cannot be counted as three. Evidence from weaker sources is weighted down. And the whole adjustment is dampened by how much history the reference class rests on — a thin base rate is moved cautiously, a deep one can bear more. The model supplies judgement about what the evidence shows; the system supplies the arithmetic, identically, every time.
And the headline is polarity-correct by construction. Whether the question is “will this hold?” or “will this escalate?”, the displayed probability always answers the question actually asked — the number never quietly reports the opposite of what the reader is reading. It sounds trivial; it is the kind of error that quietly corrupts a forecasting record, so VERA routes it deterministically rather than trusting prose to keep it straight.
An analysis moves through a fixed arc. The article is ingested and corroborated against a set of independent sources gathered around the same story. It is then read closely — through a battery of structured lenses, applied at once, that look past what the article says to how it is built and what it leaves out. The reference class is selected and calibrated into the forecast. The situation is modelled forward — who is constrained by whom, who moves first, what to watch. And the whole thing is audited before it is allowed out.
Every article is read through these lenses simultaneously:
The full stage-by-stage blueprint lives in the architecture diagram.
The political compass. Every analysis places the article's framing on two axes — economic left–right and libertarian–authoritarian — so a reader can see where a piece stands before absorbing its frame. The instrument works on one principle: the model codes, the code counts. The language model marks framing devices in the text; deterministic code counts them into a placement. And the placement carries a validation state the site respects — VERA will not present an unvalidated reading as if it were settled.
This is the line that separates the compass from a media-bias chart: it reads the framing of a text, not the politics of a person or an outlet-as-tribe. Bias-rating sites sort outlets into camps; VERA reads each article fresh, and says so.
Reliability and style, kept separate. VERA measures a source's credibility and its editorial style as two independent things, deliberately never merged. A source can be reliable and propagandistic; collapsing the two is exactly how bias charts mislead. Keeping them apart is a design decision, enforced in the system.
Surface versus shadow. Alongside what the article says, VERA reconstructs what its framing suppresses — the drivers, incentives, and mechanisms left out of the frame. A piece reporting a bilateral standoff as a clash of two leaders may quietly obscure the structural fact that a single chokepoint, and the actors who profit from pricing its risk, is the real driver. The reader sees both: the surface story as the article tells it, and the reading underneath — each labelled, so the reconstruction is never smuggled in as though the article had said it.
None of this reads a political soul. It reads devices on a page — word choice, sequencing, what is granted agency and what is rendered passive, what is asserted as natural rather than argued. That is a claim VERA can actually support from the text, and it is a narrower, more honest claim than the ones bias-rating sites make about who an outlet is.
Not every claim in an analysis carries the same weight, and VERA never pretends otherwise. Every claim on every page is graded by how it is known — and the grade travels with the claim wherever it appears. Analytical inference arrives visibly hedged; it never masquerades as a counted fact. Explained here once, and used everywhere on the site without re-explanation:
Counted — drawn from a dataset, with a citation. Article-grounded — stated in the source. Analytical inference — VERA's reasoning, marked as such.
The most dangerous forecast is one that only makes its own case. So VERA builds the opposing case into every analysis, deterministically. A falsifiability register consolidates the counter-evidence: counted findings that cut against the forecast, era-stamped; declared blind spots — what the record knows it cannot see; and tensions, where a counted historical rate opposes an analytical claim on the same population — never a mismatched dataset dressed up as a contradiction.
The distinction VERA holds here is a fine one and it matters: a tension is only recorded when a counted rate and an analytical claim genuinely bear on the same population. It is easy, and dishonest, to manufacture a contradiction by setting a number from one dataset against a claim about a different one; VERA declines to do that, and under-reports tensions rather than inventing them. A blind spot, likewise, is a specific thing the record is structurally unable to see — named, not gestured at.
Then it invites disagreement. Five schools of thought read the same set of scenarios, each from its own tradition, under one rule: a reading with too little evidence behind it abstains rather than bluffs. An unanchored school is an opinion, not a position, and it is dropped rather than counted. Where the schools that do have footing diverge, the single strongest dissent from VERA's own house view is published beside it — not buried in a footnote. And, per the thesis above, none of this argument can move the probability. It exists to inform the reader's judgement, not to quietly relitigate the number.
Two things make VERA's track record real rather than rhetorical. First, every prediction is stamped when it is made — the class, the exact resolution condition, the date, the headline probability — and written to an append-only record. It cannot be quietly revised after the world moves.
Second, VERA runs two independent layers over each forecast — a deterministic calibrator and a separate analyst pass — and it publishes their disagreement as a signed number. Zero means the analyst genuinely deferred to the calibrated figure; a signed value means it shifted, and by how much; and an explicit “no reading” is kept distinct from a quiet zero. That single number is the most direct answer VERA can give to the fair question, is this just a language model guessing? — no: there are two layers, and their gap is on the record for anyone to check. The measured public scoring of these predictions begins on the track record from August 2026.
A forecast can be numerically sound and still describe itself misleadingly. So a set of deterministic checks reads VERA's own writing and flags where it drifts from its own numbers: a scenario called the “most likely” one that isn't; a “tail risk” label on something well above tail probability; market prose asserting an impact the analysis itself recorded as unmeasured; a stated derivation whose arithmetic doesn't reconcile to the authoritative one. Each check does one thing and one thing only — it reads the number and flags the label. It never rewrites the prose. The correction is a visible flag, not a silent edit.
Nothing publishes automatically. Every analysis passes three gates first:
The adversarial review is itself AI-generated, as disclosed above — and its score, and every failure it finds, are published alongside the analysis, not filed away. A poor review is part of the record.
VERA is built to fail out loud rather than fill in silently. Identities are resolved by live entity grounding against Wikidata — the current record of who holds which office — not the model's training memory, which goes stale. The primary article is corroborated against a quorum of independent sources. Institutional feeds are polled on a clock and cited where they bear on a story: Lloyd's JWC war-risk circulars (signature-verified in the source documents), the UN Security Council, the IAEA, CENTCOM, the FCDO.
Every analysis carries a vintage plate — a stamp of what the world looked like when it was written — and if VERA verifies that an article has been overtaken by events, it declines to analyse it, with a cited reason. The refusal goes on the record. Underneath all of it runs one discipline: every degradation the system suffers is typed and shown whether flattering or not; a considered abstention is kept distinct from an error; nothing is dropped in silence.
A single analysis produces far more than a probability. A sample of what one record contains — each rendered, cited, and graded on the page itself: