OpenStampsfor industry

Calibration

Were our forecasts any good?

Wherever this product predicts something, it should be possible for anyone to check afterwards whether the predictions held. This page is that check, and it scores us — not suppliers, not countries, not industries. It is public because a report only its authors can read is no check on its authors.

Reading the outcome log…

What is being predicted

One thing, stated once so it cannot drift: a delivered request receives a complete reply within seven days of delivery.

delivered
the provider's own receipt says delivered or read, or a bot deep link was opened. Our word for it is not enough.
complete
a plot record and its stamp arrive, they bind — the published bytes are the stamped bytes — and the plausibility rules refute nothing in them.
the deadline
seven times twenty-four hours from the delivery receipt's own instant, not from when we sent anything.
which requests count
only those whose link this deployment sent itself, through a messaging provider. A link a buyer copied and sent by hand is not measured at all: there is no receipt to date it from, so there are no seven days to be inside.
when a request appears here
outcomes settle nightly at 04:00 UTC, two days after each seven-day deadline, so a request answered today appears here tomorrow at the earliest. An unexplained lag would read as a missing outcome, so it is stated.

This is a number we compute about ourselves, and we pay for every request in it: a request only enters this page once a real message has been sent to a real handset from an account's own allowance. Nothing there stops us from flattering it. What stops it being useless is that the requests we could not observe are counted too, in the open, and broken down by the reason we could not.

Three rules that decide whether a number like this is honest

An unobserved request is not a failure
A request the buyer closed, one that expired without anybody polling it, one never delivered at all — none of those is a supplier who did not reply. They are outcomes nobody observed. They are counted and shown separately, and they never enter a rate. Counting them as failures would quietly damn whoever happens to be hardest to reach.
Abstention is shown, not hidden
A request we made no prediction for is not a prediction we got right. The share of delivered requests carrying no forecast sits beside every number here, because a model that only speaks when it is confident can post a perfect score and be useless.
Nothing is published that would name somebody
Broken down far enough, "one delivered request in one country, not replied to" is a person. Any cell below the floor is withheld — visibly, with its count folded into a stated total, so the page cannot be read as though those requests never happened.

Predicted against observed

Ten bins. In each, what we said on average, and what then happened. A forecast is well calibrated when those two columns agree — which is a different thing from being useful, so the scores below report both.

We saidRequestsSaid, meanHappened

Consent

The buyer consents when the request is created: the desk says, in a sentence beside the button, that the reply's timing is measured so this page can exist. The supplier's page carries its own notice — what is measured is the timing of a reply, never the person, and never anything about them beyond the record they chose to send.

Nothing on this page names an account, a supplier, a plot, a request or a device. The report is aggregate by construction: the module that computes it is lib/calibration.js and it withholds any cell too small to publish.

When this publishes

Yearly, and the counts live at any time. No rate is published below the floor, because ten bins over a few dozen events is noise with error bars wider than the chart.

The report is served from /api/calibration, cached for an hour, and readable from any origin.