Part 6 of 7 · 6 min

The number that declines to flatter

A figure that always looks good is not a measurement. Part 6 of the Refusal Series: two numbers where the industry reports one, denominators that travel with their counts, and caps that announce where they stopped.

Part 5 left the decisions intact. A rejection recorded a year ago still stands, and the reviewer never had to make it twice.

Now look underneath those decisions. Every one of them rested on a number. Someone read a coverage figure or a completion fraction and acted on it. A figure that flatters corrupts every decision made on top of it. It does the damage quietly, because those decisions still look well founded from outside.

So ask the day-one question about RET-204. How much of the obligation is covered?

Two numbers where the dashboard wants one

The honest answer on day one is two numbers, and the first one is small. A few links between the retention obligation and the estate carry a human confirmation. A much larger set carries a machine's suggestion and nothing else.

We refuse to add them together. Our own design records state the reason without decoration: the moment they are blended, nobody can tell how much of a coverage figure is real. The blended number is bigger, and it is the number everyone prefers to see. It is also unrecoverable. No reader downstream can separate it again.

The same rule governs the empty case. A family of tables that nobody examined reads as unmeasured. It never reads as clean. Zero rows checked is not a pass, and a pass is what an unexamined zero looks like on most dashboards.

A count without its denominator

A percentage carries a hidden argument about what it is a percentage of. Change that population and the same evidence produces a different figure, with no dishonesty anywhere in the arithmetic.

Here is a measured example from our own estate. One ingested document grounds 38 of the 87 concepts in the vocabulary pack it belongs to. That reads as 44 percent, and the figure looks like a defect, as though one file held up half a model. Read the identical document against the whole embedded graph and it grounds 65 concepts out of 827. That reads as 7.9 percent, and it looks unremarkable.

Both figures are correct. Neither is a measurement on its own, because a count without its declared population is a number and nothing more. So a figure that leaves our system carries two labels beside the count: the level of the question, and the population behind it.

The figure that refuses to count our own work

There is a specific temptation in a platform that both measures a gap and helps close it.

When a team ships something that satisfies a modeled concept, the deployment writes a link back to that concept. The link is real and the record keeps it. We deliberately exclude it from the confirmed-coverage numerator. Delivery grows the substrate, and it does not move the coverage score.

The reason is that the alternative is a scoreboard we control. If our own output counted as evidence of coverage, the figure would rise whenever we worked, whether or not the estate got better understood. A number that responds to our activity measures our activity.

What a label does not promise

Every grounding link in the graph carries a second field beside its name. The name says what kind of relationship it is. The field records what kind of authority put it there, on a four-step ladder that runs from a deterministic derivation down to an unwitnessed assertion. Only the top two steps count as confirmable.

That field exists because a name is cheap. A link labeled confirmed can be a human attestation or a threshold promotion, and the two are not the same claim.

Measured on our own development estate this month: 8,719 links carry the confirmed label. Of those, 8,702 record a machine as the source of the confirmation. Nine record a person. That is the unflattering shape of a young estate, and we would rather publish it than publish a coverage figure that quietly rests on it. Part 4 argued that a data model can make dishonesty unrepresentable. This is the same move pointed at measurement.

Every cap says where it stopped

The second species is the announced bound.

A truncated list that does not announce its truncation reads as a complete list. That failure has a reliable victim. An agent that receives eight of thirty objections, with no marker, treats those eight as the whole set. It answers accordingly. A person reading a capped audit trail with no marker concludes that is all there was.

So caps announce themselves. An evidence log that drops entries counts the drops and reports the count. A context window that shows part of a portfolio appends a marker that says more exists. A bulk approval presents the complete membership of the set the signer is about to approve. The sample it shows includes the weakest member rather than the one that reassures most.

One boundary belongs here, because it is where teams get this wrong in the other direction. A display cap is presentation and never semantics. The screen shows a short list of candidate targets, and the check for unmatched targets reads the full set behind it. A column is not an orphan because it ranked below the fold.

The gate we switched off

The last case is the one we find hardest, so we state it plainly.

A governed change process wants outcome gates. A phase should advance when the business result actually arrived. Every quantity available to us for that job is weaker than it sounds. One is an attestation somebody typed into a field. The second is a telemetry composite that measures implementation completeness, and the third is a ratio computed by arithmetic on names.

Our ruling is that a manually sourced figure may render and it may trend, and it may never derive a verdict. The outcome gate ships switched off. It stays off until the value comes from a measured external system, with a provenance link to the artifact behind it. Our records put the reason in one line: the first outcome gate must not be the first dishonest gate.

We wrote that ruling. The code does not exist yet. This paragraph describes a commitment rather than a shipped feature, and this article is a poor place to blur that difference.

The cost of a small number

A buyer who compares vendors sees our two numbers beside a competitor's one number, and ours is the lower figure. That is a real cost and it lands in the demo, which is exactly where a buyer is least equipped to check the difference.

The trade is about what happens after the demo. A large blended figure survives until somebody uses it. Then it fails in a room where the decision already happened. A small confirmed figure with its denominator attached tells a team where to spend the next month.

An honest number is honest on the day somebody writes it. A year later the material under it moves. Procedures get rewritten and systems get replaced. The grounds a claim stood on stop being the grounds it stands on now. Part 7 asks what keeps a number honest after that, and what this entire posture costs the company that adopts it.

Part 7: What honesty costs, and why you would pay it. The seal that breaks itself, the claim that returns when its grounds move, and the bill stated before you find it.

Model your first domain today.

You send five documents, we model them, and the first cut comes back in days.

This site uses cookies

We use essential cookies for the site to function and analytics cookies (Google Analytics) to understand how you use it. Analytics cookies are only activated with your consent. We do not track you across other websites. Your data is stored in the EU and processed in accordance with GDPR. Read our Privacy Policy