The first decision in an honest migration map is not about any column. Part 2 of the Proof Series: why a suggestion and a decision are different objects in the model, and what one confirmed table pair opens.
Part 1 ended on a question a room full of qualified people could not answer. What must a mapping be, before somebody outside the project signs it?
The answer starts one level above the columns, in a place most mapping exercises skip.
Return to Danubia, our invented manufacturer. The table under the microscope is the supplier master: eighty-four columns on the legacy side, including the one-character status column nobody living can explain. On the target side sits a supplier object with its own field list and a validation rule behind each field. Around them, the rest of the estate: call it three hundred tables a side, which is modest for a twenty-two-year-old system.
The obvious move is to point a matcher at all of it. The tools are happy to oblige. Every column on the left gets a ranked list of candidates on the right, with scores beside each. The output is enormous, largely plausible, and structurally identical to the four spreadsheets from part 1. It is more mapping, generated faster, adjudicated by nobody.
Before any of that output can matter, one law has to hold, because every part of this series after this one relies on it.
In our model, a machine's suggestion and a person's confirmation are two different kinds of edge. They are never one edge with a status field on it.
The distinction sounds like pedantry until you ask what a status field survives. Any writer can flip a property. A migration script can default it. A bulk update can set it across ten thousand rows in one statement, and the record afterward looks exactly like ten thousand considered decisions. A reader that selects confirmed pairs by property receives whatever the last writer left behind.
A separate edge type behaves differently. A reader that does not name the confirmed type cannot receive it by accident. A writer that produces suggestions is not the writer that produces confirmations, so no single code path can promote its own guess. The split holds in four independent places in the platform: scope, evidence, field maps, and configuration. In each one, the proposal and the decision are different types, written by different hands.
The effect is the one that matters for a signature. A guess cannot leak into a sealed proof, because the seal reads only the decided type. Readers of our last series will recognize the move. This is honesty by construction rather than by policy, applied to the very first object in the chain.
Now the rule this part is named for.
Before any column can be mapped, a person confirms the table pair. At Danubia, a delivery lead states that the legacy supplier table corresponds to the supplier object in the target. The platform records that statement as a decision with her name on it. Every field mapping that will ever exist for these tables hangs under that confirmation.
The gate fails closed. No confirmed pair means no field map, for anyone, by any path. There is no provisional mode where column work runs ahead of the pair decision and gets ratified later. The order is the point: the person decides the frame, and only then does the machine work inside it.
And the pair decision is a decision in the full sense of our last series. A pair somebody rejected stays rejected. A re-run of the matcher cannot re-suggest it, and the refusal to resurrect it is recorded as a correct outcome. The twenty minutes a reviewer spends rejecting a bad pair are twenty minutes the estate never pays again.
Here is what happens in the minute after the delivery lead confirms.
Eighty-four columns become searchable that were not searchable before. The estate-wide soup, hundreds of tables against hundreds of tables, collapses into one bounded question: these columns, against that field list. Candidate ranking that was noise at estate scale becomes signal at pair scale, because the population is small enough for the scores to separate.
The economics of review change with it. A reviewer facing an estate-wide candidate list samples it, and part 1 described where sampling ends. A reviewer facing one confirmed pair can actually finish: eighty-four columns is an afternoon, and the afternoon produces decisions rather than spot checks.
The confirmation also bounds the claim. Confirming the pair says the tables correspond. It says nothing about any column, and the platform does not treat it as saying more. The column work is still ahead, and each column needs its own decision. What the pair buys is not coverage. It is a frame inside which coverage can be honestly counted.
One boundary belongs here, so the suggested tier gets its due.
Suggestions are not waste. The ranked candidates, the scores, the machine's whole view of what probably matches: all of it renders, informs the reviewer, and orders the afternoon's work. Analysis lives on the suggested tier, and it would be a worse product without it.
What the suggested tier never does is count. It does not enter coverage, it does not enter the reconciliation, and it cannot enter the seal, because the seal cannot reach its type. Our previous series spent a whole part on what blending those figures costs. Here the split carries weight. One tier exists to help a person decide, the other records what the person decided, and the model keeps them apart by construction.
That is the standing of the map at the end of this part. One pair confirmed by name. Eighty-four columns open, each with ranked candidates waiting on the suggested tier. Nothing decided yet that a seal would count.
The next question is the one the columns force. When a value crosses from a legacy column to a target field, something happens to it on the way: a cast, a decode, a default. What is the system allowed to claim about that transformation? Part 3 is about keeping the list of possible answers closed.
Part 3: A closed list of things that can happen to a value. Eight operations, a decode that is computed rather than authored, and the escape hatch that can never confirm itself.
This site uses cookies
We use essential cookies for the site to function and analytics cookies (Google Analytics) to understand how you use it. Analytics cookies are only activated with your consent. We do not track you across other websites. Your data is stored in the EU and processed in accordance with GDPR. Read our Privacy Policy