Built in the fire 10 min read
Stop Interviewing Your Experts. Replay Them.
The expert could hand over his process in ten minutes. The judgment inside it, the part you actually pay him for, had never been written down. So we stopped asking him to explain and read six years of his decisions instead.
One of the sharpest experts I have ever worked with could not write down the rules behind a decision he makes constantly. Not wouldn’t. Couldn’t.
He had made the call more times than either of us could count. Two decades of it. When I sat him down and asked him to put it on paper, the process came easily. He gave me the steps, in order, the same way he runs them every time, and it took about ten minutes.
That part is worth almost nothing, and it took me a while to understand why. The process is not what anyone pays him for. Anyone can follow steps. What he gets paid for is the judgment he applies inside them: which step actually matters this time, where the exception lives, when the obvious answer is the wrong one. That is the part he could not hand over. The hundreds of small calls that fire on their own and never rise to the front of his mind where he could catch them and say them out loud. Later he told me the whole thing had “lived largely in his head,” never as a written set of rules. He was not being evasive. He had the process. It was the judgment inside it that had never been stored as words.
Every business I have ever worked inside has a version of this person. The one who can look at a situation for four seconds and tell you it is going to be a problem, and be right, and not be able to tell you how they knew. When they are in the room, things go smoothly. When they are out sick, or on vacation, or thinking about retiring, everyone quietly holds their breath. Their judgment is the load-bearing wall nobody drew on the blueprints.
We call it tribal knowledge, or institutional memory, or “that’s just Dave.” Whatever we call it, we treat it as a person-shaped risk we can never quite close. I used to treat it that way too.
Why the usual fix fails
The standard move is to try to get the expertise out of the person and onto a page. So we interview them. We build a standard operating procedure. We stand up a wiki, hand it to the team, and tell ourselves the knowledge is now “documented.”
It never really works, and I think I finally understand why.
What an SOP captures is process. The sequence, the handoffs, the order of operations, who signs what. That is the repeatable part, and the repeatable part was never the scarce thing. You can hire for it. You can write it once and be done. What the document almost never captures is the judgment exercised at each of those steps, which is the entire reason you wanted that particular person in the first place. So we produce a document that is accurate, complete, and beside the point.
When you interview an expert, you get the rules they can articulate under the pressure of being watched. That is a thin slice. It skips everything they do on instinct, because instinct does not announce itself. Ask a great operator how they price a hard job and they will give you three factors. Watch them price forty hard jobs and you will count fifteen. The other twelve fire automatically, below the level where they can even notice them to say them out loud.
So the manual comes out thinner than the person, always. It captures the rules that were already easy to say, which are usually the rules you least needed help with. The hard, valuable judgment (the part you were actually trying to preserve) is exactly the part that does not survive the interview, because it was never verbal to begin with.
Then the document goes stale, nobody trusts it, and everyone goes back to just asking Dave. We have spent years and real money on this cycle. The problem was never effort. The problem was the format. We kept trying to read judgment as if it were stored as instructions. It is not.
The reframe: read the decisions, not the person
Here is the shift that changed how I think about this.
Judgment is not stored as instructions. It is stored as decisions. What someone actually did, over and over, across years, is the most honest record of how they think that you will ever get. It does not flatter them. It does not skip the instinctive parts. Every call they made is sitting right there, already made.
So stop asking the person to explain, and go read the decisions instead.
The best place to read them is a project that is already over. A closed, finished engagement has one enormous advantage that live work never will: you already know how it turned out. You are not guessing whether a call was good. You can look at what the expert decided and then look at what actually happened next, and grade one against the other. A finished project is not old news. It is a fully answered exam. The answer key is included.
Most organizations sit on top of years of these answered exams and treat them as dead storage. Old files. Closed matters. Archive. I started treating one of them as the richest training data in the building, because that is exactly what it is.
The extraction
So we took one closed engagement and read all of it. Six years of one expert’s decisions, start to finish: 18,511 items in all, 719 documents and nearly eighteen thousand emails. Most of the volume is the email traffic, which is exactly where the real decisions get made and argued out in real time. Not the summary. Not the highlight reel. The whole corpus, including the boring parts, because the instinct we were hunting hides in the boring parts.
Then we read it seven ways at once. Seven passes over the whole record, each one hunting a different family of decision, all of them looking for the same thing: a shape that repeats. The patterns came back as plain, reusable rules. If this and this, then he tends to do that, and here is why.
About 85 of them came back. We set aside the ones outside what we were building and brought him 73, and I walked him through every single one. He approved all 73. On 16 he added a correction or a comment, sharpening the wording or noting an exception, which is exactly what you want from someone with that much time in the seat. But not one of the 73 was wrong. They were his.
And here is the number I cannot stop thinking about. 35 of those 73 rules were things he had never written down for us anywhere. Not in the questionnaire he filled out, not in years of his own email. They were real, they were his, and they had simply been running, quietly, inside choices he made and moved past without ever stopping to name them.
Watching him read them was its own thing. He got quiet. It is a strange feeling to see your own logic laid out on a page in front of you, correctly, when you have never once been able to say it out loud. That is the moment I knew this was different from every documentation project I had ever run.
The single best thing we got did not come back as a rule at all. It came back as a correction. We had mined something about how he decides where money should come from, and we had it about half right. Instead of ticking approve, he wrote out the whole thing: six questions he asks in order, six different places the money can come from, and one line I have not stopped thinking about. Deciding which pot the money comes out of matters more than negotiating the amount. That is judgment, exactly the part that never makes it into an SOP, and it turned out to have a completely repeatable structure sitting underneath it. He had simply never had a reason to say it. Asking him to describe his process would not have produced it. Showing him a rule that was close but not quite right did.
One rule in particular stopped him, because it was so him and he had never once named it. When an outside party goes slow on something he needs, an approval he cannot force, a gate he does not control, most people lean on the slow party directly. His move is the opposite. He finds whoever sits closest to that slow party, the one with the standing relationship, and hands them the exact thing to go push for. He does not chase the blocker himself. He delegates the chase to the person best positioned to win it, and treats keeping the flame up as a job of its own. Reading it back, he half laughed. Of course that is what he does. It had just never been a sentence before.
The test that makes it trustworthy: the time machine
A page of rules that sounds right is worthless. Anyone can write plausible rules. The question is whether they actually reproduce the expert’s judgment, and the only honest way to answer that is to make them prove it against decisions we already know the outcome of.
So we built a time machine.
Here we went narrow on purpose, and that restraint is the point. We pulled 241 real instances of one specific, recurring kind of call, the same decision he makes over and over, and had the system re-decide every one of them from scratch, blind. The rule was simple and strict: it could see only what was knowable at the moment of that decision. Nothing after. No hindsight, no outcome, no idea how the situation eventually resolved. We put it back on an “as of then” clock and made it call the shot with exactly the information the expert had in his hand at the time. Then we graded its call against what actually happened.
The hard part of a test like this is not running it. It is stopping the machine from cheating, because a capable system will find every possible way to peek at the answer.
So we fenced it in. Not one door closed at a time, a whole perimeter. We walled off the future, so it could only ever see what was knowable at the moment of the decision and not a single thing that came after. We walled off the outcome, so it could not reason backward from how things ended. And we walled it off from our own work: when a rule was tested against the very case its example had been mined from, we stripped that example back out, so the machine could not simply recognize an answer we had already handed it. It only ever got to stand where the expert stood, with what he had in his hand at the time, and nothing we had learned since. A rule only counts if it works on a case it has never effectively seen.
Then the honesty control I care about most. We ran all 241 cases twice. Once with the mined rules switched on, and once with them switched off. That produced 482 drafts. If the “rules on” version was not clearly better than the “rules off” version, then the rules were decoration and I needed to know it. You do not get to claim your system works until you have watched it fail without the thing you say makes it work.
That is the difference between a demo and a test. A demo shows you the good case on purpose. A test tries to catch you lying and reports what it finds.
What it means for operators
Strip out the specifics and what is left is a playbook you can run on almost any expert-dependent part of a business.
You need four things. A closed body of past decisions, which you already have and are probably calling an archive. A way to read those decisions and pull out the patterns underneath them, which is the part that got cheap in the last two years. The expert, not to write the manual, but to validate the manual the decisions wrote for them, an afternoon of “yes, no, sharpen this” instead of weeks of blank-page interviews. And a blind replay against known outcomes, so you trust the result because it earned trust, not because it sounded good.
Notice what the expert’s job becomes. Not author. Editor. It turns out people are dramatically better at judging a rule than at generating one from nothing, which is the exact failure the old interview was built on top of. We inverted it. The machine drafts, the human corrects. That afternoon of corrections is where twenty years of instinct finally makes it onto the page.
I have started to think the competitive edge here does not go to the companies with the best-documented processes, because process was always the easy half and everyone’s is roughly as good as everyone else’s. It goes to the ones who realize the other half is documented too, that they have been storing it in the wrong format and calling it dead weight. The decisions are the documentation. They always were. Everyone has this data. Almost nobody is reading it.
Close
We spent years trying to get experts to document themselves, and it barely worked, and we blamed the experts. That was never fair. You cannot hand someone a manual for something you have never had words for.
But the record was there the whole time. Every call that person made, weighed against what actually happened, sitting in six years of files nobody thought to read as a manual.
The manual was already written. We just were not reading it.
Questions
Why do SOPs and documentation projects fail to capture expertise?
Because an SOP captures process: the sequence, the handoffs, the order of operations. That is the repeatable part, and the repeatable part was never the scarce thing. What the document almost never captures is the judgment exercised at each of those steps, which is the entire reason you wanted that particular expert in the first place. The result is a document that is accurate, complete, and beside the point.
Why can't experts explain their own judgment?
Judgment is not stored as instructions. It is stored as decisions. When you interview an expert you get the rules they can articulate under the pressure of being watched, which skips everything they do on instinct, because instinct does not announce itself. Ask a great operator how they price a hard job and they will give you three factors. Watch them price forty hard jobs and you will count fifteen.
What does it mean to read the decisions instead of interviewing the person?
It means treating a closed, finished project as the source material rather than the expert's memory. A finished project has an advantage live work never will: you already know how it turned out, so you can look at what the expert decided, look at what actually happened next, and grade one against the other. A finished project is a fully answered exam, and the answer key is included.
How do you know the extracted rules actually reproduce the expert's judgment?
You make them prove it against decisions whose outcomes you already know. Re-decide past cases blind, with the system able to see only what was knowable at the moment of that decision, then grade its call against what actually happened. The critical control is running every case twice, once with the mined rules on and once with them off. If the rules-on version is not clearly better, the rules were decoration.
What do you need to run this on your own business?
Four things. A closed body of past decisions, which you already have and are probably calling an archive. A way to read those decisions and pull the patterns out of them, which is the part that got cheap in the last two years. The expert, not to write the manual but to validate the manual the decisions wrote for them. And a blind replay against known outcomes, so you trust the result because it earned trust rather than because it sounded good.