A caller says the printer is broken. She has power-cycled it, reinstalled the driver, and swapped the cable, because a previous call taught her that is what we were going to ask. She is not guessing. From her desk the world is: I press print, paper comes out. Today no paper. Therefore printer.
That call has a hundred variations and the printer is rarely the fault. I took a lot of them, first at Gateway and then at GoDaddy. After enough of them you stop hearing the stated problem as a claim about reality. It is a report from a specific vantage point, and the shape of the report tells you where that vantage point is. That turned out to be the most transferable thing I have ever learned, and I did not understand it as organizational knowledge until much later.
The stated problem describes the reporter
At Gateway, people called about the computer. The computer was slow, the computer was broken, the computer had a virus. Almost nobody called to say they suspected their available memory was insufficient for the number of startup items some trialware bundle had installed. They called about the thing they touched. The boundary of the described problem was the boundary of what they could see and reach.
GoDaddy taught the same lesson one layer up. Someone calls and says "my site is down." Sometimes the site is down. More often the site is fine and their DNS points somewhere else, or a registrar lock is doing exactly what they asked it to do, or they are looking at a cached page, or they are looking at the wrong domain entirely because they own several. The phrase "my site is down" is doing a lot of work. The caller can see the browser window and nothing else in the stack, so the browser window is what they describe.
Here is the part that took me a while: the phrasing is worth as much as the symptom. If a caller says "my site is down" I learn they are checking from a browser. If they say "my nameservers aren't propagating" I learn they have read something, possibly something wrong, and that I now have to find the edges of a mental model before I can do anything useful. If they say "you people broke my site again" I learn there is history, and some of the next twenty minutes belongs to the history rather than to DNS.
The bad move — the move I made constantly when I was new — is to accept the framing and start working inside it. Accept "the printer is broken" and you can spend forty minutes on a printer that is not broken, then hand over a new printer and the same problem. The discipline is to treat the framing as evidence about the reporter while reserving judgment about the fault.
Departments call in the same way callers do
Sales reports that the quoting process is too slow. Finance reports that sales is quoting outside policy. Operations reports that the orders coming in are incomplete. Support reports that customers are angry about delivery dates nobody in support set. Every one of those statements is an accurate description of what that department can see from its desk, and each one names a different villain.
Often that is one problem wearing four costumes. A quoting tool requires a field operations needs. Sales cannot fill it in at quote time because the customer does not know yet, so sales puts in a placeholder to get the quote out. Finance sees the placeholder as a policy exception. Operations receives an order with a placeholder in it. Support eventually talks to a customer about a date that was invented to satisfy a required field. Every person in that chain did the sensible thing available to them. Every department hits its target and the company still misses.
These are the problems that do not belong to any one department, which is the work I do now. Sales can tell you precisely what the tool will not let them do. Operations can tell you precisely what arrives broken and how often. Finance can tell you which rule gets bent and in which direction. None of them can tell you the sequence, because nobody stands where the sequence is visible.
Finding the sequence is not mysterious, but it is not free either. On that kind of engagement I sit with the people doing the work and watch an order move end to end, several times, rather than asking anyone to describe it. I count the places where someone re-enters data that already exists somewhere else, and the places where a person pauses to decide something the system should have decided. Then I take the tickets, the quotes and the exception log for the same period and line them up on one timeline. The required field shows up as a cluster: the same workaround appearing in four systems under four different names within a few days of each order.
Sometimes four departments genuinely have four problems. That happens, and the tell is usually timing and overlap. If the four complaints started at different times, involve different customers, and stop for different reasons, they are four problems and I should stop looking for an elegant single cause. If they cluster on the same orders and move together when volume moves, one mechanism is producing all four. You can check that in two days with a sample of orders. You do not need a week to find out whether the complaints are about the same transactions.
I want to be careful about one thing. There is a body of work on organizational noise — Kahneman, Sibony and Sunstein — about how much judgments vary when they should not. It is genuinely useful, and it explains some of what I find. It does not explain most of it. The distinction I actually use is between variance that comes from judgment and variance that comes from a constraint. Noise is two underwriters reading the same file and pricing it differently. A constraint is a required field nobody can fill in honestly, so twelve people invent twelve workarounds that look exactly like inconsistent judgment and are not. The remedy is completely different. You fix noise with calibration and checklists. You fix a constraint by removing it.
What a scrapped part actually tells you
Before any of the support work I ran machines, manual first and then CNC programming. That is where the reflex comes from, though I had no language for it at the time.
A part comes off the machine out of tolerance. You can stand there and be annoyed at the operator, and occasionally that is correct. Usually it is not. The part is a physical record of the conditions that produced it, and read properly it will tell you which conditions. Oversize across the entire run is a different animal from a drift over the run, which is different again from one bad part in the middle. Consistent points at the setup or an offset. Drift points at heat, wear, or something working loose. One bad part in the middle means something happened once, and now you want to know what else was happening at that time of day.
Nobody in a machine shop argues with this. The scrapped part is measurement, not an accusation. You put it on the surface plate and let it tell you about the setup that made it.
Organizations produce scrap too, and almost nobody reads it that way. A missed handoff, a rework loop, a customer escalation, a process everyone quietly routes around — those are output from a setup. When the response is to identify who was holding the controls, you have thrown away the measurement and kept the blame, and you will make the same part again next week on the same setup.
The taxonomy carries over exactly, and this is the most useful thing I can hand anyone. A failure that shows up on every order is the setup: a rule, a required field, a form, a policy that makes the honest answer impossible. A failure that gets steadily worse across a quarter is drift: load rising against fixed headcount, a queue growing, a threshold nobody reset. A single bad escalation in an otherwise clean month is an event, and the question is what else happened that day. Sorting your complaints into those three buckets before you sort them by department will change what you look at first.
The sequence I use is not clever. Observe, measure, map, identify, experiment, measure again. The last step is the one that gets skipped, and skipping it is how organizations accumulate improvements nobody ever checked. A recommendation nobody tests is just an opinion in a nicer font.
Why I get suspicious when a problem arrives pre-diagnosed
"Our project management tool is the bottleneck" tells me someone has already decided the answer is a purchase. "The new team lead is not working out" tells me the pain became visible around the time that person arrived, which is information about timing and not necessarily about the person. "We need better communication" tells me almost nothing about the mechanism and quite a lot about how tired everyone is of discussing it. Each of those is an accurate reading from where the speaker stands. Each one also rules out everything upstream, and upstream is where the required field usually lives.
Years of assessing other companies' security programs from the outside sharpened this more than anything else I have done. You send a questionnaire, you get back a confident description of how access reviews work, and then the evidence arrives and shows the review ran once, eighteen months ago, for one system. The person who answered was not lying. They described the part of the company they could see, and what they could see was the policy document, not the practice. Multiply that by hundreds of companies and you stop treating any self-report, including a client's opening statement, as a description of the system.
A pre-diagnosed problem also arrives with a scope attached, and the scope is the giveaway. If the problem is described entirely inside one department's boundary, either it genuinely is inside that boundary, or that boundary is the edge of the reporter's vision. Those two look the same on the phone. They do not look the same after a week of watching the work actually move.
The stance has a cost and I have paid it. I once spent most of a week on a slow month-end close looking for the cross-functional mechanism I was sure had to be there, because the complaint came from three groups at once. There was no mechanism. One reconciliation genuinely was slow, for the boring reason that the data came out of a system that could only export it one way, and the other two groups were downstream of that and complaining about the wait. The stated problem was the actual problem. I lost four days to my own assumption that it could not be. Reading the framing as evidence is a habit, and like any habit it is wrong sometimes and expensive when it is.
So the thing I try to hand over is not a theory about what is wrong. It is a map of how the work actually moves, with the three or four places marked where someone has to invent something to keep going, and a small test for each one with a number attached and a date to check it. Sometimes that ends with a recommendation to buy nothing and delete a field. An MSP is paid to keep answering tickets. I am trying to work out why the tickets keep happening.
