18 min read

The Wrong Denominator

Two years ago I could have told you exactly what my engineers were working on. Jira knew, because one person worked on one thing. That assumption broke, and the number I got for free went away without anything appearing to break.
The Wrong Denominator
Photo by Antoine Dautry / Unsplash

I run about twenty engineers. Between us we have something like two hundred and fifty concurrent agentic threads of work running at any moment. Each one is a working copy with an agent moving it forward. Until recently I could not have told you how many of those threads were pointed at the thing I said mattered most. Not within a factor of two.

That is a strange thing for a CTO to admit, so let me be precise about what I mean. I knew what every project was. I knew which ones my leads were worried about. I could walk the floor and come away with a good read on how things were going.

What I could not do was say what fraction of the organization was working on any one of them. If I had told you our biggest bet should get thirty percent of us and asked whether it was getting thirty, I would have guessed.

Two years ago I could have answered it exactly, and I did not think of that as an achievement. We worked from Jira. One person worked on one thing, and when they finished they marked the ticket done. Planning was slow enough that it all made its way into the tracker before anyone started, so the tracker was a reasonable description of what engineering was doing. Count the tickets, group them by project, and there is your answer. Nobody called that instrumentation. It was just the byproduct of a work system where the unit of work and the unit of person were the same thing.

That system did not erode. It died fast. Planning got quick enough that everyone does it constantly instead of in a cycle, and once that happened, stopping to write a plan into a tracker before starting stopped making sense. One person no longer works on one thing. One person directs ten or twenty threads at once. Jira went from being how we planned to being something we no longer really used, in a matter of months.

It still runs, in zombie form. The rest of the company is not agentic and uses Jira out of habit and convenience, which is a perfectly good reason to use it. So the tracker still loads and still has things in it, and none of those things add up to a description of what my engineers are doing.

What I noticed too slowly is that losing the number was invisible. Nothing broke. No dashboard went dark. I simply stopped being able to answer a question I had never had to work to answer, and in the gap I formed opinions from whichever projects came up most in conversation. Which is a measure of what is loud rather than what is large.

I also had no good way to change the answer. I could change what the organization worked on, but only by going and moving a hundred separate levers by hand. Talk to a lead. Reprioritize a backlog. Kill a project. Each one is a conversation, each one lands weeks later, and none of them tell you whether they worked.

So I set out to build a steering wheel. What I ended up with was closer to instruments, and the difference is most of what I learned.

What it looks like

Three moments, so you know what we are talking about.

An engineer has a watch window open on the small monitor under her main one. We buy those, twelve or fourteen inches, and they sit below the display where the actual work happens. Hers lists her stacks and updates itself as they move, all day, without her asking. At the bottom of that list is a footer. Nine threads live, four of them on automation work, two on new product, three spread across everything else. Beside her numbers are the organization's: new product thirty percent, automation twenty, and so on down. She is heavy on automation and light on the thing we said mattered most, and she has not had to go look that up.

She finishes one of them and starts a new thread. The tool suggests what to point it at before she has decided, and the suggestion names a goal along with two numbers: this is what the organization said it wanted here, and this is what it is actually getting. She can ignore it with one flag and often does.

Monday morning, I open the dashboard. Two hundred and forty-seven live threads across nineteen of twenty-two engineers. New product declared at thirty percent, running at eighteen. I click that slice and get a list of who has a thread on it and who does not, and the second list is the one I take into my week.

That is the whole system. Every number in those three moments is the result of a decision that could have gone another way, and one of those decisions was much harder than it looks.

Two things you need to know about

None of what follows makes sense without knowing how work here is visible at all.

Running twenty threads at once is not something a person does out of a notebook. We have a tool for it, called Stack. Everyone here works in one monorepo. A stack is a working clone of it with an agent moving it forward, and running twenty threads means twenty of those clones open at once. Stack is how an engineer manages that: creating clones, linking each to whatever it advances, marking which are urgent, and watching all of them at once. Most people keep its watch window open all day on a dedicated screen, updating as threads move, because that is the only practical way to hold twenty of anything in your head. It is the human surface of agentic work, and it exists because directing a fleet is hard.

Which means Stack already knows almost everything: what each thread is working on, whether it is moving, who owns it. Nobody built any of that for measurement. It is there because an engineer running twenty threads needs it there for themselves.

The second tool is Project Telemetry, which is where the organizational picture lives. Stack forwards what it knows, Project Telemetry assembles it across everyone, and the result is a view of everything happening across the engineering organization rather than one person's slice of it.

That division matters more than it sounds. I did not build an observation layer and point it at my engineers. The observation was already a byproduct of a tool they use for their own benefit, and all the real work was in deciding what to do with it. Which turns out to be a question of arithmetic.

What I cannot know

Before any of the design, there is a constraint underneath it that I want to state on its own, because everything else in this essay is downstream of it. It was not a discovery. It was obvious from the start, and I have been working backwards from it the whole time.

Each engineer knows their capacity. I do not, and I cannot.

By capacity I mean how many threads a person can actually hold at once and still be doing good work on all of them. That number is real, they can feel it, and it is different for everybody. I run about twenty-two. Some of my seniors run fifteen to twenty. A second-year engineer on my team runs four, and four is right for her. The range across twenty people is wide enough that no single figure describes anyone.

It also moves, for two reasons at once. The person changes: nobody is at the same number in a week with an incident, or three days of interviews, or the days after a release. And the threads change. Some take enormous cognitive effort and some run themselves, and which is which changes on the same thread from one week to the next. A migration that was consuming someone entirely on Monday might be waiting on a test run by Thursday. Twelve threads can be a heavier week than twenty, and the person holding them knows it without doing any arithmetic, the way you know whether you are overcommitted.

The word cannot is doing specific work there. Not that it would be rude of me to know, or that I have decided not to look. There is no observation available to me that produces the number. It exists in the engineer's own read of their week, and the closest substitute I could build is watching how long someone sits at a thread, which measures presence rather than effort. It would not get me the number even if I were willing to collect it.

This is a fact about the situation rather than a policy. The shape of the ground I am building on, not a line I drew on it. I am not allocating anyone's capacity from above, because I cannot observe it precisely enough to do so. I affect it constantly through staffing, deadlines, interruptions, and what I ask for. What I cannot do is sit above it and apportion it, because the engineer with twenty-two threads is the only person who can tell whether twenty-two is right for them this week.

Some people will fight me on this, so let me meet it directly. There is a school of management that does not trust engineers with any of it. It wants the unit of work small and assigned, decisions made above and executed below, the engineer converted into a task doer whose throughput can be watched. Plenty of organizations run that way and the people who prefer it usually believe they are being responsible.

I have rejected that view for my entire career and I am not going to start now. I hire people capable of making their own decisions and then let them make them. My work is the gardening: clearing conditions, removing obstacles, and making the direction unmistakable. Point them the right way and get out of the road.

It is also the only position that scales here. When one person ran one ticket, an organization could micromanage and merely be slow and unpleasant. When one person directs twenty threads, there is no manager alive who can hold the state required to make those calls for them. The scheme that worked badly before does not work at all now. Direction is the only lever left that reaches, and a compass beats a set of instructions when you cannot see the terrain the other person is standing on.

That changes what the system is for. If I cannot allocate anyone's capacity, the useful thing I can do is make sure they are managing it against the same picture I am. Not a report I read about them. Information already in front of them, which is why that footer matters more than my dashboard does.

Everything below is downstream of that. I mention it now because the design decisions that follow will look arbitrary in isolation, and most of them are this constraint working itself out.

The obvious design

The mechanism is simple. None of the hard parts turned out to be mechanical. Write down the goals the organization actually has, and give each one a share of attention. Not a rank, a share.

Ranked priority lists are the default everywhere, and they fail in a specific way once an organization runs many things at once.

Order was a reasonable thing to communicate when work happened in sequence. One engineer, one task, finish it, take the next one off the list. A rank is exactly the right instruction for that, because the question the engineer is asking is what to do next, and a ranked list answers it.

That is not the question anymore. Nobody here is doing one thing and then doing another thing. They are running ten or twenty in parallel and dividing themselves across all of them, and the question they are asking is how much of me goes where. A ranked list has no answer to that. Taken literally it says everyone should be on item one until item one is finished, which nobody means and nobody does, so the list gets interpreted privately in twenty different heads. Some read it as strict order, some read the top three as equal, some read their own project's position as a statement about their standing. The list looked like a decision and was a prompt for twenty separate decisions.

The mismatch is not that rank is imprecise. It is that rank is a sequential instrument, and the work stopped being sequential.

A share says something a rank cannot. Thirty percent means that of everything the organization has in flight, roughly a third should be pointed here. That is checkable: you can look at what is running and see whether it is a third. There is nothing you can look at to see whether an order is being respected.

Each share also carries an articulation, one paragraph saying what the goal is and why it holds the share it holds. Our largest slice states in plain language why it is the largest and what we gave up to make it so. That does work a number cannot: it hands an engineer the reasoning, so they can apply it to a case I never anticipated. A rank gives an instruction. A share plus a reason gives a heuristic somebody can run without me in the room.

The shares live in one file, one goal per section, each with its percentage in the heading. We have a handful of goals, and their percentages add up to eighty.

Not a hundred. Eighty, with the remaining twenty not assigned to any goal, on purpose. I would defend that gap harder than anything else in the design.

Those twenty points are the discretionary share. Not slack, not rounding. A fifth of everything we have running is expected to be work an engineer picked because they thought it was worth doing, and no part of the system asks them to justify it.

Writing it as a number does something that saying it never did. Every organization I have worked in claimed to want engineers following their instincts, and in every one of them that work was first to go when a quarter got tight, because it was the only work with nobody assigned to argue for it. A goal has a sponsor. Autonomy has nobody. Give it a share and the arithmetic defends it: a new goal's points come from somewhere, and if they come from this slice the diff says so and a reviewer watches the number go down. It cannot be eroded quietly, and quietly is the only way slices like that ever get eroded.

Claiming that share is something you do on purpose. Every thread names the goal it advances, and to put one in the discretionary slice you name discretionary. You are saying, on the record, I chose this myself.

So what about a thread that names no goal at all? It could have counted as discretionary. Somebody is working on something they picked, after all. It does not. Those threads land in a separate pile of work nobody has labeled, and that pile is a worklist. Each entry resolves two ways: either it does advance a goal and nobody said so, or the person picked it themselves and it is discretionary. Both take seconds.

The part that matters is who owns that chore. It is mine. If a handful of threads are unlabeled, somebody was moving fast and did not stop to categorize. That is nothing. If a lot of them are, the honest reading is that my goals do not describe what my organization is doing, and no amount of asking engineers to label things more carefully fixes that. The fix is to go change the goals. So the size of that pile is the one number in this system that measures me rather than the people working for me. It needs to stay its own number for that reason alone.

Which is the reason the distinction is worth the pedantry. If unlabeled work silently counted as autonomy, a goal list slowly going stale would read as a thriving culture of engineer initiative. The number that should have been telling me my goals were wrong would instead be telling me my engineers were flourishing. The failure would look exactly like the success.

Keeping them apart costs one word from whoever creates the thread. Autonomy is a thing a person chooses. It should have to be said out loud.

Changing any of this is a pull request.

That constraint is the entire point, and it is the part most organizations skip. A wall of aspirations is free. Everybody has one. You can add a priority to a list of priorities at no cost, which is exactly why those lists grow until they mean nothing. A pie cannot work that way. Raising one slice requires lowering another in the same diff, in front of a reviewer, with your name on it. The document stops being a statement of what you care about and becomes a record of what you were willing to give up. Those are very different artifacts, and only one of them is a decision.

It also means the trade is never invisible. When a new goal takes ten points from somewhere, the diff says where. I did not appreciate how much of the value was in that until I watched the first few of these get reviewed.

Then approximate the other side. Take the distribution of live threads as an observable proxy for where attention actually went, and subtract it from the declared distribution. Two distributions subtract. Simple enough, except that the word approximate is carrying more weight than it looks like it is.

I still think that part is right. The measurement half is what taught me something.

The denominator

Measuring the other side sounds like counting. Before you can count anything, though, you have to decide what you are dividing by.

Here is the choice, and it is the only interesting decision in the whole system.

Take my twenty-two threads and that second-year engineer's four. When I roll those up into one organizational picture, I can do it two ways. I can give each person equal say, dividing my twenty-two by twenty-two and her four by four, so we each contribute one unit. Or I can count threads, so my twenty-two are twenty-two and her four are four, and I move the number roughly five times as much as she does.

The first one looks obviously right. Fair, statistically careful, and proof against anybody whose tooling habits would otherwise swamp the picture. One person, one vote. I would guess most people reaching for this problem land there without much deliberation, and I understand why.

It is also the wrong denominator for what I was building, and not because it is unfair in the other direction. It answers a different question. Divide each person to one and you are asking where the average engineer's portfolio is pointed, which is a perfectly coherent thing to want to know and might be exactly right for a different purpose. Count every thread and you are asking where the live fleet is pointed. Two real questions. The denominator is where you choose between them, silently, usually without noticing you chose.

I need the second one, and the reason has nothing to do with fleets mattering more than people. It is that the thing I am building acts on threads, one at a time, at the moment somebody starts one. If the population you measure is not the population you act on, the number cannot tell you whether the action worked. Twenty live threads are twenty objects in the fleet I am trying to steer. Four are four. Whether those twenty consume five times the human effort of the four is a separate question, and my instrument cannot answer it and does not try.

So: normalize once, over the population you are measuring, and never per person inside it. The organization divides by every live thread. A single engineer's view divides by their own threads. Both answer the same question at different sizes, which is what makes them comparable. Per-person normalization inside the organizational view divides once by each person's thread count and again by headcount. That cleanly answers the average-engineer question. It does not answer the fleet question, and the fleet question is the one this system acts upon.

One consequence of that looks like a bug. High-concurrency engineers move the organizational picture more than low-concurrency engineers do.

That is not a position I am arguing for. It is what is already true. Somebody running twenty threads is directing more of the live objects this instrument can steer than somebody running four, on any given Tuesday, whether or not I ever build the instrument. Whether those objects consume five times the human effort is a separate question and the instrument makes no claim about it. All a per-person denominator would do is decline to show the part that is observable.

I should say the obvious thing here. I am the person with twenty-two threads. The rule I chose gives me more weight in my own measurement than the rule I passed over would have. I noticed that, and I sat with it for a while, and I still think the fleet denominator is the right one for an instrument that steers threads. But you should know I had a stake in the answer, and you should weigh the argument accordingly rather than taking my word for it.

What else the unit decides

Once you count threads, a few other choices that look independent turn out to be the same choice again.

The obvious next move is to weight them. Our threads carry a marker for whether they are urgent, the ones with a deadline and somebody waiting, versus the ones being advanced whenever there is room. Count an urgent thread as two and the rest as one, and the number gets more realistic. Urgent work does absorb more of a person.

That weight asserts a ratio, and I cannot observe the ratio. Is an urgent thread worth twice a background one, or one and a half, or four? It varies by person, by week, and by the thread. There is no measurement I can take that settles it, so any number I pick is a number I made up and then dressed as arithmetic. The marker is also self-declared, which means a weight attached to it is a weight anyone can inflate by relabeling.

Threads count equally, then. The argument is not that urgency does not matter, because it plainly does. The argument is that I know which threads exist and I do not know what fraction of a person each one takes. Counting them equally measures something I know. Counting them unequally measures a guess.

Liveness is the more interesting case. The simple approach is to have people mark threads dormant when they park them, and it fails for a reason that has nothing to do with discipline. Nobody is going to do bookkeeping across twenty-two threads to keep a dashboard honest, and I would not either. Any count that depends on that maintenance is a count that quietly drifts. So dormancy has to be derived from whether the work actually moved. That sounds obvious until you notice the trap: if you derive liveness from telemetry heartbeats, a terminal somebody left open in a forgotten tab keeps a dead thread alive forever. Presence is not movement. The system now looks at whether the working copy changed, not whether the tooling phoned home. That distinction is the single most important line in the implementation.

There is an asymmetry underneath both of those decisions that I keep coming back to. A thread wrongly marked dormant fixes itself the instant somebody touches it. A thread wrongly counted as alive stays wrong indefinitely. When your errors have that shape, be aggressive. Prefer the mistake that self-corrects.

What I refused to build

Instrumenting people is where this kind of project goes bad, so here is what I would not do.

The data is open. Every engineer can see every other engineer's distribution, including mine. There is no privacy guarantee and I did not offer one. I think open metrics are correct and I think pretending otherwise is worse than being direct about it.

What the system does not do is compute a verdict. No scores. No rankings. No thresholds. No alerts about individuals.

That is easy to write in a design document and harder to hold, because the violations do not arrive labeled. Here is the one that nearly got past me.

Go back to that Monday morning list of who has a thread on the underserved goal and who does not. The obvious version of that view is sorted, worst first, every engineer with how far their distribution sits from the target. Useful. Actionable. Exactly what somebody in that situation is asking for.

It is also a ranking. Sorting people by deviation means computing a per-person number that says how far off each of them is, and then presenting those numbers in order. Nothing about that changes because the number lives in a sort order instead of a column. A rank reconstructed in a sort is still a rank, and I would have shipped the thing I said I would not build while believing I had not.

So the view does not sort. It puts people in two groups, those with a thread on the goal and those without, alphabetical inside each. That answers the question a manager actually has, which is who has no exposure to this, and it answers it better. The sorted version would have looked precise while quietly encoding every measurement problem I have already admitted to. Somebody eleven points from the target is not measurably different from somebody four points away, and an ordinal list insists they are.

The deeper constraint is epistemic rather than merely ethical, and it is the capacity problem again. The measurement sees breadth and not depth. A goal with fifteen threads on it might have five people committed to nearly nothing else, or fifteen people giving it a corner of their week, and my aggregate cannot tell you which. Depth is the part that lives with the engineer, and I decided at the outset not to go after it. Having made that choice, I do not then get to take a number with that much slack in it and turn it into a rank ordering of human beings.

The wager

Here is the thing I am least sure about, stated as plainly as I can.

I called this system attention allocation. It does not measure attention. It measures the direction of the live fleet, and I am betting those two things track each other closely enough to steer by. An engineer with twenty threads at five percent of themselves each counts five times an engineer with four threads at a quarter each, and nothing in my system can tell those two situations apart.

The bet is that in an organization where the domain depth has moved into the agents and the humans are directing fleets, direction of the fleet is the best available proxy for where the organization's effort is going. That may be wrong. If it is wrong, I know what the failure will look like: the distribution will read healthy while a handful of high-concurrency people determine it and everyone else's work barely registers.

There is a second way this fails, and it is worse, because nobody has to be acting in bad faith for it to happen. If our tooling, or the nature of the work itself, causes some work to decompose into smaller threads, the proxy moves underneath me. Infrastructure work may naturally produce fewer and larger threads than product work does, which would tilt the picture toward product without anyone deciding to tilt it. A person or a goal that produces more and smaller threads gains weight in the measurement without directing more work. That is not a metric that is merely imperfect. That is a metric that can change meaning as the work system it measures changes, which is how something becomes gameable without anyone choosing to game it.

So I am watching for both of those rather than assuming them away. I do not have a solution to either. I have a name for each of them and an intention to look, which is the most anyone can honestly claim about a system built on an assumption they cannot yet verify.

I did not build a measure of human attention. I built a measure of direction, weighted by volume. Thread count carries some magnitude signal, and that is why the denominator matters at all: if threads were pure direction with no size behind them, counting the fleet rather than counting people would be a much smaller decision than I have made it out to be. But the magnitude is coarse and I cannot calibrate it. It tells me twenty is more than four. It does not tell me how much more.

The instrument cannot tell me how hard the organization is pulling. It can tell me which way the fleet is pointed, roughly how much of the fleet is pointed each way, and whether that is the way I chose.

For the question I opened with, that turns out to be enough.