An operator who does only what a frontier model would have done is worthless. Anyone could have asked the model and gotten the same result.
To measure the rest, replay their sessions and ask, at each move, what the frontier model would have done instead: the moves the model cannot match are the operator’s skill.
I built a tool that runs this measurement on agent transcripts, and I ran it on my own sessions first. The report opened with a line I did not write about myself: “He treats the model as a drafting surface, not an authority.” I posted it where my colleagues could see it.
One of them read it and was sure it was about him. “I related to the interaction style,” he wrote, and after learning whose session it was: “I had to learn it was you before being sure it wasn’t myself they were talking about.” He had never seen the session. A machine-written portrait of one person’s prompting was vivid enough that a reader mistook it for himself.
Everyone who works with these models senses that operating them well is a real skill. Almost no one has seen it measured. The hunger for the measurement is strong enough that people project themselves into it.
The model is the constant; you are the variable
A frontier model is a public utility, identical for you, for me, and for your competitor. When everyone holds the same instrument, the instrument stops being the differentiator.
Call what the model produces for anyone who simply asks the baseline: the average answer, the plausible plan, the menu of options it hands to every comer. Call everything in your output that the baseline does not explain the surplus. Your agentic skill is the surplus, and it has a sign. Match the baseline and your skill is zero, whatever your title says. Fall below it, by truncating the model’s reasoning or pasting its first draft into production, and your skill is negative. Your employer is paying for worse than a subscription.
Most operators are middleware with a keyboard
Most people’s surplus is near zero. Watch the average session. A request, a skim, a paste. The model asks a question, they answer it. The model offers three options, they pick one. The model says impossible, they believe it. They are not operating the model. They are middleware with a keyboard, moving text between the model and the ticket tracker, and everything they produce was already in the model, waiting for anyone to ask.
A small class of operators routinely leaves a session carrying results the model called impossible. Same model, same tokens, different species of result. Nobody measures the difference. The industry hires, pays, and promotes on vibes, while the one quantity that captures an operator’s value goes unrecorded.
The surplus is judgment the model cannot supply
What does the surplus look like from inside a session? My report gives four answers. When the model offered a menu of options, the operator dissolved the menu by reframing the question at a different altitude. When the model called a capability impossible, the operator declined the verdict, and he was right. When the model’s own design had holes, variance across re-runs, files vanishing from output sets, the operator found them.
The mistakes matter most. Wrong turns and abandoned approaches mark the points where human judgment and machine judgment diverged and the human lost. Correct answers mark the divergences the human won. Both are the skill. Both are visible only in the session record.
Skill is divergence from the frontier baseline
The measurement is one comparison, repeated. Take the transcript. At each human turn, give a frontier model the same context and ask what it would do. Then ask two questions of the operator’s actual move: did it diverge from the model’s move, and was the divergence better? Random flailing diverges and adds nothing. Skill is valuable divergence.
Score it against today’s frontier, and be honest about what that means. As models improve, yesterday’s surplus becomes today’s baseline, and an old score decays. That is the measure telling the truth about a moving target.
If this sounds like Newcomb’s paradox, look closer. Newcomb needs a near-perfect predictor of you, with the payoff fixed before you act. A benchmark predicts nobody and pays nothing. It is the same public model for every operator, and comparing a performance against a benchmark is the oldest measurement there is.
You cannot see the operator in the artifact
Why has nobody measured this? Because you cannot see the working in the finished work. A finished prompt or design shows the destination. You cannot tell whether the model produced it in one shot or the operator extracted it across forty corrections. The surplus sits in the session record, and the session record is what everyone throws away.
A reader of this post’s first draft dismissed it on the grounds that he could not tell how much of it was the model’s average and how much was mine. He was right to demand the distinction and wrong only about the remedy. When the work cannot be shown, the safe assumption is the average. Surplus that cannot be shown will not be paid for.
So keep the record. Append-only, immutable, everything saved. The record grows more valuable untouched, because each new generation of models reads more technique and intention out of the same frozen session. It is also process-supervision training data, the rarest kind, because it teaches how a skilled operator arrives at an answer rather than only what the answer is.
The measurement already exists
None of this is a proposal. Skillgate reads every human turn in your transcripts and judges one thing: how each prompt treats the model’s previous reply. Each turn lands in a lane. Fabrication, when a turn asserts what the transcript contradicts. Reactive, when the reply demanded engagement, graded from ignoring it to acting on its specifics. Proactive, when the turn opens new work. Every turn gets one sentence on its inferred motive, and the sentences compress into a short, unsparing portrait.
Here is the portrait from my own session, unedited - the one a colleague mistook for himself:
He treats the model as a drafting surface, not an authority. When the reply carries substance - a feasibility estimate, an architecture option, a framing of the problem - he almost always engages it, but on his terms: overriding the model’s pricing (”impossible”), rejecting its taxonomies, collapsing its menus into a single principle, or spotting loopholes it missed. His highest-engagement turns consistently do the same thing: the model offers a structured choice, and he dissolves the choice by reframing the question at a different altitude. The lowest-engagement turns cluster around implementation details (naming, config paths, protocol questions, build setup) where the answer is a preference, not a decision - and there the bare approvals and terse redirects are appropriate delegation, not failures to engage.
Reactive: The dominant pattern is correction-by-elevation. When the model presents options, he rarely picks one - he reframes the problem so the options become irrelevant, then supplies the principle the model should have started from. His strongest turns find structural holes the model missed entirely: stochastic variance in re-runs, files vanishing from output sets, the impossibility of edit-time replay detection, mid-pipeline resume requiring Lua VM state. When he does pick from a menu, he always adds a constraint or design principle the model didn’t offer. The weakest reactive turns are concentrated in the implementation tail of the session (config file naming, crate prefixes, Whisper model sizes, Rust enum syntax) where shallow acknowledgment is the right move because the model’s answers need routing, not rethinking.
Proactive: He initiates new threads at pivotal moments - the pipeline architecture for one-off reports, the “oops” pattern for lazy divergence detection, the status-bar observer threaded through every subsystem, the full UI surface enumeration that forced acknowledgment of IDE-scale scope. His proactive turns also serve as session steering: invoking the architect tool when exploration reached critical mass, bringing in external red-team feedback for independent evaluation, volunteering cross-references to his own lesson files. The training-data one-liner - dropped without elaboration after the model’s four-payout summary - is the sharpest proactive turn in the set: one sentence that closed an economic loop the model hadn’t seen.
The model got democratized. Skill didn’t. The instrument is public, the record is valuable, and the question stands: why is no one measuring or talking about this? Run Skillgate on your own transcripts and read your portrait. Find out whether you are the variable or the middleware.
2026-08-26 11:09 - kimi-k3

