Why the numbers look real, why better change management doesn't fix them, and what would make a pilot worth trusting.

Twelve accounts, three months, a pilot program that looked like a clear win. Adoption climbed. CSAT ticked up. The team presenting results to leadership had a strong case, and leadership approved the wider rollout.

Eight months later, at full scale, none of those numbers showed up. Adoption stalled. The CSAT lift never materialized.

The postmortem blamed execution. The rollout was rushed. Training wasn't thorough enough. Some sponsors weren't fully bought in.

Nobody asked whether the original number was ever measuring what everyone assumed it was measuring.

Pilot premium: the gap that was never a mistake

The twelve pilot accounts weren't a preview of the next twelve hundred. They were a different population. Customers who knew they were part of a pilot. Customers who'd been asked for feedback along the way. Customers who got meaningfully more attention from the team running the program than anyone at scale ever will.

I call this gap pilot premium: the inflated performance a program shows in pilot, driven by conditions that disappear the moment the program stops being watched.

The pilot number wasn't wrong. It was accurate for the population it measured. It was never going to be accurate for the population it gets used to justify.

This isn't unique to customer programs. Software teams learned the same lesson with beta releases years ago. Beta cohorts are self-selected volunteers, closely supported, motivated enough to opt in before anything is proven. Nobody is surprised anymore when general availability conversion looks nothing like beta conversion. The gap has a name and an expected shape. Most GTM programs still treat their pilot number as a forecast instead of a beta metric, and the same category of mistake shows up in how most teams read attribution data: a number that's technically real, read as if it means something it was never built to mean.

It isn't one effect. It's four.

Treating pilot premium as one vague bias is why most fixes for it don't work. I think it's actually four separate effects, stacked on top of each other, and each one needs a different fix.

Awareness effect. People behave differently when they know they're being observed and asked for feedback. A pilot participant who knows someone's watching engages more, tries harder, reports more favorably, not because the program is working, but because attention itself changes behavior. This fades the moment observation stops. Fix: build in a quiet cohort. A slice of the pilot population that gets the program but isn't asked for regular feedback or told it's being closely tracked. If results hold there, the effect is real.

Selection effect. Pilot participants are rarely a random sample. They're usually volunteers, early adopters, or the accounts your team already had the best relationship with. That population was already more likely to succeed before the program touched them. Fix: include a deliberately unenthusiastic slice. Accounts that didn't ask to be included, assigned rather than recruited. If the program still works there, the lift is coming from the program, not the volunteers.

Support effect. Pilots get more hands-on attention than any team can sustain at scale. Faster response times, more senior people in the room, more willingness to solve edge cases manually. None of that travels to account twelve hundred. Fix: hold support constant between pilot and rollout, or if that's genuinely not possible, measure what happens to the results as support is deliberately reduced during the pilot itself, before the business case gets built on numbers that assumed white-glove treatment forever.

Competition effect. The first three effects distort who's in the sample or how they behave. This one distorts the conditions themselves. A pilot's touchpoint typically runs where little else is competing for that same moment of the customer's attention. At scale, it has to hold up against everything else already competing for attention at that exact stage of the journey. This is an external validity problem, not a sample problem: the effect size was real, measured under conditions that don't generalize. Fix: before rollout, identify what else already competes for attention at that stage, and measure the pilot's effect against that same level of competition, not a protected one.

Four different mechanisms, four different fixes. A single generic warning to be careful with pilot data doesn't tell anyone which one they're dealing with, which means it doesn't tell them what to actually change.

Change management and measurement are different jobs

The standard advice for a pilot that doesn't scale is to manage the rollout better. Communicate more clearly. Get stronger sponsor alignment. Phase the deployment instead of doing it all at once.

I don't think that advice is wrong. I think it's answering the wrong question.

Better change management makes a real result travel further through an organization. It does nothing to tell you whether the original result was real in the way it's being used. If the pilot number was inflated by awareness, selection, and support effects that don't exist at scale, the most disciplined rollout in the world still won't reproduce it, because there was never a program-level effect that size to begin with.

Change management fixes how a number spreads. It can't fix a number that was never going to hold still.

Treating this as a change management problem doesn't just fail to fix it. It hides the real question long enough for a second, equally expensive rollout to make the same mistake with better execution.

What actually makes a pilot worth trusting

None of this means pilots are useless. It means the number a pilot produces needs to be treated as a beta metric, not a forecast, until it's been checked against the four effects that inflate it.

A pilot worth trusting has a quiet cohort alongside the closely watched one. It includes accounts that didn't volunteer, not just the ones that did. It either holds support constant on the way to scale, or it tests what happens when support is deliberately pulled back before the business case gets written. It measures the result against the same level of competition for attention the touchpoint will actually face at scale, not a protected moment nothing else is contesting. And whoever presents the results to leadership says so explicitly: this is what happened under pilot conditions, and here's what we don't yet know about what happens without them.

That's a smaller, less impressive number to bring into a leadership meeting. It's also the only one that survives contact with the population you're actually about to scale to.


A useful signal: if nobody on the team can say whether the next thousand accounts will need more support than the first twelve, not less, treat that itself as the answer.