Skip to content
What the Record Proves

Home / Designing it out

Measuring Whether It Worked

The five measures that show whether an intervention worked, why one of them should get worse, and how to run a before-and-after that means something.

Designing it out · Reference

One site, before and after a second terminal was installed

Punches in the final 90 seconds41% before, 12% after✓ Supports it

Measured the same way, same shifts.

Median queue at shift start3m 50s before, 55s after✓ Supports it

Timed with a watch, three mornings each.

Several badges per terminal per minute11 a week before, 1 after✓ Supports it

From the terminal log.

Recorded lateness3 a month before, 9 after△ Consistent with it

People now punch when they arrive.

Investigations opened2 that quarter before, 0 after✓ Supports it

Small numbers; one quarter.

Four measures improved and one got worse, which is the one that shows the first four are real. Recorded lateness rose because the avoidance stopped. This is one employer's own data.

Measure the same things before and after, or the change is an opinion. Every intervention in this collection — a second terminal, a softened rule, a staggered start, a control — can be assessed within a month with measures the organisation already has, and almost none of them ever is.

The practical lesson in “Measuring Whether It Worked” is that visibility is not certainty. For teams researching task switching cost, read the operational guidance can add time and project context to the operational record, provided the purpose is explained, access is restricted and any material inference is checked through conversation and proportionate human review.

The most useful signal is counterintuitive: one measure should get worse. If the behaviour being removed was avoidance, removing it makes the thing being avoided visible.

The Slack Trust Center offers another lens on the issue raised in “Measuring Whether It Worked”. Compare its principles with the actual record, ownership model and review route rather than importing a generic checklist unchanged.

The five measures

Distribution of punches around the shift start, by minute.

Queue time at the terminal, timed by hand on three mornings.

Multiple-badge events per terminal per period, from the log.

Recorded lateness, which should rise if avoidance falls.

Matters opened, as a slow indicator over quarters rather than weeks.

Why one should get worse

Where people were passing badges to avoid being marked late, stopping that produces more recorded lateness — not because anybody got later, but because lateness is now being recorded.

An intervention where every number improves, including that one, is usually an intervention where the avoidance moved rather than stopped. That is the single most useful diagnostic available here.

Running it properly

  1. Choose the measures before making the change, and write them down.
  2. Take a baseline over at least two weeks.
  3. Make one change, not three.
  4. Wait at least a month before measuring.
  5. Measure the same way, at the same shifts, by the same method.
  6. Record the result whichever way it went.

Step three is the one that gets broken. An organisation that installs a terminal, staggers the starts and softens the rule in the same week has learned nothing about any of them.

The small-numbers problem

Investigations opened is a real measure and a slow one. Two a quarter becoming none is not evidence of anything on its own.

Say so rather than claiming it. The fast measures — distribution, queue time, multiple-badge events — move within weeks and carry the finding; the slow one is confirmation a year later.

Measuring a control as well

The same discipline applies to the expensive interventions. A biometric rollout should be measured on queue time, failure rate and the distribution, and the failure rate is the one that predicts whether the problem returns in another form.

Organisations measure controls on whether they were installed, which is not a measure of anything.

Publishing the result

Internally, to whoever decided and to the people affected. A second terminal that visibly worked makes the next operational fix much easier to get approved.

It also makes the honest version of a failure possible. An intervention that did not work is worth knowing about, and the only way that knowledge survives is if somebody wrote the number down.

Keeping the baseline

The baseline measurements themselves, with dates and method.

Those are what make every subsequent comparison possible, including ones nobody anticipated — and they are the thing that is never kept, because at the time they were taken they looked like a description of a problem rather than a reference point.

Who measures

Not the person who proposed the change, for the same reason nobody checks their own arithmetic well. The measurement is twenty minutes of work and the independence is free.

In practice that means somebody from a different function takes the baseline and the follow-up, using a written method. It also means the result is believed when it is good, which is the point of the exercise.

Writing the method down first

The method matters more than the instrument. Three mornings, the same shifts, the same two minutes, counted the same way.

A follow-up measured differently from the baseline produces a number that cannot be compared, and the difference is usually discovered only when somebody questions the result. Writing the method on the same page as the baseline removes it entirely.

What a failed change teaches

An intervention that did not work is a finding about the diagnosis rather than about the intervention, and it is worth saying so.

A second terminal that changed nothing means the queue was not the cause, which narrows the problem usefully and is worth more than the terminal cost. That only happens if somebody recorded the result when it was disappointing.

What the measurement is for

To move these decisions from argument to evidence. Whether a queue causes badge-passing is an empirical question, it is answerable in a month for almost nothing, and answering it is what stops the next conversation being about opinions.