Skip to content
What the Record Proves

Home / Measurement

Productivity Scores as Evidence

What a productivity score is made of, why nobody in the organisation can usually describe the formula, and the two conditions for using one at all.

Measurement · Reference

A productivity score of 54%, and what it was offered to support

The claimWhat the record showsWeight
The score was 54% for the weekThe platform's own report✓ Supports it
The employee was less active than the teamTeam median 71% on the same measure✓ Supports it
The employee worked lessThe measure counts input in listed applications✕ Does not show it
The employee was not workingOutput for the week was in line with normal✕ Does not show it
The score is comparable between rolesThe weighting was set once, for a different team· No record

The score is a true statement about a formula nobody in the organisation can describe. This is one employer's own file, not a statement of what any rule permits about monitoring.

A productivity score is a formula, and in most organisations nobody can say what the formula is. It combines inputs, application categories, active time and a weighting, all chosen by the vendor or by whoever configured the platform, and it produces a percentage that looks like a measurement of a person.

The evidential discipline in “Productivity Scores as Evidence” should also govern workforce technology. A team evaluating check this official Monitask resource for attendance sheet template can use time and project records as operational context, but should preserve the original record, document access and let the employee correct a misleading entry before it supports a conclusion.

The number is a true statement about the formula. Whether it is a statement about work depends entirely on whether the formula models the work, and that question is answerable only if somebody can describe it.

The New Jersey wage and hour resources offers another lens on the issue raised in “Productivity Scores as Evidence”. Compare its principles with the actual record, ownership model and review route rather than importing a generic checklist unchanged.

What they are usually made of

Time in applications classified as productive, divided by time in applications classified as anything, adjusted for idle time and weighted by categories somebody set.

Each of those three components carries a judgement: which applications count, what idle means, how the categories are weighted. Each judgement was made by somebody, often a vendor, usually years ago, for a generic office.

The classification problem

An application classified as unproductive is frequently where some roles do their work: a browser for research, a messaging tool for coordination, a spreadsheet for everything.

So the score systematically penalises roles whose work lives in the wrong category, and the categories were not built with those roles in mind. That is not a flaw in the person being measured.

The two conditions for using one

First, somebody in the organisation can describe the formula, including the classification and the weights, well enough to explain a score to the person who received it.

Second, the score has been checked against something independent — output, delivery, quality — for the roles it is applied to, and the relationship has been looked at rather than assumed.

Without both, the score is a number with a percentage sign, and treating it as evidence about an individual is not defensible on any view.

Comparing people with it

Comparison within one role, on one configuration, over a reasonable period, is the only use that holds. Comparison between roles compares classifications rather than people.

That is the error that produces the most unfairness, and it happens by default because the platform presents a single leaderboard.

The measure that changes the work

Any score that people can see and are judged on will be optimised, and scores built on input are unusually easy to optimise without doing anything.

That is worth anticipating rather than policing. A measure that can be satisfied without producing anything will be, and the response is a better measure rather than a rule against gaming it.

What it is reasonable to do with one

Look at aggregates. A team whose score falls by a third after a system change is telling the organisation something about the system.

Look at trends for an individual over time, as a prompt for a conversation. Do not use a single period's score as a finding about a person, and do not put one in a disciplinary file as evidence of hours.

Explaining a score to the person

If that cannot be done — if nobody can say why a particular week scored 54 — the score should not be used for anything that affects the person.

That test is simple, it is applied almost nowhere, and it disposes of most of the uses these numbers are put to. It also tends to be the point at which somebody discovers the weights have not been reviewed since the platform was installed.

What the score is made of, shown

Where a score is used at all, the components should be visible to the person receiving it — not as a formula, but as the handful of things that moved it.

  1. The applications counted as productive for this role.
  2. The idle threshold applied.
  3. The hours the score was computed over.
  4. The weighting, if more than one component is involved.
  5. The comparison group the percentage is relative to.

Five lines next to the number. If they cannot be produced, the number should not be shown to anybody, because nothing can be done with it except resented.

What to hold on file

The formula, the classification list, the weights, the date each was last reviewed, and whatever check has been done against independent output.

Five items. An organisation holding them can defend using the score. One that cannot produce them is relying on a vendor's model of work and calling it a measurement of its own people.