Noise
Daniel Kahneman, Olivier Sibony, Cass R. Sunstein
Why judgments vary too much and how institutions can reduce unwanted variance.
Nuxflo may earn a small commission from qualifying Amazon purchases, at no extra cost to you.
Examines unwanted variation when different people judge comparable cases.
Noise audits
decision hygiene
consistency protocols
Daniel Kahneman, Olivier Sibony, and Cass R. Sunstein examine unwanted variation in judgments that should be comparable. Their publisher's description places this problem across institutional decisions. This makes Noise especially relevant when several people use the same scoring process.
Imagine three volunteers rating the same grant application on a ten-point scale. They give 4, 7, and 9. The average is about 6.7, but it conceals a five-point range. Before using that number, ask what each reviewer thought the scale meant. One may have scored the clarity of the application, another its likely impact, and another the applicant's enthusiasm.
This is a Nuxflo example, not a reported experiment. To investigate it, give the reviewers the same written criteria and several sample applications. Have them score independently before discussion. Keep each component score and the reason for it. You can then see whether disagreement concerns facts, interpretation of a criterion, or how much weight to give different goals.
Suppose all three now give every application a 7. Disagreement has vanished, but the scores might be useless. Reviewers could be following an overly vague rubric or sharing the same mistaken assumption. Agreement does not establish accuracy or fairness.
Check two things separately: whether similar evidence receives similar treatment, and whether the criteria actually serve the decision. For example, rewarding polished language could disadvantage an applicant whose project is strong but whose writing is less fluent. A consistent rule can still measure the wrong thing.
For a low-stakes review process, take five anonymized examples, record independent ratings, and discuss the largest disagreements. Rewrite ambiguous criteria, then try a different set of examples. Do not silently change old scores and claim that the process was always consistent. Preserve the original ratings so the effect of the change remains visible.
Choose this book when your problem is variation across a system. If your question is why one person's vivid impression overwhelms statistical context, begin with Thinking, Fast and Slow.