Cohen's d
How big, not just whether. Cohen's d states a difference between means in standard-deviation units, measuring practical size.
- Term
- Cohen's d
- Is
- A standardized effect size
- Formula
- Mean difference ÷ pooled SD
- Measures
- Practical size of a difference
Parts of speech & senses
- Cohen's d is a standardized effect size equal to the difference between two means divided by their pooled standard deviation, measuring how large a difference is in standard-deviation units. "The lift was significant, but Cohen's d was tiny."
What Cohen's d is
Cohen's d is a standardized effect size that tells you how large the difference between two group means is, measured in standard-deviation units. You take the difference between the two means and divide it by their pooled standard deviation — a combined measure of how spread out the data are. The result is a pure number with no units: a Cohen's d of 0.5 means the two means sit half a standard deviation apart, whatever the original scale. That standardization is the point. Because the difference is expressed relative to the natural variation in the data, you can compare effect sizes across studies, metrics, and units — a lift in seconds, in dollars, or in clicks all reduce to the same scale. Cohen's d answers the question a raw difference cannot on its own: is this gap big or small relative to how much the numbers ordinarily vary?
Jacob Cohen, the psychologist who introduced the measure, offered rough benchmarks that are still widely quoted: around 0.2 is a small effect, 0.5 medium, and 0.8 large. These are conventions, not laws, and they should be read against the context — in some fields a d of 0.2 is meaningful, in others trivial. The sign of d shows direction; its magnitude shows size. Because it uses the spread of the data as its yardstick, Cohen's d stays stable as sample size grows, unlike a p-value, which shrinks with more data regardless of how big the underlying effect is. That stability is exactly why analysts report d alongside significance tests: it describes the effect itself, cleanly separated from how much data happened to be collected. A large study and a small one measuring the same true difference should report similar d values.
Cohen's d versus the p-value
The most important distinction is between Cohen's d and the p-value, because they answer completely different questions and are constantly confused. A p-value measures statistical significance: how surprising the data would be if there were truly no difference. Cohen's d measures practical significance: how big the difference actually is. The two can diverge sharply. With a huge sample, a difference so tiny it barely matters can produce a dazzlingly small p-value, tempting you to call it important when d reveals it is negligible. With a small sample, a genuinely large effect — a big Cohen's d — may fail to reach significance simply because there were not enough observations to rule out chance. Significance tells you whether an effect is likely real; effect size tells you whether it is likely to matter.
That is why reporting both is now standard practice, and why a p-value alone is treated as incomplete. Imagine a test where a variant lifts a metric by a hair across millions of users: p may be minuscule, yet Cohen's d is close to zero, and rolling out the change would be effort spent on nothing. Now imagine a pilot with forty users showing a substantial improvement that just misses significance: d is large, and the sensible move is to gather more data, not to dismiss the effect. A low p-value with a small d is statistically real but practically hollow; a large d with a high p-value is promising but unproven. Read together, they keep you from two opposite mistakes — celebrating trivial effects and discarding important ones — which is the whole reason effect sizes earn their place beside significance.
Using Cohen's d well
Using Cohen's d well means reporting it as a matter of course, not as an afterthought, and pairing it with a confidence interval so readers see both the estimated size and its uncertainty. Interpret the number in context rather than clinging to the 0.2, 0.5, 0.8 labels: a small standardized effect on a metric that touches millions of dollars can matter more than a large one on something trivial, so translate d back into real-world consequences before deciding it is important. Be careful about which standard deviation you pool, especially when groups have unequal spreads or sizes, since that choice shifts d. For paired designs, compute the effect size on the within-pair differences, matching the way a paired t-test analyzes the data. The goal is a number that honestly conveys magnitude.
The failures start with omitting effect size entirely and leaning on the p-value, which hides whether a significant result is large or laughably small. Treating Cohen's benchmarks as strict cutoffs is another trap — they are rough guides, and blindly labeling a d of 0.49 medium and 0.51 large ignores context. Interpreting d without its confidence interval overstates precision, particularly in small samples where the estimate is shaky. And confusing effect size with importance in the other direction — assuming a big d automatically justifies action regardless of cost — skips the business judgment that should follow. Used properly, Cohen's d restores the question significance testing leaves out: not merely whether a difference exists, but whether it is large enough to care about. That single addition sharpens far more decisions than an extra decimal place on a p-value ever will.
Synonyms & antonyms
Synonyms
Antonyms
Origin & history
Cohen's d is named for the psychologist Jacob Cohen, who formalized standardized effect sizes in his work on statistical power analysis in the behavioral sciences.
Etymology: source.
Usage trends
Search interest for this term over the last five years:
Common questions
- What is Cohen's d?
- Cohen's d is a standardized effect size equal to the difference between two means divided by their pooled standard deviation. It expresses how large a difference is in standard-deviation units, letting effect sizes be compared across studies, metrics, and scales.
- How is Cohen's d different from a p-value?
- A p-value measures statistical significance — whether a difference is likely real. Cohen's d measures practical significance — how big it is. A tiny effect can be highly significant in a large sample, while a large effect can miss significance in a small one.
- What counts as a large Cohen's d?
- By Jacob Cohen's rough convention, about 0.2 is small, 0.5 medium, and 0.8 large. These are guides, not rules — interpret the number against the context and the real-world stakes rather than applying the cutoffs mechanically.
Resources & people to follow
- referenceRGM analysis — definitions, senses, and usage verified per term
Curated, non-competitor resources verified per term.
Related training
Disciplines
Areas of marketing where cohen's d is a core concern: