Saral Shiksha Yojna
Courses/Behavioral Research: Statistical Methods

Behavioral Research: Statistical Methods

CG3.402
Vinoo AlluriMonsoon 2025-264 credits

FWER vs FDR; Bonferroni, Holm, BH

NotesStoryCaveman
Unit 8 — Multiple Comparisons (FWER, FDR)

Many spear-throw, one always hit by luck

Ugg tell you by fire. Ugg throw ONE spear at mammoth-shadow in dark — Ugg miss almost always. Good. But Ugg throw TWENTY spear at twenty shadow? One spear hit SOMETHING by pure luck. Ugg not great hunter. Ugg just throw many time. This the trap.

Multiple comparisons problem — when tribe run many test on data, the chance of one FALSE "yes!" grow big, way past the small chance Ugg promise for one test.

Story: CEO-man bring memory-berry to Maya. Test say — no good. CEO grumble "try concentration!" . "Reaction time? Verbal?" Nothing. After twenty test, test twenty-one show . CEO shout "miracle berry, give crore!" NO. That just a luck-spear hitting shadow. Ugg not fooled.

The luck-math

One test use α (alpha) = 0.05 — Ugg allow 5-in-100 chance of crying "yes!" when truly no. That false shout is Type I error (false positive, FP).

One test: chance of NO false yes = 0.95. For tests that not touch each other:

Ugg read this: take clean-chance (0.95), multiply by self times (one per test), take away from one. That leave the chance at least one lie sneak in.

Plug : . SIXTY-FOUR in hundred! Ugg carve this number in cave wall. More test, worse: , . Coin-story same: one coin flip 9-head look strange; but clone coin twenty time, ONE showing 9-head is just normal luck.

Two ways to count mistake: FWER and FDR

Let = number test. = number Ugg shout "yes!" (reject). Among them, some true (TP), some lie (FP).

FWER (Family-Wise Error Rate) — chance of ANY false yes in whole hunt: . Ugg ask "did I make ANY mistake?" It count events. Very careful, very strict.

FDR (False Discovery Rate) — the average SHARE of Ugg's "yes!" that are rotten: . Ugg ask "how many of my catch are bad?" It count proportion. More relaxed, more berries found.

Ugg trick: when EVERY null is true, FWER = FDR. They split apart only when some real effect hide in the data.

Use FWER when even ONE lie is bad — costly hunt, medicine going in tribe body, confirming old finding. Use FDR when Ugg dig THOUSAND rock for shiny — genes, brain-picture — where a few lie okay so long as Ugg not miss real shiny.

Bonferroni — chop the gate small

Bonferroni correction — simplest fix. Chop α by number of test:

Fifty test at 0.05? New gate = 0.001 each. Only smaller than that counts.

Why work? Union bound (Boole's inequality). Ugg words: chance of "this OR that OR other" never bigger than all the single chances added up. Add up gates of and you get back . So FWER stay under α. Solid rock.

But Ugg WARN: Bonferroni too tight! It miss real mammoth (raise Type II error, β, the false "no"). And it pretend all test not touch — but brain-region touch its neighbor, survey-question touch its cousin. Then Bonferroni over-chop and kill good findings.

Holm — Bonferroni's smarter cousin

Holm's correction — same safety, more catch. Line up small-to-big: . Compare each to a gate that OPENS wider as Ugg walk down:

Smallest vs (same as Bonferroni). Pass? Reject, move on. Next vs — bigger gate! Keep going till one FAIL, then STOP.

Example: , α = 0.05, = 0.005, 0.012, 0.018, 0.030, 0.080. vs yes. vs yes. vs NO, stop. Reject first two. Holm always catch everything Bonferroni catch, sometimes more. Ugg say: no reason use plain Bonferroni — use Holm.

Benjamini-Hochberg — the FDR rock

Benjamini-Hochberg (BH) — control FDR, find more shiny. Sort small-to-big. Each rank gets a gate , where = the FDR Ugg allow (say 0.05). **Find the BIGGEST rank where . Reject that one AND all below it** — even if some below dipped over their own gate. The biggest passing rank sets the line.

Example: , , = 0.001, 0.008, 0.039, 0.041, 0.250. Gates = 0.01, 0.02, 0.03, 0.04, 0.05. Rank 1: yes. Rank 2: yes. Rank 3: no. Biggest passing = rank 2 → reject test 1 and 2. BH beat Bonferroni when many null are truly false — that why gene-tribe with thousand test use it.

Permutation — let the shuffle tell truth

When test touch each other (brain-voxel, repeat-measure), even BH over-chop. Permutation test fix it:(1) run test, keep raw .(2) SHUFFLE the group-labels many time (1000+) — now all null true by force.(3) each shuffle, run all test, write down the MOST extreme result.(4) the 95th-biggest of those extremes becomes your true gate.(5) judge real data against it.

Beauty: it FEELS the correlation by self. Tests touch a lot → fewer real independent test → gate more kind. Tests not touch → gate match Bonferroni. Ugg get power back, keep FWER safe. Brain-picture tribe use this always.

Ugg warn: no do these!

  • No run twenty test and show only the one shiny. That the drug-berry lie.
  • No slap Bonferroni on THOUSAND test — it murder real effect. Use BH.
  • No pretend subgroup-splits (by gender, age, tribe) are free. Each split a test — six subgroup push FWER near 26%.
  • No peek at after each new hunter and stop when happy (optional stopping) — that inflate error past α.
  • No report a BH-catch as if it FWER-safe. BH promise only FDR , not the stricter α.
  • No confuse the two: FWER = chance of ANY lie; FDR = SHARE of lie. Different scale.

Ugg remember

  • Many test inflate false yes. ; , α = 0.05 → 64%. Carve it in the wall.
  • FWER , "any mistake?", strict — for few costly confirmatory test. FDR , "what share rotten?", roomier — for many exploratory test.
  • Bonferroni : simple, safe by union bound, but too tight and assume tests independent.
  • Holm = sequential : same safety as Bonferroni but more power — always prefer it.
  • BH = sort, reject up to biggest with : controls FDR. Permutation for touching tests. And pre-register your hunt-plan to block the garden-of-forking-paths cheating.
End of storyUnit 8 — Multiple Comparisons (FWER, FDR) · FWER vs FDR; Bonferroni, Holm, BH