Two footprints walk together — but who push who?
Ugg watch two things move. Sun hot, more berry Ugg eat. Sun hot, more river-swim. Two things move together! Ugg call this correlation — a number saying how much two things move as one. But first, deep fire-truth.
Together-walk is NOT push-cause
Biggest rule. Carve in cave wall. Correlation is not causation — two things moving together does NOT mean one pushes the other.
Two things walk together for four reasons:
1. A push B. Smoke-in-lung, lung get sick. 2. B push A. Lung sick, so man smoke more for calm. 3. Third thing C push both. Ice-treat sales and drown-deaths rise together — but ice-treat drown nobody. Hot summer push BOTH. 4. Dumb luck. Small hunt, few tracks — two unrelated things look joined by pure chance.
Ugg warn: never say "A cause B" from together-walk alone! To know real cause, do experiment — split tribe by coin-flip (randomise), poke one group, watch. Or careful cause-thinking that rules out sneaky third thing.
Pearson stone — measure straight-line together-walk
When both things are smooth-number and bell-shaped (Normal), Ugg grab Pearson r — number for how much two things move in a straight line together.
Caveman-read: for each track see how far x sits from its middle, how far y from its middle, multiply, add all up, then divide by the wiggle-size so the number stays tidy. Same beast as standardised covariance, — the together-wiggle squashed by each thing's own wiggle, so it come out unit-less.
r lives between −1 and +1. Sign = direction (plus: both climb; minus: one up, other down). Size = strength. Cut-off rocks: tiny, – small, – middle, big.
Square the stone — r²
Square it, get r² (coefficient of determination) — how much of Y's wiggle is shared with X. → "49% of Y's wiggle explained by X." Lives 0 to 1. Ugg warn: r² is already a piece of the whole — no report "r² times 100%" as if fresh; multiply by 100 only to speak as percent.
What Pearson stone NO see
Pearson stone is blind to curve! A perfect U-shape gives , yet the two things are TIGHT joined. So ** does not mean unrelated** — only no straight line. Always draw the dot-picture (scatter plot) first.
Pearson stone also fears the outlier — one crazy-far track drags r wild. Ten tracks hug the line, r ≈ 0.99; drop in one monster point far off, r crashes to 0.5. (Anscombe's quartet: same r, four different stories.) Ugg warn: no trust r on wild dots — LOOK.
Spearman stone — rank the tracks
When data is order-only (ordinal), or the join is bend-but-always-climb (monotonic, not straight), or outliers wreck Pearson — grab Spearman ρ. The trick: Spearman is just Pearson r on the RANKS, not raw numbers. Line up x smallest-to-big, hand out rank 1, 2, 3…; same for y; then Pearson those ranks.
Why Spearman no fear outlier? Rank squashes the monster. A point at one-million becomes only "rank n" — same as if it were a whisker above the next. Extremeness thrown away, order kept.
Kendall τ — pair-counting cousin. For each pair, same-direction = concordant, opposite = discordant. . Even steadier than Spearman; good for tiny hunt or many ties.
Strong is not the same as sure
Ugg compute r = 0.85. Real or luck? Depends on hunt-size n. Test if true-r is zero: with df. Small hunt (n=10): even r = 0.99 can be noise. Huge hunt (n=1000): even puny r = 0.10 can pass as "significant." So strength = size of r; sureness = p-value, which needs r AND n. Ugg warn: never show p alone — always report r beside it.
Partial stone — strip the sneaky third thing
Maya wants the join of X and Y but fears third thing Z fools her. **Partial correlation ** = join of X and Y AFTER Z's straight effect is stripped from BOTH. Recipe:(1) fit X from Z, keep leftovers (residuals);(2) fit Y from Z, keep leftovers;(3) Pearson the two leftover-piles.
Worked: r(GPA,IQ)=0.75, r(GPA,study)=0.56, r(IQ,study)=0.46 → partial = (0.75 − 0.56·0.46)/√((1−0.56²)(1−0.46²)) = 0.492/0.735 ≈ 0.67. GPA–IQ join still strong after stripping study-time.
Semi-partial (part) correlation strips Z from only ONE side. Ugg warn: no mix them up — partial strips both, semi-partial strips one.
Reliability — how much the tribe agree
Cohen's κ — two counters drop items into name-buckets (nominal). They agree some by luck, so κ fixes for luck:
= agree-seen, = agree-by-chance. κ=0 is only chance, 1 is perfect, negative is worse than chance (counters fight). Ugg warn: κ is smaller than raw agree because it removes luck. Cousins: Fleiss' κ for more than two counters (nominal); Kendall's W for many counters ranking items (0 to 1); Krippendorff's α the do-everything stone (any counters, missing data, any scale). Same counter twice over time on smooth data → plain Pearson r (intra-rater).
Cronbach's α — do all test-questions measure the same one thing?
k = number of questions. If questions are tight-joined, the total wiggle towers over the sum of piece-wiggles → α high. α ≥ 0.7 okay, ≥ 0.9 great — but Ugg warn: α ≥ 0.95 maybe TOO high, questions just copy each other; drop the twins.
Outlier — the crazy-far track
Outliers come from measure-mistakes, write-mistakes (someone typed "999" for missing), process-mistakes, mixed sources, or a true rare beast (real genius). Spot them with dot-pictures, box-plots (1.5 × IQR rule), or points 2–3 SD from the middle, Grubbs' test. Five ways to deal: omit (only with proof of error, write down why), replace (winsorise), use rank/robust methods, value it (sometimes the best find), or transform (log, sqrt). Ugg warn LOUD: the big student mistake is quiet-dropping the weird point. NO. Detect, investigate the cause, decide, document.
Ugg remember
- Together-walk is NOT push-cause — four reasons: A→B, B→A, C→both, dumb luck. Never claim cause from correlation alone.
- Pearson r = straight-line join in , scared of outliers, blind to curve; r² = shared wiggle; and does NOT mean unrelated.
- Spearman ρ = Pearson on ranks — for ordinal, bendy-climb, or outliers; Kendall τ is the pair-count cousin.
- Strength (size of r) is not sureness (p, needs r AND n). Small n + big r = suspicious. Always report r beside p.
- Partial strips Z from both X and Y; semi-partial from one. Cohen's κ = chance-fixed agreement (κ<0 is bad); Cronbach's α = internal consistency (>0.7 ok, >0.95 redundant).