Saral Shiksha Yojna
Courses/Behavioral Research: Statistical Methods

Behavioral Research: Statistical Methods

CG3.402
Vinoo AlluriMonsoon 2025-264 credits

Scales, Reliability, Validity

NotesStoryCaveman
Unit 2 — Research Design & Measurement

Wrong measure make whole hunt waste

Before Ugg hunt beast, name it and say how to count it. You cannot study "sad-spirit" or "smart" until you write an exact rule to measure it. That rule is the operational definition — a working rule for HOW a researcher measures a fuzzy thing. Ugg say "sad = how many times friend sit by fire with tribe in one moon." Maybe a bad rule (sad friend still sit sometimes), but at least Ugg can COUNT it. No count-rule, no science — only loud opinion in fur coat.

Four kind of counting-rock — NOIR

Ugg sort all measure into four rock-piles. Remember the order: N-O-I-R, small-info to big-info. Ugg write — each pile can do every trick of the pile before it, plus one more trick.

Nominal — name-only rock. Just groups, no bigger-or-smaller. Eye colour, blood type, which-cave-you-from. Ugg WARN: no make average! "Average eye colour" is grunt-nonsense. Only count heads, find most-common (mode), do .

Ordinal — line-up rock. Now there IS order, but the GAPS not equal: 1st beat 2nd by a mammoth-length while 2nd and 3rd finish nose-to-nose. Use middle-one (median), percentiles, rank-friendship (Spearman, Kendall).

Interval — even-step rock, but NO true zero. Fire-heat in Celsius: 0°C not mean "no heat," just mean water-turn-ice. Gaps trustworthy (30° − 20° = 20° − 10°), BUT Ugg WARN: no do ratio! 20°C NOT "twice hot" as 10°C. Now allowed: mean, spread (SD), Pearson, t-test, ANOVA.

Ratio — even-step rock WITH true zero. Reaction-time, weight, height, count-of-mammoth. Here zero mean truly nothing, so "Ugg twice as fast as you" is REAL. All tricks allowed.

Second question, a SEPARATE knife: is the number continuous (any value in between, like weight) or discrete (only jump-steps, like year-of-birth)? Reaction-time = ratio + continuous. Year-of-birth = interval + discrete. Cave-of-transport = nominal + discrete.

Likert scale — the "strongly-disagree → strongly-agree" 1-to-5 rock. Strictly it is ordinal. But tribe usually TREAT it as interval, since people space the steps kinda-even. Exam trick answer: technically no, in practice yes, depends. Treat-as-interval → t-test/ANOVA/Pearson; treat-strict → Mann-Whitney/Spearman.

Same-answer magic — reliability

Reliability — does the measure give the SAME answer when nothing changed? Consistency. If the stick says a new number each time for the same rock, it is garbage. Four flavours:

Test-retest reliability — same over TIME. Same person take smart-test twice, three moons apart, should score close.

Inter-rater reliability — same across PEOPLE-who-judge. Two wise-elders each judge the same 50 tribe-mates — do they agree? Count with Cohen's Kappa (2 judges, name-group data), Fleiss' Kappa (more than 2 judges), Kendall's W (ordinal), Krippendorff's Alpha (any). Kappa: — take the agree-you-saw () minus agree-by-dumb-luck (), then stretch so 0 = luck-only, 1 = perfect.

Parallel forms reliability — same across two EQUAL versions. Two weighing-rocks should give the same weight.

Internal consistency reliability — same across ITEMS inside ONE test. If 10 questions all measure the same smart-thing, they move together. Count with Cronbach's α: — high α (near 1) mean items hunt as one pack. Above 0.7 okay, 0.8 good; but above 0.95 WARN — items maybe just copy each other. Also split-half and KR-20/21.

Split-half trick: cut test into odd-items vs even-items and correlate: . But a half-test is short so it looks weak — fix with Spearman-Brown correction: , which guesses the full-length strength.

Ugg warn — things that break reliability: measure-error, apparatus drift, practice-effect, mood and time-of-day, faker-participant, tired or biased researcher.

Right-answer magic — validity

Validity — DIFFERENT from reliability! Not "same answer" but "RIGHT answer, on target." Five flavours:

Internal validity — can Ugg truly say CAUSE made effect? If the sick-cave group and healthy group differ in many ways (food, healer, shiny-box internet), you cannot blame the sickness alone — those are the confounds. Random-assign is the cure.

External validity — does the finding spread BEYOND your little sample? A study only on young cave-scholars at one cave does not speak for all humans.

Construct validity — does the measure grab the REAL fuzzy-thing? Measure sadness by "like my tree-carving post" is terrible — likes come from who-follows-you, not from sadness. Prove it with convergent validity (agrees HIGH, r = 0.78, with a trusted sad-measure like PHQ-9) AND discriminant validity (disagrees, r = 0.10, with an unrelated thing like outgoing-ness). Need BOTH.

Face validity — does the test just LOOK right on the surface? Weakest kind. Wise-tribe not care much, but it helps convince the chief and the crowd.

Ecological validity — does the test-cave resemble the real world? An eyewitness test in a calm lab misses real-world fear and rush. Nice to have, not must-have — many lab findings still travel.

Scale that lie the same way every time

Big lesson: reliability and validity are SEPARATE. A bathroom-rock always 5 kg too heavy is reliable (same answer) but NOT valid (wrong target). A stopped sun-dial always says "high-noon" — reliable, but right only twice a day by luck. So: you CAN be reliable without being valid; you CANNOT be valid without being reliable — if the answers jump random, they cannot track any true thing.

Trickster spirits that fool Ugg

Regression to the mean — pick the most-extreme rock today, and tomorrow it drifts back toward the middle all by itself. Famous fool: sky-warrior trainers saw pilots PRAISED after a great landing do worse next time, and pilots SCOLDED after a bad landing do better. They cried "punishment work!" WRONG — extreme scores just drifted to average. A real statistic-ghost mistaken for a cause.

Confound — a sneaky third-thing tied to BOTH your cause AND your effect, making a fake link. Threatens internal validity. Best fix: random-assign (breaks the link), or add it as a covariate in the model.

Double-blind — neither the tribe-mate NOR the researcher knows who got real medicine vs fake. This kills both researcher-bias AND participant-reaction. Add a fake-pill control plus random-assign and you have the gold-standard RCT.

Ugg remember

  • NOIR: Nominal (name) → Ordinal (rank) → Interval (even step, no true zero) → Ratio (even step, true zero). Continuous-vs-discrete is a SEPARATE knife.
  • Reliability = same answer. Validity = right answer. Reliable-but-not-valid can happen; valid-but-not-reliable CANNOT.
  • Reliability flavours: test-retest (time), inter-rater (people, Cohen's κ), parallel-forms (versions), internal-consistency (items, Cronbach's α).
  • Validity flavours: internal (cause), external (spread), construct (right thing — need convergent AND discriminant), face (looks right), ecological (real-world).
  • Ugg warn: no average a name-group; 0°C is NOT "no heat"; high Cronbach's α proves items AGREE, not that they measure the right thing; never mistake regression-to-the-mean for a real cure.
End of storyUnit 2 — Research Design & Measurement · Scales, Reliability, Validity