Attack rate, relative risk & odds ratio: the math that turns hypotheses into evidence
This is where Step 7, evaluating a hypothesis epidemiologically, actually happens. Everything starts with sorting people into a 2×2 table: exposed or not, ill or not.
🥔 Worked Example: The Cedarwood Picnic (Cohort Study)
Investigators suspect the potato salad at a class picnic made people sick. Because everyone who attended is known (a defined cohort), they can survey all 120 attendees about what they ate and whether they got sick:
Ill
Not Ill
Total
Ate potato salad
60
20
80
Didn't eat potato salad
8
32
40
Total
68
52
120
Attack Rate
Cases in groupTotal in that group
Exposed: 60 ÷ 80 = 75% Unexposed: 8 ÷ 40 = 20%
A special "incidence" just for outbreaks
Relative Risk
Attack rate, exposedAttack rate, unexposed
75% ÷ 20% = 3.75
The exposed group got sick 3.75× as often
Odds Ratio
(a × d)(b × c)
Used for case-control studies, where relative risk can't be calculated directly
See worked example below ↓
Reading the results: A relative risk of 3.75 is strong evidence. It means people who ate the potato salad were nearly 4 times as likely to get sick as people who didn't. (Say "times as likely": some graders take points off for "times more likely.") An RR (or OR) of exactly 1 would mean no difference in risk at all between the two groups. Above 1, the exposure goes with more illness; below 1, with less (it's protective, like a vaccine).
The "everybody ate it" trap: a food that 95% of guests ate will show a high attack rate among eaters, simply because nearly every sick person ate it too. Don't stop at the exposed attack rate: compare it to the unexposed attack rate. If the few people who skipped the dish got sick just as often, the RR is near 1 and that dish isn't the source.
Risk difference (attributable risk): relative risk says how many times as likely; the risk difference says how much extra risk. Risk difference = attack rate exposed − attack rate unexposed. At the Cedarwood Picnic: 0.75 − 0.20 = 0.55, so 55 more cases per 100 people who ate the potato salad.
Turning an RR into a percent: the exposed group's risk is (RR − 1) × 100% higher. RR = 1.8 means 80% higher risk (not 1.8% or 180%); RR = 3.75 means 275% higher. Below 1, the risk is (1 − RR) × 100% lower: RR = 0.6 means 40% lower. For an odds ratio, say "odds" instead of "risk."
How much illness is due to the exposure? State/NatsAttributable risk percent (among the exposed) = (AR exposed − AR unexposed) ÷ AR exposed = (RR − 1) ÷ RR. At the Cedarwood Picnic: (3.75 − 1) ÷ 3.75 = 73%, so about three-quarters of the illness among salad eaters came from the salad. The population attributable fraction asks the same about everyone: (attack rate in the whole group − attack rate unexposed) ÷ attack rate in the whole group. It depends on how common the exposure is as well as the RR, so a strong but rare exposure causes only a small share of a population's cases. Number needed to treat (NNT) = 1 ÷ the risk difference (the absolute risk reduction). If a drug lowers heart-attack risk from 2.0% to 1.0%, NNT = 1 ÷ 0.010 = 100: treat 100 people to prevent one heart attack.
Confidence interval: a study only samples some people, so a result like RR = 3.75 is an estimate. A 95% confidence interval gives a range likely to contain the true value (for example 2.1 to 6.7). If the interval for an RR or OR includes 1, the study can't rule out "no difference." A bigger study gives a narrower (more precise) interval. State/Nats For an odds ratio the interval is worked out on the log scale: 95% CI = eln(OR) ± 1.96 × SE.
🧫 Worked Example: A Case-Control Study
Now imagine the outbreak was community-wide, so investigators can't identify every exposed person. Instead, they compare 50 cases (confirmed ill) to 50 controls (healthy people, similar in age and location) and ask both groups about past exposures:
Cases
Controls
Exposed (a, b)
40
15
Not exposed (c, d)
10
35
Total
50
50
Odds ratio = (40 × 35) ÷ (10 × 15) = 1,400 ÷ 150 = ≈ 9.3. Because a case-control study starts with the outcome (already-diagnosed cases) rather than following a group forward in time, investigators can't calculate a true attack rate or relative risk, but the odds ratio approximates it well, especially when the disease is rare.
More than one exposure in the table? Work out a separate 2×2 for each exposure. For exposure A, "exposed" is everyone who had A and "not exposed" is everyone else, including people who had exposure B. Then compare the odds ratios: the largest one (with a sensible story) points to the likely source. Say it the right way round: "cases had 12 times the odds of having been exposed," not "exposed people had 12 times the risk."
🤔 Which Formula Do I Use?
Did the study start with an exposed group, followed forward to see who got sick? (like the picnic: a cohort)
→ Use Relative Risk
Did the study start by picking already-sick vs. healthy people, then look backward at exposure? (a case-control)
→ Use Odds Ratio
How Good Is the Test? Sensitivity & Specificity
Lab tests confirm cases, but no test is perfect. Two numbers describe how good a test is. Learn them by meaning, not by letter, because tests put the table together in different ways:
Sensitivity = true positives ÷ everyone who really has the disease. "Of the sick, how many did the test catch?"
Specificity = true negatives ÷ everyone who really doesn't have the disease. "Of the healthy, how many did the test correctly clear?"
Example: a quick test is tried on 40 people who have the disease and 160 who don't. It comes back positive for 36 of the sick people and 32 of the healthy ones.
Layout 1: disease in the rows
Test +
Test −
Has disease
36 (TP)
4 (FN)
No disease
32 (FP)
128 (TN)
Layout 2: disease in the columns
Has disease
No disease
Test +
36 (TP)
32 (FP)
Test −
4 (FN)
128 (TN)
Sensitivity
TPTP + FN
36 ÷ (36 + 4) = 90%
Misses 10% of real cases
Specificity
TNTN + FP
128 ÷ (128 + 32) = 80%
Wrongly flags 20% of healthy people
Positive predictive value
TPTP + FP
36 ÷ (36 + 32) = 53%
Only half of positives are really sick
Same numbers, same answers, either layout. Before plugging in, circle the two cells where the test was right (true positive, true negative). Then ask: sick people go on the bottom of sensitivity, healthy people on the bottom of specificity.
What low values cost:low sensitivity means false negatives: sick people are told they're fine, go untreated and keep spreading the disease. Low specificity means false positives: healthy people get unneeded treatment, worry and cost, and case counts get inflated.
Predictive values depend on how common the disease is.Positive predictive value (PPV) = true positives ÷ everyone who tested positive: "if I test positive, how likely am I to really have it?" Negative predictive value (NPV) = true negatives ÷ everyone who tested negative. Sensitivity and specificity belong to the test, but PPV falls when a disease becomes rarer: with few real cases, the false positives from the many healthy people make up a bigger share of the positives. That's why screening a low-risk population produces lots of false alarms. Screening vs. confirming: a screening test should be highly sensitive, so it misses as few sick people as possible (a negative result then really means "probably not sick"). Anyone who screens positive then gets a second, highly specific test to confirm it and weed out the false positives. To raise a test's sensitivity: lower the cutoff for a positive, test again, or collect the sample at a better time. (Why not always use the best "gold standard" test? It's usually slower, costlier or more invasive.)
Vaccine Effectiveness State/Nats
Vaccine effectiveness (VE, also called vaccine efficacy in a trial) uses the same cohort table as relative risk, with vaccinated as the "exposed" row. It answers: by what percent did the vaccine cut the risk of getting sick?
Step 1: Attack rates
Vaccinated: 10 of 200 ill = 5% Unvaccinated: 50 of 200 ill = 25%
Step 2: Relative risk
AR vaccinatedAR unvaccinated
5% ÷ 25% = 0.20
Step 3: VE = 1 − RR
1 − 0.20 = 0.80 = 80%
Vaccinated people had 80% lower risk
Say it in context: "Vaccinated people had an 80% lower risk of getting the disease than unvaccinated people." An RR below 1 means the "exposure" (the vaccine) is protective. The best study for testing a new vaccine is a randomized controlled trial: volunteers are randomly assigned to the vaccine or a placebo, then followed to see who gets sick. The main concerns are ethics (withholding a working vaccine from the placebo group) and cost.
Try Your Own Numbers
Build your own 2×2 table and watch the math update live: a good way to build a feel for how each number moves the result.
Ill
Not Ill
Exposed
Unexposed
Attack Rate, Exposed
N/A
Attack Rate, Unexposed
N/A
Relative Risk
N/A
Odds Ratio
N/A
Try setting exposed and unexposed attack rates equal: watch the relative risk head toward 1, meaning no difference in risk at all.
✓ Check Yourself
✓ Complete
Q8Using the Cedarwood Picnic data (attack rate 75% exposed vs. 20% unexposed), what does a relative risk of 3.75 mean in plain language?
People who ate the potato salad were 3.75 times as likely to get sick as people who didn't eat it.
Q9Why can't investigators calculate a true relative risk from a case-control study?
A case-control study starts by selecting people based on their outcome (already sick or not), not by following a defined population forward from exposure, so there's no true "population at risk" to calculate a rate from. An odds ratio is used instead.
Q10A test is given to 23 people with a disease (18 test positive) and 88 people without it (52 test negative). What are its sensitivity and specificity?
Sensitivity = 18 ÷ 23 = 78%. Specificity = 52 ÷ 88 = 59%. Divide by everyone who truly has (or doesn't have) the disease, whichever way the table is drawn.
Q11State/NatsOf 30 unvaccinated birds, 18 got sick; of 30 vaccinated birds, 4 got sick. What is the vaccine effectiveness?
AR vaccinated = 4/30 = 13.3%; AR unvaccinated = 18/30 = 60%. RR = 0.222. VE = 1 − 0.222 = 77.8%: vaccinated birds had about 78% lower risk.