← Disease Detectives
0/8 sections complete done☰ All chapters
Chasing the Curve/Chapter 1
01

Two Ways to Test an Idea

Experimental vs. Observational studies, and the four observational designs

Once field investigators have a hunch about what's causing an outbreak, they need to actually test it. Epidemiologic studies fall into two big families, and the difference between them comes down to one question: does the researcher control who gets exposed?

๐Ÿงช Experimental Studies

Who controls exposure?
The researcher does: they assign it
Key question
"If I give this group the treatment/exposure, what happens?"
Examples
Randomized controlled trials (a new vaccine), quasi-experiments (a policy change without random assignment)
Strength
Best at proving cause and effect: randomization spreads out unknown factors evenly

๐Ÿ”Ž Observational Studies

Who controls exposure?
Nobody: the researcher just watches and records what already happened
Key question
"Among people who happened to be exposed, what happened to them?"
Examples
Case-control, cohort, cross-sectional, and ecological studies
Strength
Practical and ethical for outbreaks: you can't randomly assign people to eat contaminated food!
Why field epidemiology leans observational: You can't ethically expose people to a suspected pathogen on purpose just to study it. During a real outbreak, almost every study a Disease Detective runs is observational, which makes picking the right kind of observational study a critical skill.

Here are the four observational designs your team should be able to recognize, along with when each one shines and where it struggles:

Case-Control Study

Start with people who already have the disease (cases) and people who don't (controls), then look backward to compare their past exposures.

Controls have to be chosen carefully. Good controls come from the same source population as the cases and had the same opportunity for exposure, they just didn't get sick. Sloppy control selection is a major source of bias in case-control studies.

Best forRare diseases, or outbreaks where you need an answer fast
Pros
  • Fast & cheap
  • Great for rare diseases
  • Small sample needed
Cons
  • Relies on memory (recall bias)
  • Can't calculate incidence directly
Cohort Study

Start with a group defined by exposure (exposed vs. unexposed) and follow them forward to see who develops disease.

Best forA defined group with a known exposure, like everyone who attended one banquet
Pros
  • Can calculate attack rate & relative risk directly
  • Shows clear time order
Cons
  • Needs a well-defined group
  • Can take a long time for rare or slow diseases
Cross-Sectional Study

Measure exposure and disease at the same single point in time: like a snapshot of a population.

Best forQuickly estimating how common a condition is right now
Pros
  • Quick & inexpensive
  • Good for measuring prevalence
Cons
  • Can't tell what came first: the exposure or the disease
Ecological Study

Compare data about entire groups or populations (like whole towns or countries), not individual people.

Best forSpotting broad patterns when individual-level data isn't available
Pros
  • Uses data that's often already collected
  • Good first clue-finder
Cons
  • Risk of the "ecological fallacy": a group-level pattern may not hold for individuals
Retrospective vs. prospective cohort: a prospective cohort follows people forward from today and waits for disease to happen (Framingham). A retrospective cohort starts after everything already happened: investigators identify everyone who was exposed at a past event, like a party or a banquet, and look back through interviews or records to see who got sick. Most outbreak cohorts are retrospective. Pros: quick and cheap, since the data already exist; gives attack rates and relative risk. Cons: depends on memory and records (recall and interviewer bias), and like any observational study it can show association but can't prove causation.

A couple of terms from above, unpacked:

Randomized Controlled Trial (RCT)

The gold-standard experimental study: the researcher randomly assigns who gets the treatment/exposure and who doesn't, which helps make the two groups similar overall. It reduces the chance that another difference between the groups explains the results, but it does not guarantee that the groups are exactly the same. That is why an RCT is the strongest way to show the treatment caused the difference in outcome.

Quasi-experiment

An experimental-style study where the researcher still controls who's exposed, but without random assignment: like comparing two schools where only one adopted a new handwashing policy. Useful when randomizing people isn't practical or ethical, but weaker evidence than a true RCT since the groups might differ in other hidden ways.

Double-blind study

A trial in which neither the participants nor the researchers who measure the outcome know who got the real treatment and who got the placebo. It keeps expectations from shaping what people report or what investigators record. (Single-blind: only the participants don't know.)

Recall bias

A distortion that creeps into case-control studies because people are asked to remember past exposures from memory. Someone who got sick often searches their memory harder for "what did I eat?" than someone who stayed healthy, which can make an exposure look more strongly linked to disease than it really is.

Placebo effect (and Hawthorne effect)

The placebo effect: people feel better, or report fewer symptoms, just because they expect a treatment to work, even a sugar pill. Placebo groups and blinding exist to cancel it out. The Hawthorne effect: people change how they act because they know they're being watched.

Case series

A descriptive report on a group of patients with the same new illness: their symptoms, tests and outcomes. There's no comparison group, so it can raise a hypothesis but not test one. The first reports of AIDS were a case series.

Community intervention trial

An experiment where whole communities, not individual people, get an intervention (like adding fluoride to the water), and are compared with similar communities that didn't.

Cluster randomized trial

Groups (schools, villages, clinics) are randomly assigned instead of individuals. Used when the intervention works on a whole group, or when people in the same group would share it anyway.

Crossover trial

Every participant gets both treatments, one after the other in random order, with a "washout" break in between. Each person serves as their own control.

When Studies Go Wrong: Bias & Confounding State/Nats

Bias is a systematic error: something about how a study was run pushes the answer the wrong way, and a bigger sample won't fix it. Tests group biases into two families, plus confounding:

Selection Bias

The people studied differ from the population the study is about.

  • Volunteer (participation) bias: people who sign up are healthier or more motivated than average.
  • Berkson's bias: cases or controls drawn from a hospital, whose patients have other illnesses linked to the exposure.
  • Attrition bias (loss to follow-up): in a long cohort, the people who drop out differ from those who stay.
  • Poor control choice in a case-control study (controls who couldn't have been exposed).
  • Healthy worker effect: people with jobs are healthier than the general public (the very sick can't work), so comparing workers with everyone else hides a job's real risk.
  • Survivorship bias: studying only the people who survived or stayed, and missing those who died or left.
  • Convenience sampling: studying whoever is easiest to reach, so the sample may not represent anyone else.
Information (Measurement) Bias

Exposure or disease is measured wrongly, often differently between groups.

  • Recall bias: sick people remember past exposures better.
  • Reporting (social desirability) bias: people give the answer that sounds better.
  • Interviewer / observer bias: investigators ask or record differently for cases and controls.
  • Misclassification: a test or definition puts people in the wrong group.
  • Panel effect: being interviewed again and again changes how people answer or behave.
  • Detection (surveillance) bias: one group is tested or examined more closely, so more disease is found there even if the true risk is the same.
Confounding

A third factor is linked to both the exposure and the outcome, so it creates or hides an association. Coffee drinkers seem to get more lung cancer only because more of them smoke.

Fix it withRandomization, restriction, matching, or adjusting for the confounder in the analysis
How studies fight bias: choose controls from the same population as the cases; randomize who gets an intervention; blind participants and investigators; use the same standard questionnaire for everyone; check records instead of memory where possible; and keep following everyone in a cohort. When exposure and outcome are measured differently in the two groups, ask which way it pushes the result. For example, if vaccinated people get tested earlier, when a test misses more infections, the vaccine looks better than it really is.
Confounding vs. effect modification: to check a suspected confounder, split the data by it and work out the RR in each group (a stratified analysis). If the association disappears in every group, the overall result was confounded. If it's real but clearly different between groups (a vaccine that protects children much better than older adults), that's effect modification: a true finding to report group by group, not a bias to remove. Interaction is the related idea that two exposures together do more than their separate effects added up (smoking plus asbestos and lung cancer).
Type I vs. Type II error: a Type I error (α) is a false positive: concluding there's an effect when there really isn't. It's usually capped at 5% (p < 0.05). A Type II error (β) is a false negative: missing an effect that's really there, most likely in a small study. Power = 1 − β, the chance of finding a real effect. Very small studies give imprecise results that chance can swing either way.
Internal vs. external validity: a study has internal validity when its results are true for the people it studied (free of bias and confounding). It has external validity (generalizability) when those results also apply to other people. A careful study of high-school girls may be internally valid but tell you little about older adults.

โœ“ Check Yourself

Q1Investigators identify everyone who attended a wedding, then follow them forward to see who got sick. What kind of study design is this?
A cohort study: it starts from a defined group based on exposure and follows them forward in time.