DIP Episode 197 - Bias in Biostatistics
Topic
Biostatistical biases; study design flaws (selection, measurement); confounding vs. effect modification; survival analysis pitfalls.
Key Takeaway
To accurately interpret research findings, one must systematically identify the type of bias present—whether it is selection, measurement, or statistical—and apply appropriate methodological solutions such as randomization, blinding, stratification, or intention-to-treat analysis.
Episode Notes
Source / episode info
- Episode: 197
- Title: Divine Intervention Episode 197 – Bias in Biostatistics.
- Published: 2019-12-31
- Source: Episode page
One-liner
Episode 197 provides a comprehensive review of common biostatistical biases encountered on board exams, covering selection bias (e.g., Berkson's), measurement bias (e.g., Hawthorne effect), and statistical pitfalls like confounding, effect modification, lead time vs. length time bias, and recall bias.
High-yield summary
- Selection Bias: Occurs when the study sample is not representative of the target population (e.g., using only hospital patients). Solutions include randomization and ensuring generalizability.
- Measurement Bias: Arises from flaws in data collection or analysis, such as leading questions or participant awareness (Hawthorne effect). Mitigation requires blinding and objective measurement tools.
- Confounding vs. Effect Modification: Confounding occurs when a third variable distorts the observed relationship; stratification eliminates this difference. Effect modification occurs when the magnitude of the association differs across levels of a third variable, even after controlling for it.
- Lead Time Bias: Mistakenly concluding that early detection increases survival when the true endpoint (e.g., age at death) is fixed and unaffected by diagnosis timing.
- Length Time Bias: Only studying slowly progressive cases of a disease because rapidly progressing or severe cases die too quickly to be captured in screening/surveys.
Learning objectives
- Differentiate between various types of selection bias (e.g., Berkson's, loss to follow-up).
- Identify and correct for measurement biases using blinding techniques (single vs. double/triple blind).
- Distinguish clinically and statistically between confounding variables and effect modification.
- Recognize the pitfalls of survival analysis, specifically lead time bias versus true improved survival.
- Apply appropriate study design solutions (e.g., Intention-to-Treat) to minimize observed biases.
Board exam buzzwords
| Condition | Key Finding | Association | Board Exam Tip |
| Selection Bias | Sample is not representative of the population | Berkson's bias (using hospital/sick populations) or Loss to Follow-up | Solution: Intention-to-Treat analysis; Randomization. |
| Measurement Bias | Data collection process distorts results | Hawthorne effect (knowing you are observed); Leading questions | Mitigation: Blinding (double/triple blind studies). |
| Confounding | Third variable influences both exposure and outcome | Lifestyle factors (smoking, obesity) often confound outcomes. | Solution: Stratification or multivariable regression analysis. |
| Lead Time Bias | Apparent increase in survival time after early detection | Fixed endpoint (e.g., age at death); Detection timing is spurious. | Remember: Early diagnosis does not necessarily mean longer life. |
Rapid review table
| Topic | Key Point | Context | Exam Relevance |
| Selection Bias | Sample population must mirror the target population. | Using a specialized group (e.g., ICU patients) to generalize results to the community. | Always question generalizability when the sample is restricted. |
| Intention-to-Treat (ITT) | Analyze all participants according to their initial assignment, regardless of adherence. | Used to correct for loss to follow-up or non-compliance in a clinical trial. | Best method for minimizing selection bias due to dropouts. |
| Confounding | A variable that causes both the exposure and the outcome. | Example: Smoking (Exposure) -> Lung Cancer (Outcome); but SES is the confounder. | Solution: Stratify by the suspected confounder. |
| Blinding | Keeping participants, data collectors, and analysts unaware of group assignment. | Essential for minimizing observer expectancy bias or placebo effects. | Triple-blind studies are ideal; double-blinding is common in drug trials. |
Board-speak -> diagnosis
| Board-speak / Vignette phrase | Diagnosis / Concept | Why it fits |
| A study comparing outcomes only among patients admitted to the ICU versus the general ward population is likely flawed because... | Berkson's Bias (Selection Bias) | Hospital/ICU populations are inherently sicker and different from the general community, limiting external validity. |
| When assessing the efficacy of a new screening tool for cancer, researchers find that early detection appears to extend life span; however, this finding is likely due to... | Lead Time Bias | The true endpoint (e.g., age at death) remains fixed regardless of when the disease was detected. |
| A study investigating the link between smoking and lung cancer finds a strong association, but fails to account for the fact that smokers also tend to have lower socioeconomic status, which is independently linked to poor health outcomes. This flaw represents... | Confounding Variable | Socioeconomic status (SES) is the third variable influencing both smoking habits and cancer risk, masking the true relationship. |
| A study comparing two drug regimens finds that the effect of Drug X on blood pressure only appears strong in obese patients but disappears when stratifying by BMI. This suggests... | Effect Modification | The factor (obesity/BMI) changes the magnitude or nature of the observed association between the drug and the outcome. |
| Researchers ask participants, "Do you have any allergies?" rather than "Are there any allergies you are aware of?" to minimize suggestion bias. | Avoiding Leading Questions (Measurement Bias) | Leading questions introduce systematic error by guiding the participant's response toward a specific answer. |
| A study tracking patient outcomes after surgery finds that patients who were highly motivated and actively participated in physical therapy had better results, but this difference may be due to their inherent motivation level rather than the surgery itself. | Observer Expectancy Bias / Selection Bias | The researcher/participant's expectation or pre-existing characteristic (motivation) influences the outcome measurement. |
Differential diagnosis / distinguishing features
Lead Time Bias vs. Length Time Bias
| Key Features | Distinguishing Findings | Next Step |
| Lead Time Bias: Early detection appears to extend survival time. | The true endpoint is fixed (e.g., age at death); the diagnosis timing is spurious. | Focus on whether the intervention changes the natural history of the disease, not just when it's found. |
| Length Time Bias: Only capturing slowly progressive cases of a disease. | Rapidly progressing or severe forms of the disease die too quickly to be captured in screening/surveys. | Stratify by disease severity (e.g., grade of astrocytoma) to ensure all stages are represented. |
Management pearls
- Addressing Selection Bias: Use Intention-to-Treat analysis when analyzing data from patients who drop out or fail to adhere to the protocol, as this preserves randomization integrity.
- Minimizing Measurement Bias: Implement a triple-blind study design , where participants, data collectors, and outcome analysts are all unaware of group assignments.
- Controlling for Confounders: When designing an observational study, collect data on potential confounders (e.g., age, BMI, SES) to allow for statistical adjustment or stratification.
- Interpreting Survival Data: Always assume that apparent improvements in survival due to early detection are Lead Time Bias unless the intervention can demonstrably alter the natural history of the disease.
Don't miss
Integration & clinical reasoning
- Study Design: The ideal research design is a randomized, double-blind, controlled trial that uses an intention-to-treat analysis to minimize bias from non-compliance or dropouts.
- Epidemiology: When analyzing population health data (e.g., cancer screening), the risk of lead time and length time biases must be considered to avoid overestimating survival benefits.
- Clinical Practice: Recognizing potential biases helps clinicians interpret patient histories and research literature critically, preventing premature conclusions about causality based on flawed study designs.
Concept connections / cross-references
- No explicit cross-references.
High-yield association table
| Condition | Association | Mechanism | Clinical Significance |
| Berkson's Bias | Restricted population sample (e.g., hospital) | The characteristics of the restricted group are fundamentally different from the general population. | Limits external validity; results cannot be generalized to the community. |
| Hawthorne Effect | Behavior changes when observed by researchers/public health officials. | Participants modify behavior because they know they are being monitored (e.g., Jiko effect). | Requires blinding or unobserved measurements to obtain true baseline data. |
| Confounding | A third variable influences both the exposure and the outcome. | Example: Smoking -> Cancer; but SES is the confounder. | Must be identified and controlled for statistically (stratification) or by design. |
| Lead Time Bias | Apparent extension of survival time due to early detection. | The true endpoint (e.g., age at death) is fixed, regardless of when the disease was found. | Leads to overestimation of screening benefits; requires careful statistical modeling. |
Key terms glossary
| Term | Definition | Context | Example |
| Selection Bias | Systematic error due to non-representative sampling or loss of participants. | Any study where the sample group differs systematically from the population it intends to represent. | Using only patients who can afford specialized care in a private clinic. |
| Intention-to-Treat (ITT) | Analyzing all subjects according to their initial randomized assignment, regardless of whether they completed treatment or adhered to protocol. | Used primarily in Randomized Controlled Trials (RC Ts) with dropouts. | A patient assigned to Drug A but who only takes the placebo is still analyzed as if they were taking Drug A. |
| Confounder | A variable that influences both the exposure and the outcome, creating a spurious association. | When an unmeasured factor causes two variables to appear related when they are not causally linked. | Weight (confounder) -> High Blood Pressure (outcome); Obesity (exposure). |
| Effect Modification | The phenomenon where the effect of an exposure on an outcome differs across levels of a third variable. | When the risk ratio changes depending on whether the patient is male vs. female, or smoker vs. non-smoker. | A drug may be highly effective in one ethnic group but ineffective in another. |
Study optimization
| Topic | Study Approach | Priority | Resources |
| Bias Identification | Systematically ask: 1) Is the sample representative? 2) Was data collection flawed? 3) Is there a third variable at play? | High (Must be done for every study design question). | Review definitions of all major biases (Berkson's, Hawthorne, etc.). |
| Confounding vs. Effect Modification | Use the "stratification test": If stratifying removes the difference -> Confounder. If stratifying keeps the difference -> Modifier. | High (Common trap question). | Practice applying stratification logic to vignettes involving lifestyle factors. |
| Survival Analysis | Always assume fixed endpoints unless proven otherwise; distinguish between detection timing and true survival benefit. | Medium-High (Crucial for oncology/cardiology questions). | Memorize the difference between Lead Time Bias and Length Time Bias. |
Question pattern recognition
- The "Why" Question: When faced with a correlation, always ask: What else could be causing this? This prompts consideration of confounding variables.
- Design Flaw Recognition: USMLE loves to present study designs that are inherently flawed (e.g., case-control studies relying on memory).
- The "Is it the same?" Trap: Be prepared to differentiate subtle biases like lead time vs. length time, or confounder vs. effect modifier.
Test yourself
Common mistakes to avoid
Common traps
Original transcript with highlights
Original transcript with highlights
Okay, welcome. My name is Divine, I'm a resident. This is episode 197 of the Divine Intervention Podcasts. In this podcast I'm going to be talking about something I call bias and bios statistics. Bias is kind of a high ill topic, especially for step one, step two, second, step three. It's high ill for the US semily stance to confuse a lot of people. So I'm just gonna go ahead and talk about it and just devote this podcast to it. So you can maybe think of this as like the Bios stats but two podcasts. We are essentially just talking about talking about the biases that we're finding in research studies. So what are the big things you want to keep at the back of your mind with bias? Well, I'll tell you that the very first thing right off the bat is you want to be able to first identify the type of bias, right? And then secondly, you need to know what the solution to the bias is, right? The NV Me pride itself in asking you questions not just about the type of bias, but it also the solutions to that bias, right? So that's something you want to keep at the back of your mind. Now, the thing is these questions people tend to find them to be challenging, right? But let me suggest an approach. Think about it when you were like when you're in trauma medicine or rotation, right? Usually the smart thing you would do is you would go you listen to a patient's chief conflict, right? You listen to all your problems. You write it down, you know, as the patient is talking, write it down, right?
And then as you write down all those problems, you then begin to kind of put it, then after you've written down all the accomplices, you then put a name on it, like, oh, you know, this is the problem the patient has, right? And then after you put a problem on it, right? Then you then match it up with the knowledge you have to suggest the treatment strategy, right? That is really the way you should approach these bias questions on the USMLA's and to be honest with you, that is actually the way you should approach testing general, right? Basically taking a USMLA exam, right? I mean, again, that's why the national body kind of like rewards people that go through this iterative process is something that needs practice and guidance for you to become very good at, right? But if you go through this iterative process of writing down problems in a presentation, putting a name on those problems, and then matching it up with your knowledge, like playing a matching game in the third step to pick out an answer will usually be very successful on exams. At least I certainly have been thankful and successful by doing all those things. So that's kind of like the big thing when I keep on the back of your mind, you know, so you know, kind of learn this biases, learn the knowledge, right? But again, use this approach when you're approaching when you're getting through a bias questions, right? So the first one let's talk about, right? So let's go ahead and talk about selection bias, right?
So selection bias is the first bias. Basically, this bias is that people in your study are bad representations of the population, right? So people in your study are bad representations of the population. That's election bias, right? So basically the question you want to ask yourself, when you're with a, when you're learning about selection bias, or you see an acute status, can I generalize the results of this study to the rest of the world? Usually the answer is no. Okay? So that's something you want to keep at the back of your mind with selection bias, right? So there are many ways selection bias shows up on the USMN exams, right? So like one week can be like if you select a predominant like hospital population, right? For your study. Obviously people that go to hospitals are very different from people that don't, right? People that go to hospitals are probably sicker than people that don't go to hospitals, right? So you cannot take the results of a study that examined only hospital patients and then say, oh, this is representative of the entire population. That's that's not pretty. So that's an example of selection bias. If you want a specific info that that's what's known as, Berksons bias. So B-E-R-K-S-O-M, right? When you use like a hospital population to make those, to make those determinations, that's that's Berksons, that's Berksons bias, right? Another classic one on NBM exams is let's say for example you do a study, right? And in one study arm, right?
There's like a hundred people in one study arm and then in the other study arm, there's like 50 people, right? And then you say, oh, let's say the people in the 100, the group with 100 people are like, oh yeah, you know, they have this advantage, they have this significant difference. Yeah, yeah, yeah, the thing is that's another example of selection bias, right? Because again, in that selection bias, which sometimes actually people call sampling biases, if you don't have equal numbers of people in both groups, right? Then you may be detecting a difference. You may not be, you may not, let me put it this way, you may not be detecting a difference in the 50-person group, right? That actually does exist because the sample size is just too small in that group, right? So the thing is you may actually be making like some kind of statistical error, either like a type one error or a type two error because you have a smaller number, right? But I'll say maybe more so a type two error because the thing that's happening is by having like fewer people in a study group, right? Maybe a difference really does exist, but you're not detecting it, right? That's a power error. That's a type two error. So the thing that's essentially happening is you're incorrectly accepting, sorry, you're incur, you know, you're incorrectly accepting the non-hypothesis, right? So that's something you absolutely want to keep in mind for the purposes of the USML is, right?
Another cause of a, another example of selection biases, let's say a ton of people, right? Drop out of a study, right? Obviously, the people that stay in the study may have characteristics, right? So let's say for example, you're doing a study of like a drug, right? And let's say like 30% of the study population drops out from that study. It may be that the people that drop out from the study, maybe they get very sick, right? And they're like, oh, this drug does not seem to be helping me, but people that stay in the study, maybe they are people that are not very sick, they're like, oh, you know what? This drug seems to be helping, even if it's really not, right? So you're observing differences. So there are like key differences between the people that are in the study and who that have dropped out of the study, right? So like that loss to follow up, that's an example of selection, that's an example of selection bias, okay? That's an example of selection bias. And the problem with that is that you can, you know, you can again kind of mess up the results of your study, right? So ultimately, the way you fix that problem is with an intention to treat analysis. So again, that's a solution to this loss to follow up or kind of business, right? So, I mean, like again, if you want to, let's say, you want to do a study in California, right? And then there's selecting people that live in Beverly Hills, right?
From people that live in like, you know, like some non, I don't want to get in flock for this. So I'm going to be as diplomatic as possible. But you know, like a non-Berwillie Hills part of California, right? Or the people that live in Beverly Hills are very rich, right? So again, you'll not be able to make generalizations to the entire state of California just by looking at people from Beverly Hills, right? That's probably not very pretty, right? So again, that's why you have, um, um, that's why again, that's why, again, you want to make sure that your study population, um, you want your study population to as, want as closely as possible, approximately, the general population, right? Or another classic one mission, an exam is like, uh, they call it like volunteer bias or like non-respondent bias, right? Because think about it, right? Let's say if you're doing like a survey, obviously, people that participate in a survey, right? Maybe people that have more awareness of the health condition or something, right? Versus people that choose to know respond, right? So the thing is by having those differences in people that respond or not respond to a survey, right? Again, people that respond may have characteristics that are different from people that do not respond, right? So that will obviously introduce a selection bias to your study, right? So again, in general, the way you fix election bias is, you know, you're randomized people are in the study, right?
You try to go as much randomization as possible. You try to match the characteristics of people because the thing is when you're doing a research study, right? Especially like a control trial. The thing you want to accomplish is you want the people that you want the only like variable, right? To be the intervention that you're deploying, right? You don't want other variables to come into play because that can begin to introduce confounding factors as I'll talk about in a bit, right? So you want to avoid all those problems, right? So in general, you try to match people, you know, similar characteristics, similar age, similar race, similar everything, right? So that you don't have, you don't have these kinds of a selection or sampling bias problems, right? Again, like I said, like the intention to treat analysis is good for the loss to follow problem, right? And again, you want to try to keep the people in the study groups versus the control groups. You want to keep them as equal in number as possible, right? That's that's an important thing to keep at the back of your mind. And then you'll probably, let me just maybe talk about this open label trial term, right? So an open label trial is a trial where you just enrol, you know what arm you're involved in. I mean, some trials employ that, but that's probably not a good trial, right?
Because again, you'll see as I mentioned in a bit that, you know, when you have a blinded trial, that usually brings some benefit that you may not that's maybe absent in an open label trial, right? So now the next general, so if you notice, there are many types of selection bias, right? So selection bias is like a big dog, right? And then there are many small dogs under like small kids under it. Now, the next kind of bias is a measurement bias, right? The next kind of bias is measurement bias. And so measurement bias, basically it's the way you obtain your data distorts the information you get from the study. I'll see that again. The way you obtain your data, right? Distorts the information that you get from the study, right? So basically in this case, like the researcher may have a bias of his own, right? That affects like his data collection, his data analysis, right? I mean, it's kind of like a common error that people make on step 2 CS, right? So like I talk about this with people all the time when I'm touring or step 2 CS, right? And you ask people about allergies, you ask them, do you have any allergies? That's it, right? But if you see, do you have any drug allergies, right? That you're essentially introducing measurement bias because you're using a leading question, right? You don't. In general, in clinical medicine, you don't use leading questions because if you use leading questions, patients will tell you what you think you want to hear, right?
So again, don't ask leading questions. But that's usually a bad idea on step 2 CS. In fact, that's probably one of the things that makes people like lose ice points on the exam, right? So just something to keep, something to keep at the back of your mind, right? Something to definitely keep at the back of your mind, for example. So what are some examples of this measurement bias? So one is like the Hawthorne effect. I feel like the easiest way to understand the Hawthorne effect, think about it, right? When you were a kid, when you're a parent, you're home, right? You were more likely to do the right thing, right? But once they go out, they're like, oh, kids, we're going out for a trip this weekend, right? You then, you then kind of loosen up, you know, maybe poor parties and home and everything, that you would ordinarily not doing your parents were there, right? So because you knew you were being observed by your parents, you know, you did the right thing, but once your parents were gone, you started doing the wrong thing, right? That's basically the Hawthorne effect, right? So if people know, or another classic example that maybe drives this home, right? So I mean, you probably, everyone in medicine has probably heard of Jiko, right? So they are like a buddy that comes and, you know, kind of sets up some very, let's call them unique requirements that certain parts of hospitals, you know, need to meet, right?
So the thing is when people know that the Department of Health or Jiko is around, right? People are going to do the right thing, right? Like, you know, there will not be computer monitors in the hallways, everything will be perfect in the hospital, right? Because they know Jiko is coming, right? But once Jiko goes, right? Everything returns back to normal, right? So again, the nurses, the physicians, the med students, the healthcare professionals, they know that, oh, Jiko is watching them, right? Big brothers watching them, so they need to do the right thing, right? So that's an example of the Hawthorne effect. The participants in the study, they know they are being observed. So because they're being observed, then they stop a, they start acting in ways that may skew the results of the study, right? That's an example of the Hawthorne effect. Now, and one thing that really helps is blinding, right? If you're blind the study, if you don't know what area you're, if you don't know, as either like, like as a study participant, you don't know what group you're in, the person analyzing data doesn't know what group you're in, the person collecting data doesn't know what group you're in, right? That's like a triple blind study. So the analysts, the participants and the study coordinators do not know anyone's group, right? That really does help. That's triple blind.
Double blinding is just the patients and this study like the principal investigators, not knowing who, what assignments the patient has, either that, he doesn't know if they're in the control group or in the experimental group, right? So again, those are all big things, those are all big things to keep in mind for, yeah, those are all big things to keep in mind, for example, right? So blinding, blinding absolutely positively helps with this, with this type, with this type of air, right? Another one you may see is like the Pagmallion effect, the Pagmallion effect is an example of a measurement bias, right? So it's where like, you know, like if it's like, oh, you have this expertise as a researcher, right? Then you're more likely to collect your data in a certain way versus a person that is non-biased, right? So it's like if, for example, you have like big, big, big expertise and a subject or you're like friends with some people in the study, right? Then you may collect data on them differently from then a person that, oh, you know, let's say this person, I mean, let me think of, let me give a classic example. So let's say, for example, you're a PI, right? And you're going, you're doing an Indian PI, right? And let's say you grew up in India for decades, right? And then you're collecting data like Indian subjects, chances are, you know, you probably collect your data in a skewed way, whether you, and it's not something you're doing intentionally.
These are just unconscious biases that people have. Even me, I know I have my own unconscious biases, right? So, so like an Indian PI will be more likely to make this kind of Pagmallion effect a problem than a person that let's say it's someone that doesn't know anything about India, doesn't know anything about Indians and just goes and collects data on them, right? So that's who, that's an example of the Pagmallion effect is just like some pre-existing knowledge you have or some pre-existing characteristic you have that skews the way you collect data, right? So, let me make sure that you understand the difference between the whore-thorn effect and the Pagmallion effect. And the whore-thorn effect, it's a research participant characteristic that kind of messes up the way you, like they know they're being looked at, right? They're being observed, so that modifies the whore-thorn effect. Pagmallion effect is almost like the reverse. It's like the person that's collecting the data has certain characteristics within himself that modifies the way he collects data, right? So those are all examples of measurement bias. Another classic example you'll be seeing examples is something called like the observer expectancy bias, right? Observer expectancy bias, right? So it's like, oh, whatever you expect from the study, right? Like you're like, oh, going in, you're like, oh, these are the results I expect from this study. And then that kind of affects the outcomes that you get, right?
Because the way I think about it is, if, for example, you expect certain results from a study, let's say, oh, you're like, okay, you know what? By, I'm going to go ahead and do a research study of US Emily scores of students at Johns Hopkins, right? And you're like, oh, okay, Johns Hopkins, these are very smart students. It's a top ranked school in the US, right? So I expect that these people from Johns Hopkins should all have very high US Emily scores, right? So the thing is because you expect that people from Hopkins should do well, again, I'm using Hopkins because I was a Hopkins med student, but that's a different conversation. But you know, because you expect people at Hopkins to do well, you know, you may give them more TLC, right? So like more tender love and care, like if they don't understand something the first time, you may try to find other ways to explain it and everything, just to kind of force the results that you want, right? So that's an example of an observer expectancy bias, right? So again, that's an example of measurement bias because, again, the way you've interacted with those people, right, has kind of affected basically the way I think of measuring biases. It's a bias that arises from interactions. Like the way a study participant interacts with the PI, I mean, affects the kinds of results you get from that study, right? That's an example of a measurement bias. And again, one of the best ways to deal with measurement biases blinded, right?
It's blind as much as possible. Blind, blind, blind, blind, blind, blind. And also, when you have like control groups, right? Control groups also helps with eliminating some of those some of those biases, right? Now, one bias that is very, very, very common is understood amongst people is lead time bias, right? I mean, this is something that you know, you see many resources, not really explain very well, right? So let me, because many times people memorize the buzzword, right? I mean, I'm sure everyone of you listening to this post can probably hear this that, oh, confusing early detection with improved survival, right? But the thing is, I mean, I'm, again, thankfully I have 200 tons of people in my life, right? And the thing is, there are many people that does the thing that kind of kills people on these, on these US Emily exams. They know all the buzzwords, they know all the buzz phrases, but they're wondering, and how am I doing very poorly on tests, right? And this is kind of like my, my quep with with Anki, right? Many people, Anki is great. Anki is actually phenomenal, but many people use Anki wrong. You see many people, they just use Anki to like, even they are ready threads on this before those me to Disney. You see people, they use Anki, they're like, oh, yeah, divine, I'm learning this material real well. And you see them clicking through, hitting the ones that choose the three, they're like, yes, I am crushing this knowledge.
But the thing is, they think you didn't write an exam question and they just completely fall apart, right? The thing is, if memorizing, trust me, if memorizing was the key to doing well on US Emily exams, most people do extremely well. Well, that's not the case, right? The thing is, you need to take context and understanding into consideration, because many people, they understand lead time bias, right? But how you give them like a well-written question on lead time bias, and they're like, what's going on here? They don't recognize it, right? So again, that's something you want to be careful about. Always try to approach something from a standpoint of understanding and from a standpoint of context, like it is not like one thing that is kind of annoying when I tutor, sometimes is you hear people, they say, oh, so divine, if I see this, that means it must be this, right? No, that's not the way it is. You also have to take the context of a question into consideration, right? So again, that is where the understanding comes from. That's where the understanding comes from, right? So lead time bias, the buzzword is, oh, confusing early detection would increase survival. Fine. Well, what in the world does that mean? Let me give you an example, right? So think of it this way. Let's assume you have this disease, right? Let's say it's a particular kind of cancer, let's call it cancer X.
And let's say between when a person gets the first genetic mutation that causes that cancer, and when they ultimately die is 10 years. So from like first mutation to death, 10 years, right? So let's see, that's like a fixed thing, like there's nothing changing that, you know, mutation, death, 10 year period. Obviously, if you detect that cancer at year five, the person is going to live for five years after the diagnosis was made. But if you detect that cancer at year three, the person is going to live for seven years after the diagnosis was made. That is an example of lead time bias. The thing is if we could extrapolate to when the genetic mutation initially occurred in that person and measure to when the die you see that the time, if we cannot wait whatever data we have for those who will see that, oh, they don't actually end up dying at the same time, right? That's an example of lead time bias, right? That's lead time bias, right? It's like the time you die is fixed, what you just like is like, okay, let's say everyone that gets again, this is a spurious thing. It's a terrible example, but you'll probably drive the point home. You know that everyone that gets this disease dies at each 40, right? If you find the disease when they are age 35 versus you find the disease when they are age 20, it'll look like the person that you found the disease at each 20 lived longer, but no, they didn't live longer. You just found it earlier, right?
But it's still died at the same fixed age of 40, right? So that's an example of that's an example of that's an example of lead time bias, right? But the thing is sometimes your friends at the MVMDM are questions where it is not lead time bias, but you'll write the question to make it look like a lead time bias question, right? So let's go with this 10 year mutation to death sequence I talked about, right? So let's say for example, right? You know, you detect the disease by year three, right? And by detecting the disease, you know, maybe you have like a treatment strategy you employ, and you notice that, oh, this person actually lives for like 17 years after the detection was made. Then obviously that is not lead time bias, that is actually a true legitimate improved survival, right? That is a true legitimate improved survival, right? So that is not lead time bias, because by actually finding the disease earlier, you're able to do something that made the person not die at that fixed 10 year point, right? That is not lead time bias, right? So again, make sure you understand lead time bias. This thing just pops up on pretty much every USM elite exam, right? So it's not something you want to get, it's not something you want to get broke. Now, here's another quan, people have with lead time bias, right? Some people confuse lead time bias with something called length time bias, right? So people confuse lead time bias with something called length time bias.
In fact, another name for length time bias is something called lead look bias. Basically, the thing that happens here is you essentially never come in contact with the worst cases of a given disease, right? So say for example, right? Let's say you're collecting data on people that have brain cancer, right? And you're like, oh, let me collect data on the kinds of symptoms that you have, right? And you know, you see people bring cancer and they just have headaches, and you know, they don't seem to have any severe neurologic deficits. They're still able to accomplish the activities of daily living. Well, what if those people are getting that data from a just people that have like meningiomus? Then obviously, right? That kind of skews your data, right? Because the thing is, for example, you may be less likely to catch people that have glioblastoma because many of them just die like the time between diagnosis and death is so short, right? So like the rapidly progressive cases of that disease, you do not catch, right? So it may be let us stray, right? So basically, it's like the construct here with this length time bias that tends to pop up on the USML exams. And the thing is they love to do it in terms of cancer screening as well because they know that when people see cancer screening, the only thing that comes to their is lead time bias, right? So don't confuse this on your exam, right?
Length time bias is just, you're not catching the people that have like a rapidly progressive form of disease. You're only catching people that have like slowly progressive disease, right? So that again, me skewed results of your study because you may be saying, oh, so everyone that has brain cancer, they don't have this, they don't have that, they don't have this, they don't have that. But if you found people that actually have GBM, you're like, ooh crap, they do actually have this really nasty symptoms that are pop up, right? So the thing that can happen to fix this kind of problem is stratify people by disease severity, right? So like if you are checking symptoms and people that have brain cancer, like stratify them by, oh, these are the symptoms that are found in people that have glial blastoma motif for me. These are the symptoms that are found in people that have many genomes, people that have a panepangimomas, people that have hemangioblastoma, because think about it again, if you had a study that just, oh, only sample people that have hemangiobl stoma, you think that, oh, everyone that has brain cancer has elevated hematocrate from increased levels of epa, but that's not true. That is just something that's unique to hemangioblastoma, right? So again, or if you are comparing like astrocytomas, right? Again, astrocytomas come in different flavors, right? Like GBM is a grade four astrocytoma, right?
It's a grade four astrocytoma, but they are grade one astrocytomas that are not as virulent if you may, right? They're not as rapidly fetal, right? So again, you want to stratify by disease severity to sort of deal with this length time bias, right? And then another common bias mission is exam, right? Record bias, right? So record bias, this is a big one that shows up in case control studies, essentially the thing that happens is people have hematocritic record, right? And again, it's not like these people are trying to lie to you, right? It can just be hematocritic record for many reasons, right? So because think about it, right? Let me let me ask you this. Do you remember exactly the food you do remember exactly what you ate? I mean, this is December 31st. Thank God for bringing us to the end of the year. But do you remember exactly what you ate on December like on January 17th? This year, do you remember exactly what you ate to the little crop? Like, oh, this is what I ate on December 17th. It was this and this and this and that. No, right? You can't, right? So obviously you can already see that this is something that would happen in a case control study because the case control study are you telling people you take cases of disease and then take controls that do not have disease and then instead asking them about things from like backing the day, right? It's very hard.
I mean, not everyone has a good memory and even people that have good memories, they make a little mistake, right? So again, in accurate record, that's record bias, right? So or I mean, think about it, people that have a certain disease that have had to manage for a long period of time, you know, they probably have better recall because they're usually much better at tracking those kinds of things compared to people that have just never had that kind of disease, right? So that's one thing you want to keep at the back of your mind for, for example. So, how does some ways you can fix recall bias? Well, one way you can fix it is just try to reduce the length of time between when you when the event happens and when you're trying to cause, like, when the outcome happens and how far behind you're trying to get people to do recall, right? So, that's one thing you want to do. And other things, just try to find some means of like corroborating the information that the person is giving you, right? That's one, those are two nice ways of dealing with dealing with recall bias. Now, the last, I guess, two biases I want to talk about is, and I may be able to compare and contrast, right? But let me talk about confounding, right? Let me talk about confounding and then subsequently talk about something called effect modification. And that's something that just screws people over like a lot on exams, right? Again, this is not something that people seem to understand very well, right?
So, let's say, like, let me maybe start with this construct. Think about it. If you're doing a research study, right? You want as many things as possible to be constant, right? With one exception. The thing the exception to that rule is the intervention, right? Or the exposure you're employing in that study, right? You want like, oh, okay, this group is getting this, this group is not getting this, because whatever outcome you get, if every other thing is held constant, whatever outcome you get is likely linked to the intervention or the exposure that you're through to play, right? So, but if, for example, there is some other factor in the study that is explaining the differences in outcome, that is a confounder. Because you like your results, no one is arguing with the results of your research study. You're like, yeah, you know, this, this, I saw that there was a bigger decrease in blood pressure in this group compared to the control group. But if you're not sure of what was responsible for causing that decrease in blood pressure, you're scratching your head, you're like, wait, was it the drug I gave? Was it something else? Was it this? Was it that? When you have that kind, when you read an MDM exam question and you're kind of asking yourself those kinds of questions, then you are dealing with confounder, okay? You're asking yourself, why in the world did I get these results, right? That's a confounder, okay? That's a confounder.
And really, how do you fix a confounder problem? Well, the way you fix it is you just stratify people, right? You stratify people. So, let me, let me think of a, let me think of a nice, let me think of it, of a nice example I can use, I can use here, right? So, let's say for example, you say that, oh, you know, there's a, you notice that there is a high incidence of, you don't think of a good example, okay? Let's say, there's a high incidence of pancreatic cancer in people that drink a lot of soda, right? People that drink soda tend to get pancreatic cancer, right? And you're like, oh, that's it. But let's say, you then notice that, oh, wait, a lot of people, if you then say, okay, wow, soda, like soda seems to, you know, cause a lot of pancreatic cancer. And then let's say you're like, okay, let me stratify, by, okay, let me use a more classy example. So, let's say a lot of, coma, let's see, you do the study and you find out that, oh, coma nurse tend to have long cancer, right? Tom, like, oh, coma name was basically a long cancer. But it may be that weight. That coma name may not actually be the risk factor for long cancer. It may not really just be that, people that are coma nurse smoke a lot, right? So, smoking, basically, confounding is something where, like, you're like, you think something is the cause of the outcomes you're observing in the study, when there is, in fact, something behind the curtain that's actually causing those problems, right?
So, it may be that, oh, wait, these co-minors, yes, coma name has that relationship, right? But it's actually not the co-mining that causes the long cancer. It's actually the smoking that the co-minors do that causes the long cancer, right? So, if you have one variable, essentially, masking the effect of an order of variable, that is something you want to think about as, that is something you want to think about as, it doesn't really want to think about as a confounder. But let's say, for example, you're like, oh, okay, let me see if this is actually true. Let me test and see if the smoking is actually a confounder. And then let's say, you know, you compare long cancer rates in, in co-minors that smoke, and then you compare long cancer rates in co-minors that don't, like, you look at long cancer rates in co-minors that smoke, and in co-minors that don't smoke. And then you notice that, oh, wait, there are actually, there are actually differences observed here. Then you know for sure that you're dealing with a confounder because when you have, when you have confounding, when you stratify the patients in those two groups, that eliminates whatever effects that you find. Or, I feel like, okay, in fact, now that I think more about it, I feel like this coworker long cancer example is something that people see all the time, right? But it still confuses them a little, maybe like, oh, divide, this is a characteristic big thing.
Why don't you use an example that involves an intervention, right? So let me use an example that involves an intervention. So let's say you make a new blood pressure drop, right? You make a new blood pressure drop, and then you have like, you compare that a new blood pressure drop to like the current, let's see, compare it to like placebo, right? Or you're comparing like two groups, right? And you notice that, oh, in this group, let's say like, okay, you give the drop to two groups, and let's say in group, you notice that, man, these people's blood pressure went down by like 40 points. But in this other control group, this, they're blood pressure went down only by like, five points, and you're like, what's going on? And then let's say you, you then figure out, you're like, examining your data, you're looking at the characteristics, and you're like, wait, wait, hold up here. The people in group A, most of them had, most of them were like, not obese, but the people in group B, most of them were obese. So could it be that obesity is the thing that's actually responsible for these decrements in blood pressure, not necessarily the drug and given, right? So you say, okay, let me figure out if obesity is a confounding factor. So you go ahead and then say, okay, let me just take obese people, just obese people, right? And then break them up into two study groups, give them this drug, and then give them a placebo, right?
If you notice that, oh, wait, when you actually stratify people by like their BMI, and you notice that, oh, wait, they actually no differences in blood pressure lowering between these groups is like, oh, by cutting people or by being obese or not obese or just cutting people or by BMI, the effects that you saw that, oh, my drug reduced blood pressure, if it evaporates, then that tells you that obesity is likely a confounding factor. But let's say for example, you do that second thing where you're like, oh, you know, let me, let me, you know, kind of break people up by being a stratify people by obesity. But you notice that in those obese groups, you are still observing decreases in blood pressure, then you're dealing with something known as effect modification, because stratification did not erase those effects that you were studying. If you see that kind of thing, then you know that you are likely dealing with effect modification. And the thing is that actually, you may say, oh, divine, that's a bad thing that has screwed up the results of my study. No, it actually gives you more information that, oh, okay, you know what? There's probably some factor that some of the obese people have that some of the others do not have that may be more defined the effects that I'm seeing from my intervention. That's why it's known as an effect modification, right? But again, I know that people may not have the time to think this deeply about this. So let me give you what?
Let me give you a trick that may help you on USMLA exams to differentiate effect modification from confounder, right? The thing is if stratification eliminates the differences that you observe between two groups in your research study, you're dealing with a confounder. If stratification does not eliminate the differences you observe between two groups in your research study, that's effect modification. That's one way to tell them apart. But sometimes it may not be that clear on a USMLA exam. So here's on that trick you can use. The thing is if the key stem that is being referred to if the question you're reading is talking about like a lifestyle thing, right? Like smoking or obesity or blah blah blah, right? Or alcoholism. That is very likely to be a confounding question. Okay? That is very likely to be a confounding question. So if it's again related to like a lifestyle thing that you do is very likely to be a confounding question versus if, for example, they are talking about like an actual intervention like a medication or a kind of surgery or something, right? In the key stem, then it's likely to be an effect modification style of question. This trick probably will get you like 99% of your questions right involving either of these either of these concepts. So I think I'm going to go ahead and stop here. I mean, I feel like I've talked about pretty much everything I want to talk about.
I mean, like they could I would have to investigate this further, but I feel like this whole thing about ACE inhibitors not being like first line for hypertension in African Americans is probably a good example of effect modification because it's like why do ACE inhibitors not work as well in African Americans as they do in Caucasians, right? It's probably an example. I can actually imagine I mean again, it will take a lot more thought on my part to do this, but I can already imagine writing a being able to write a question that relates that factoid to effect modification as a that will actually be a phenomenal endemic exam question. So again, those are just again, things to things to keep at the back of your mind for exams. So as I do, I do I do every podcast. I do all for one or one tutoring for many exams, both one on one and actually large group tutoring for step one, step two CK, step two CS, step three, preclinical medical exams, 30-ish-off exams, if you're a medicine resident, like the medicine-intering exam or the internal medicine, IBA, IBA, IBA exam, or if you're a college student, like general chemistry, organic chemistry, physics, biochemistry, histology, physiology, most basically most of the MCAT subjects I do offer tutoring for those things. And again, I do large group, I do one or one.
If you want to find out about any of these things, you can just send me an email, divine intervention podcasts with an S at the end at gmail.com, or where you can reach out to me through the website. And then if you need like, I also offer this booster course for the USMN Es, it's like 10 hours for step two CK, step three, 20 hours for step one. Basically, if you are at the end of your dedicated period, or you know, you feel like a knowledge business is pretty good, and you want to in a very short period of time, in an integrated format, review like the most knows for the exam. Again, reach out to me, I've done this with many people, many people done this with, they found it to be like, wildly successful and very helpful for them. And then there's also this thing I do, but it's for a very unique subset of people. So this unique subset of people are basically people that, like there are certain people that learn things, like they learn by question review, right, where I literally we just work together, and you know, we take a Q bank, and essentially go through that Q bank together. It's very time consuming, but it's very effective for certain kind of learn, right? The thing is those kinds of things, like those kinds of determinations, you can make them after tutoring a present for like one or two sessions, right? They are certain people where like I've worked with where I literally take like a Q bank, and then we go through like 40 random questions each time.
And the thing is, like it's almost like a guided discussion, where like I see the thought process in picking out an answer, and then I explain my thought process. And then as they explain, the thought processes and detecting errors in the way they think, right? So it's like at the same time I'm teaching them how to like take tests better, but I'm also teaching them concepts, and then I'm also doing like a compare and contrast of the different answers for them, right? Many people find it to be like a pretty phenomenal learning experience, but again, it's very time consuming, but it's a very effective method for actually certain groups, certain groups of people. So if that's something you're interested in, again, it will be something that will be done ideally over like a five, maybe six week period, like five six week dedicated study period, but again, it's something that has been very effective for a bunch of people I've actually worked with. So if you're interested in any of that, or if you need coaching for like, if you're like a met student applying to residency, so like an era's application, or a college student applying to a met school, so like an Amcass application, I do offer like again, one-on-one coaching, one-on-one advising like rec letters, personally submitted in editing applications, mock interviews, again, I've worked with tons of people.
Many of people have worked with, they've much done their first choices, and I've also I have admissions committee experience. I've been on the admissions committee of a top two met school for like a year, right? So again, I've reviewed thousands of high, very high quality applications, so again, if any of that expertise reached out to me as I've mentioned earlier, so have a wonderful rest of your day, and I wish you a happy new year, I'm praying that it's a year of blessings, of peace, and of your hard desires being fulfilled. So I'll see you in the next podcast, thank you, and God bless you.
Practice questions — USMLE style
Question 1 — Biostatistics/Selection Bias
A research team conducts a study investigating risk factors for cardiovascular disease by recruiting participants exclusively from a large urban hospital's emergency department over a six-month period. The researchers find a strong correlation between elevated blood pressure and poor diet among the admitted patients. They conclude that poor diet is a major, generalizable risk factor for CVD in the entire metropolitan population. Which type of bias most likely compromises the external validity (generalizability) of this study?
- A) Observer expectancy bias
- B) Length time bias
- C) Berkson's bias
- D) Confounding by indication
Answer: C. The study is subject to Berkson's bias, a specific form of selection bias. This occurs when the sample population (hospitalized patients) is inherently different from the general population because they are already sicker or presenting with acute conditions. Therefore, the results cannot be reliably generalized to healthy individuals in the community.
Question 2 — Biostatistics/Measurement Bias
A clinical research team designs a study to test a novel physical therapy regimen for chronic back pain. The participants are informed that the study is designed to identify the most effective treatment and that their adherence to the protocol will be closely monitored by the researchers, who provide frequent feedback and encouragement throughout the trial period. After six months, the researchers observe significantly better outcomes in the intervention group compared to the control group. Which type of measurement bias is most likely responsible for artificially inflating the observed positive results?
- A) Observer expectancy bias
- B) Hawthorne effect
- C) Recall bias
- D) Pagmallion effect
Answer: B. The Hawthorne effect describes a phenomenon where participants modify their behavior simply because they know they are being observed or studied. In this scenario, the awareness of monitoring and the active involvement of the researchers likely caused the participants to improve their adherence and self-care, leading to artificially inflated results that do not reflect true long-term efficacy.
Question 3 — Biostatistics/Confounding vs. Effect Modification
A study is conducted to determine if a new medication (Drug X) reduces blood pressure. The researchers compare two groups: those receiving Drug X and those receiving placebo. They observe that the reduction in blood pressure is significantly greater in the Drug X group. However, upon analyzing the data, they find that most participants in the Drug X group were also highly physically active, while the control group had a lower average level of physical activity. To determine if physical activity (a potential confounder) was responsible for the observed difference rather than the drug itself, the researchers stratify the patients by their BMI status and then compare the blood pressure reduction between groups within each stratum. They find that when comparing only obese individuals, the difference in blood pressure reduction between Drug X and placebo is still statistically significant. What does this finding suggest regarding the relationship between physical activity, medication, and blood pressure?
- A) The initial observation was due to a confounder, meaning the drug has no true effect.
- B) The observed difference represents confounding, which must be controlled for by statistical adjustment.
- C) The observed difference suggests effect modification, indicating that the drug's efficacy is dependent on an underlying characteristic (obesity).
- D) The study suffers from Berkson's bias because the sample was drawn from a specific clinical setting.
Answer: C. Effect modification occurs when the relationship between an exposure and an outcome changes depending on the level of a third variable (the effect modifier). Since stratifying by BMI did not eliminate the observed difference in blood pressure reduction, it suggests that obesity is modifying or defining the drug's true effect. If the differences had been eliminated upon stratification, it would suggest confounding.
Question 4 — Biostatistics/Lead Time vs. Length Time Bias
A large-scale cancer screening program detects a rare type of lung cancer in asymptomatic individuals years before symptoms typically appear. Researchers calculate that early detection significantly increases the estimated survival time for these patients compared to historical data. However, subsequent analysis reveals that all patients ultimately succumbed to the same underlying biological process and died at roughly the same age range, regardless of when the diagnosis was made. This discrepancy between the observed increase in survival time due to early detection versus the fixed natural history of the disease is best explained by which bias?
- A) Length time bias
- B) Berkson's bias
- C) Lead time bias
- D) Observer expectancy bias
Answer: C. This scenario describes lead time bias. Lead time bias occurs when measuring the duration between diagnosis and death (lead time). If the true biological endpoint (the fixed age of death or mutation onset) is constant, detecting the disease earlier will artificially inflate the calculated survival period, making it appear that early detection improves overall survival when, in reality, it only changes the timing of the diagnosis.
Quick fire review
What is selection bias?
When the study participants are not representative of the general population from which results are generalized.
What specific type of selection bias occurs when a study uses only hospital patients to draw conclusions about the community?
Berkson's bias.
How can loss to follow-up (attrition) introduce selection bias in a drug trial?
The people who drop out may systematically differ from those who remain, potentially skewing the observed treatment effect. This is best addressed by Intention-to-Treat (ITT) analysis.
What is the primary difference between the Hawthorne effect and observer expectancy bias?
The Hawthorne effect relates to participants changing behavior because they know they are being observed. Observer expectancy bias relates to the researcher's expectations influencing data collection or interpretation.
In a case-control study, what type of bias is most common due to poor memory recall over time?
Recall bias.
What is the key difference between confounding and effect modification?
Confounding is eliminated by stratification (the relationship disappears). Effect modification occurs when stratification does not eliminate the difference, suggesting a genuine interaction between variables.
Definition of Lead Time Bias?
Confusing earlier detection of a disease with improved survival time; the total lifespan remains fixed regardless of diagnosis timing.
What is the primary method to mitigate confounding bias in observational studies?
Stratification (or matching) by potential confounders.
Which type of bias arises because participants change their behavior simply knowing they are being studied or observed?
Hawthorne effect (a form of measurement bias).
If a study finds that the association between two variables disappears when stratifying by a third variable, what is this third variable likely doing?
Confounding.
What must be done to minimize selection bias related to participants dropping out of a long-term study?
Use Intention-to-Treat (ITT) analysis.
Name the three components that must all be blinded in an ideal trial design to prevent measurement bias.
The study participant, the data collector, and the data analyst.
Quick recall / Anki-style questions
Definition of Lead Time Bias?
Confusing earlier detection of a disease with improved survival time; the total lifespan remains fixed regardless of diagnosis timing.
What is the primary method to mitigate confounding bias in observational studies?
Stratification (or matching) by potential confounders.
Which type of bias arises because participants change their behavior simply knowing they are being studied or observed?
Hawthorne effect (a form of measurement bias).
If a study finds that the association between two variables disappears when stratifying by a third variable, what is this third variable likely doing?
Confounding.
What must be done to minimize selection bias related to participants dropping out of a long-term study?
Use Intention-to-Treat (ITT) analysis.
Name the three components that must all be blinded in an ideal trial design to prevent measurement bias.
The study participant, the data collector, and the data analyst.