Skip to content

Episode Notes

USMLE purpose: Separate confounding from effect modification and choose the right statistical test for means/tables.

How to use this page

  1. Start with the confounding vs effect modification table.
  2. Review the statistical test guide until you can match the data type to the test.
  3. Use the rapid recall toggles before looking at the practice explanations.
  4. Use the transcript only if you want Randy’s original wording.
Episode metadata
FieldDetails
EpisodeRandy Neil Biostats 09
TopicEffect modification vs confounding; t-test/ANOVA
Runtime14 min
Published2020-07-26
SourceOpen YouTube video

One-liner

Confounding affects both exposure and outcome; effect modification changes the outcome relationship in a subgroup and should be reported, not “controlled away.”

Confounding vs effect modification

ConceptWhat it affectsWhat to do
ConfoundingExposure and outcomeControl/adjust for it
Effect modificationOutcome relationship differs by subgroupReport stratified effects

Statistical test quick guide

Data / comparisonTestShortcut
Two meansTwo-sample t-testTwo groups, continuous outcome
Three or more meansANOVA“A-NO-VA” = think multiple means
Proportions / 2×2 tableChi-squareBox/table comparison
Pooled data from studiesMeta-analysisCombines studies; increases power
NNT1 / ARRHow many treated to prevent one event?

Exam pattern recognition

  • Smoking changes medication → DVT relationship but not medication exposure → effect modification.
  • Smoking relates to alcohol use and pancreatic cancer → confounding.
  • Compare vitamin D means in two groups → t-test.
  • Compare categories in a 2×2 table → chi-square.
  • RR/OR confidence interval crosses 1 → not significant.

Common trap

⚠️ Trap: Treating effect modification as bias.
Fix: Effect modification is real subgroup variation; confounding is distortion.

Board Exam Buzzwords

BuzzwordWhat it should triggerClinical / exam context
Third variable affects exposure and outcomeConfoundingDistorted relationship
Third variable changes outcome relationship by subgroupEffect modificationReport subgroup-specific effects
Two meanst-testContinuous outcome, two groups
Three or more meansANOVAContinuous outcome, multiple groups
2×2 categorical tableChi-squareProportions / association
Pooled studiesMeta-analysisIncreases power

Anki-style rapid recall

How do you recognize confounding?

The third variable is associated with both the exposure and the outcome.

How do you recognize effect modification?

The effect differs by subgroup. It is a real interaction, not a bias to remove.

Which test compares two means?

Two-sample t-test .

Which test compares three or more means?

ANOVA .

Practice Questions

Question 1

Smoking changes the medication → DVT relationship, but smoking does not affect whether the patient received the medication. What is this?

  • A) Confounding
  • B) Effect modification
  • C) Recall bias
  • D) Selection bias
Reveal answer & explanation

Answer: B) Effect modification

The effect differs by smoking subgroup. It modifies the outcome relationship rather than distorting both exposure and outcome.

Question 2

Which statistical test compares vitamin D means between two groups?

  • A) Chi-square
  • B) ANOVA
  • C) Two-sample t-test
  • D) Meta-analysis
Reveal answer & explanation

Answer: C) Two-sample t-test

Two groups + continuous outcome/mean comparison = t-test.

Quick Reference Summary

TopicKey pointUSMLE buzzword
ConfoundingAffects exposure and outcomeStratify / adjust
Effect modificationEffect differs by subgroupInteraction
t-testTwo meansContinuous data, 2 groups
ANOVAThree or more meansMultiple group means
Chi-squareCategorical table2×2 / proportions

📌 Transcript note: The transcript is kept below for reference only. Use it if you want to verify the original video wording; the high-yield study content above is the main review tool.
Full Transcript

Alright guys, so in this video we're going to talk about the confounding and versus effect modulation and we'll also kind of touch on the difference between the ANOVA Chise Square and the T-test and such. So this will kind of round out the rest of the bios stats that I'm aware of and hope you like to video. Alright guys, so question number one, it says a new medication is aimed to treat schizophrenia works on the M&MDA receptor as an antagonist, a prospective study shows that women who smoke in women who smoke that the drug increases the risk of deep brain thrombosis. The relative risk associated with this phenomenon is 1.5 with a p-value of 0.02. In treated non-smokers, no risk of DBT is evident relative risk of 0.95 and p-value of 0.45. This finding is an example of which of the following. So long story short, we didn't see your answer choices and it's like they did a study and then when they kind of threw in this extra little piece here, it obviously changed the result because the relative risk looks like here is 1.5, meaning it's 1.5 times greater and then as p-value, remember we want to p-value less than 0.05 to make it significant, right? This is like saying this is the chance of it occurring by chance. But then when they said look, if we had non-treated non-smokers of this, then everything kind of went different, right? Went from 1.5 to reverse, right? Because if it's less than 1, then it was like a reduction. But p-values of pretty stinks, it's way too big. But when you see the difference of when you see con-founding and effect modulation in the answer choices, you know that they're probably leaning in that direction. But you got to know the difference between the two right off the right of the gate. So with a con-founder and effect modulation, and actually it's pretty easy that a work for step 1 step 2, is that a con-founder, okay, I think that like this,

you draw it out and then have exposure and outcome, exposure and outcome, okay? And this is a simplified version of this. Is that a con-founder affects both the exposure and the outcome, whereas if it's an effect modulation, it only affects the outcome. Now in this question, you know, we started talking about there's this new medication, and that it showed that this medication, okay? So that's the exposure and they were looking at an outcome of DVT. But then they did, when they kind of did a follow-up on this, they said, well look, what about smoking, okay? Because they treated us as an intruded non-smoker's at a different reaction. So smoking, we have to determine whether that was a con-founder or an effect modulation, and it's real simple. If it affects the, you know, the astrosel, does smoking affect DVT, and we say yes, well, we know that, even in that's kind of a given, and plus it was kind of proven in this. But does smoking in theory affect medication? We don't know, okay? And there's no evidence of it. So the smoking only affected the outcome, so then we say, oh, that's effect modulation. Now, think of it like this, in a con-founder, the most common one they like to use when they do smoking is think of, think of they did a study that said, if alcohol increases the risk of pancreatic cancer, okay? So if you did that study and then you teased out smoking, could you say that smoking affects alcohol? And you say, well, yeah, the chances are that if someone who's an alcohol, the chances are there's an increased level of smoking, so smoking does affect alcohol check, so that's the exposure. And the smoking also affect pancreatic cancer is like a higher association with that, and you're going to say yes, it is, because it'll be in the in the in the question, or it's even known. So if it affects just the outcome, you go with effect modulation. If it affects both the

exposure and the outcome, you'll go with confounding. Now in this situation, you know, it's a medication, so the smoking, or the non-smoking is getting affected by that, but it does affect the DVT, so it only affects one, and then we go with the effect modulation. And you know, just like on the step exam, you're going to put chance of choice A is going to be the one that is going to dangle that in front of you, hoping that you jump on it. But you got to understand, confounder of effects both. Effect modulation just affects the outcome, one piece of it, okay? This question says, a group of researchers perform a prospective study that reveals an association between G, so should done this one first. Between alcohol consumption and prostate cancer, the relative risk is associated with this finding is 1.7 with the p-value of 0.02, okay? So these are good. Then they split the group into smokers and non-smokers, and then re-examine the findings of alcohol and take in prostate cancer. The result was follows. Yeah, I should have had these backwards of them, and been a lot easier. So at first, when they do the study, relative risk is 1.7, so it's 1.7 times greater, okay? Unless just, let's just write this out, exposure. And you know, there's going to be the outcome or disease. And the exposure on this one is going to be alcohol, all right? And then the disease is pancreatic cancer. And we have to determine smoking. Okay? Now, is this smoking going to be an effect modulation? We're only effects the outcome or is it going to affect both? And of course, we just talked about this. So, and if we looked at it, again, is smoking effect the alcohol? Yes, is smoking effect the pancreatic cancer? Yes, because when they really teased this out, you could see that it went from a 1.7 with a nice p-value to total kind of garbage on the p-value. And then, of course, these relative risk change

significantly. So, as I say, initially, the data looked good. Okay? It looked good up here, until they teased out smoking. And then everything went south. But once again, if this is associated with both the exposure and the outcome, you got to go with confounding. I think when I typed these up, my regio intent was to have this question go first, because it was almost like a giveaway thus far. Now, if you see the word confounding, and that's what your choice is, because it affects both exposure and the outcome, they might do a follow-up question saying, how do you compensate for that? Or how do you associate that? And that's where you're going to say, you're going to do stratified analysis. Okay? And just remember confounding is a bias. Okay? It's a bias. You're going to stratify analysis. Again, confounding affects both exposure and the outcome. Effect modulation just affects one of them. Okay? Now, this one says, plasma vitamin D25 evaluation. Levels are measured in patients with coronary artery disease. The mean vitamin D levels of this group was measured at 19. Here's normal with a standard deviation of 1.1. In a separate group of patients with similar demographics, the mean vitamin D levels are 17.5. Okay? So they're talking mean here, mean here. With a standard deviation of 1.4, which of the following tests should be used to compare the mean of vitamin D levels? Okay? So this is just a basic definition because you're looking at your choice. You've got analysis of variants or a Nova, a two sample T test. But analysis was herbivisant or chisquared. To get through this step exams, really all you have to do on this one is no couple little facts. In this question, they're comparing two sets, right? This set and that set. It's the mean. So they're comparing two means. So if you're going to compare two means as simple as it looks, it's going to be a two sample T test.

You might see some that says a Z test, but we're not good. That's got more global per se. So a two sample T test compares two means. Now that's what this question says. But you got to know the definition of these other guys. If you were comparing two or more or always just think of three. Okay? Three or more. Sorry, not two or more. It's three or more. You're going to think analysis of variants or a Nova. How many syllables are in this? A Nova. Three. So three syllables. So you got at least three means to compare. If you're comparing just two means, you're going to go two sample T test. And then if you're comparing three or more means, you're going to go analysis of variants. You know meta analysis is just that when you're pulling more data. Okay? You're pulling data. And what's that going to do? The more data that you pull into a study, that's going to increase what? It's going to increase power. Okay? So meta analysis and they like to use this as an answer choice a lot when they do, I've seen it multiple times. Those drug ads because they're always pulling in extra data or something like something to that effect. And meta analysis somehow built in there. You know, observer bias is just more of a distractor in this type of question. That's really just a cognitive bias. You know, by the observer that might influence the participants. Now, chi-square, think of it like this. Chise-square, intimacy of box and going with chi-square. Okay? Yeah, this basically measures, you know, proportional data. Okay? Proportional data. Okay? You know, categorical or proportional. So any long story short, if I see a box and they tell me to compare stuff in there, I'm going with chi-square. If I compare two means, I'm going to two sample test. If I compare three or more means, I'm going to analysis of variants. If they talk about pulling data, I'm going with meta analysis. Now, in this

one, a study was performed to assess the association between light. Yeah, actually, say light box therapy and vitamin D levels in patients who underwent treatment for major depression with the ACT. The data from the study was given below. Okay? We have light box therapy, no light box therapy, normal vitamin D levels. We're in 45 people with light box, 35 with no, low vitamin D, low vitamin D with light box therapy. I probably should have, you know, if I would have made this more accurate, I should probably should have went like that. Right? Because you would think the light box, if they light box therapy, you should have more people with low, I mean, with those right. Light box therapy should have normal. So yeah, they'd be a lower number here. So I had to write them in the first time. So anyways, it's a chart. The following is the best statistical method to assess the association between light therapy and vitamin D levels. So again, if I see a two by two table and they ask me to compare that data and, you know, they're not going to add, if they're not asking you about sensitivity of these specific obesity, probably all that kind of stuff, because that's easy, right? If my answer choices are these, and I see a box, I'm going with triscord. Okay? Now, there might be a fancier way to do this, but if you just want to get through step one, step two, step three, this will work. Analysis of variance. What does that remember? Three or more means to compare. Two sample T test, you're comparing two means. Now, you know, you also have to have, you know, just as a technical piece, you know, you got an as simple size, they got to give you standard deviation, any ad or portion you two means. So anyways, analysis of variance, three means, two sample T test, two means. Meta analysis, fooled data. Okay? Increases power. Chisequare test, any time I see a box, it's categoric category.

You're comparing categories, or I was like, think, proportion data. Okay? But if I see a two by two table or a box, I'm going with Chisequare. All right. This one says, a prospective cohort study was performed to assess the role of daily vitamin D intake and occurrence of kidney failure. The study revealed a seven year relative risk of 1.2 for people who consumed 500 international units compared to those who do not. Okay? So that's saying there's 1.2 greater times chance of kidney failure if you take a bunch of vitamin D or a 500 units daily of vitamin D. The 95% confidence interval was from 1.05 to 1.4. Now, without me telling you anything, is that good or is that not good? This is where you're going to tell me it's good because why? It doesn't cross 1.0. And I can't stress that concept enough. If at any point that you see a confidence interval that crosses 0, you're going to say the data is kind of void. And I'm telling you, that's how you're going to get through those drug ads because you're going to give you a lot of this kind of stuff. And you're automatically going to say there's no association because at some point, it was equal. The exposure. Yeah, there was no difference between exposure and non-exposure. So I'm going to say that one that on the results. So just the fact that my confidence interval stays above that, then I know, hey, this is a good study. It's got a relative risk of 1.2, meaning it's a 1.2 times greater chance. And my confidence interval looks good, which of the following p-value is as most consistent with this. Okay, so if everything checks out, which of these is a good p-value. Now, you know that a good p-value must be what? It must be less than 0.05. Kind of a given. So which is the one that we can use here? I don't see it as 0.05. Well, we said it's less than that. So you got to go with answer choice E. Okay? And remember, this is just,

you know, when we did the whole null hypothesis, alpha, error, all sent in, it's p-value, and we want it less than 0.05. Okay? Very important. Because we want this low, low, low, low, because that's saying, what is the possibility of something happening by chance? And we don't like that. We want that really small. Now, if you notice on all these questions, I use vitamin D. You know, it doesn't matter what it used. The point to learn here is the concept behind all these questions is to understand what's the con- what's a con- founder versus effect modulation. You know, what are the differences between those tests analysis of variance? Two sample T tests, two means meta-analysis, pull data, and try squares when I see a table. Okay? And it's all you need to know, and that'll get you through the exam. This question says, you are asked to evaluate the efficacy of a hair loss medication that is being marketed as low one. There results of your randomized control trial at seven years as shown below. I got a little box here. So okay, I'm not seeing any of this try square stuff. Good. What is the number to be treated to prevent one person with endogenic LAPI show with this new medication? Okay. And this was kind of requested by a person who left it in the comment box, and who just wanted me to work another problem with the number needed to be treated. And I could have just said that I should have said what is the, you know, how many, you know, if I were to treat people, I could have just gave the guess definition for number need to be treated long story short, but I was nice and put it like that. So we know that the number needed to treat is one over the absolute risk reduction, absolute value. So absolute risk reduction is the vent rate control minus the event rate with the treatment, or vice versa, it doesn't matter to me because it's an absolute value. So this is the

treated group with the medication, this is treated with placebo. So hair loss, I probably should have swapped that right because we want to, how many to prevent. So it should have been no hair loss. I should have done that. Let's do it like that. We'll put no hair loss on top and hair loss on the bottom. No hair loss. Twenty people had no hair loss, so vice kind of pretty bad data then. What we could do it the other way, but I don't, I think the answer choices were, but I made this for for that. So if no hair loss was twenty and hair loss was still 80 on this one, but if they were treated with placebo, no hair loss was 40, so it actually, you know, it would have done better. This is pretty pretty messed up data, but for the most, just roll with it for right now, is if you're in the treatment group, you would say it's 20 out of how many people had no hair loss, so they had a success. Twenty out of how many, it's not 80 because how many people total, 20 had no hair loss, 80 had hair loss, so there's a hundred people in this group, and 20 had no hair loss, so really it's 20 out of a hundred. Now in this one, it's going to be, no hair loss was 40 out of what, not 40 out of 60, but how many people were in this group that were treated with placebo, 40 plus 60, so it's a hundred, and you know how to know it is 40 of them had no hair loss. Okay, so it's 40 out of a hundred. Don't make them a stake of putting the 80 or the 60 in a denominator, it was out of a hundred. Okay, now, and again, don't worry about the technicalities of this drug. So with that being said, I keep my denominator, it's a hundred, 20 minus 40 or 40 minus 20, it doesn't matter, absolutely, you're going to get that, which is basically one fifth. Okay, so when I do all the simple math, however you learn it, I would get a number needed to treat a five. Okay, again, the key with this number needed to treat,

you gotta know the formula. Okay, is one of the absolute reduction, and make sure it's an absolute value. And then all you gotta do is find the treatment versus non-treatment group, and then make it a fraction. Okay, make it a fraction. Don't fall for the mistake of putting it over just how many over the other group. It's how many are in the total of that section. And then if you just saw, you'll get a nice little number like that. Okay, so again, guys, I'd say you have to know that the perfect takeaway of this is again, confounding versus effect modulation, and you have exposure outcome, and if it just affects the outcome, and I see it then we're going to go with the effect modulation. And I'll be honest with you a lot of the questions that I see, and it doesn't, the rule doesn't always stand, but if usually the exposure is some type of drug or medication, I would see it lean toward effect modulation. Now confounding, that's usually like you're smoking, drinking, it can all be, I've seen examples of like when they do a study of birth order and birth order with down syndrome and stuff like that, and then they have to take in consideration the age and things like that. So multiple different examples, but if it affects both the exposure and the outcome, you're going confounding. If it just affects the outcome, you're going with the effect modulation, analysis of variants, you're going to say three means or more, and if it's a two sample T test, you're going to go two means comparison, and meta analysis, you're going to say pulled data, increases power, and then chi square, you are going to be thinking if I see a box of table, anything that says you got to compare proportions. Okay, hope to see you guys. Thank you.