Skip to content

Episode Notes

Randy Neil Biostats 10 — Quick review
Runtime: 26 min · Published: 2020-10-21 · Content type: Core Concepts
USMLE purpose: Rapidly rehearse the highest-yield calculation patterns before test day.

Source / episode info

FieldDetails
EpisodeRandy Neil Biostats 10
TopicBiostats quick review
Runtime26 min
Published2020-10-21
Sourcehttps://www.youtube.com/watch?v=oR06YvSibTA

One-liner

This is a speed-run of the formulas and visual patterns: standard error, confidence intervals, specificity, correlation, cohort/case-control, cutoffs, accuracy/precision, power, RR/OR, and standard deviation.

Quick-review checklist

TopicExam shortcutTrap to avoid
Standard errorIncreasing sample size lowers SEMDo not confuse SEM with SD
Confidence intervalMore sample size → tighter CIFor RR/OR, crossing 1 weakens significance
CorrelationDirection = sign; tightness = magnitudePositive/negative depends on slope
Case-controlOdds ratioStarts with outcome/disease
CohortRelative riskStarts with exposure
Cutoff movementOnly FP/FN regions changeDo not focus on TP/TN regions
Power1 − βLower n → lower power → more type II error
Normal distribution68–95–99.7Remember to split tails

Exam pattern recognition

  • “More participants added” → confidence interval narrows; SEM decreases.
  • “Older compared with younger” → the first named group goes on top.
  • “Greater than 240” with mean/SD → calculate tail, then split tails if needed.
  • “Instrument readings tightly grouped but far from gold standard” → precise, not accurate.

Common trap

⚠️ Trap: Memorizing a single RR/OR table orientation.
Fix: Put whatever group is named first in the numerator, then apply OR vs RR structure.

Step-focused notes and complete transcript for Randy Neil Biostats video 10.

Video: https://www.youtube.com/watch?v=oR06YvSibTA

Exam-Ready Notes

Organized Notes

High-Yield Biostatistics

Study Designs:

  • Case-control vs cohort
  • Randomized controlled trials
  • Cross-sectional studies
  • Ecological studies

Statistical Measures:

  • Relative risk
  • Odds ratio
  • Hazard ratio
  • Number needed to treat

Quick Question Strategies

Pattern Recognition:

  • Identify question type quickly
  • Choose appropriate study design
  • Select correct statistical test
  • Interpret results accurately

Time Management:

  • Read efficiently
  • Eliminate wrong answers
  • Focus on key concepts
  • Use systematic approach

Essential Formulas

Relative Risk:

  • RR = Incidence in exposed / Incidence in unexposed
  • RR > 1: Increased risk
  • RR < 1: Decreased risk
  • RR = 1: No difference

Odds Ratio:

  • OR = Odds in exposed / Odds in unexposed
  • Approximates RR when disease is rare
  • Used in case-control studies

Board Exam Buzzwords

Classic USMLE Phrases & Keywords:

| Buzzword | Associated Concept | Clinical Context |

|----------|-------------------|------------------|

| Odds ratio | Case-control study | Retrospective analysis |

| Relative risk | Cohort study | Prospective analysis |

| NNT | Treatment benefit | Clinical significance |

| p-value | Statistical significance | Hypothesis testing |

Differential Diagnosis Tables

Study Design Selection

| Question Type | Best Study | Key Measure |

|---------------|------------|-------------|

| Rare disease | Case-control | Odds ratio |

| Common exposure | Cohort | Relative risk |

| Treatment efficacy | RCT | Relative risk |

| Disease prevalence | Cross-sectional | Prevalence |

High-Yield Facts

  • Case-control: Odds ratio, retrospective
  • Cohort: Relative risk, prospective
  • RCT: Gold standard, randomization
  • p-value: Statistical significance threshold
  • Confidence interval: Precision of estimate

Clinical Pearls

  • Study design: Match to research question
  • Statistical tests: Choose based on data type
  • Interpretation: Clinical vs statistical significance
  • Time management: Quick pattern recognition
  • Practice: Repetition improves speed

Practice Questions

Question 1:

Which study design is most appropriate for investigating a rare disease?

A) Case-control

B) Cohort

C) Randomized trial

D) Cross-sectional

Answer: A) Case-control

Explanation: Case-control studies are efficient for rare diseases as they start with cases and work backwards.

Quick Reference Summary

| Topic | Key Point | USMLE Buzzword |

|-------|-----------|----------------|

| Case-control | Odds ratio | Retrospective |

| Cohort | Relative risk | Prospective |

| RCT | Randomization | Gold standard |

| p-value | Significance | Hypothesis testing |

Full Transcript

Alright guys in this video it's a quick review of the Biostats questions. Most of these were taken from the original video that was done five years ago. I think right before I went into residency and it's basically I've already got to worked out. It's mainly meant so that you can see the problem and actually just work it in your head more so then getting a blank problem and trying to figure out where do I start to pick up speed as you're studying through this. If you already had seen these before it's just a matter of being sure you have the have the concept down and that's what I'm trying to accomplish with this is just to go fast and to review and that's pretty much how you can study pretty much any topic but especially math and I'll save you a lot of time. The first question is a new one. It was sent in by one of the viewers per se and so that one is something you've seen for the first time with me but other than that they're just review questions mainly from the old videos just set a very up tempo pace. So hopefully I'd like to be you. Alright guys so we got one new problem in the rest you're going to be review very quick review but the first one says which of the following statistical changes would most likely be most likely if more as a matter of children were included in the study. It gets all of the standard error, the mean, upper confidence interval and lower confidence interval. So our main end. So alright so you got to understand what this is basically but you're going to see a lot of this kind of stuff on this step exam. It says a recent study addresses the role of air pollution and asthma development. 200 children diagnosed with asthma and 400 children without asthma are asked a series of questions regarding their homes. An air pollution index ranging from 0 to 10 is then calculated based on each child's response. The mean air pollution index for

children with asthma is calculated at 4.5 with a 95% confidence interval between 3.2 and 5.8. Which of the following statistical changes would be most likely if more as a matter of children were included in the study. Okay so you got to know a couple things here and we in confidence intervals key. The key with the confidence interval is that it's like saying look when they got all the data most of that data, most of that data fell within this range. The key with this is that you never want to use and we talked about this on the drug add is that we never want to use any confidence interval that crosses zero. Any time you see something like I don't know negative 1 to 2.5 across zero right and so it's no good. But this is good. Anyways the standard air the mean the upper confidence interval is going to be this guy and the lower confidence interval points going to be that guy. But the change in this is what happens to this confidence interval where things should fall within that range if you put more asmatic children in this study or you can increase the study with the asmatic children. Well that means this air pollution index for children as well calculated would be it seems to be tighter right because it would be more people more asmo kids you're going to get more dots in the same spot that are going to make that confidence interval a lot tighter. Okay because you get more the data should be more concise. But what's this whole standard air the mean stuff? Well you got to pray you don't get this one stuff one but that's just basically it's a formula standard air the mean which is like center deviation. I guess I'm using that term over end. Now the end is the key here and is the number. Okay now a standard air the mean it kind of implies that the air kind of goes away as the sample size goes up. Okay so if you put more people in the study what happens the bottom number

the denominator gets bigger and if the denominator gets bigger then the whole number gets smaller. So if you're going to put more kids in this study the standard air the mean and you have to memorize this. I mean this is one of those really obscure ones. This guy is going to get smaller so it could be this one this one or that one. Okay but what we say about the confidence interval that if you put more people in the study that confidence interval should get a little bit tighter. So the upper confidence interval should go down. Okay and the lower confidence interval should go up right because it should go get to a higher number over here and that should go up. So really the only answer that has all of them is going to be answer choice B. And the key with this just knowing this and I told you that when I'm especially when we did the formulas for pharmacology I don't get lost in trying to plug and chug. It's all about can I put in do I understand what happens if I increase this guy what happens to that guy. That's really where you get to put your money. Okay the confidence interval will get tighter and standard air the mean goes down. Now when it comes to you get to fly through these this is how I would study for early time. Any topic if you got the questions but especially by the stats. You just got to understand the concept don't work problems when you study. You work on one time you write them out and then just tell yourself do you understand what you did here. Specificity why no sensitivity goes down specificity goes up. So it should be 180 I put that on top because I circled it and it goes up. So whatever boxes I combine 180 plus 20. So if you 180 here 180 plus 20 on the bottom and that goes and it should be 90. 90% I mean what it and sort of that so that's going to be the specificity but make sure you understand sensitivity specificity, pot of dye, negative value

and how you do that. So again you just look at this and then mentally go through it you don't have to work the problem each time. It's a lot quicker this way. The scatter you see one that has a scatter plot you're looking for the correlations right the correlation of age of income to closest which is the following. So the below scatter diagram shows correlation income in age of 90 states and you can see it goes in this direction. Well it goes in that direction if you remember from middle school math you know it's like if you go to the positive in the x direction it's a positive number and it ups positive. So positive and positive this is going to be in the positive direction. If it went down that way of course it would be in the negative slope. Which is you know negative and even though this is a positive direction negative and positive is going to be a negative. Okay it's like this is something bad happened to a good person that's a bad thing. It's something good happened to a good person that's a good thing. And I'll tell you guys you get the positive number now. This stuff is how tightly these dots are correlated to that line right. So you know that if if those dots are basically right on that line then that's going to be like a line. Okay positive one but these are just off that not too bad not not not upscurly off right they're not big time off. They're just a little bit off so I say we it's called like most like most nicely with .9 okay but understand the directions for those. You know when you look at this problem you say okay they're going to ask about case control and cohort. And so again five year studies plan to assess the incidence in the allergy of respiratory disease 600. In 600 individuals greater than 50 years age is study consists of two groups. One group cares for pet dog and seven group does not care for animals in the household. At the onset of

respiratory symptoms cultures and psychological studies will be performed which the following best describes the study. What was the key to this question? Is it at the beginning of the study did anybody have any symptoms? No so both groups both groups did not have symptoms remember when it's the one in San A and no one group has it. One group has it one group does not when it's a cohort right because this A and O stands for case control the owner is cohort. So when at the beginning of the study nobody had disease so it's going to be a cohort. Now they and what is associated cohort the word relative risk. If it was case control one group has it one group doesn't as associated to word odds ratio so I say case control odds ratio one number over. One number we get that later cohort study as it oh so sort of relative billet risk one number over to. This question a study designed to diagnose prostate cancers being evaluated the sensitivity of the test is seventy percent. The specificity is ninety in the studies there's a hundred patients who truly have utilized 200 to do not. How many faults negatives now they could change that anyways but you know to any of the any of the four types. You know true positive false positive true negative so. Always draw your box when you see these always draw your box reality test you got a label it positive negative. It said remember there's a hundred patients who truly have UTI so in reality a hundred do. And then in 200 truly do not. But if we knew that the sensitivity was seventy. Sensitivity is only going to be on this side right sensitivity is this box going down and so if that box going down. Is this we know there's a hundred people in that category so a hundred people would be this guy plus this guy because there's a hundred total so he goes on bottom. We know it's seventy percent we don't like percent we like decibels so now we can

calculate the true positive right. And if I know this guy is seventy the whole thing is a hundred that I know this guy must be thirty okay and doesn't matter what they asked on this one you could easily. Easily calculate each one so again when you study this little you look at it and say do I understand the concept of trying to teach me. I know need to work it just kind of mentally go through this. This is the question that says when they move the marker whether it goes up or the marker goes down. By a marker with this people 500 volunteers 120 patients. Changing the cut off value of the biomarker from point c to point A so we're going from c to A or we're going down. And so what we say we take that middle line we take the middle line and we smash it to the left okay. Now you have to be able to label this right true positive true negative those are always on the outside do we care about them you know we don't. So but the inside we do care about that's all we care about when it comes to this move the marker thing so. Everybody keeps her last name so on the right side it's it's the true positive so this will be a false positive on the left side they were negative so this one's going to be a keeps this last name false negative. All I care about is whether those things get changed and since I'm going from c to A I take this line I move to the left I smash it. So what's going to happen in this guy he goes he goes um he goes down and what happens this guy he goes up so now I take that information and I say you know I work the box sensitivity well sensitivity I know goes down false negative went down and a false negative goes down the whole number goes up so no it's not lower it would be higher more true negatives. No we don't remember we don't care about the outside so that nothing to do with it more false negatives no we said false negatives went down in this one. Higher

negative predictive value well negative predictive value will be this guy going that way so that's true negative over a negative plus false negative. And if the false negative went down the negative pretty value went up okay so again snapshot that. Very good. The following graph shows a distribution values of a healthy group of a group of healthy and disease people points a through e blah blah blah blah blah. What cut off would determine is sensitivity of 100 percent well if we know sensitivity is this formula right based on our box. This formula to get this guy 100 percent meaning if I had a 10 over 10's 100 percent right so I need these to be the same so I need to get rid of this guy I need to make him what I need to make him zero. So where does this guy become zero well if I look at this where can I smash the false negative into oblivion into zero I would take it and go to the left. So where I make the false negative zero to make my sensitivity my sensitivity 100 percent right to point see. Now they said specificity and don't memorize this okay and then you say oh I just memorized it but understand the concept in case I messed with you a little bit. Where would I make if to make it a specificity. It would be the false positive has to be zero and I would move it this way to point e okay where he would be zero. Let it in this case sensitivity 100 percent see. Okay what snapshot that good. Sorry and then this one what will happen same concept but in this one they're saying and what it what it what it what will happen to sensitivity and specificity of a test when the markers are moved from the solid. That's outside to the dash one so again all we care about is just that center piece right all we care about him. So we label it true positive true negative we don't care about those guys but we care about the center this center piece is false positive and fall negative

but when we move from the solid to this to this one what happens. That area in here got. It got from here to here got smaller so what happens if false positive area smaller what happens if false negative area got smaller so these guys got smaller. Well if this guy gets smaller sensitivity if the denominator gets smaller sensitivity gets bigger. If to true negative get smaller right the specificity gets bigger and that's why it's transfer choice. Okay okay again snapshot mentally play that through your head. New instrument purchased by the hospital it checks here and levels of X the published value for the standard is 40 the technology runs the the test on patient getting the reading of this. What can be concluded well if you got to you got to know where you're shooting so to get the gold standard or the bull's eye is 40 but then you're getting 70 68672775. You're not hitting the bull's eye but you're being pretty you know you're you're over there hitting all close close marks together right. So if you hit the bull's eye you'd be considered accurate okay to hit 40 you're accurate but they're not accurate. But they're pretty consistent over here you know getting all these things so they're like this right so in this situation with this data they were. But they were precise right they were precise but they were not accurate they never got close to bull's eye. Okay. I'm sure you kind of understand that this one says based on the data what is the relative risk for the development of prostate cancer and man who had no children compared to men who had children. The key with this and I can't stress this enough this is regular lot of the most questions from this is they say well I don't understand first date has it written a certain way and I'm saying you can't go by first day because that only works. In one way if you understand this and we read math problems talk to bottom

left to right when I say left to right it's whatever they mentioned first goes on top. And first of all relative risk right if I said you know case control is associated with odds ratio and that's one number over one number. If I said a cohort study that's also associated with relative risk and that's one number over two. Know the difference you know be able to recite that like the back of your hand. And so in this situation relative risk is one number over two but what goes on top that's the key what goes on top. It's whatever was said first in this situation. It's the development prostate cancer of man who had no children compared to many of your had children. So who goes on top? No children compared to man who had children over a yes. Okay, now they could have said had children versus no children then this then the had children would go on top. But in this situation it's no children it's no children so what are you going to go here no. So what goes on top 80 what goes on the denominator that 80 plus 920 because it's relative risk I had put two of them on there. Because on the bottom whatever I said second way had children that's 220 but I put one number or two numbers on bottom. Oh I put two numbers why because it's relative risk and it's 220 plus one two 80. The key with this it's whatever they said first goes on top okay. Which got to understand this piece here case control odds ratio one number over one number cohort relative risk one number over two. This one member it says new study show that the mean HDL level non diabetic was 42 and the mean. HDL of an diabetic was 35 the probability this was due to chance was 0.5 point zero five there's a 15% probability that conclude there's no difference in HDL measurement when you're actually was one. I like this question says what is the p value well p value I think of is something my chance well we they even tell it

all this. This point zero five okay and we want our p value to be less than 0.05 okay 0.05 or better. So anyways long story short p value of the study 0.05 percent piece cake. What is the power of the study well you got to know the null hypothesis box right it's the same type of box you put reality up top test up here h1 and the null h1 and the null. Top right box is the alpha error type one also knows p value well we knew that one I was 0.05 alright that's no brainer but we got to get they're asking for the power we need this box. So here's how they write the questions there's a 15% probability of concluding that there is no difference in each HDL measurement when there is one. So they're saying that the test is going to say the test would say that there's no association when a reality there is one and that fits right there there's a 15% chance of that so that's why you get the 0.15 in that box okay this whole sentence describes that box right there in words. So then you put him in there one minus beta power so then you go one minus 0.15 and get 0.85 and then that's your answer. This is really that if they ask any null hypothesis questions they can't ask any more than what you just learned here they're going to be the could ask you alpha error type one p value but it's out to give it to you obviously. And they're either going to ask you well I'm telling you they're going to ask you power and they're going to give you your you find beta and it's probably going to be something just like that there's only so many ways they can ask this. A study of 200 patients hospitalized with complaints related to pneumonia show that their serum cholesterol levels are normally distributed variable with a mean up to 110 standard deviation of 15 based on the study results have many patients would you. Expect to have cholesterol greater than 240 well it's greater than 240 is what

they're asking for right and that's the kicker it's greater than 240 so again if the mean was 210 you put that dead center they got to give you standard deviation of 15. And so you know one standard deviation 68 to is 95 and 399 is 7 you got to know these to at least okay you got to have that memorized unfortunately. So if two tens here and you went one standard deviation out that'll take you to one 95 because all that is same one it one center deviations 15 so you went down 15 and then you're going to add 15 this way so that's one standard deviation is from. One 95 to 225 so that's like saying 68% of the people fall within that range 68% of that well let's go out another standard deviation. Minus 15 this way at 15 that way get 180 and 240 and what's what's that saying that. That saying that 95% of all the people in the study fall in that range 95% of 200 well 95% of 200 it's a hundred 90 so they're saying a hundred and ninety people fall within this and there's our 240 because the question said how many would you expect a greater than 240. Now they didn't say how many are on the outsides in remaining they said how many is above 240 well if I got. A hundred ninety people in here. And there's 200 people in the study that's 10 people on account for there's 10 people outside. Five of them are down here and five of them are up there and that's why the answer is going to be. Five because how many was greater than 240. Good question. Okay. This one says. A test being connected to the tremendous older students have a lower US and less step one is compared to younger students a group of older students. A greater than 50 you took step one compared to a group younger students. He took step one data show them a little blah blah blah blah. Here's the box okay. What type of test is being conducted well. You know in this situation one group has a low test score and run group does

not in theory not a great word question but the end of the day if one group has it one group doesn't. It's going to be a case control okay and then our case control studies like if I know one group has it one group doesn't have it or whatever do I care what happens moving forward for the most part no. If I want to know with the answer this I got to look backwards and that's why case control also associated with odds ratio. Is retrospective okay cohort. I don't know nobody has it okay nobody would have it yet so I can look backwards to get some data and I can follow these people moving forward so it's retro and forward pro perspective as well okay and that's associated with relative risk when number over two. So in this situation one group had it the low test score one group did not one group has it one group does not case control they could have put what what what what word is associated with the type of test in this one and they could have put like. on odds ratio relative risk here and you would have jumped all over odds ratio because that's what's associated with. With this type of study case control. And then this one this goes back it's the same worded question but this one goes back what is the odds of older students having a lower exam score compared to younger students now think about that what goes on top. This is not this is a case control so it's odds ratio it's one number over one number. But which what goes on top and you better repeat this back to me because on top whatever I said first in the sentence what is the odds ratio older students. Having a lower exam score compared to this is where this is like your big division sign compared to younger students so who goes on top. older because on bottom younger. So the older here's my older group so happens to be the right box could have been the left one it happens to be the right one so I got to take this

data 60 it's odds ratio so it's one number over one number so it's 60 over 20. If this was a relative risk I'd put 60 over 60 plus 200. Okay 60 plus 200. Again I don't even repeat that because I'm not a student too fast or I'm not a student running a second ago. Case control odds ratio one number over one number so in this situation it's 60 over 200. If this was a relative risk it'd be one number over two and it would be 60 over 60 plus 200. Okay there's a difference. We're back to case control odds ratio one number over one number old on top. Young on bottom why because I said older people first compared to my division younger people on bottom because I said it second and then young less than 50. Younger is 40 over 160 and that's why we got to answer twice see okay but you don't memorize the formula you got to know whichever one which everyone they said first. So again guys this is just kind of a quick way to get through the by of stats and I say I look at these when I study this I look at them and say what was the concept oh standard error I get into that. I go through this best of specificity okay I got that I got that I got that there's no need to work these problems out over and over you just got to look at them and then you run through them. This is how I would study pretty much each night with questions I did know in any topic and I would just run through these I'd have a book and I would just kind of keep running through the book. Of all the questions that I didn't know so anyways guys I hope this was helpful. And yeah it's just kind of a re-visitation of the by of stats and this is the common stuff that you kind of see on your step exam. I try to upload these online it may take a while it always takes a while it seems but again we're here to help here to move forward and you know I think the day that we made this video people are starting to put their

residency applications in and get interviews so it's pretty exciting. I was happy to be part of that with some people that I'm working with on that so hope to see you all guys. Thank you.