Skip to content

Episode Notes

USMLE purpose: Use the 2×2 table to calculate PPV, specificity, standard deviations, and correlation direction quickly.

How to use this page

  1. Start with the diagnostic-test box and make sure you can label TP, FP, FN, and TN.
  2. Review the core formulas until you can choose the right formula from the wording of the stem.
  3. Work through the PSA example without looking at the answer.
  4. Do the practice questions before opening the answer toggles.
  5. Use the transcript only if you want to verify the original video wording.
Episode metadata
FieldDetails
EpisodeRandy Neil Biostats 01
TopicBiostatistics extra review
Runtime9 min
Published2020-03-09
SourceOpen YouTube video

One-liner

This episode is a rapid drill on the highest-yield “box math” moves: build the 2×2 table from prevalence, sensitivity, and specificity; calculate PPV and specificity; apply standard deviation cutoffs; and interpret correlation direction.

Why this matters for USMLE

🎯 If the stem gives sensitivity, specificity, prevalence, or a 2×2 table , slow down and draw the box. These questions are usually free points if the table is labeled correctly.

Core formulas / rules

ConceptFormula / ruleExam trigger
PPVTP / (TP + FP)“Positive test; probability patient truly has disease?”
SpecificityTN / (TN + FP)“Correctly identifies disease-free patients”
Prevalence effectHigher prevalence → higher PPV; lower prevalence → lower PPVScreening test performance in different populations
Standard deviationMean ±1 SD = 68%; ±2 SD = 95%; ±3 SD = 99.7%Normally distributed value above/below a cutoff
Correlation coefficientDirection = sign; tightness = magnitudeScatter plot interpretation

High-yield visual: diagnostic-test box

Disease +Disease −
Test +True positiveFalse positive
Test −False negativeTrue negative

Episode-specific notes

1. Positive predictive value from prevalence, sensitivity, and specificity

  • The patient already has a positive PSA , so the stem is asking: “Given a positive test, what is the chance the patient truly has disease?”
  • That wording points to PPV , not sensitivity.
  • Build the 2×2 table from the information given:
    • Prevalence tells you how many people are disease positive.
    • Sensitivity fills the disease-positive column.
    • Specificity fills the disease-negative column.

2. Specificity calculation

  • Specificity is the ability to correctly identify people without disease.
  • Formula: TN / (TN + FP)
  • In a table, specificity lives in the disease-negative column .

3. Standard deviation cutoffs

  • If the mean is 200 lb and 1 SD is 60 lb, then:
    • 1 SD above = 260 lb
    • 2 SD above = 320 lb
  • For most USMLE questions, the calculation is simple once you identify the mean and SD.

4. Correlation coefficient interpretation

  • Downward trend = negative correlation .
  • Upward trend = positive correlation .
  • Tight points near the line = magnitude closer to 1.
  • Scattered points far from the line = magnitude closer to 0.

Step-by-step worked example: PSA positive predictive value

🧠 Try first: The patient’s PSA is positive. Sensitivity = 80%, specificity = 90%, prevalence = 10%. What is the probability the patient truly has prostate cancer?
StepMoveResult
1Prevalence = disease-positive proportionDisease + = 0.10; Disease − = 0.90
2Sensitivity = TP / all disease-positiveTP = 0.80 × 0.10 = 0.08
3Specificity = TN / all disease-negativeTN = 0.90 × 0.90 = 0.81
4False positives are the rest of disease-negative patientsFP = 0.90 − 0.81 = 0.09
5PPV = TP / (TP + FP)0.08 / (0.08 + 0.09) = 0.08 / 0.17 ≈ 47%
Takeaway: Once the test is already positive, the question is asking for predictive value , not sensitivity.

Exam pattern recognition

Stem wordingWhat to think
“PSA is positive; what is the chance the patient has cancer?”PPV
“Correctly identifies patients without disease”Specificity
“Two standard deviations above the mean”Mean + 2 × SD
“Dots trend downward on scatter plot”Negative correlation
“Dots are far from the best-fit line”Weak correlation; magnitude closer to 0

Common traps

⚠️ Trap: Using sensitivity when the question asks for PPV. Fix: If the patient’s test result is already known, think predictive value.
⚠️ Trap: Forgetting that PPV changes with prevalence. Fix: Use prevalence to fill the disease-positive and disease-negative columns first.
⚠️ Trap: Reading the whole long stem before identifying the ask. Fix: Read the final sentence first, then extract only the numbers needed.

Board Exam Buzzwords

BuzzwordWhat it should triggerClinical / exam context
Case-controlOdds ratioRetrospective study; starts with disease status
CohortRelative riskProspective or retrospective exposure-based study
Randomized trialIntervention / treatment effectExperimental design; strongest for causality
p-valueStatistical significanceProbability results occurred by chance if null is true

Study design selection

Study typeBest use caseKey feature
Case-controlRare diseaseRetrospective; starts with cases and controls
CohortCommon exposureFollows exposed vs unexposed groups
Randomized controlled trialTreatment efficacyRandomization supports causal inference
Cross-sectionalPrevalenceSingle time-point snapshot

Anki-style rapid recall

What is the positive predictive value formula?

PPV = TP / (TP + FP)

Use it when the test is already positive and the question asks whether the patient truly has disease.

What is the specificity formula?

Specificity = TN / (TN + FP)

It measures how well the test correctly identifies people without disease.

What happens to PPV when prevalence increases?

PPV increases.

When disease is more common, a positive test is more likely to be a true positive.

How do you interpret a negative correlation?

The variables move in opposite directions. As one increases, the other decreases.

What does a correlation coefficient near 0 mean?

The relationship is weak. The scatter plot points are far from the best-fit line.

Practice Questions

Question 1

A disease is rare. Researchers identify patients who already have the disease and compare them with controls to look backward for exposures. Which study design is this?

  • A) Case-control
  • B) Cohort
  • C) Randomized controlled trial
  • D) Cross-sectional
Reveal answer & explanation

Answer: A) Case-control

Case-control studies are efficient for rare diseases because they start with people who already have the disease and compare them with controls, then look backward for exposures.

Test-taking takeaway: Rare disease + looking backward = case-control → odds ratio.

Question 2

A patient’s screening test is positive. The question asks for the probability that the patient truly has the disease. Which metric is being tested?

  • A) Sensitivity
  • B) Specificity
  • C) Positive predictive value
  • D) Negative predictive value
Reveal answer & explanation

Answer: C) Positive predictive value

PPV asks: “Given a positive test, what is the probability of true disease?”

Formula: TP / (TP + FP)

Question 3

A test has 23 true negatives and 5 false positives. What is the specificity?

  • A) 23 / 28
  • B) 5 / 28
  • C) 23 / 23
  • D) 5 / 23
Reveal answer & explanation

Answer: A) 23 / 28

Specificity = TN / (TN + FP) = 23 / (23 + 5) = 23 / 28 ≈ 82%.

Question 4

A normal distribution has a mean of 200 lb and a standard deviation of 60 lb. What value is 2 standard deviations above the mean?

  • A) 260 lb
  • B) 300 lb
  • C) 320 lb
  • D) 360 lb
Reveal answer & explanation

Answer: C) 320 lb

Two standard deviations above the mean = 200 + (2 × 60) = 320 lb.

Question 5

A scatter plot slopes downward from left to right, and the points are widely spread around the line. Which correlation is most likely?

  • A) Strong positive correlation
  • B) Weak positive correlation
  • C) Strong negative correlation
  • D) Weak negative correlation
Reveal answer & explanation

Answer: D) Weak negative correlation

Downward slope = negative. Widely scattered points = weak relationship, so the coefficient is closer to 0 than to −1.

Quick Reference Summary

TopicKey pointUSMLE buzzword
PPVTP / (TP + FP)Positive test → true disease?
SpecificityTN / (TN + FP)Correctly identifies disease-free patients
Case-controlOdds ratioRare disease; retrospective
CohortRelative riskExposure groups followed over time
CorrelationSign = direction; magnitude = tightnessScatter plot

📌 Transcript note: The transcript is kept below for reference only. Use it if you want to verify the original video wording; the high-yield study content above is the main review tool.
Full Transcript

All right guys, so we're going to have four problems here and let's get started. It says 55-year-old man visits his primary care physician with the complaint of your neary frequency. Examination finds one centimeter nodule on his prostate gland. If the physician orders a prostate specific energy to your amtest, by common standards, the PSY level greater than four is considered abnormal. Using the standard test, using the standard, this test has a sensitivity of 80% and in specificity and non-early, recently published if epidemiological article found that this cross-sectional study, 10% of the men of this age have prostate cancer. The result of the PSA of the patient's PSA is 7. What is your best estimate as likely as this man actually has prostate cancer? They're saying that the guy actually has a positive PSA, so the test is positive and they're saying what is your best estimate that the guy actually has prostate cancer? They're asking positive predictive value. Anytime you see a problem with positive predictive value, you make for value or sensitivity specificity, I want you draw the box and I want you to label it. Reality, positive negative and the test, positive negative. Now we've just got to fill in the box. Now they're asking for positive predictive value. We know positive values, top left, going to the right, so we've got to figure just really these two boxes and then we're home free. What I'd recommend going to label these anything you want, ABC and D, and then what do we know? Well, we know the sensitivity is 80%. So actually this box going down would be 80% or 0.8, that would be a over a plus c. And then it says specificity is 0.90% or 0.9. Well, we know that is D going this way, so that would be D over D plus B. And then what else do we know? Because 10% of men of this age have prostate cancer, so in reality, 10% of the men already have it. So if 10% have

it, then how many don't? Well, that's just what's left, 0.9. So in theory, 0.1 equals a plus c, right? Because that's how many people have the disease at the entire column. So we would have 0.8 equals 0.1 bottom of, so it's a over 0.1. Now, let's just go ahead and solve for A, and then 0.8 times 0.1, which would give us a 0.08. So A equals 0.08. Now, we're halfway there, right? We just got to get what B is. And then we'll have our positive predictive formula, which is what? Right? A positive value is top left, which would be A over the two boxes going to the right, A plus B. Now, we know that 0.9 is going to be B plus D, right? It's the entire column. And so we put that there, and then that would give us D. And so we know that D is going to be 0.81. Well, if D is 0.81, then D plus B must equal 0.9. So that's got to be 0.09. So now we have everything we need, right? We'll have positive values A, which is 0.08 over A plus B, which is 0.08 plus 0.09. So now we have 0.08 over 0.17. And then we should do basically 0.08 divided by 0.17 with the decimal spot 2. That would give us right here. And then we'd have to say, let me see, A, 2.28, 20.30. I always get this mess. So it's 8. And 2, that'd be 6. I'm sorry. And then 2, 1, 0. And now we'll give 9. And 4, as you love it, right? So it gets this right about there. So we're looking at 0.47 or 47%. The key with this one is very long problem. But you have to know that you're looking for the positive value. And when you're looking for that draw your box. If you draw your box, they got to give you this stuff in the middle, just to fill it out. And then you should have it home free. All right. This one says, a 50 year old male presents the office for a routine health check. Managing his weight has been his focus. Has been his focus at this time. And the previous overall health. The doctor discusses his weight loss goals and

overall health benefit from weight loss including better blood pressure, management, blah, blah, blah, blah. The national average weight for males is between 15.59 is 90 kilos or 200 pounds with this inner deviation of 27. So one standard deviation equals 27 kilos. What would be the most likely expected value of his weight had two standard deviations? Well, this is like a no-brainer, right? If the mean or the average is 200 pounds, well, one standard deviation is just going to be an extra 60, right? So that's one standard deviation is going to be 260. And then, so two standard deviations is going to be 320. So two standard deviations is going to be 320. Okay? Pretty easy on that one. All right. This one. So again, long problem looks like they're getting the whole box thing going for, at least they were nice to draw it out. Specificity for breast-to-gdominations traditionally rather high among community practitioners. If a team of new researchers set forth a goal to increase the specificity and detection of breast cancer from the previously born national average is 74, based on the following results has a team achieved their goal. So they're saying that this is the old, say, based on what do we want to call it? They're saying, well, did they meet their goal getting better than that? Well, if they're telling us this is their results and we're looking for specificity, we know specificity is going to be the bottom, right? So we're going to go in up. So then we'll have 23 over 23 plus 5. So then we got 23 over 28 and then we go 23 divided by 28. We would get 0.8 for 16. Let's see, 16 and 622 and go down. That's only 16, 1, 6, and 0. That would give us a 2. So that would give us 0.82. So the new specificity is going to be 0.82, the old one has 0.74, so yeah, they increased it. So it's a yes, they increased it and they have achieved it with an increase of what's a

difference between 0.82 and 74. That's going to be 8% right? That's going to be an increase in the size of the old one. So it's going to be a specificity, a price per value or a negative value. So this one is just kind of a review from those old videos catching a lot of flag from it. So hopefully this is this will clarify what I was trying to get across. But all that kind of stuff, we're looking for the correlation coefficient essentially. In long story short, the correlation to the test scores, okay, the correlation between these test scores is most indicative of which one of these, okay? Well long story short, you know, you just want to draw a line that matches where these things trend, right? Well this one obviously trends in this direction, okay? And that's in the negative, okay? That's in the negative direction, right? That's the old high school middle school math. If it trended upward, you know, that would be in the positive direction. And if it goes in the down direction like this, left right, it's in the negative. Now the real question is, if the dots are way away from the line, then it's going to be more of a small number. Now when I say small number, these things are only going to go between essentially say say zero and one. So if they're far away from that line, it's going to be say like a zero point two, a point three, but if they're really tight right on that line, that's going to be a positive one, but it's going to be closer to one, okay? That's what you really have to know from this. My old point was look, it goes in that direction, it's positive, it goes down, it's going to be a negative, okay? Well this one obviously goes down, so we know it's going to be a negative number. But then all these dots right are they tight along the line? No, they're not tight along the line. They're really tight along this line, it would be a negative one, but these

guys are far away from the line. Therefore it's going to be a much like say small number, such as negative point two zero point two zero point. Okay, but the p is, it's up like this, it's going to be a positive number, and if they're tied on the lines, we close for one, if they're way apart from the line, you're going to give it a positive number that's going to be a lot smaller. Okay, same thing holds, it goes in that direction. Hope to self guys. Okay.