Methodology

Every number on this site, and how it was computed.

Rezumab estimates things that matter: whether a signal will get you an interview, whether an away rotation moves a program, what your odds actually are. Estimates like that are only worth anything if you can audit them. This page is the audit trail. Every model below is described with its data source, its estimator, its sample size, and the specific thing it cannot tell you.

Applicant records
921,967
applicant x program rows
Programs
5,869
in the directory
Modeled
23
specialties with signal lift
Reported null
45%
of credential estimates

Last reviewed August 28, 2026.

Standards of evidence

Five rules the models are held to

These are not aspirations. They are enforced in the build scripts that produce the data files this site reads, and every one of them costs us a number we would otherwise get to display.

  1. 1Published sources only. Every input is either a published dataset, a peer-reviewed paper, or data that applicants and programs reported themselves. Nothing is simulated, imputed across years, or filled in to make a page look complete.
  2. 2Effect sizes with intervals, not p-values. Each estimate ships with a confidence interval. p-values are not used anywhere on this site. An interval tells you how much the number could move; a p-value only tells you whether someone should be surprised.
  3. 3Small samples get pulled toward the mean. A program with nine reports does not get to claim a 90 percent signal conversion rate. Estimates are shrunk toward the specialty average with a weight set by sample size, so thin cells regress instead of shouting.
  4. 4A null result is reported as a null result. Where a credential's interval crosses no effect, the site says "no measurable independent effect" and the number does not move. Across every specialty and stratum modeled, 145 of 322 credential estimates (45 percent) come back that way, and all of them are labeled.
  5. 5Observational data is never called causal. Applicants who signal a program differ from applicants who do not, in ways no dataset records. Every lift on this site is an association measured under those conditions. The word "causes" does not appear in any model output.
Provenance

Where the data comes from

Six source families feed everything on the site. Each one is used for the questions it can actually answer, and none of them is asked to carry a claim it cannot support.

SourceVintageWhat it answers
Applicant-reported outcome records2023 to 2026 cycles921,967 applicant x program rows, de-identified and reported by applicants after their cycle. One row per program an applicant touched, carrying the signal sent, the interview outcome, the match outcome, and the applicant's credentials. This is the only source that links a specific applicant to a specific program, so every per-program invite model is built on it.
NRMP Charting Outcomes2026 edition, plus pooled 2020 to 2024Population match rates by Step 2 score, contiguous ranks, and applicant characteristics. Drives the match-probability calculator. The 2026 book publishes only score marginals, so the joint score-by-ranks structure comes from the pooled earlier vintage.
Specialty match reportsMost recent published cyclePositions offered, applicants, and match rates per specialty. Sets the population baseline every personalized estimate is anchored to.
SF Match and AUPO reports2026 summary reportOphthalmology runs its own match outside the main algorithm, on its own timeline and signal cap, so it gets its own engine fit on its own published data.
Program directory and FREIDA recordsCurrent accreditation year5,869 accredited programs: identity, location, specialty, positions offered, and self-reported interview counts. Source of truth for which programs exist and for the per-program funnel where a program reports its own numbers.
Community away-rotation reportsRolling, resynced hourlyApplicant-maintained away-rotation records: where people applied, when they heard back, and what the decision was. 12 specialties have historical benchmarks; 7 resync from live sheets every hour.
Model 01

Match probability

/calculator
Inputs
Step 2 CK, applicant type, contiguous ranks, research output, honor society, school characteristics, and the specialty you are applying to.
Method
A score-conditional grid engine rather than a single fitted curve. Published match rates by Step 2 band give the level; the joint distribution of score and contiguous ranks from the pooled earlier vintage gives the structure. The engine puts current levels under pooled ratios: your anchor is the current bin marginal, and the ranks effect is applied on top as a logit-scale ratio, because ratios age better than absolute rates. A bi-monotone projection reconciles the two vintages so probability never decreases as your score or your rank list grows. Cell counts are carried through the build so the engine can shrink thin cells itself. Individual factors move the estimate through pre-shrunk logit-scale lifts, each capped so no single credential can dominate, and the final probability is clamped to a band around your applicant type’s base rate.
Limits
Population-level. It knows your numbers, not your letters, your interviews, or your rank list strategy. Below 30 matched applicants in a cell, adjustments are damped toward the base rate on purpose.
Model 02

The ophthalmology engine

/calculator

Ophthalmology does not run through the main match algorithm, so running it through the main model would be wrong in a way that is invisible to the user. It gets a separate engine.

Method
A logistic model fit on the published 2026 summary report for the specialty, with coefficients stored as data rather than hard-coded in the app. Signals enter at the published cap for the cycle, and the coefficient threshold reflects where the published data actually separates.
Limits
The top-program sub-model still rests on a 2014 paper because no newer top-tier breakdown has been published. It is left as-is rather than refit on data that does not exist, and it is the weakest estimate on the site.
Model 03

Signal lift

/explore/signal-strategy/match-planprogram pages

The central question of the modern application: does signaling this program actually change your odds of an interview there? For every program in 23 specialties, the model estimates the invite rate for applicants who signaled and for applicants who did not, and reports the difference.

Estimator
Lift is P(invite given signaled) minus P(invite given not signaled), computed per program and per stratum, where a stratum is applicant type crossed with Step 2 band. Each cell is then shrunk toward the specialty grand mean with weight n / (n + K), K = 20, and n is the smaller of the two arms. A program with three signaled reports ends up close to the specialty average; a program with two hundred barely moves. Where a stratum’s limiting arm falls below 50 rows, the estimate falls back to the specialty-wide target rather than pretending the cell is informative.
Era
Fit on the 2024 to 2026 cycles. The 2023 cycle is dropped because it predates tiered signaling in the specialties that now tier. The baseline invite rate is recency-weighted so it tracks the most recent cycle instead of being dragged down by an older, larger pool.
Definition
A row counts as signaled if it carries any signal of any kind, flat or gold or silver. This matters more than it sounds: the legacy flat flag reads zero for every specialty-year that switched to tiers, so a flat-only definition would silently drop essentially every real signaler in those specialties. Where a specialty tiers, gold and silver also get their own per-program cells.
Intervals
Reported on the raw, unshrunk lift, using a normal approximation for the difference of two proportions. The interval belongs to the evidence, not to the smoothed display value.
Caveats shipped inside the data file
  • Lift is observational; signaled applicants may differ in unmeasured ways (geographic ties, mentorship, prior aways). Treat as directional, not causal.
  • Programs with low N have wide CIs; shrinkage pulls them toward the mean.
  • Fit on the 2024-2026 signaling era; 2023 predates tiered signaling and is dropped. The unsignaled baseline is recency-weighted so it tracks the most recent cycle rather than a stale 2023-24-heavy pool.
  • Signal is the union of the flat and gold/silver flags. The flat-only column reads 0 for tiered specialty-years, so the union is required to count real signalers.
  • Rollup uses min(n_signaled, n_unsignaled) for the shrinkage weight.
  • n_total / n_invited cover ALL slug-resolved rows; rollup.n_signaled + rollup.n_unsignaled covers the stratification-eligible subset only (rows with classifiable applicant_type and known Step 2 band).
  • Gold/silver tier cells are fit on the tiered era only (IM: 2025-2026, its 2024 rows are flat and pool into the union signaled arm, not a tier).
Model 04

Signal, geography, and aways together

/explore/signal-strategy

A signal does not act alone. Applicants who signal a program are also more likely to have trained near it or rotated there, and if you attribute all of that to the signal you will overestimate what a signal buys a stranger.

Method
A raw eight-cell cross-tabulation of the invite rate across signal status, geographic connection, and away rotation, computed per specialty on the applicant-level records. No model, no smoothing: just the observed rate in each of the eight cells, so you can see the decomposition rather than trust a coefficient.
Guardrail
Each cell carries the smaller of its two arm sizes. When that number is thin, the interface flags the cell rather than rendering a confident-looking bar.
Model 05

Credential lift

/explore/signal-strategy

How much does one more first-author publication, or an honor society election, independently change your odds of an interview, holding the program and everything else fixed? This is the most heavily defended model on the site, because it is the one people most want to over-read.

Estimator
Ridge-penalized logistic regression of the interview outcome at the applicant-by-program row level, fit separately within the signaled and unsignaled strata. Ridge lambda = 25. The credential block is collinear by construction (publications track research, quartile tracks clerkship honors, all of them track Step 2), and the L2 penalty is what keeps each coefficient stable and conservatively shrunk instead of ping-ponging between correlated predictors.
Program fixed effects
One dummy per program with at least 30 rows in the stratum, so every credential coefficient is a within-program comparison rather than a comparison across programs of different selectivity. Thinner programs collapse into a single bucket. The fixed-effect block carries only a light penalty (1); it is a nuisance parameter, not an estimate of interest.
Centering
Credentials are mean-centered within specialty, so a coefficient reads as the effect of being one unit above that specialty’s norm, not above zero. Three publications means something different in dermatology than in family medicine.
Intervals
Cluster bootstrap. Rows from the same applicant are not independent, so each replicate resamples applicants rather than rows and refits the whole model. Intervals are the 2.5 and 97.5 percentiles across replicates. Replicate counts drop for thin strata, and that is reported rather than hidden.
Missing data
First-author publication counts are entirely absent from the 2023 and 2024 records, so that slope is estimated on 2025 and 2026 rows only, with no cross-year imputation. Its entry carries the years it was actually fit on.
Deliberate exclusion
Couples matching is excluded from the model because programs do not screen interview invitations on it. Its near-null association is computed and reported separately as a check, not folded into anyone’s estimate.
Calibration
Predicted against actual invite rate by predicted-probability decile, plus expected calibration error, computed for every fit.
The null rule

When a credential’s bootstrap interval crosses an odds ratio of 1.0, it is recorded as no measurable independent effect and the credential rail refuses to move your number. Across 23 specialties, 145 of 322 credential estimates land there. Noise is not dressed up as signal.

Caveats shipped inside the data file
  • Observational. Reported numbers are PARTIAL ASSOCIATIONS, not causal effects. A credential's coefficient is its association with invites holding the other modeled covariates fixed - it is not a controlled lift.
  • The signal split is itself contaminated: applicants signal programs they already fit, so the signaled stratum over-represents strong-fit pairings.
  • A credential whose odds-ratio 95% CI includes 1.0 has 'no measurable independent effect' in that cell - it is reported as such, not as a small positive number.
  • The academic-cluster credentials (quartile, clerkship honors, final-A) overlap Step 2 CK; once Step 2 is in the model their independent effect is expected to be small.
  • First-author pubs are fit on 2025-2026 rows only (the field is 100% missing 2023-2024); the pubs entry carries fit_years and n_rows_observed. All other credential slopes are pooled across 2023-2026 with cycle-year intercept dummies.
  • The signaled stratum pools every signal type (gold, silver, and the legacy pre-2024 flat untiered signal): gold/silver tiering only began in the 2024 cycle, so a cross-cycle gold stratum would mix real gold with legacy flat signals. The strata are kept binary (signaled vs not) to avoid that.
  • The signaled stratum is thinner than no-signal in most specialties; check the reliability band on each cell before quoting a number.
  • Couples Match is excluded from the model (not a credential programs screen invites on); see couples_match_check.
Model 06

Program funnel

program pages/prescreen

How many people applied, how many were interviewed, how many seats exist. Two paths, tried in order, and the page tells you which one produced the number.

Census path
Where a full cycle census is published for a program, the counts come straight from it: applicants, interviews offered, first-year positions, and signal composition. One recent cycle, no estimation.
Self-report path
For specialties outside the standard application platform, the program’s own reported interview history and position count are used instead, including the multi-year history where a program publishes one.
Neither
Programs with no census and no self-report are skipped rather than estimated. Their pages fall back to the applicant-level model and say so.
Model 07

Rotator advantage

program pages/my-fit

Whether rotating at a program changes your odds of an interview there. The honest answer has three different qualities depending on what a given program has published, so the model carries three states and labels which one you are looking at.

State A
The program or its specialty organization published both rotator and non-rotator interview counts directly. Highest fidelity, and rare.
State B
Rotator counts come from away-rotation reports; the non-rotator baseline is derived from that program’s own funnel, netting out the rotators from both numerator and denominator.
State C
The program has rotator counts but no usable program-specific baseline, so the comparison runs against the specialty-wide non-rotator interview rate. Directional only, and marked as such.
Model 08

Away rotations

/away-tracker/explore/away-rotations

Away-rotation decisions are not published by anyone. The only way this data exists is that applicants maintain it themselves, which makes provenance and hygiene the entire methodology.

Benchmarks
Historical acceptance rates and median days-to-decision per program, built from cycle-spanning survey exports. Site names are normalized to canonical program names by fuzzy match, and rows that do not resolve to a known program are dropped rather than guessed at. 12 specialties have benchmark files.
Live sync
7 specialties resync from community spreadsheets every hour. The parser stops at an explicit end-of-cycle marker so a prior cycle cannot leak into the current one, and each row is fingerprinted so a status change updates the existing record instead of creating a duplicate.
Hygiene
Rows whose source record disappears from the sheet are flagged and excluded from every aggregate. Rows dated to a prior cycle are re-filed. Community aggregates never include flagged rows.
Limits
Self-selected. People who get accepted report more reliably than people who do not, which biases raw acceptance rates upward. Treat program-to-program comparisons as more meaningful than absolute levels.
Model 09

Match plan simulation

/match-plan/explore/signal-strategy

Your portfolio is not a list of independent coin flips, so the plan is simulated rather than summed.

Decomposition
For each program, P(match there) = P(interview there given your signal) x P(match given interviewed). The first term is the signal-lift model. The second has no published per-program value, so it is built from the specialty baseline conditional on having interviewed anywhere, scaled by how your Step 2 compares to that program’s invited cohort, through a logistic curve capped at [0.3, 1.8] to prevent implausible swings. Where a specialty publishes real per-program post-interview conversion, that is used instead.
Simulation
Monte Carlo down your rank order: draw the interview, then the match, and take the first program where both fire. This collapses to the closed-form probability of matching somewhere, but it also produces the destination distribution and the interview-count distribution, which are the numbers you actually want.
Unmodeled risk
Every simulation carries a probability that the cycle fails for reasons no model sees: a late submission, a letter that never arrives, a rank-list error, a family emergency. It is scaled by specialty competitiveness, from 1 percent in the most accessible specialties to 12 percent in the most competitive, which is why the site will never show you a 99 percent chance in dermatology.
Allocation
Signal swap suggestions are evaluated in log space, comparing the probability of matching somewhere under each candidate reallocation. The recommendation is whichever swap raises that number most, not whichever program is most prestigious.
Model 10

Cohort comparison

/share-data
Method
Applicant-level records are collapsed to one row per applicant per cycle, filtered to those who matched, and reduced to percentile cuts (10th, 25th, 50th, 75th, 90th) for Step 2, interviews received, and signals sent. Files carry both a per-cycle series across four cycles and a pooled distribution, plus a separate breakdown by applicant type where the cell is large enough to support one.
Reliability
Every percentile cut carries the cohort count behind it and a reliability band derived from that count. Where Step 2 is reported in bands rather than points, the bin midpoint is imputed and the imputation is disclosed.
Audit

Coverage by specialty

Model coverage is uneven, and pretending otherwise would be the easiest way to mislead you. This table is generated from the model files themselves at build time, so it cannot drift from what the site is actually serving.

SpecialtyProgramsSignaledUnsignaledCredential rowsTiers
Anesthesiology17914,94219,03863,587Gold / silver
Child Neurology803342,8885,443Gold / silver
Dermatology1406,6633,49423,839Gold / silver
Emergency Medicine2856,48437,83172,029Single
Family Medicine7425,53821,72853,249Single
General Surgery3529,42738,77577,177Single
Internal Medicine64729,01551,112135,226Gold / silver
Neurological Surgery1183,5404,93414,520Single
Neurology1903,06314,31426,974Single
Obstetrics and Gynecology29220,31335,35690,905Gold / silver
Ophthalmology1241,33918,95035,976Single
Orthopaedic Surgery20718,0628,98849,930Single
Otolaryngology - Head and Neck Surgery1307,9206,13928,194Single
Pathology-Anatomic and Clinical1401,1955,74710,503Single
Pediatrics2116,98632,71270,045Single
Physical Medicine and Rehabilitation1102,5657,88118,619Single
Plastic Surgery - Integrated881,4589,20519,147Single
Psychiatry3367,49139,03272,515Single
Radiation Oncology862052,9584,750Single
Radiology-Diagnostic1935,61921,10447,770Gold / silver
Thoracic Surgery - Integrated361241,4642,345Single
Transitional Year1813,7756,67915,930Single
Urology1517,0064,81321,511Single

Signaled and unsignaled are the stratification-eligible row counts behind each specialty’s lift model: rows with a classifiable applicant type and a known Step 2 band. Credential rows are the applicant-by-program rows in that specialty’s credential fit. Tiers records whether the invite model distinguishes gold from silver signals for that specialty.

Honesty

What these numbers cannot tell you

  • They are not causal. Applicants who signal, rotate, or apply to a given program differ from those who do not in ways no dataset records: mentorship, geography, family, advice they were given. Every lift here is an association measured in the presence of those differences.
  • The sample is self-selected. Applicant-reported data over-represents people who finished a cycle and chose to report it. Away-rotation data over-represents accepted rotations. Both skews push reported rates upward, and neither is correctable from within the data.
  • Programs change faster than data is published. A new program director, a changed signal policy, a class-size shift: none of it appears in a dataset until at least a cycle later. The most recent cycle is always the least well measured.
  • Nothing here models the parts that decide interviews. Letters, personal statements, interview performance, away-rotation impressions, and program-specific priorities are not in any of these datasets. They are also, by most program directors’ own accounts, where the decisions actually get made.
  • A probability is not a prediction about you. A 40 percent estimate does not mean you will match 40 percent of the way. It means that among applicants who look like you on the measured axes, roughly four in ten matched. You are one draw, and the measured axes are not all of you.
  • This is not medical or advisory guidance. Use it as one input alongside your program director, your specialty advisor, and mentors who know your file. Where this site and your advisor disagree, your advisor has information this site does not.
Data handling

What happens to what you enter

  • Your own rows are yours. Every per-user table is row-level-secured on your account id. Nobody else can read your tracker, worksheet, profile, or match plan, and no aggregate query returns a user id.
  • Community numbers are aggregates only. Away and signal community views are served by database functions that return counts and rates, never rows. Where an aggregate would be thin enough to identify someone, the interface flags it instead of rendering it.
  • Anonymous submissions stay write-only. Rotator reports, post-cycle data submissions, and feedback can be written by anyone and read back by no one, including through the public interface.
  • The most sensitive fields never leave your browser. Parts of the application worksheet are kept in local storage on your own device and are never written to the server.
  • Nothing is sold, ever. There is no paywall, no premium tier, and no data broker on the other end of this.
Maintenance

Corrections and versioning

Every model file carries a method version and the exact method string it was built with, and this page renders those strings directly rather than a summary of them. When a model changes, its version increments and this page changes with it. The signal-lift files on this build are method version 3; the credential-lift files are method version 2.

If a number here contradicts something you know to be true, that is worth an email. Program corrections, data errors, and methodology challenges all go to the same place and all get read.

rezumab.med@gmail.comSupport and FAQLast reviewed August 28, 2026