False discovery rate
In one sentence
The false discovery rate is the expected fraction of false findings among a procedure’s selected findings, with that fraction defined as zero when none are selected.
The intuition
Suppose a search returns a basket of candidates. You care how many selected candidates could be false finds, rather than only how many searches were performed. False discovery rate (FDR) concerns the basket produced by a defined procedure. It does not certify each item in it.
How it works
Imagine repeating the entire procedure on comparable data. In each repeat, divide false selections by all selections, using zero if there are no selections. FDR is the average of that fraction across repeats under the statistical model. The false labels are generally unknown in a real dataset.
An FDR-controlling procedure therefore targets an expected proportion. It is different from controlling the probability of even one false selection, called the family-wise error rate. The original Benjamini–Hochberg procedure establishes its control under independent test statistics; dependence conditions and the method used need attention. Benjamini and Hochberg, 1995.
For that procedure, sort the p-values. With m tests and target q, compare the p-value at rank i with (i/m) × q. Find the largest rank meeting its threshold, then select all findings up to that rank. This step-up rule is a teaching description, not a substitute for choosing a method appropriate to the data.
Why it matters in cancer
Gene-expression screens and mass spectrometry can make many comparisons or identifications. Retain the tested family, procedure and reporting level. A protein-level error estimate and a peptide-level estimate describe different sets. A controlled selection error still does not establish the biological importance of a change.
Worked example
Five fictional tests give sorted p-values of 0.001, 0.008, 0.026, 0.080 and 0.200. At q = 0.05, the rank thresholds are 0.010, 0.020, 0.030, 0.040 and 0.050. The largest passing rank is three, so the step-up procedure selects the first three findings.
You do not know which, if any, of those three are false. The calculation does not say each has a 5% error probability, or that exactly 5% of this particular basket is false. It describes the procedure's expected behavior when its assumptions hold.
Common confusions
- An FDR control target does not supply a patient's false-positive probability. Clinical positive predictive value needs its population and test-performance information.
- An FDR target of 5% is not a guarantee about a particular gene or patient.
- Screening more hypotheses and reporting only selected ones without the method obscures the error claim.
- Different adjusted p-values or q-value methods should not be treated as interchangeable labels without checking their definitions.
Try it
A researcher reports 20 selected genes at FDR 5%. Must exactly one be false?
Answer: No. The realized number is unknown. The target concerns an expected fraction across repetitions under the procedure's assumptions.
Related concepts
Sources and scope
Source check: October 9, 2026. Invented p-values; expert and learner review remain pending.
- Benjamini and Hochberg, 1995: the original step-up procedure.
- Benjamini and Yekutieli, 2001: dependence conditions.
Used in
Browse the concept index for related learning paths.