Skip to lesson
OncoGuideeducationDiana’s wiki
THE EDUCATION LIBRARY

Read the validation behind the claim

Validation is strongest when the evidence tests the exact claim on suitable data that were not used to build or tune the method. A promising chart is a starting point for questions about how the chart was made.

Before you start: Analytical validity, clinical validity and clinical utility identifies the question being tested. Trial phases and randomization explains what a comparison can establish.

Where this step sits

You have written the claim. Now inspect whether the study supports it. The final lesson will connect the evidence to a decision.

Follow the sample through the study

Imagine that the impressive chart is the last page of a story. To understand it, go back to the people and samples that entered the study.

Who was eligible? How were participants chosen? Which samples failed, and where did those failures go? Were researchers given the outcomes before choosing a threshold? Were patients receiving the same care as the population for whom the claim is now made?

These questions are practical. A test that performs well only when a perfect sample is available may struggle in the setting where you need it. A model tested on the same patients used to train it can reward memorization and tuning. Independent testing asks whether the result travels beyond those examples.

External validation means evaluating the method on data from outside its development dataset. Look for relevant differences in sites, sample handling, time period and patient care. “Independent” should describe the data and evaluation, rather than simply a new logo on the slide.

Match the evidence to the proposed use

Use a small evidence table while you read:

Study featureWhat to write downWhy it matters
Population and settingCancer type, stage, treatment context and selectionA study may not match the proposed patient group
Specimen and methodMaterial, handling, test version and quality criteriaA different input can change performance
Reference or comparatorWhat counts as truth, or what care strategy is comparedA weak reference can distort the apparent result
Outcome and timingWhat was measured and whenResponse, recurrence and survival answer different questions
Analysis planPrespecified rules, tuning and independent test dataSearching many choices can produce a flattering result
Missing informationFailed samples, dropout, unavailable outcomesThe visible chart may omit difficult cases
UncertaintyIntervals, sample size and remaining limitsA point estimate is not a promise for a person

A prespecified rule is written before seeing the relevant results. An exploratory analysis generates ideas after looking at data. Exploration is valuable, but the resulting claim needs confirmation that did not reuse the same search.

If the evidence begins in cells or animals, identify what the preclinical model actually tested. If it comes from patient care without assigned treatment, ask about confounding. Neither a useful laboratory result nor an outcome association automatically validates the proposed clinical action.

Worked example: two fictional validation slides

Slide A shows the performance of a risk score on the same patients used to choose its variables and threshold. Slide B evaluates a locked version on a separate hospital's patients, with a defined outcome and failed tests reported.

Slide B answers a more useful question about performance outside development. It still needs careful reading. Perhaps that hospital treated advanced disease while the current request concerns patients after surgery. Perhaps the score ranks patients well but overestimates their risk.

Ranking and risk accuracy are different. A method can put higher-risk patients above lower-risk patients while its reported probabilities are poorly calibrated. Calibration asks whether stated probabilities agree with observed outcomes in comparable groups. There is no single performance cutoff that makes every clinical use appropriate.

Even a well-calibrated risk score does not establish benefit from the treatment suggested in the report. The prognostic versus predictive distinction remains in place.

What can go wrong at this step

  • Successful samples become the whole denominator. Request the number attempted, the number reported and why others failed.
  • A chosen threshold travels silently. Ask whether it was set in advance, tuned here, and evaluated again independently.
  • A cross-disease dataset supports a same-disease claim. Keep the mismatch attached to the conclusion.
  • Laboratory status replaces clinical evidence. Certification and authorization answer another question.
  • A headline metric hides the decision. The cost of false positives and false negatives depends on what action follows.

Try it

A fictional model was tuned using outcomes from Hospital A. The developers then report “validation” after rerunning the unchanged Hospital A samples. What would you request?

Answer: Ask for performance on suitable data held out from every tuning choice, ideally from an independent setting. Also request the locked method, target population, outcome, failures and uncertainty. Repeat analysis of the development data can check reproducibility, but it does not provide independent validation.

Explain it back

Complete: “This study supports ____ in ____. It does not yet support ____.”

One possible answer: “It supports risk ranking in treated advanced Cancer X. It does not yet support changing postoperative treatment in early Cancer X.”

Takeaway

Keep the population, input, comparison, outcome and independence of the study attached to every performance claim.

Next: Connect evidence to a decision.

Sources and scope

Source check: October 9, 2026. The slides are fictional. This is a reading method, not a universal validation rulebook. Expert and learner review remain pending.