3556 min

Evidence-Based Neurophysiology

Diagnostic accuracy, prognosis literature, trial design, and AI

Learning objectives
01Interpret sensitivity, specificity, predictive values, and kappa, and identify the biases that inflate reported accuracy.
02Critically appraise prognostic EEG literature for spectrum, verification, incorporation, and outcome-blinding bias, including the self-fulfilling prophecy.
03Recognize the design challenges of clinical trials in epilepsy and seizure monitoring, including blinding, endpoint reliability, and the placebo response.
04Appraise AI-assisted EEG interpretation in 2026 and explain automation bias and how to guard against it.

Why metrics need appraisal, not just citation

Every claim about what an EEG finding means rests on a body of evidence, and that evidence is of wildly uneven quality. A number such as a sensitivity or a hazard ratio is only as trustworthy as the study design that produced it, and EEG research is unusually prone to a specific set of methodological traps, because the reference standard is often itself uncertain, the readers are human and frequently unblinded, and the populations are selected in ways that flatter the test. The competent neurophysiologist therefore reads the literature the way they read an EEG: with artifact as the default hypothesis and field logic as the discipline. The question is never simply what was the reported accuracy but what biases could have manufactured that number, and whether the study population resembles the patient in front of you closely enough for the estimate to transfer at all. This module is, in effect, the cognitive-bias module applied one level up: the same skepticism that treats a single-electrode transient as artifact until it proves it has a field treats a published statistic as inflated until the study proves it controlled for the relevant biases.

Diagnostic accuracy and its inflations

The core metrics of a diagnostic test are sensitivity, the probability of a positive test given disease, and specificity, the probability of a negative test given no disease. These are relatively stable properties of the test across populations, but they are not what a clinician needs at the bedside; that role belongs to the predictive values, which depend on prevalence and answer the actual question - given this result, what is the probability of disease? As the cognitive-bias module established with worked arithmetic, a positive interictal EEG has a far lower positive predictive value in a low-prevalence setting than in a high one, even though sensitivity and specificity are unchanged. Confusing the prevalence-independent metrics with the prevalence-dependent ones is among the commonest errors in citing EEG accuracy, and it is exactly the error that fuels overdiagnosis when a sensitivity quoted from a high-prevalence cohort is applied implicitly to a low-prevalence clinic.

Reported accuracy is, moreover, systematically inflated by a recognizable family of biases that the appraiser must hunt for deliberately. Spectrum bias arises when the study population is enriched with severe, classic cases and clearly normal controls; the test looks superb on this easy spectrum but degrades sharply on the ambiguous patients who actually generate referrals - the difference between distinguishing a textbook spike from a flat line and distinguishing a borderline sharp transient from a wicket wave. Verification or work-up bias occurs when only test-positive patients go on to receive the reference standard, which distorts both sensitivity and specificity because the test-negative patients never have their true disease status confirmed. Incorporation bias arises when the EEG result is itself part of the reference standard used to define the disease - rampant in epilepsy, where the clinical diagnosis may partly rest on the very EEG being evaluated, guaranteeing spuriously high apparent agreement by circular reasoning. And lack of blinding between the index test and the reference standard inflates accuracy whenever a human reader who already knows the answer interprets the trace, importing the expectation bias of the previous module directly into the study design. A reported sensitivity is meaningful only after each of these questions has been asked and answered.

BiasMechanismEffect on reported accuracy
Spectrum biasPopulation enriched with classic cases and clear controlsAccuracy overstated; degrades on real, ambiguous patients
Verification biasReference standard applied selectively to test-positivesSensitivity inflated, specificity distorted
Incorporation biasEEG is itself part of the reference standard defining diseaseSpuriously high apparent agreement by circular reasoning
Unblinded interpretationReader of the index test knows the reference resultAccuracy inflated by expectation bias

Interrater reliability and kappa

Because EEG interpretation is fundamentally a human judgment, interrater reliability - the degree to which independent readers agree - is a first-class measure of the evidence base, and often a sobering one. Raw percent agreement is misleading because two readers will agree a certain amount by chance alone, especially when one category is common; if ninety percent of records are normal, two readers who both call everything normal will agree ninety percent of the time while demonstrating no real skill. Cohen's kappa corrects for this chance agreement: it expresses agreement beyond chance as a fraction of the maximum possible beyond-chance agreement, running from 0 (chance only) to 1 (perfect), with negative values possible when agreement is worse than chance. Cohen's kappa is defined for exactly two raters; when three or more readers are compared, the analogous chance-corrected statistic is Fleiss' kappa, and multi-reader EEG studies typically report Fleiss' kappa or the mean of all pairwise Cohen's kappas. The values are most commonly glossed by the Landis and Koch bands, in ascending order, as slight (up to about 0.20), fair (about 0.21 to 0.40), moderate (about 0.41 to 0.60), substantial (about 0.61 to 0.80), and almost perfect (above about 0.80) - though these labels are an arbitrary convention, not a measured property, and different authors draw the lines slightly differently. The empirical reality, established by large multicenter studies, is that agreement is excellent for unambiguous findings - a clearly normal record, a frank evolving seizure - and frequently only fair to moderate for exactly the judgments that matter most and are hardest: whether a sharp transient is epileptiform, and where a pattern sits on the ictal-interictal continuum. In the best-known study of expert annotation of individual interictal discharges, chance-corrected agreement on whether a given transient was epileptiform sat at the fair-to-moderate border and was characterized by the investigators as only fair, even though agreement on whether an entire record contained any epileptiform activity was substantial - a pattern replicated in subsequent work. This is one of the strongest empirical motivations for the standardized, criterion-based terminology introduced earlier: explicit definitions measurably raise kappa relative to free-text impressions, converting hidden disagreement into something that can be seen and managed.

Key concept

Kappa, not raw percent agreement, is the honest measure of interrater reliability because it removes the agreement expected by chance. EEG kappa is typically high for clear findings and only fair to moderate for borderline epileptiform and ictal-interictal judgments - precisely where structured terminology and double reading earn their keep.

Appraising the prognostic literature

Prognostic EEG studies - does this pattern predict outcome - carry their own biases beyond those of diagnostic accuracy, and the most important is outcome ascertainment that is not blinded to the EEG. If the treating team sees a malignant-looking pattern and withdraws life-sustaining therapy, and the study then counts the resulting death as evidence that the pattern predicts death, the prophecy is self-fulfilling and proves nothing about biology. This self-fulfilling prophecy bias has demonstrably contaminated parts of the post-cardiac-arrest prognostication literature, where patterns such as a suppressed or burst-suppression background and certain periodic discharges were treated as predictors of poor outcome in cohorts where those same patterns may themselves have driven the decision to withdraw care. The concern is sufficiently well recognized that international resuscitation guidelines have for years called for blinding the treating team to the results of any predictor under formal investigation, precisely to break this loop. Recent evaluations of malignant EEG patterns continue to caution that unblinded withdrawal practices may overestimate the specificity of those patterns. Rigorous prognostic studies therefore either blind the treating team to the EEG or, at minimum, document and statistically account for withdrawal practices. Beyond blinding, the appraiser checks for an adequately defined and homogeneous inception cohort assembled at a common point in the disease course, sufficient and complete follow-up, blinded and pre-specified outcomes, and validation in a population separate from the one used to derive the predictor - because a rule that has only ever been tested on the data that generated it is a hypothesis, not a prognostic tool.

Watch out

Beware the self-fulfilling prophecy in EEG prognostication: if clinicians withdraw care because of an ominous-looking pattern, a study that then links that pattern to death has proven nothing about biology. Demand outcome assessment blinded to the EEG, or explicit accounting for withdrawal practices, before believing a prognostic claim.

Trial design in epilepsy and seizure monitoring

Clinical trials that use EEG or seizures as endpoints face design problems the appraiser must understand to read them honestly. Seizure counts are an imperfect endpoint: patient and caregiver diaries notoriously under-report, especially for nocturnal, brief, subtle, or amnestic events, so a trial relying on self-reported seizure frequency may be measuring reporting behavior as much as underlying biology - one of the central motivations for objective ambulatory and inpatient EEG monitoring as an outcome measure. Blinding is genuinely hard in antiseizure-drug trials, because the drugs produce recognizable side effects, such as sedation or dizziness, that can unblind both patient and clinician and thereby threaten the integrity of a placebo-controlled comparison. Regression to the mean and the placebo response are unusually large in epilepsy, because patients typically enroll when their seizures are at a peak and would have improved somewhat regardless of any intervention, so uncontrolled before-after comparisons systematically overstate benefit - the high placebo-response rates seen in epilepsy trials are a standing warning against uncontrolled designs. Where the EEG itself is the outcome, central blinded reading by reviewers unaware of treatment allocation is essential, precisely because of the interrater-reliability and expectation-bias problems established above; a trial whose endpoint is an unblinded local read inherits all the cognitive bias of the previous module. Finally, the population and the comparator must be scrutinized for transferability: a drug shown superior to placebo as add-on therapy in refractory focal epilepsy has not thereby been shown useful as monotherapy in new-onset generalized epilepsy, and conflating the two is a common and consequential interpretive overreach.

AI-assisted reading and automation bias in 2026

By 2026, machine-learning systems for EEG - automated spike and seizure detection, background classification, and decision support - have moved from research demonstrations into commercial and regulatory reality, and appraising them is now part of evidence-based neurophysiology rather than a speculative addendum. The same critical framework applies, with a few sharpened edges. Most reported AI performance comes from datasets that are small, pre-selected, and enriched for the very findings being detected - a form of spectrum bias - so impressive headline accuracy on curated, discharge-rich records routinely degrades on the unselected, artifact-laden recordings of real clinical practice. Reported failures are not hypothetical: published evaluations have described detection systems that miss a substantial fraction of epileptiform discharges in certain clinical contexts, and real-world validations on routine recordings have tempered the earlier non-inferiority claims drawn from idealized data. A crucial regulatory subtlety is that clearance is not the same as clinical effectiveness: regulatory authorization typically confirms analytical performance under defined conditions, not generalizability or improved patient outcomes, and the prudent reader does not mistake a clearance for proof that a tool helps real patients. Training-data bias is a further hazard, since models trained on non-representative populations can systematically underperform for groups underrepresented in the data.

The most important cognitive development is automation bias - the well-documented human tendency to over-trust automated output, both by accepting a machine's assertion uncritically (commission error) and by failing to notice what the machine missed (omission error). Automation bias is the gestalt over-reliance of the cognitive-bias module wearing a new costume: instead of anchoring on a clinical history, the reader anchors on the algorithm's label, and confirmation bias does the rest. A detector that highlights a transient as a spike supplies a powerful, authoritative-seeming prior that can override the reader's own field logic; a detector that stays silent can lull the reader into satisfaction of search. The defense is the same disciplined inference applied to a new source of priors: treat the algorithm's output as one more piece of evidence with its own likelihood ratio, not as ground truth; preserve the human verification steps - field, morphology, evolution - rather than delegating them; demand validation on representative, unselected populations before trusting a tool; and remember that a confident, well-designed interface is not evidence of a well-validated model. The clinician who applies the skepticism of this entire pillar to the machine, rather than exempting it, is the one who will use AI well.

Clinical pearl

Treat an AI EEG tool exactly as you treat a clinical history: a source of priors to be weighed, not obeyed. Beware automation bias in both directions - accepting a machine flag uncritically and missing what the machine ignored. Regulatory clearance confirms analytical performance under defined conditions, not real-world effectiveness or generalizability.

The through-line of this pillar is that critical appraisal and bedside interpretation are the same discipline applied at different scales. The Bayesian reasoning that calibrates a single read against a base rate is the reasoning that turns a published sensitivity into a predictive value for your patient. The skepticism that treats a single-electrode transient as artifact until it proves it has a field is the skepticism that treats a reported accuracy as inflated until the study proves it controlled for spectrum, verification, incorporation, and blinding - and that treats an algorithm's confident label as a hypothesis until it proves itself on representative data. Reading the EEG, reading the evidence about the EEG, and reading the machine that reads the EEG are, in the end, one habit: disciplined inference under irreducible uncertainty, in which humility about the limits of the data is the method itself.

Check your understanding

1. Why is Cohen's kappa preferred over raw percent agreement for reporting EEG interrater reliability?

2. A study reports that a suppressed EEG pattern after cardiac arrest predicts death, but the treating team saw the EEG and used it in deciding to withdraw life-sustaining therapy. The chief threat to this prognostic claim is:

3. In appraising an AI-assisted EEG tool in 2026, which statement reflects sound evidence-based reasoning?

Assessment →