MORGOTH
Toward unified and comprehensive automated electroencephalogram interpretation: a multicentre development and validation of an electroencephalogram foundation model
Clinical Question
Can a single AI foundation model deliver expert-level, comprehensive automated EEG interpretation across all major clinical settings (routine outpatient, epilepsy monitoring units, ICU, sleep labs) and across the full spectrum of clinically relevant patterns (seizures, epileptiform discharges, rhythmic and periodic patterns of the ictal–interictal–injury continuum, pathological slowing, sleep staging)?
Study Overview
Objective
Multicentre development and external validation of MORGOTH, a foundation model for comprehensive automated EEG interpretation across 7 event-level and 17 EEG-level clinical tasks, benchmarked against human experts and state-of-the-art models.
Study Summary
- MORGOTH achieved AUC-ROC 0.86–0.98 across 17 EEG findings and exceeded ≥20% of experts on every task; outperformed ≥90% of experts on 3 of 7 multi-expert-annotated datasets.
- Event-level detection was especially strong for seizure/ictal–interictal–injury continuum (EUC 96.6%) and spike detection (EUC 100%); on IIIC-Test AUC-ROC 0.990 and on SN2-Test spike detection AUC-ROC/PR 1.00.
- Generalisation drop from internal to external test sets was modest (event-level AUC −1.21%, EUC −3.33%; EEG-level AUC −2.12%, EUC −9.52%) and MORGOTH showed lower age sensitivity (20.90% vs 30.90%) and fewer sex-related differences (33.33% vs 44.00%) than SPaRCNet.
Intervention
MORGOTH — a 19-channel EEG transformer foundation model (tokeniser with 8192 tokens via contrastive learning + vector quantisation; 12 encoder blocks with multihead attention; fully connected event-level heads and convolutional/attention EEG-level heads) — vs board-certified human experts and SOTA models (SPaRCNet, SpikeNet 2, U-Sleep, SCORE-AI, LaBraM, EEGFormer, BrainBERT, Kaggle competition winner).
Patients per Arm
Development 18 677 patients (14 500 pretraining, 18 677 fine-tuning); internal validation 13 334 patients / 13 618 EEGs; external validation 1573 patients / 1800 EEGs from 48 institutions. Total 12 datasets, 34 602 patients, 52 hospitals, 8 countries.
Bottom Line
MORGOTH is the first EEG foundation model to deliver expert-level, generalisable performance across the full range of clinically relevant EEG tasks and settings. It matched or exceeded expert consensus and outperformed SOTA task-specific models on internal and external multi-institutional data, with only modest generalisation losses and superior robustness across age, sex, and moderate channel loss, providing a scalable path to expand EEG diagnostic capacity in low-resource and high-volume settings.
Major Points
- Foundation model trained on 14 500 patients (HEEDB pre-training) and fine-tuned on 18 677 patients from 4 Boston-area hospitals (MGH, BWH, BIDMC, BCH), then internally validated on 13 334 patients / 13 618 EEGs and externally validated on 1573 patients / 1800 EEGs from 48 institutions across 8 countries (52 hospitals total, 12 datasets, 34 602 patients).
- Performs 7 event-level tasks (3 binary — normal vs abnormal, burst suppression, spike detection; two 3-class — focal vs generalised vs none for slowing and spike localisation; one 5-class sleep staging; one 6-class seizure/IIIC — seizure, LPDs, GPDs, LRDA, GRDA, other) and 17 EEG-level binary tasks (one per finding).
- Internal validation: average AUC-ROC 0.921 (95% CI 0.780–0.981) across 7 event-level and 17 EEG-level tasks; outperformed on average 81.0% (46.9–100; 17/21) of experts on 3 multi-expert internal datasets.
- External validation: average AUC-ROC 0.908 (0.705–0.980) across 7 event-level and 8 EEG-level tasks; outperformed on average 50.8% (14.3–98.0; 18/35) of experts on 4 multi-expert external datasets; exceeded ≥20% of experts on every one of the 17 evaluated tasks.
- MoE dataset (21 experts; MoE-Internal 1841 patients/2125 EEGs and MoE-External 409 patients/636 EEGs, 14 categories): AUC-ROC 0.945 (0.925–0.962), AUC-PR 0.803 (0.731–0.860), outperforming automated burst-suppression detector by 0.084 and 0.428 (statistically significant). IIIC-Test set (30 experts): outperformed 96.6% (82.1–100; 29/30) of experts, surpassing all in LPD, GPD, LRDA; AUC-ROC 0.990 (0.988–0.991), AUC-PR 0.940 (0.926–0.953), exceeding Kaggle winner by 0.041/0.168 and SPaRCNet by 0.126/0.387 (per Figure 2B per-class averages).
- SN2-Test (24 experts): outperformed all 24 experts in spike detection; MORGOTH AUC-ROC and AUC-PR both 1.00 vs SpikeNet 2 (0.975 / 0.990).
- External generalisation drop was small — event-level AUC −1.21% (0.00–3.94), EUC −3.33% (0.00–33.33); EEG-level AUC −2.12% (0.00–5.83), EUC −9.52% (0.00–75.00). Overfitting minimal: mean train–test accuracy difference 2.75%, never >5%.
- Sleep staging: on external UPenn (6 experts) AUC-ROC 0.930 (0.927–0.933) vs U-Sleep 0.917 (0.914–0.921); AUC-PR 0.739 vs 0.706, despite U-Sleep using an added EOG channel. On MASS dataset outperformed U-Sleep by average AUC-ROC 0.030 (0.022–0.038). N1 remained hardest (AUC-PR 0.262).
- Reliability (Cohen's κ): matched or exceeded expert consensus on 7 multi-expert datasets. On IIIC-Test outperformed experts on all IRR metrics with substantial LPD lead. On SN2-Test achieved moderate agreement with individual experts and near-perfect agreement with consensus. On TELEEEG (23 LMICs, 8099 EEGs) AUC-ROC >0.7 in 9 of 12 tasks.
- Robustness across demographics: on IIIC-Test lower age sensitivity (20.90% vs 30.90%) and fewer sex-related differences (33.33% vs 44.00% SPaRCNet; 48.00% Kaggle winner). Local features more affected by missing channels than global patterns.
Design
Study Type: Multicentre retrospective development and external validation of an AI foundation model for comprehensive automated EEG interpretation (diagnostic accuracy study; not a randomised interventional trial)
Randomization:
Blinding: Expert annotators blinded to model outputs; multi-expert labels obtained independently and used as consensus ground truth. IRR analysis compared expert–expert with expert–model pairs.
Enrollment Period: Jan 1, 2003 – Feb 1, 2025 (EEG data collection window)
Follow-up Duration: Not applicable (diagnostic study; no longitudinal patient follow-up)
Centers: 52
Countries: USA, Canada, Belgium, Switzerland, Brazil, UK, China, Greece
Sample Size: 34602
Power Calculation: Not reported (diagnostic model validation; 95% CIs derived by 10 000 bootstrap iterations; differences considered significant when intervals did not overlap)
Analysis: Performance evaluated with ROC and precision-recall curves. Expert benchmarking used experts' operating points under the curve (EUC) — fraction of expert operating points below the model's ROC/PR curve. Cohen's κ for reliability. Calibration indices derived by parametric-model fitting (range −1 undercalibrated to +1 overcalibrated). Platt scaling and isotonic regression for post-hoc calibration. Training: focal loss with curriculum learning; AdamW optimiser with learning-rate scheduling and early stopping. Two-stage training — self-supervised pretraining on 14 500 HEEDB patients, then task-specific fine-tuning heads. Bootstrap 10 000 iterations for 95% CIs.
Inclusion Criteria
- Clinical EEG recordings from routine outpatient clinics, epilepsy monitoring units, intensive care units, or sleep laboratories
- All ages (0 to >90 years) accepted
- Expert-labelled event-level and/or EEG-level annotations available
- Recordings compatible with 19-channel 10–20 system input; other sampling rates resampled to 200 Hz; missing channels imputed by mean of available channels
- Institutional review board approval at contributing sites; waiver of informed consent for retrospective analysis (BIDMC IRB protocols #2022P000481 and #2022P000417)
Exclusion Criteria
- No explicit patient or EEG exclusion criteria reported by the authors; the study retrospectively included the available expert-annotated recordings from participating institutions
Arms
| Field | MORGOTH (AI foundation model) | Control | Control |
|---|---|---|---|
| Intervention | 19-channel EEG transformer foundation model: tokeniser (8192 tokens via contrastive learning + vector quantisation) + 12 encoder blocks with multihead attention and temporal/spatial positional encoding + task heads. Event-level heads use fully connected layers on 10-s or 1-s segments; EEG-level heads use convolutional/attention on 10-min or full recordings. Supports variable-length input at inference. | Board-certified neurologists, epileptologists, and sleep medicine specialists — 6 to 30 per test set (MoE 21; IIIC 30; SN2 24; SAI 14; ON 15; UPenn 6; HEP 14). Multi-expert consensus used as reference; individual experts also compared via EUC and Cohen's κ. | SpikeNet 2 (spike detection); SPaRCNet and Kaggle competition winner (seizure/IIIC classification); U-Sleep (sleep staging); an automated burst-suppression detector; SCORE-AI (EEG-level tasks); and EEG foundation models LaBraM, EEGFormer, BrainBERT on TUH benchmarks. |
| N |
Outcomes
| Outcome | Type | Control | Intervention | HR / OR / RR | P-value |
|---|---|---|---|---|---|
| Overall diagnostic performance across 7 event-level and 17 EEG-level tasks, measured by AUC-ROC, AUC-PR, and experts' operating points under the curve (EUC), on internal (HEEDB-Test, IIIC-Test, SN2-Test, MGH-PSG-Test, BCH-PSG, MoE-Internal) and external (HEP, SAI, ON, TUH-Test, UPenn, MASS, MoE-External, TELEEEG) validation datasets | Primary | MORGOTH internal validation (average across 7 event-level + 17 EEG-level tasks): AUC-ROC 0.921 · MORGOTH external validation (average across 7 event-level + 8 EEG-level tasks): AUC-ROC 0.908 · AUC-ROC 95% CI (internal): 0.780–0.981 · AUC-ROC 95% CI (external): 0.705–0.980 | |||
| IIIC-Test (30 experts) AUC-ROC — average across 6 classes (per Figure 2B) | Secondary | MORGOTH: 0.990 (95% CI 0.988–0.991) · Kaggle winner: 0.949 · SPaRCNet: 0.864 · Notes: MORGOTH exceeded Kaggle winner by 0.041 and SPaRCNet by 0.126 (Figure 2B per-class averages) | |||
| IIIC-Test AUC-PR — average across 6 classes | Secondary | MORGOTH: 0.940 (0.926–0.953) · Kaggle winner: 0.772 · SPaRCNet: 0.553 · Notes: MORGOTH exceeded Kaggle winner by 0.168 and SPaRCNet by 0.387 | |||
| IIIC-Test EUC — % of experts outperformed | Secondary | MORGOTH: 96.6% (82.1–100; 29/30) · Notes: surpassed all experts in LPD, GPD, LRDA | |||
| MoE dataset (21 experts, 14 categories; MoE-Internal 1841 patients/2125 EEGs and MoE-External 409 patients/636 EEGs) AUC-ROC | Secondary | MORGOTH: 0.945 (0.925–0.962) · Automated burst suppression detector delta: +0.084 (significant) | |||
| MoE AUC-PR | Secondary | MORGOTH: 0.803 (0.731–0.860) · Delta vs burst-suppression detector: +0.428 (significant) | |||
| MoE experts outperformed | Secondary | MORGOTH: 81.0% (46.9–100; 17/21) on 13 of 14 tasks | |||
| SN2-Test spike detection (24 experts) AUC-ROC / AUC-PR | Secondary | MORGOTH: 1.00 / 1.00 · SpikeNet 2: 0.975 / 0.990 · Notes: outperformed all 24 experts (SpikeNet 2 values from Figure 2C) | |||
| Sleep staging — UPenn dataset (6 experts) AUC-ROC | Secondary | MORGOTH: 0.930 (0.927–0.933) · U-Sleep: 0.917 (0.914–0.921) | |||
| Sleep staging — UPenn AUC-PR | Secondary | MORGOTH: 0.739 (0.730–0.748) · U-Sleep: 0.706 (0.698–0.715) · Notes: U-Sleep additionally used an EOG channel; N1 AUC-PR 0.262 (0.248–0.278) — MORGOTH still outperformed 52.7% of experts; REM AUC-PR 0.857 (0.849–0.864) — did not surpass any | |||
| MASS sleep dataset — average AUC-ROC advantage over U-Sleep | Secondary | MORGOTH: +0.030 (0.022–0.038) | |||
| External single-scored corpora — AUC-ROC | Secondary | TUH Seizure Corpus: 0.883 (0.881–0.885) · TUH EEG Events Corpus: 0.729 (0.695–0.763) · TUH Slowing Corpus: 0.747 (0.715–0.779) · TUH Abnormal EEG Corpus: 0.902 (0.900–0.904) | |||
| SAI dataset (14 experts, 5 tasks) — average EUC across tasks | Secondary | MORGOTH ROC EUC: 50.62% (24.16–77.14) · MORGOTH PR EUC: 53.50% (26.36–77.58) · Notes: outperformed half of experts | |||
| ON dataset (15 experts) — EUC-ROC / EUC-PR | Secondary | MORGOTH ROC EUC: 43.8% (14.3–78.6) · MORGOTH PR EUC: 39.0% (8.9–75.0) | |||
| HEP dataset — event-level AUC-ROC / AUC-PR | Secondary | MORGOTH: 0.999 / 0.999 · Notes: near-perfect; surpassed SpikeNet 2 at both event and EEG level | |||
| Generalisation drop (internal → external) | Secondary | Event-level AUC drop: −1.21% (0.00–3.94) · Event-level EUC drop: −3.33% (0.00–33.33) · EEG-level AUC drop: −2.12% (0.00–5.83) · EEG-level EUC drop: −9.52% (0.00–75.00) | |||
| IRR on IIIC-Test (Cohen's κ) | Secondary | MORGOTH: outperformed experts on all IRR metrics with substantial lead for LPDs | |||
| IRR on SN2-Test | Secondary | MORGOTH: moderate agreement with individual experts; near-perfect agreement with consensus | |||
| TELEEEG (8099 EEGs from 23 LMICs) | Secondary | MORGOTH: AUC-ROC >0.7 in 9 of 12 tasks; some decline attributed to missing channels and label noise | |||
| Not applicable — diagnostic model, no patient interventions or adverse events reported. | Safety | Notes: Overfitting monitored as internal-to-external accuracy delta: mean 2.75%, never exceeded 5%. | |||
Subgroup Analysis
Prespecified subgroup analyses across age and sex on IIIC-Test showed MORGOTH's demographic robustness exceeded baselines: age sensitivity 20.90% vs 30.90% for SPaRCNet; sex-related differences 33.33% (MORGOTH) vs 44.00% (SPaRCNet) and 48.00% (Kaggle winner). Local features (e.g., spike detection) most stable across age; slowing detection showed greater variability. Local features more affected by missing channels than global patterns.
Criticisms
- Uses only the standard 19 EEG channels; does not incorporate EMG, EOG, or other polysomnography signals — limits sleep-related task performance and detection of arousals, apnoeas, and limb movements.
- Training data drawn mainly from tertiary US epilepsy centres (MGH, BWH, BIDMC, BCH) — may reflect referral-centre case-mix biases.
- Does not explicitly detect certain normal variants or artifacts (breach rhythm, photoparoxysmal responses) that even experts interpret variably.
- Model distinguishes focal vs generalised activity but does not localise focal activity or distinguish between seizure subtypes — limits presurgical utility.
- Does not use clinical context (e.g., treatment response) — focuses on pattern recognition rather than outcome prediction (epilepsy risk or syndrome classification).
- External validation drop for EEG-level tasks larger than for event-level (EUC −9.52% vs −3.33%), suggesting some site-specific overfitting for whole-recording classifications.
- N1 sleep stage detection remained hard (AUC-PR 0.262; no improvement post-calibration), reflecting expert-level ambiguity.
- Human–AI workflow (preliminary reports generated for clinician review) not prospectively validated — prospective trials still needed for regulatory approval.
- TELEEEG (LMIC) performance declined (some tasks AUC-ROC <0.7) due to missing channels and label noise, indicating need for further expert review in resource-limited real deployment.
- Study is retrospective diagnostic accuracy work — not a randomised interventional trial; does not directly measure impact on patient care or outcomes.
Funding
US National Institutes of Health (multiple grants including RFG064312, RF1NS120947, R01AG073410, R01HL161253, R01NS126282, R01AG073598, R01NS117904, R01NS131347, R01NS130119, R01NS111022, R01NS131967, K23AG063899, NS121559). Additional support from: Swebilius Foundation; CONDA Award; DHHS LB606 Nebraska Stem Cell Grant; Fonds National pour la Recherche Scientifique; Brussels Institute for Research and Innovation (INNOVIRIS); Fonds Erasme pour la Recherche Médicale; Fonds Jaumotte; Department of Defense (HT9425-23-1-0242, HT9425-25-1-0170, W81XWH-19-1-0861, W81XWH-21-C-0075); AMFDP 843457 (20CDA35310297 and 24DIVSUP1274116); Regents of the University of California; Cures Within Reach (2022CAL-Amorim); Zoll Foundation; Hellman Foundation; NINDS (K23NS124656, K23NS112596, R01NS117904, SR21NS137117); Brain Aneurysm Foundation; and Veterans Affairs Office of Research and Development (I01HX003107-01A2).
Based on: MORGOTH (Lancet Digital Health, 2026)
Authors: Sun C, Karakis I, Herlopian A, ..., Jing J)
Citation: Lancet Digit Health 2026; published online. DOI: 10.1016/j.landig.2026.101039
Content summarized and formatted by NeuroTrials.ai.