← Personalized BCI research

Personalized BCI
for mental health

Research synthesis · 8 September 2026 · Scalp EEG, anxiety and longitudinal personalization

The research answer

The most useful place to attack is the gap between an EEG signal that predicts a laboratory label and a measurement that tracks meaningful changes in one person over time. Foundation models and rapid personalization are promising tools for crossing that gap. They are not yet a demonstrated end-to-end solution for everyday anxiety monitoring.

The working hypothesis should be: a shared representation, a carefully estimated personal baseline, explicit session and context information, and controlled adaptation can outperform a fixed population decoder. Each component must earn its place against simpler alternatives. In particular, do not assume a frozen foundation model with a small output layer is sufficient: the 2026 EEG-FM-Compass benchmark found linear probing frequently inadequate and specialist models competitive. Liu et al., EEG-FM-Compass, version 3, August 2026.

For a mental-health BCI, four questions require separate answers: Is the recording reliable? Does it measure the intended psychological state? Does prediction survive new people and future days? Does acting on the prediction improve an outcome? A good answer to one does not establish the others.

This report focuses on noninvasive scalp EEG for monitoring anxiety-related states and, eventually, informing feedback. Motor imagery, emotion recognition, sleep and invasive depression studies are included only where their methods illuminate the problem. They are not treated as clinical anxiety validation. Recommendations and the study design below are research proposals, not established treatment protocols.

1. Define what the decoder is supposed to measure

“High anxiety” is not a single interchangeable label. A diagnosis, a person's enduring anxiety tendency, momentary worry, bodily arousal, exposure to a stressor and a need for intervention answer different questions. Decide whether the system estimates a current state, forecasts deterioration, or selects an intervention before collecting training labels.

A useful initial target is a person's momentary anxiety severity, measured with a prespecified repeated assessment, together with change from that person's initial baseline. Evaluate sustained worsening separately. Predicting that someone was assigned a stressful task is a weaker endpoint than predicting how anxious they actually felt.

DASPS illustrates the scope problem: its original report describes 23 participants exposed to psychological stimulation. This is a useful elicitation dataset, but it cannot by itself establish months-long monitoring of psychiatric patients. Its task and rating definitions must travel with any reported result. Baghdadi et al., DASPS, 2019.

Anxiety also has separable dimensions. In a study of 130 participants, anxious apprehension and anxious arousal showed different relationships with frontal alpha asymmetry. This argues for measuring worry and somatic arousal separately rather than assuming every symptom has one EEG signature. Diverging patterns of EEG alpha asymmetry, Biological Psychology, 2021.

Self-reports are informative measurements, not perfectly precise ground truth. Align their recall period with the EEG window, retain continuous ratings where possible, record uncertainty and disagreement, and use clinician assessments for slower clinical outcomes. Do not label every second of a recording using a questionnaire that summarizes the previous two weeks. That would ask the model to solve a different problem from momentary decoding.

2. Separate recording drift from genuine clinical change

A useful conceptual model is that measured EEG combines neural activity, person-specific anatomy, the electrode and reference configuration, physiological artifacts and instrument noise. The neural activity itself depends on the target state, sleep, attention, tasks, treatment and other context. This decomposition is an organizing hypothesis, not a claim that these components can always be uniquely recovered from scalp recordings.

Between-person differences change the mapping from neural sources to sensors and the relationship between features and labels. Between-session differences can include cap placement, contact quality, reference changes and context. Within-session changes add movement, fatigue and evolving states. A model may encounter several of these shifts at once.

Covariate shift means the distribution of inputs changes while the input-to-label relationship is assumed stable. Label shift means the frequency of the target state changes. Concept drift means the relationship between inputs and the target changes. An alignment method that helps with the first does not automatically solve the other two. If anxiety becomes more frequent, forcing predictions back to an old class balance can hide deterioration.

Sleep is especially revealing because it can change both EEG and anxiety. A 2025 counterbalanced sleep-deprivation study analyzed 17 healthy participants and reported changes in state anxiety and EEG measures. Its small sample does not establish a universal biomarker; it does show why treating all sleep-associated variation as removable noise is a questionable assumption. Hao, Xie and Bu, Frontiers in Psychology, July 2025.

The research target is selective stability: remain insensitive to irrelevant recording changes while remaining sensitive to real symptom changes. A drifting personal baseline can otherwise absorb persistent anxiety until a worsening person looks “normal” relative to their newly adapted baseline. Preserve an initial anchor and report both absolute severity and recent deviation.

3. The bottlenecks, why they persist, and how to attack them

Bottleneck 1 — The target is weakly identified

Barrier: a classifier can succeed by detecting task demands, distress, sleepiness or muscle tension without measuring anxiety specifically. Changing the model architecture does not resolve an ambiguous label.

Why it persists: labeled clinical episodes are harder to collect than recordings from a convenient laboratory manipulation. A task label is cheap and precise; the psychological response to that task is neither uniform nor perfectly observed.

Attack: collect anxiety, worry, arousal, mood, sleepiness and context together. Include comparison conditions that separate physical arousal from anxiety, and anxious worry from movement. Estimate within-person associations as well as between-person differences. Require predictive benefit after accounting for non-EEG information, and interpret that adjustment carefully when sleep or treatment is part of the causal pathway.

Bottleneck 2 — Artifacts can masquerade as brain biomarkers

Barrier: eye and facial-muscle activity can covary with distress. In a two-person neuromuscular-blockade experiment, scalp power above 20 Hz changed substantially when muscle activity was suppressed. This is strong evidence for a contamination mechanism, but not an estimate of contamination in every wearable or population. Whitham et al., Clinical Neurophysiology, 2007.

Why it persists: comfortable daily-use devices trade spatial coverage and preparation time for usability. More aggressive cleaning can also remove signal, while motion-related missingness can preferentially discard the moments that matter.

Attack: record contact quality, channel locations and reference, plus motion and eye or muscle channels where feasible. Test deliberate cap reapplication and channel dropout. Compare raw, cleaned and artifact-only predictors. Track usable coverage alongside accuracy. If an artifact-only model matches the EEG model, a claim of neural specificity needs stronger evidence.

Bottleneck 3 — We lack the right repeated observations

Barrier: more windows from the same few people do not provide the same evidence as more independent people, sessions and clinical transitions. Generic clinical EEG can support pretraining while lacking time-aligned everyday anxiety labels.

Existing resource: Wang and colleagues released 60 participants with three EEG sessions, including short repeats and a roughly one-month follow-up, alongside behavioral measures. It is valuable for reliability and session-shift work. Participants were young adults and current psychiatric disorders were excluded; this is not a dense clinical anxiety trajectory dataset. Wang et al., Scientific Data, September 2022.

Why it persists: repeated sensor setup, synchronized symptom assessments, participant retention and clinical follow-up require sustained effort. Privacy and data-use agreements complicate pooling. A large number of recording hours is easier to advertise than coverage of clinically meaningful transitions.

Attack: prioritize repeated weeks of EEG plus contemporaneous assessments over an indiscriminate increase in unlabeled hours. Document which devices, demographic groups, treatments and symptom ranges are represented. Retain timestamps, medication timing and changes, sleep, missingness and device metadata. Use consented data with explicit reuse terms. Federated training may help collaboration, but does not create missing labels or fix incompatible protocols.

Bottleneck 4 — Benchmarks can answer the wrong deployment question

Barrier: random EEG-window splits can put the same person's signature, recording or event into both training and testing. Brookshire and colleagues experimentally showed that segment-based evaluation inflated performance relative to subject holdout in translational EEG. Their diseases were Alzheimer's and seizures, so this is a methodological warning rather than an anxiety effect estimate. Brookshire et al., Frontiers in Neuroscience, 2024.

Why it persists: window-level examples appear plentiful, and convenient splits reward models that recognize people or recordings. Publication scores can be hard to compare when target-data access, preprocessing and adaptation budgets differ.

Attack: assign participants, sessions and events to partitions before making windows. Keep overlapping windows together. Fit learned preprocessing and hyperparameters inside training folds. Separate inductive evaluation, where target data are unavailable during fitting, from transductive evaluation, where an unlabeled target batch is used. Online systems must not normalize using future recordings. Audit pretraining overlap with the evaluation cohort.

Bottleneck 5 — Population invariance can remove useful personal information

Barrier: a representation that suppresses every difference between people may remove meaningful symptom-related differences or individual physiological responses. Conversely, retaining person identity can enable shortcuts.

Why it persists: an adversarial objective can reward indistinguishable subjects without establishing that the retained information is clinically useful. A shared representation is easier to optimize than a validated decomposition of biology, context and measurement.

Attack: compare shared-only, personal-only and hierarchical models. A hierarchical model shares statistical strength across people while allowing a personal intercept or response pattern. Keep session information explicit. Test whether adaptation improves symptom prediction, not merely whether distributions become similar. Use identity-prediction controls as diagnostics; identity information alone neither proves leakage nor proves clinical invalidity.

Bottleneck 6 — Adaptation lacks trustworthy feedback

Barrier: updating on the model's own predictions can reinforce mistakes. Updating on a new day's entire distribution can absorb a genuine shift in symptom prevalence. Frequent adaptation also risks forgetting previously useful behavior.

Why it persists: prompted motor tasks can provide trial labels during operation. Everyday anxiety rarely provides equally frequent, unambiguous labels. An unsupervised objective does not tell us whether a changed signal reflects a bad electrode or clinical worsening.

Attack: separate rapid recording-quality adjustment from slower symptom-model updates. Use sparse verified assessments, a bounded replay set, small candidate updates and a frozen fallback. Evaluate before updating on each new labeled event. Reject changes that harm an anchored validation set. A distribution alarm should trigger a quality check or targeted assessment, not automatic acceptance of a new baseline.

Bottleneck 7 — Reliability and usefulness are different

Barrier: a stable EEG measure may have little relationship to the outcome. In 79 depressed adults assessed at intake and three months, frontal alpha asymmetry and frontal theta showed sufficient reliability but mostly weak or nonsignificant psychiatric correlations. This finding concerns particular features and a particular cohort; it does not rule out all EEG representations. Validity and reliability of EEG biomarkers for depression, 2013.

Why it persists: demonstrating an association, a classification score and a beneficial intervention requires progressively different studies. A feedback system can also change behavior and thus its future input distribution.

Attack: establish prospective prediction in silent monitoring before letting predictions trigger feedback. Then test the feedback strategy against an appropriate control, including participant burden and meaningful outcomes. Personalized closed-loop treatment in one person with treatment-resistant depression is an important invasive proof of concept, but cannot validate a scalp-EEG anxiety product. Scangos et al., Nature Medicine, 2021.

4. What existing solutions actually establish

Alignment: useful, inexpensive, and limited by assumptions

Euclidean alignment adjusts EEG trials using covariance information to reduce distribution differences. The original work evaluated motor imagery and event-related potentials using offline and simulated-online experiments. It is an important low-cost baseline; it is not prospective mental-health validation. He and Wu, IEEE Transactions on Biomedical Engineering, 2020.

Riemannian Procrustes analysis matches covariance distributions through geometry-aware transformations and was evaluated across eight public BCI datasets covering three paradigms and 243 subjects. Methods and transformations differ in their target-label requirements. Compare them under the same permitted calibration information, and do not assume distribution matching preserves a psychiatric endpoint. Rodrigues, Jutten and Congedo, 2019.

Foundation models: credible transfer, incomplete deployment evidence

BENDR demonstrated large-scale self-supervised EEG pretraining and transfer across several downstream tasks. Learning structure from unfamiliar recordings is valuable, but does not itself show that the required clinical labels can be decoded without adaptation. Kostas et al., BENDR, 2021.

LaBraM pretrained on about 2,500 hours from around 20 datasets using channel patches and masked neural-token prediction. A critical qualification is that its SEED-V emotion experiment divides trials while merging subjects and sessions across the partitions. That result is not a held-out-person or future-session demonstration. Jiang et al., LaBraM, ICLR 2024, Appendix F.

EEGPT uses representation alignment and masked reconstruction and reports multiple downstream evaluations. Its adaptation includes learned spatial filters, and some tasks use a more substantial classifier. “Universal representation” should not be interpreted as proof that one tiny personalized layer will solve all settings. Wang et al., EEGPT, NeurIPS 2024.

Recent comparisons are mixed rather than uniformly negative. The ICLR 2026 benchmark Are EEG Foundation Models Worth It? finds advantages depend on the evaluation setting and data regime. A CHIL 2026 study across six datasets reports gains on longer-context tasks but comparable compact-model performance on short-window BCI tasks and limited robustness under reduced sensor coverage. Model choice must match the deployed headset and time window. Yang et al., ICLR 2026; Kommineni et al., CHIL 2026.

An August 2026 preprint found that several encoders strongly retained dataset identity and that model rankings depended on comparators and evaluation units. This is a useful negative-control proposal, not proof that all foundation models fail or that geography caused the differences. Zare, A Negative-Control Protocol for Clinical EEG Foundation-Model Benchmarks, version 3.

Continual learning: promising, but read the label assumptions

EDAPT, published in April 2026, combines population training with supervised trial-by-trial personalization and optional unsupervised alignment. It evaluates nine datasets and three BCI paradigms. Its key implication is that calibration can be integrated into operation when feedback exists. It does not eliminate the need for informative labels in passive anxiety monitoring. Haxel et al., Journal of Neural Engineering, 2026.

T-TIME instead adapts to arriving unlabeled trials, using ensemble prediction and information-maximization objectives; its evidence comes from three motor-imagery datasets. It supports testing online adaptation, not assuming that the same objectives remain valid when real anxiety prevalence changes. Li et al., IEEE Transactions on Biomedical Engineering, 2024.

The synthesis is therefore conditional: pretraining may reduce sample demands; alignment may reduce nuisance variation; personalization may learn a person's mapping; temporal modeling may distinguish persistent from transient change. None makes the other components or clinical evaluation unnecessary.

5. A system worth testing

The following architecture is a proposed research design. Its defining feature is that different kinds of change have different update rules.

Multimodal data serve two purposes: improve prediction and identify alternative explanations. In a real-world wrist-sensor study of 83 students, examination-week stress did not map simply onto increased aggregate physiological arousal. This is not EEG evidence, but it illustrates why context matters and why a peripheral-sensor comparator is necessary. Detecting Prolonged Stress in Real Life, JMIR, 2023.

If the same self-report supplies both the model input and the target, apparent success can be circular. Specify when each input becomes available. For forecasting, every feature must precede the prediction horizon; for contemporaneous estimation, explain what the system adds beyond simply asking the person.

6. Where Ravia should attack first

Priority 1 — Build a decisive longitudinal benchmark

The first deliverable should be a benchmark for within-person anxiety change across days, with truly unseen participants reserved for final evaluation. Use one defined recording workflow initially. Include cap reapplication and natural changes in context. This isolates whether the central problem is measurement, labels, personalization or inadequate signal.

A feasible discovery plan could recruit approximately 40–60 participants for 6–8 weeks, with clinical and comparison participants, repeated brief EEG measurements and time-aligned assessments. These numbers are illustrative planning assumptions, not a power calculation or an efficacy-study recommendation. Size the confirmatory study from observed participant-level variance, event frequency, attrition and a prespecified useful effect.

Begin with adherence and signal-quality feasibility. Schedule a manageable number of assessments and measure the burden. Maintain naturally occurring medication and sleep records rather than deliberately manipulating clinical treatment for model training. A clinical research partner should shape the endpoint and participant protocol.

Priority 2 — Prove incremental value before scaling pretraining

Run the same future-day evaluation for: the person's historical mean; their most recent available symptom report; sleep and context; peripheral physiology; simple EEG features; a compact EEG network; and pretrained EEG models with several adaptation levels. Evaluate EEG alone and its incremental contribution to multimodal models.

The decisive question is whether adding EEG improves meaningful prediction at an acceptable cost in preparation time, missingness, discomfort and false alerts. If it does not, investigate measurement and label design before paying for a larger encoder. A null result here can save an entire research program from optimizing an irrelevant benchmark.

Priority 3 — Make personalization resistant to false adaptation

Use controlled recording perturbations and naturally observed clinical transitions as separate tests. A cap reapplication should not create a symptom alarm; persistent symptom worsening should not disappear through baseline updates. Compare fixed, scheduled supervised and continuously adaptive models, and report how often adaptation helps or harms individual participants.

Only after these tests show a reproducible advantage should Ravia expand pretraining, devices and sites. The defensible long-term asset is a well-characterized longitudinal dataset plus a trusted evaluation and adaptation system. A large model alone is less informative about the actual deployment problem.

7. The evaluation contract

For continuous severity, report error, calibration of uncertainty intervals and within-person change sensitivity. Correlation alone can conceal a biased scale or differences between people. For an alert endpoint, prespecify its clinical definition and report precision–recall performance, sensitivity at an acceptable false-alert rate, false alerts per day and detection delay.

For every setting, report usable recording coverage, abstention, dropout, per-person results and relevant subgroup results with uncertainty. Resample at the participant or session level rather than treating correlated windows as independent patients. Choose thresholds on validation data; leave the final evaluation set untouched.

Progression criteria should be set before looking at the final results: a prespecified incremental improvement over the strongest non-EEG comparator, acceptable participant burden and alert load, and no unacceptable degradation in the groups the device is intended to serve. Numeric thresholds require the intended use and pilot estimates; inventing a universal accuracy target would be misleading.

8. What would change this conclusion?

Evidence in favor of the foundation-model hypothesis would be a replicated improvement over compact personalized models on held-out people and chronological future sessions, with fewer labels, acceptable sensor coverage and preserved detection of genuine worsening. Stronger evidence would include an external site and then a controlled study of the proposed feedback intervention.

Evidence against the proposed approach would include no incremental EEG value after honest comparison, reliance on artifacts or task identity, adaptation that degrades vulnerable users, or burden that makes clinically important periods systematically unobserved. In those cases, change the measurement, endpoint or sensing approach rather than assuming more pretraining will rescue it.

The open question is not whether personalized EEG decoding can work in any experiment. It is whether it can deliver a reliable, specific and useful measurement for the intended person, device, context and clinical decision over time. The highest-value experiment is the one that can falsify that claim early.

Research scope and limitations

This is a targeted evidence synthesis, not a preregistered systematic review or meta-analysis. Searches covered foundational alignment and pretraining papers, 2026 model comparisons, continual adaptation, translational leakage, anxiety labels, reliability and repeated-measure datasets. Original papers, author abstracts and official data records were prioritized. Evidence was checked through 8 September 2026; versioned preprints are identified as such where used.

Full original text was examined where accessible, including the LaBraM evaluation protocol and the longitudinal dataset design. Some comparisons and clinical studies were available through author abstracts or indexed primary records only. No numerical effects are pooled across incompatible tasks. The selected literature does not establish prospective, multi-month scalp-EEG anxiety monitoring with a universally effective small personalization layer. That is a bounded assessment of the reviewed evidence, not proof that no relevant study exists.

Research stopped after the consequential claims had primary support or an explicit limitation, including contrary evidence on stability and foundation-model performance. The ranking of bottlenecks, architecture, priorities and study plan are this report's synthesis and should be tested rather than treated as settled findings.