Benchmark catalogue

Scientific questions, made measurable.

BCI-Bench is organized into focused tracks that test generalization and deployment behavior under standardized conditions.

v0.1 planning

Track definitions and dataset selections are under community development.

Cross-Subject Generalization

v0.1

Can a BCI algorithm work on unseen users without subject-specific training?

Evaluation

Train on one participant group and evaluate on held-out participants excluded from training and validation.

Planned data

Multi-subject motor imagery and ERP datasets.

Cross-Session Stability

Planned

Does an algorithm remain reliable across future sessions and neural drift?

Evaluation

Develop on earlier sessions and test chronologically on unseen future recordings.

Planned data

Longitudinal and repeated-session datasets.

Cross-Device Transfer

Planned

Can performance transfer across acquisition systems, channel layouts, and sites?

Evaluation

Train on one device or center and evaluate on recordings from another.

Planned data

Compatible multi-device and multi-center cohorts.

Few-Shot Calibration

Planned

How much new-user data is required before a model becomes useful?

Evaluation

Report zero-shot, 1-shot, 5-shot, and calibration-curve performance.

Planned data

Datasets with sufficient trials per participant.

Missing-Channel Robustness

Planned

How gracefully does a model degrade when channels are noisy or unavailable?

Evaluation

Controlled channel dropout, signal corruption, and montage reduction.

Planned data

High-density datasets with channel metadata.

Foundation Model Evaluation

Experimental

Do pretrained neural models transfer fairly across BCI tasks and populations?

Evaluation

Frozen, adapted, and fully tuned protocols with controlled data access.

Planned data

Multi-task datasets spanning common BCI paradigms.

Deployment and Latency Evaluation

Planned

Can an algorithm satisfy practical compute, memory, and real-time constraints?

Evaluation

Standardized runtime profiles, end-to-end latency, memory use, and throughput under declared hardware conditions.

Planned data

Representative benchmark inputs with repeatable runtime harnesses.

Reporting standard

No single number tells the whole story.

Track reports will pair central performance with subject-level distributions, calibration cost, robustness, and resource usage where relevant.

Detailed benchmark definitions, eligible datasets, split files, and baseline implementations will be versioned publicly before each track opens.
Read the metric framework