Documentation
How BCI-Bench will work.
Early technical notes for researchers, dataset stewards, benchmark contributors, and future participants.
Interfaces and policies may change before the first benchmark release.
01 / Overview
Platform overview
BCI-Bench is an open evaluation platform for benchmarking generalizable and deployment-ready brain-computer interface algorithms. It is designed as evaluation infrastructure rather than a dataset repository or model toolbox.
Benchmark definitions will specify data roles, participant and session splits, allowed preprocessing, calibration settings, primary and secondary metrics, resource reporting, and verification level.
02 / Workflow
Benchmark workflow
- Browse the benchmark catalogue and select a task.
- Download the public data, definition, and starter kit.
- Develop and validate the algorithm locally.
- Submit predictions or a packaged model.
- Run standardized evaluation on the designated test data.
- Receive a private, multi-dimensional result report.
- Choose whether to publish the verified result.
- Appear on the public leaderboard with provenance metadata.
03 / Submission types
Progressive verification levels
Submission modes will increase in reproducibility and infrastructure requirements. Each published result will state how it was evaluated.
Prediction-file submission
Participants submit outputs in a validated schema. This is the simplest planned mode for public or server-side labels.
Packaged model submission
A model artifact and inference interface are evaluated under a standardized runtime.
Docker-based hidden evaluation
A container runs against hidden data without exposing the recordings or target labels.
Full training-pipeline verification
Training code, environment, and declared data access are reproduced to verify the complete result chain.
04 / Evaluation metrics
Performance, reliability, and cost
Every track will name a primary metric, but benchmark reports will include the additional dimensions needed to interpret deployment readiness.
- Overall task performance with confidence intervals
- Cross-subject and cross-session performance
- Worst-subject and subject-level outcome distributions
- Calibration curves and few-shot performance
- Robustness under missing channels and signal noise
- Inference latency, memory use, and throughput
- Verification level and reproducibility status
05 / Data policy
Separation of development and evaluation data
Public training data, validation data, and hidden evaluation data will have explicit roles. Benchmark releases will preserve source licenses and document any access requirements.
- Training sets support local method development.
- Validation sets support controlled model selection.
- Hidden test sets support independent final evaluation.
- Registered datasets retain their original access controls.
Final data-use, retention, privacy, and submission policies will be published before evaluation services open.
06 / Roadmap
Release stages
- v0.1: public website, scientific direction, and draft documentation.
- v0.2: benchmark definitions, starter kits, baselines, and validation splits.
- v0.3: prediction submission, automatic scoring, and private reports.
- v1.0: hidden container evaluation and verified readiness metrics.
07 / Citation
Citation
A canonical paper and machine-readable citation file will be released with the first benchmark definition. Until then, please reference the project name and website.