Documentation

How BCI-Bench will work.

Early technical notes for researchers, dataset stewards, benchmark contributors, and future participants.

Draft documentation

Interfaces and policies may change before the first benchmark release.

01 / Overview

Platform overview

BCI-Bench is an open evaluation platform for benchmarking generalizable and deployment-ready brain-computer interface algorithms. It is designed as evaluation infrastructure rather than a dataset repository or model toolbox.

Benchmark definitions will specify data roles, participant and session splits, allowed preprocessing, calibration settings, primary and secondary metrics, resource reporting, and verification level.

The current release is a static website MVP. Submission, scoring, and hidden-evaluation services are planned but not yet active.

02 / Workflow

Benchmark workflow

  1. Browse the benchmark catalogue and select a task.
  2. Download the public data, definition, and starter kit.
  3. Develop and validate the algorithm locally.
  4. Submit predictions or a packaged model.
  5. Run standardized evaluation on the designated test data.
  6. Receive a private, multi-dimensional result report.
  7. Choose whether to publish the verified result.
  8. Appear on the public leaderboard with provenance metadata.

03 / Submission types

Progressive verification levels

Submission modes will increase in reproducibility and infrastructure requirements. Each published result will state how it was evaluated.

Prediction-file submission

Participants submit outputs in a validated schema. This is the simplest planned mode for public or server-side labels.

Packaged model submission

A model artifact and inference interface are evaluated under a standardized runtime.

Docker-based hidden evaluation

A container runs against hidden data without exposing the recordings or target labels.

Full training-pipeline verification

Training code, environment, and declared data access are reproduced to verify the complete result chain.

04 / Evaluation metrics

Performance, reliability, and cost

Every track will name a primary metric, but benchmark reports will include the additional dimensions needed to interpret deployment readiness.

  • Overall task performance with confidence intervals
  • Cross-subject and cross-session performance
  • Worst-subject and subject-level outcome distributions
  • Calibration curves and few-shot performance
  • Robustness under missing channels and signal noise
  • Inference latency, memory use, and throughput
  • Verification level and reproducibility status

05 / Data policy

Separation of development and evaluation data

Public training data, validation data, and hidden evaluation data will have explicit roles. Benchmark releases will preserve source licenses and document any access requirements.

  • Training sets support local method development.
  • Validation sets support controlled model selection.
  • Hidden test sets support independent final evaluation.
  • Registered datasets retain their original access controls.

Final data-use, retention, privacy, and submission policies will be published before evaluation services open.

06 / Roadmap

Release stages

  1. v0.1: public website, scientific direction, and draft documentation.
  2. v0.2: benchmark definitions, starter kits, baselines, and validation splits.
  3. v0.3: prediction submission, automatic scoring, and private reports.
  4. v1.0: hidden container evaluation and verified readiness metrics.

07 / Citation

Citation

A canonical paper and machine-readable citation file will be released with the first benchmark definition. Until then, please reference the project name and website.

BCI-Bench. Open evaluation infrastructure for generalizable and deployment-ready brain-computer interfaces. https://bcibench.com