Biocomputing papers often demonstrate different tasks on different biological preparations with different digital readouts. Direct ranking is therefore usually unjustified. A useful benchmark has to define the complete hybrid system and compare it with a strong electronic baseline performing the same task.
Biocomputers.co.uk comparison checklistThis is an editorial framework for reading experiments consistently. It is not an adopted field standard.
Count the whole system
For energy and cost comparisons, include culture maintenance, incubation, pumps, stimulation, recording electronics, data acquisition and the digital model used to encode or decode signals. Reporting only neuronal metabolic power can make an experimental platform look far more efficient than the apparatus operating it.
Separate biological and digital learning
Many hybrid experiments train a digital readout around a biological reservoir. Others deliberately alter the biological state through feedback. Both can be useful, but they answer different questions. A benchmark should say which parameters changed during training and where the retained information resides.
Biological replication matters
Repeated trials on one culture estimate task variability. Independent cultures estimate whether the result survives biological variation. Both numbers are useful and should be reported separately.
| Metric | What to report | Comparison note |
|---|
| Task performance | Accuracy, reward, error or task-specific score | Compare against a strong electronic baseline under matched conditions |
| Learning speed | Samples, episodes or wall-clock time to target performance | Report both biological adaptation and digital readout training |
| Data efficiency | Training examples or interactions required | Separate pretraining and online adaptation |
| Stability | Performance drift across hours or days | Show mean, variance and failure modes |
| Culture-to-culture variance | Spread across independent biological preparations | Report biological n, not only repeated trials on one culture |
| Retention | Performance after a defined rest interval | State whether readout weights, stimulation protocol or biological state are retained |
| I/O bandwidth | Independent input and output channels plus update rate | Include effective usable channels, not electrode count alone |
| Whole-system energy | Culture support + stimulation + recording + digital compute | Measure the full experimental stack at the wall where possible |
| Preparation cost | Cell culture, maturation, consumables and specialist labour | State amortisation assumptions |
| Lifetime | Useful operating period at specified performance | Distinguish tissue viability from computational usefulness |
| Replication | Independent repeats and external laboratory replication | Label company-only, preprint and peer-reviewed evidence separately |
| Recoverability | Ability to restart, recalibrate or transfer function after failure | Record whether learned state can be copied or reconstructed |
Download the checklistCSV What would be convincing?
A strong result would use a pre-registered or clearly fixed task, multiple independent biological preparations, a competitive electronic baseline, transparent digital preprocessing and whole-system resource measurements. Independent replication would then tell us whether the advantage generalises beyond one laboratory and one culture protocol.
The evidence database deliberately records task and evidence status without pretending that unlike tasks form a leaderboard.