Performance qualification method
At least 2× improvement over SQLite is an acceptance target, not an achieved guarantee. No result in this checkout establishes a general database performance advantage. Earlier isolated read or batch experiments are not a finished engine or reproducible product benchmark.
Compare equivalent work
Section titled “Compare equivalent work”Before claiming a win, define and publish the workload matrix and retain every passing and failing cell. Compare equivalent operations, returned values, transaction boundaries, constraints, and durability acknowledgments against prepared SQLite and the best-tested SQLite configuration for each workload.
Measure three layers separately:
- Engine: native RefPot operations versus prepared native SQLite.
- Python binding: equivalent public operations and materialized results versus Python SQLite interfaces.
- ORM: complete application operations versus optimized ORM baselines, alongside raw-engine measurements.
Include synchronization, commit, checkpoint, compaction, reclamation, and final maintenance costs. Report tail latency, memory, disk growth, recovery, and concurrency separately. Keep native macOS, Linux VM, and bare-metal evidence distinct. Repeat complete campaigns when results are unstable.
Matrix and reproducibility gates
Section titled “Matrix and reproducibility gates”The canonical roadmap requires correctness cases, successful CRUD, misses, repeated mutations, rollback, mixed workloads, point/range reads, batch and dataset sizes, key locality, memory limits, and concurrency. Persist raw trials, environment identity, source hashes, configurations, the exact matrix, expected-state validation, and a reproduction command. Retain failed cells; do not narrow the matrix after results are known.
A faster lookup or favorable large batch does not establish that RefPot is faster as a database. Dataset scaling, recovery, durability and small-transaction outcomes matter. Claim only the exact measured scope with its limitations.