// Evidence
Headline benchmark results, Standard RAG comparison, and cross-benchmark summaries drawn from the public black-box report.
Five formal test sets, 1,461 total samples. Comparison systems: standard_rag, hybrid_rag, mem0_platform. All test sets use locked data versions with reproducible results.
Test sets
5
Total samples
1,461
Comparison systems
3
LHCSB test split
291 samples
SPM-Core leads all comparison systems on every formal test set. The benchmark is presented as a direct comparison against standard_rag, hybrid_rag, and mem0_platform.
| Benchmark | SPM-Core | Best baseline | System | Lead |
|---|---|---|---|---|
| LHCSB | 85.57% | 41.58% | standard_rag | +43.99pp |
| LoCoMo 100q | 86.47% | 73.33% | standard_rag | +13.14pp |
| BABILong full-250 | 99.60% | 46.99% | hybrid_rag | +52.61pp |
| LongMemEval 100q | 73.00% | 52.00% | hybrid_rag | +21.00pp |
| RULER 720-case | 86.47% | 81.08% | standard_rag | +5.39pp |
The public report uses black-box visuals to show operating-point trade-offs and domain-level behavior without disclosing implementation details. Charts are available on the Evaluation page.
Three practical takeaways: SPM-Core leads all comparison systems on every test set; T5 abstention accuracy is 100% while all baselines score below 2.5%; BABILong full-250 shows the largest absolute lead at +52.61pp over hybrid_rag.