// Datasets
The public test surfaces used to evaluate SPM across long-horizon conversational state, factual retention, confusion control, and long-context stress.
The public benchmark suite is split into multiple surfaces so readers can distinguish long-horizon conversational state, multi-session recall, synthetic state tracking, and long-context validation.
| Dataset | Cases | What it tests |
|---|---|---|
| LHCSB | 291 (test split) | Long-horizon conversational state: fact tracking, abstention, temporal reasoning |
| LoCoMo 100q | 100 | Multi-session long conversation recall |
| BABILong full-250 | 250 | Synthetic state tracking at 16k context (QA1–QA5) |
| LongMemEval 100q | 100 | Long-term memory question answering |
| RULER 720-case | 720 | Long-context stress across full category coverage |
| Ultra-long | 4 | Extreme-length validation above 0.8M baseline tokens |
Publicly described cases follow a common shape: target facts or constraints are injected early, unrelated history is added, and a fact-sensitive probe is asked at the end.
early fact injection
topic switches or noise flooding
similar-fact distractors where appropriate
fact-sensitive end question
The suite turns SPM claims into something auditable. It gives readers a way to understand both everyday and extreme evaluation surfaces without requiring access to internal implementation detail.