K-Bench Leaderboard
Evaluating LLM unlearning across memory substrates and observable channels in agentic deployments — without misleading cross-substrate averages.
The Substrate Matrix
A single method earns fundamentally different verdicts on different memory substrates. Comparing rows reveals that unlearning is substrate-dependent.
| Method | P Parametric (Weights) | C In-Context Memory | R-text Unstructured Retrieval | R-struct Structured Retrieval |
|---|
Linked method names open their source paper from the paper's own bibliography; NOISE and None are controls, not published methods, and carry no reference.
Per-Substrate Ranking
Rankings are computed strictly within each individual substrate and base model. Because each base model's K-Score floor depends on its own no-intervention baseline, absolute values are compared within a base model rather than across models, and are never combined across substrates.
| # | Method | OR(all) | Retain | Δsel | Degeneration % | K-class | K-Score |
|---|
Twenty published methods, three base models
How to read this chart
Observation-Budget Recovery
Baseline forget-set recovery rate across nested attacker observation budgets (A1 ⊆ … ⊆ A5). Secrets suppressed in final answers frequently leak when reasoning traces or tool execution channels are observed.
How to Submit
Every score on K-Bench is backed by a hash-verified bundle to ensure reproducible evaluation.
Submissions go through a pull request against github.com/OniReimu/kbench under leaderboard/submissions/.
Each submission ships a hash-verified bundle built by kbench bundle and validated by kbench report against the frozen evaluation harness.
The Space displays merged results only. This strict verification gate guarantees that every row and verdict here can be independently checked and audited.