121 self-audit
Same rubric. Same receipts. Same correction burden.
121 does not get to grade vendors without grading its own products first. Current scores are provisional where directly observed and TBD where a measured run has not happened.
First-run scores
| Product | State | Model identity disclosure | Reroute and fallback visibility | Drift disclosure | Data retention clarity | Pricing, quota, and effort clarity | Correction and redress path | Note |
|---|---|---|---|---|---|---|---|---|
| Eleanor | Pending instantiation | N/A | N/A | N/A | N/A | N/A | N/A | Named but not live; scoring begins only when an actual Eleanor runtime exists. |
| Quill | Scaffold v0.1 | TBD | TBD | TBD | TBD | TBD | TBD | Public pages exist, but no measured Quill runtime is available for scoring. |
| Companion | Scaffold v0.1 | TBD | TBD | TBD | TBD | TBD | TBD | Foundation surfaces are live; companion behavior remains incomplete. |
| Switchboard | Routing layer | 2/5 provisional | 1/5 provisional | TBD | 3/5 provisional | 2/5 provisional | 1/5 provisional | The routing layer exposes status and route disclosures, but signed receipts and public correction flow are not complete. |
Substrate continuity
Reference implementation target
Eleanor-on-Hermes is the intended reference case, but Eleanor is not live. The current measured target is Switchboard route disclosure and recovery behavior.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Status: TBD until the substrate-continuity canaries run against a real route-change event.
Known limitations
What this page cannot claim yet
- No vendor score is published until a benchmark run has a receipt.
- No certification or trust badge exists in v0.1.
- The Trust Wire is curated and review-gated; raw scraper output is not a product.
- External vendor data retention, internal model parity, and training corpus truth remain only partially observable.
- Private canaries may exist to reduce benchmark gaming, but public scores need public methodology and preserved history.
- The Anthropic Fable 5 incident is a trigger for this category, not the target or the whole story.
Incident and correction log
No public Observatory corrections yet.
The correction ledger starts with this v0.1 scaffold. Future entries must preserve the original claim, source labels, correction date, corrected text, and reason for change.