121 self-audit

Same rubric. Same receipts. Same correction burden.

121 does not get to grade vendors without grading its own products first. Current scores are provisional where directly observed and TBD where a measured run has not happened.

v0.1 pilot. Public self-application first; no vendor score is treated as meaningful without this page staying current.

First-run scores

ProductStateModel identity disclosureReroute and fallback visibilityDrift disclosureData retention clarityPricing, quota, and effort clarityCorrection and redress pathNote
EleanorPending instantiationN/AN/AN/AN/AN/AN/ANamed but not live; scoring begins only when an actual Eleanor runtime exists.
QuillScaffold v0.1TBDTBDTBDTBDTBDTBDPublic pages exist, but no measured Quill runtime is available for scoring.
CompanionScaffold v0.1TBDTBDTBDTBDTBDTBDFoundation surfaces are live; companion behavior remains incomplete.
SwitchboardRouting layer2/5 provisional1/5 provisionalTBD3/5 provisional2/5 provisional1/5 provisionalThe routing layer exposes status and route disclosures, but signed receipts and public correction flow are not complete.

Substrate continuity

Reference implementation target

Eleanor-on-Hermes is the intended reference case, but Eleanor is not live. The current measured target is Switchboard route disclosure and recovery behavior.

Identity continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track
Memory continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track
Tool continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track
Capability continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track
Disclosure continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track
Recovery continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track
Audit continuity

Status: TBD until the substrate-continuity canaries run against a real route-change event.

Research Track

Known limitations

What this page cannot claim yet

  • No vendor score is published until a benchmark run has a receipt.
  • No certification or trust badge exists in v0.1.
  • The Trust Wire is curated and review-gated; raw scraper output is not a product.
  • External vendor data retention, internal model parity, and training corpus truth remain only partially observable.
  • Private canaries may exist to reduce benchmark gaming, but public scores need public methodology and preserved history.
  • The Anthropic Fable 5 incident is a trigger for this category, not the target or the whole story.

Incident and correction log

No public Observatory corrections yet.

The correction ledger starts with this v0.1 scaffold. Future entries must preserve the original claim, source labels, correction date, corrected text, and reason for change.