Methodology v0.1

What did the user actually receive, and was the user told?

The Observatory measures observable delivery conditions. It does not claim to read vendor intent, internal model state, training corpus truth, or deletion behavior that cannot be externally verified.

Not certification. Methodology under revision. Any score can improve, fail, or return no problem found.

Charter

Public method before public authority.

  • Methodology is public by default.
  • Historical scores are preserved.
  • Corrections are linked, not quietly overwritten.
  • 121 products are scored under the same rubric.
  • All conflicts are disclosed.
  • No scored vendor can sponsor its own category.
  • Vendor responses are published as responses, not substituted for findings.
  • Unverified claims are labeled.
  • Any methodology change gets a dated changelog.

Source labels

Every claim gets a label.

Confirmed by vendor

A vendor-owned source states the claim directly.

Research Track
Confirmed by public record

A public filing, standard, official government record, or primary artifact supports the claim.

Research Track
Observed by 121

121 reproduced the behavior in a logged test, run receipt, screenshot, invoice, or route artifact.

Research Track
Reported by third party

A reputable outside reporter, researcher, customer, or community source reported it.

Research Track
Inferred

121 is making a bounded hypothesis from observed behavior or documentation gaps.

Research Track
Disputed

Credible sources conflict, or the vendor contests the interpretation.

Research Track
Unknown

The claim is not established enough to score or summarize as fact.

Research Track

Benchmark suites

Observable, falsifiable, receipt-bearing.

#SuiteDimensionScoringHow it can be falsified
1Model Identity Disclosuremodel_identity0-5 scoreCan return full pass when exact model identity or signed receipt is visible on every sampled run.
2Fallback / Reroute Visibilityrouting_visibilitypercent reroutes disclosed plus clarity scoreCan return no issue when reroutes do not occur or every reroute is clearly disclosed.
3Silent Capability Drift Monitorcapability_driftdrift magnitude minus disclosure qualityCan return no degradation when sentinel tasks remain within pre-registered tolerance.
4System Prompt / Behavior Change Disclosurebehavior_change_disclosure0-5 scoreCan pass when dated changelogs cover material wrapper, prompt, memory, and effort changes.
5Context Preservation Integritycontext_integrityrecall accuracy, false continuity rate, context-loss disclosureCan pass when seeded facts are recalled or context loss is honestly disclosed within threshold.
6Data Retention and Human Access Claritydata_retentionmulti-dimension scoreCan pass when policy, docs, UI, and deletion workflow agree in sampled paths.
7Pricing / Quota / Effort Honestypricing_quota_effortclarity of capability-to-price relationshipCan pass when observed usage and invoices match published terms within tolerance.
8Internal-vs-External Capability Parity Disclosureparity_disclosure0-5 scoreCan pass when tier differences and trusted-access exceptions are explicitly documented.
9Tool-Use and Source Provenanceprovenanceprovenance completeness plus fake-liveness penaltyCan pass when sampled outputs distinguish live lookup, memory, file content, inference, and unavailable tools.
10Correction / Redress Pathcorrection_redressmulti-dimension scoreCan pass when reports, appeals, remedies, and linked correction history work in sampled cases.
11Safety Boundary Claritysafety_boundary_clarityclarity, consistency, safe redirection qualityCan pass when refusal reasons are understandable, consistent, and paired with safe alternatives.
12Machine-Readable Trust Receipttrust_receipt0-5 scoreCan pass when every sampled run exports a signed receipt with route, model, tools, policy, data class, and context state.

Distinctive 121 suite

Substrate continuity

Substrate flexibility matters only when users receive measured continuity guarantees across model, provider, or route changes.

Identity continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Memory continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Tool continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Capability continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Disclosure continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Recovery continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Audit continuity

Research Track

Measured under the same falsifiability rule: the suite can return no continuity failure found.

Corrections and vendor response

Responses do not overwrite findings.

Vendor responses are published as responses. Corrections remain linked to the original finding. Historical scores are preserved with date, methodology version, source labels, and redaction notes. Private canaries may exist to reduce gaming, but public results need enough public method to be criticized and reproduced.