The report today is a well-cited narrative with a vibes number (74/100) at the top. FAR AI just shipped the first credible, re-run, production-grade measurement of the misuse safeguards frontier labs actually ship — not what their system cards claim.1 This proposal folds that evidence into the report as a live Shipped Safeguards Scorecard tile, then stages three moves toward a measurement surface Byron can credibly own alone.
Rendered as it would ship — placed directly below Signal Alert, above the Global Safety Index, because empirical evidence outranks synthesized narrative. This is the first data in the report that no LLM touched: it renders from a versioned, dated JSON snapshot, never from the synthesis pass.
| Model (FAR) | Universal jailbreaks | Est. cost / jailbreak | Operator read — how to file this evidence |
|---|---|---|---|
| Grok 4.5far.ai 07-29 ↗ | 448baseline | ~$58 | Highest universal-jailbreak count in the suite (448); FAR reports cyber/chem lag across models — treat as an EU AI Act Art. 55 GPAI systemic-risk open item; map to the NIST AI RMF MEASURE function in the vendor risk register. |
| Gemini 3.1 Profar.ai 07-29 ↗ | 249baseline | ~$278 | Mid-band cost-to-break; same Art. 55 / NIST MEASURE flag. Cost/jailbreak is not a safety guarantee — Gleave: don't rely on the "jailbreak tax."4 |
| Claude Fable 5far.ai 07-29 ↗ | 0baseline | none >$14,200 | Strongest shipped-safeguard evidence in this suite; still verify against your own threat model — this is a minimum standard, not adaptive/white-box. |
| GPT-5.6 Solfar.ai 07-29 ↗ | 0baseline | none >$14,200 | As above: no universal jailbreak found at >$14,200 spend. File as corroborating evidence, re-check on next model release. |
Data pipe — permission first, then pixels. Step one is one email to Adam Gleave: "may I mirror your data with attribution? here's the tile." That converts a fragile scrape of an undocumented SPA endpoint into a sanctioned data partnership — and converts Gleave from scrape-victim into the ally the strategy needs anyway. Only then snapshot: the numbers are committed to a versioned far-snapshot-YYYY-MM-DD.json alongside the retrieval script and a SHA-256 checksum, so "snapshot verified" links to a reproducible object, not a mood. Refresh is event-driven — a watcher cron trips on FAR re-runs (major model releases); the tile shows two honest timestamps ("FAR last ran" / "snapshot verified"). Every number carries FAR attribution in the component: a cited fact, never Byron's ranking of the models.
| Component | Weight | Source | Empirical? |
|---|---|---|---|
| Safeguard Robustness | 25% | FAR leaderboard: normalized cost-per-jailbreak across covered models, domain coverage | Hard data |
| Claims Integrity | 20% | Byron's rubric: do third-party findings corroborate system-card claims? Scored per lab per quarter | Semi-empirical |
| Disclosure Quality | 15% | Rubric on system cards: eval methodology published? red-team results? severity taxonomy? | Rubric'd |
| Governance Momentum | 20% | Regulatory Watch, rubric'd: enforcement actions, standard adoption, framework maturation | Structured |
| Incident Pressure | 10% | Public incident / misuse reports, direction of travel | Structured |
| Defense Research Velocity | 10% | Are defenses improving faster than attacks? (Gleave's defense-dominant finding moves this up) | Structured |
Different object of measurement, zero duplication of FAR. FAR measures the models. Byron measures the labs' epistemics and the ecosystem's accountability. That makes FAR's data more valuable, not competitive — and makes Adam Gleave a natural ally. Email him in week one.
Ship the tile. Write one Operator's Read: "What FAR's leaderboard means for your model risk register." Map FAR's six domains to NIST AI RMF functions and EU AI Act GPAI systemic-risk obligations.
A persistent table: each lab, each safety claim in system cards, matched against independent evidence (FAR, AISI, academic red teams, incidents) with status Corroborated / Contradicted / Untested. Weekly. First version = ten rows.
A quarterly, rubric-based, publicly-methodologied index of (a) disclosure quality per lab, (b) claims-vs-evidence corroboration rate, (c) framework-crosswalk completeness. Plus a Minimal Standard for Safeguard Disclosure — the complement to FAR's Minimal Standard for Safeguards.
The risk section Byron will scrutinize hardest. The line is bright, and the career target and the wall constraint point at the same position — not a compromise, the strategy. Principal-scale RAI roles are governance-of-measurement roles, not red-teaming roles.
Byron is the standards-and-accountability layer above the evaluators. He defines what good measurement disclosure looks like, tracks whether claims survive contact with independent evidence, and translates it into operator decisions. Hygiene item: run this through AWS's outside-activity / conflicts process proactively and keep the approval on file before anything goes live. Audit-clean means paper, not vibes.
Gleave flagged that labs rate the same jailbreak wildly differently — P0 at one lab, P2 at another — and explicitly called for a shared severity standard.4 Nobody has written it. Byron drafts a severity taxonomy crosswalked to NIST AI RMF, EU AI Act systemic-risk tiers, and ISO 42001, published as an open RFC, with FAR, labs, and AISI folks invited to comment.
It answers a named call from the field's most credible evaluator; it is pure governance work (fully wall-safe); and if even one lab or evaluator adopts language from it, Byron authored a piece of the industry's measurement infrastructure. That is the referee's chair by construction.
FAR's dashboard shows current state; it structurally cannot show change over time. A public, hash-verified, dated archive of every FAR run — with a per-model changelog of what regressed and what improved on each release — is something FAR doesn't offer, agents can consume, journalists will cite, and no one can call derivative, because the object of measurement is the diff, not the snapshot. It costs one email and a cron job, it makes Byron infrastructural to FAR's data rather than parasitic on it, and it is the "Trend vs prior run" column weaponized into a product. Pair it with the signed machine-readable JSON feed so rai.arnao.ai is a data surface agents consume, not just a page humans read.
The earlier version of this very section claimed a Fable audit "had already happened" — written before any pass ran. A manufactured audit trail on a page selling audit integrity is the brand's kill shot.
Deleted the pre-written pass. This section now records the actual 2026-08-04 critique and only changes that were really made in response. If it reads less tidy, that is the point.
The red/amber/green bands were Byron's undisclosed judgment, not FAR's — a rank-ordering of partner models in his voice. Worse optics: the AWS-partner model (Anthropic) glows green while xAI/Google go danger-red.
Bands are now a published mechanical transform of FAR's count (0/1–99/≥100), stated in a key beside the table, shown next to the raw count. The colour is a function, not a verdict.
"The artifact on this page is derivative reselling." The tile's columns were FAR's columns; the crosswalk value existed only as roadmap prose. A CISO gains nothing over bookmarking leaderboard.far.ai.
Added an Operator read column to the hero tile — one cell per model mapping the finding to EU AI Act Art. 55 / NIST AI RMF MEASURE. The translation is now in the artifact, not deferred.
The data pipe was scrape-then-email-Gleave and "if ToS permits" hand-waved republishing an undocumented endpoint. "EVIDENCE VERIFIED" had no verifiable object — "the 74/100 problem in monospace."
Re-ordered to permission first (email Gleave to mirror with attribution), added SHA-256 checksum + retrieval script behind "snapshot verified," and demoted the topbar stamp to an honest "mockup" label.
| What | Why | Link |
|---|---|---|
| FAR AI Security Leaderboard | The live dashboard — the source every number in the tile snapshots from. | leaderboard.far.ai ↗ |
| Current RAI report | The live report this tile folds into — see the 74/100 headline being replaced. | rai.arnao.ai ↗ |
| FAR press release | "Hundredfold gap in frontier AI model safeguards" — the launch citation. | PR Newswire ↗ |
| The Cognitive Revolution podcast | Adam Gleave on the leaderboard — source for the defense-dominant + severity-standard claims. | cognitiverevolution.ai ↗ |