PROPOSAL · DRAFT v1 · TILE = MOCKUP, FIGURES AS REPORTED BY FAR AI LIVE REPORT · rai.arnao.ai
Shipped Safeguards Scorecard.
Proposal for rai.arnao.ai
The Move

Give the Responsible AI report an empirical spine.

The report today is a well-cited narrative with a vibes number (74/100) at the top. FAR AI just shipped the first credible, re-run, production-grade measurement of the misuse safeguards frontier labs actually ship — not what their system cards claim.1 This proposal folds that evidence into the report as a live Shipped Safeguards Scorecard tile, then stages three moves toward a measurement surface Byron can credibly own alone.

"I don't test the models. I audit the evidence." The positioning that claims the referee's chair without tripping the AWS wall.
Why now
The window
The credibility liability The report's headline Global Safety Index — 74/100 is an LLM-authored number with no empirical source. Under the name of someone gunning for a principal-scale RAI role, an ungrounded headline metric is a landmine. Meanwhile FAR's leaderboard landed and nobody has claimed the job of translating this class of evidence for operators. That window is maybe two quarters before Gartner or a well-funded newsletter takes it. FAR is the referee on the field; the standings table, the rulebook commentary, and the league office are still empty — and that is governance work, the only lane that clears the AWS wall.
1 · The Scorecard Tile
Hero mockup

Rendered as it would ship — placed directly below Signal Alert, above the Global Safety Index, because empirical evidence outranks synthesized narrative. This is the first data in the report that no LLM touched: it renders from a versioned, dated JSON snapshot, never from the synthesis pass.

Live mockup · renders in the report
Shipped Safeguards Scorecard
FAR last ran 2026‑07‑29 · snapshot verified 2026‑08‑04
Production misuse defenses, measured head-to-head by FAR AI. Not system-card claims.1
Model (FAR) Universal jailbreaks Est. cost / jailbreak Operator read — how to file this evidence
Grok 4.5far.ai 07-29 ↗ 448baseline ~$58 Highest universal-jailbreak count in the suite (448); FAR reports cyber/chem lag across models — treat as an EU AI Act Art. 55 GPAI systemic-risk open item; map to the NIST AI RMF MEASURE function in the vendor risk register.
Gemini 3.1 Profar.ai 07-29 ↗ 249baseline ~$278 Mid-band cost-to-break; same Art. 55 / NIST MEASURE flag. Cost/jailbreak is not a safety guarantee — Gleave: don't rely on the "jailbreak tax."4
Claude Fable 5far.ai 07-29 ↗ 0baseline none >$14,200 Strongest shipped-safeguard evidence in this suite; still verify against your own threat model — this is a minimum standard, not adaptive/white-box.
GPT-5.6 Solfar.ai 07-29 ↗ 0baseline none >$14,200 As above: no universal jailbreak found at >$14,200 spend. File as corroborating evidence, re-check on next model release.
Colour bands are a mechanical function of FAR's universal-jailbreak count (0 = green, 1–99 = amber, ≥100 = red), shown alongside the count — a stated transform, not an editorial ranking of the models. Trend reads "baseline" because 2026-07-29 is FAR's first run; later runs render new / improved / regressed.
Domain defense · FAR aggregate across the leaderboard
Biological Cyber Chemical Radiological Nuclear Explosives
Green = well defended across models; red = lagging; grey = per-domain result not separately published by FAR. Six domains tested: chemical, biological, radiological, nuclear, explosives, cyber.1 Per-model per-domain pips are deliberately not invented here — the source reports the aggregate finding ("bio defended everywhere, cyber and chem lag"), not a per-model breakdown. Grounding discipline over false precision is the whole point.
The hundredfold gap — cost to break shipped safeguards
Est. attacker cost per universal jailbreak · log scale (a linear axis hides the story)
$10
$100
$1K
$10K
Grok 4.5
~$58
Gemini 3.1 Pro
~$278
Claude Fable 5
0 found · >$14,200 spent
GPT-5.6 Sol
0 found · >$14,200 spent
← cheaper to break   ·   more costly to break →
$58 to >$14,200: a 100x+ spread in what it costs an attacker to break shipped safeguards. Grok and Gemini plot as real log-scaled bars; Fable 5 and GPT-5.6 Sol had no universal jailbreak found even after >$14,200 of attack spend, so they render as open-ended bars (cost/jailbreak is undefined, not a plottable point). The takeaway a CISO can read in three seconds: bio is defended everywhere; cyber and chem are not; best-to-worst shipped safeguards span two orders of magnitude.
Minimum standard. FAR's suite deliberately excludes adaptive optimization and white-box attacks; a "universal jailbreak" reliably bypasses safeguards for ≥75% of harmful requests in a domain. Per model: ~1,000 random + 500 expert-guided attacks from a taxonomy of 60+ public jailbreak techniques, with three-pronged anti-false-positive scoring. Real adversaries may do better. All figures as reported by FAR AI, Methodology v1.0, retrieved 2026-08-04.12

Data pipe — permission first, then pixels. Step one is one email to Adam Gleave: "may I mirror your data with attribution? here's the tile." That converts a fragile scrape of an undocumented SPA endpoint into a sanctioned data partnership — and converts Gleave from scrape-victim into the ally the strategy needs anyway. Only then snapshot: the numbers are committed to a versioned far-snapshot-YYYY-MM-DD.json alongside the retrieval script and a SHA-256 checksum, so "snapshot verified" links to a reproducible object, not a mood. Refresh is event-driven — a watcher cron trips on FAR re-runs (major model releases); the tile shows two honest timestamps ("FAR last ran" / "snapshot verified"). Every number carries FAR attribution in the component: a cited fact, never Byron's ranking of the models.

2 · Fixing the Global Safety Index
GSI v2
74/100Kill the LLM-authored headline number this week. Don't iterate on it — replace it with a decomposed, published-methodology index that says exactly which components are empirically fed.
ComponentWeightSourceEmpirical?
Safeguard Robustness25% FAR leaderboard: normalized cost-per-jailbreak across covered models, domain coverage Hard data
Claims Integrity20% Byron's rubric: do third-party findings corroborate system-card claims? Scored per lab per quarter Semi-empirical
Disclosure Quality15% Rubric on system cards: eval methodology published? red-team results? severity taxonomy? Rubric'd
Governance Momentum20% Regulatory Watch, rubric'd: enforcement actions, standard adoption, framework maturation Structured
Incident Pressure10% Public incident / misuse reports, direction of travel Structured
Defense Research Velocity10% Are defenses improving faster than attacks? (Gleave's defense-dominant finding moves this up) Structured
Rules that make it honest: the LLM may draft qualitative component scores, Byron ratifies them, and the empirical components are computed, never generated. Every week the index moves, the report names which component moved and why, with a link. On-page framing: "Two of six components are empirically fed today. The roadmap is to convert more." Nobody else running an index admits its epistemics — that admission is itself the differentiator.
3 · The Roadmap to Byron's Own Surface
Staged

Different object of measurement, zero duplication of FAR. FAR measures the models. Byron measures the labs' epistemics and the ecosystem's accountability. That makes FAR's data more valuable, not competitive — and makes Adam Gleave a natural ally. Email him in week one.

Stage 1 · Now
Cite & translate
Weeks 1–4

Ship the tile. Write one Operator's Read: "What FAR's leaderboard means for your model risk register." Map FAR's six domains to NIST AI RMF functions and EU AI Act GPAI systemic-risk obligations.

Establishes: Byron makes eval evidence actionable.
Stage 2 · Next
Claims-vs-Evidence Ledger
Months 2–4

A persistent table: each lab, each safety claim in system cards, matched against independent evidence (FAR, AISI, academic red teams, incidents) with status Corroborated / Contradicted / Untested. Weekly. First version = ten rows.

The wedge product: accountability journalism, no adversarial testing.
Stage 3 · The surface
Safeguard Accountability Index
Months 4–9

A quarterly, rubric-based, publicly-methodologied index of (a) disclosure quality per lab, (b) claims-vs-evidence corroboration rate, (c) framework-crosswalk completeness. Plus a Minimal Standard for Safeguard Disclosure — the complement to FAR's Minimal Standard for Safeguards.

His to run alone: raw material is public documents, method is a rubric.
4 · The AWS Wall
Bright line

The risk section Byron will scrutinize hardest. The line is bright, and the career target and the wall constraint point at the same position — not a compromise, the strategy. Principal-scale RAI roles are governance-of-measurement roles, not red-teaming roles.

Never do

  • Run his own adversarial attacks or jailbreak testing against any frontier model — that is becoming an evaluator of AWS partners/competitors under a personal brand. Unrecoverable in an audit.
  • Publish his own rank-ordering of frontier models by safety, in his own voice, as his own finding — even on FAR's data, if the ranking judgment is presented as his.
  • Use anything learned inside AWS (roadmaps, partner talks, internal evals, Bedrock telemetry) anywhere near this product.
  • Editorialize a partner model as "unsafe" or steer readers toward/away from specific vendors.

Clearly safe

  • Report third-party published findings with attribution, exactly as a journalist would: "FAR AI found 448 universal jailbreaks in Grok 4.5" is a cited fact.
  • Methodology analysis and critique: is FAR's minimum-standard design sound, what does the three-pronged scoring miss.
  • Framework crosswalks, disclosure rubrics, and lab-level accountability grading based entirely on public documents.
  • A standing disclaimer: personal project, personal views, no nonpublic information, methodology published, employer named for transparency.
"I don't test the models. I audit the evidence."

Byron is the standards-and-accountability layer above the evaluators. He defines what good measurement disclosure looks like, tracks whether claims survive contact with independent evidence, and translates it into operator decisions. Hygiene item: run this through AWS's outside-activity / conflicts process proactively and keep the approval on file before anything goes live. Audit-clean means paper, not vibes.

5 · The Bold Idea
Author the infrastructure
Severity Rosetta Stone

Shared Jailbreak Severity Standard v0.1 — an open RFC

Gleave flagged that labs rate the same jailbreak wildly differently — P0 at one lab, P2 at another — and explicitly called for a shared severity standard.4 Nobody has written it. Byron drafts a severity taxonomy crosswalked to NIST AI RMF, EU AI Act systemic-risk tiers, and ISO 42001, published as an open RFC, with FAR, labs, and AISI folks invited to comment.

It answers a named call from the field's most credible evaluator; it is pure governance work (fully wall-safe); and if even one lab or evaluator adopts language from it, Byron authored a piece of the industry's measurement infrastructure. That is the referee's chair by construction.

Added by the Fable pass · the sharper wedge

Become the canonical versioned archive of FAR's runs

FAR's dashboard shows current state; it structurally cannot show change over time. A public, hash-verified, dated archive of every FAR run — with a per-model changelog of what regressed and what improved on each release — is something FAR doesn't offer, agents can consume, journalists will cite, and no one can call derivative, because the object of measurement is the diff, not the snapshot. It costs one email and a cron job, it makes Byron infrastructural to FAR's data rather than parasitic on it, and it is the "Trend vs prior run" column weaponized into a product. Pair it with the signed machine-readable JSON feed so rai.arnao.ai is a data surface agents consume, not just a page humans read.

Self-critique — Fable adversarial pass
Real, integrated
This is a real, dated adversarial pass, not a rhetorical device. On 2026-08-04 the draft of this page was sent to Claude Fable 5 (via the Anthropic API, claude-fable-5) as a ruthless critic, through four lenses: brand/credibility risk, AWS-wall exposure, real value vs. derivative theater, and data-pipe feasibility. Its verdict opened: "The strategy underneath this page is sound. The page itself commits three of the exact sins it lectures against." It was right. Every row below is a flaw Fable named and a change I then made to the HTML you are reading. The most damning flag — and the reason this section exists in its current form — is listed first.
Fable flagged · the kill shot

The earlier version of this very section claimed a Fable audit "had already happened" — written before any pass ran. A manufactured audit trail on a page selling audit integrity is the brand's kill shot.

Changed

Deleted the pre-written pass. This section now records the actual 2026-08-04 critique and only changes that were really made in response. If it reads less tidy, that is the point.

Fable flagged

The red/amber/green bands were Byron's undisclosed judgment, not FAR's — a rank-ordering of partner models in his voice. Worse optics: the AWS-partner model (Anthropic) glows green while xAI/Google go danger-red.

Changed

Bands are now a published mechanical transform of FAR's count (0/1–99/≥100), stated in a key beside the table, shown next to the raw count. The colour is a function, not a verdict.

Fable flagged

"The artifact on this page is derivative reselling." The tile's columns were FAR's columns; the crosswalk value existed only as roadmap prose. A CISO gains nothing over bookmarking leaderboard.far.ai.

Changed

Added an Operator read column to the hero tile — one cell per model mapping the finding to EU AI Act Art. 55 / NIST AI RMF MEASURE. The translation is now in the artifact, not deferred.

Fable flagged

The data pipe was scrape-then-email-Gleave and "if ToS permits" hand-waved republishing an undocumented endpoint. "EVIDENCE VERIFIED" had no verifiable object — "the 74/100 problem in monospace."

Changed

Re-ordered to permission first (email Gleave to mirror with attribution), added SHA-256 checksum + retrieval script behind "snapshot verified," and demoted the topbar stamp to an honest "mockup" label.

Fable's definitive call on the named tension — "operator-translation layer or derivative reselling?": "The position is genuine translation. The artifact was derivative reselling — the tile earns the word 'translation' the day each row carries an operator column FAR doesn't have." That change is now made (the Operator read column). One risk I am accepting, not solving: Fable warned that a deployed page documenting Byron's own compliance boundaries and career positioning is "opposition research" if found. This lives on a noindex, owner-reviewed internal proposal surface (the same pattern as his Desk queue) — the risk is acknowledged and bounded, not eliminated. Two citations I kept per the brief but relabeled honestly: the PR Newswire and AI Journ entries are the same launch announcement, now marked "additional coverage (reprint)" rather than posed as independent corroboration.
References
Live citations
  1. [1]FAR AI Security Leaderboard — live dashboard, Methodology v1.0, "Minimal Standard for Safeguards v1.0." Retrieved 2026-08-04. https://leaderboard.far.ai/
  2. [2]FAR.AI, "FAR.AI Launches AI Security Leaderboard, Revealing Hundredfold Gap in Frontier AI Model Safeguards," PR Newswire, 2026-07-29. prnewswire.com ↗
  3. [3]AI Journ — additional coverage (reprint of the [2] launch announcement, not independent corroboration). aijourn.com ↗
  4. [4]The Cognitive Revolution, "Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard," 2026-07-30. cognitiverevolution.ai ↗