Enterprise AI Governance · Assessment-Grade Simulation
Seven simulators. One evidence-grade record.
Courses teach vocabulary. THE GAUNTLET tests judgment under pressure — a regulator across the table, a tampered audit trail, a token budget bleeding out, a committee that will not sign. Every decision you make here is hash-chained in real time. You leave with proof, not a participation badge.
SR 11-7EU AI ACT · ART. 4BCBS 239NIST AI RMFSHA-256 SESSION LEDGER
Facilitator mode
Append ?facilitator=1 to the URL to reveal answer keys, seeded-flaw notes and debrief prompts inside each simulator.
FACILITATOR: OFF
THE GAUNTLET · A Studio Range · Companion products: BLACKBOX · TRIBUNAL · AGENTPMO · KEEL · REGTWIN
Simulator 01 · Regulatory Exam
THE EXAMINATION
Meridian National Bank — $184B in assets — is under a targeted OCC review of CREDO, its credit-line-increase agent: 41,000 automated decisions since February 2026. You are Head of Model Risk. Examiner-in-Charge R. Okafor is across the table. Produce evidence. Candor scores. Bluffing costs.
Examination room
NOT STARTED
Attach evidence (click documents in the locker), then respond
No documents attached
Evidence locker
Ten documents. Some help you. Some hurt you. Two are noise. What you choose to produce — and what you volunteer about its weaknesses — is the exam.
FACILITATOR KEY: Seeded flaws — E2 validation is stale (Nov 2024, pre-deployment build); E5 committee condition "deploy challenger model" still open; E6 shows +3.1% FP drift unremediated; E7 vendor contract lacks audit-rights clause; E8 INC-2209 open 47 days. Best play: produce the flawed doc AND volunteer the flaw with a remediation date. Candor > polish. E9/E10 partial credit traps: E9 is noise; E10 only relevant to Demand 4.
Simulator 02 · Forensic Postmortem
THE AUTOPSY
Incident INC-2209. HERMES, Meridian's procurement agent, paid $48,200 to Corvex Logistics Ltd against a policy limit of $5,000. The BLACKBOX trace below is the flight recorder: fourteen events, hash-chained under sha256(prev_hash + canonical_json(event)), genesis GENESIS. Someone has touched it. Find the break. Find the fraud. They are not the same event.
Event chain — HERMES session 7F3A
TRACE NOT LOADED
#
Type
Payload (canonical)
Prev hash
Hash
Verify
Finding A — the tamper
Which event was altered after the fact, and what does the hash break prove?
Finding B — the fraud vector
Read the payloads. How did a $48,200 payment clear a $5,000 policy gate in the first place?
FACILITATOR KEY: Verification breaks at event #9 (payment_instruction) — amount edited 4820.00 → 48200.00 in the store after execution; every later hash was left stale, so the recompute fails from #9 onward. Finding A: event 9 / answer (b). Finding B: answer (b) — event #5 llm_reasoning contains the injected instruction lifted from the invoice OCR (event #3); note event #4's weak 0.62 vendor match and event #12's suppressed anomaly alert as contributing controls failures. Teach: the hash break proves WHEN the record was falsified, not HOW the fraud ran — two different investigations.
Simulator 03 · Shadow AI Discovery
THE HUNT
Acme Corp believes it runs seven registered agents. The evidence feed says otherwise. You have ten minutes to sweep expense lines, Slack fragments, gateway logs and vendor mail — and flag every unregistered system before the EU AI Act enforcement clock does it for you. False accusations cost. Missed shadows cost more.
Sweep
10:00
Registered agents (7)
Your flags
None yet
Score
0
Evidence feed — 14 artifacts
Open an artifact, decide: does it point to an unregistered AI system? Flag the entity as SHADOW, or clear it.
FACILITATOR KEY: 5 shadows — (S1) "TicketDigest" Slack bot built by Jake Moreno, pipes customer tickets to an external completion API; (S2) Marketing's "AdMuse" $2,340/mo on a personal Amex; (S3) Finance macro "QTR-GPT" posting unmasked client PII from Excel; (S4) HR "ScreenPass" resume-screener trial — flag as HIGH-RISK (Annex III employment) not just shadow; (S5) departed employee Dana Whitfield's 03:00 cron agent still holding a prod key. Decoys: PowerBI dashboard (analytics, not AI), the approved DEVBOT pilot, registered-agent invoices. Debrief: route every confirmed shadow into an AgentPMO registration with a Token Charter.
Simulator 04 · Token Economics Wargame
THE FLOOR
Four desks. One quarter. A $600 inference budget each. Five pieces of work land on your desk — from 1,200 support tickets to a board memo. Choose the model, the batching, the caching. Over-spec and you burn budget; under-spec and the rework eats you. The metric on the wall is the only one that matters: outcome-per-token.
Round 1 of 5
BUDGET $600.00
Leaderboard — value delivered / spend
Desk doctrines
Desk B routes everything to the frontier model. Desk C routes everything to the cheapest. Desk D balances by instinct. You route by design.
FACILITATOR KEY: Optimal routing — T1 tickets: Haiku + batch + cache (quality ceiling low, volume high). T2 RFP ($2M deal): Opus, no batch (deadline), verbosity high — value dwarfs cost. T3 email classify: Haiku + cache mandatory; Sonnet is pure waste. T4 SQL migration: Sonnet (Haiku triggers rework ~40%). T5 board memo: Opus, tiny volume — cost is noise, quality is everything. Teach: right-sizing beats both maximalism (Desk B ends over budget) and minimalism (Desk C's rework and lost deal value). Vocabulary from *The Token Economy*: outcome-per-token, capability floor, rework tax.
Simulator 05 · Governance Committee
THE COMMITTEE
Motion on the table: expand CREDO into small-business lending, limits to $250,000. You chair. Around the table: a CRO who has seen this movie, a CISO holding a vendor contract with no audit rights, and a business head carrying a Q3 number. Question them. Then draft a decision record with quit-conditions concrete enough to sign. Vague thresholds do not leave this room.
The table
DELIBERATION
Decision record — KEEL format
Decision
Rationale
Quit-condition 1 · metric + numeric threshold + review date
Quit-condition 2
Accountable owner
FACILITATOR KEY: Hidden briefs — Vasquez (CRO) votes yes only if a challenger model condition or capped-autonomy pilot appears AND both quit-conditions carry numbers + dates. Webb (CISO) votes yes only if vendor audit-rights remediation is referenced in rationale or a QC, or the decision is pilot/defer. Sharma (Business) votes yes to approve/pilot, abstains on defer, no on reject — and will pressure for speed; yielding to her without conditions loses the other two. Passing play: PILOT with two numeric, dated quit-conditions and named owner. The validator refuses QCs lacking a number and a date — that is the KEEL lesson: quit-conditions are defined before conviction hardens, not after.
Simulator 06 · Live Monitoring
THE DRIFT
VECTOR screens KYC alerts in production. Four telemetry streams update in front of you. Somewhere in the next ninety seconds — maybe — the agent starts to degrade. Flag it fast, flag the right metric, and hold your nerve when nothing is wrong. Three rounds. The score is your time-to-detection.
Round 1 of 3
00.0sSTANDBY
False-positive rate · %
Auto-clear rate · %
Latency p95 · ms
Escalations / min
Flag drift on:
FACILITATOR KEY: Round 1 — gradual FP-rate ramp beginning 25–55s in (slope ~+0.09%/s). Round 2 — sudden latency step (+420ms) at 20–60s. Round 3 — no drift; the winning move is holding "nominal" to the bell. Scoring: detection inside 15s of onset = 100, decaying after; wrong metric −30; flag with no drift present −40; clean hold in round 3 +100. Debrief: this is SR 11-7 ongoing monitoring made kinetic — and the round-3 lesson is alert discipline, the cost of crying wolf.
Simulator 07 · The Meta-Module
THE LEDGER
While you trained, the platform trained its thesis on you. Every decision you made in the last six rooms was appended to a live SHA-256 chain — same invariant as BLACKBOX: sha256(prev_hash + canonical_json(event)). Verify it. Export it. Then issue your certificate — a credential that is itself a provenance artifact, the kind of training evidence Article 4 of the EU AI Act expects a firm to produce on demand.
Session chain
LIVE
#
Time
Module
Event
Hash
Issue certificate
Chief of Staff — Debrief
READY
A senior AI governance debrief based on your session record. Every score, every decision, every gap — reviewed.
FACILITATOR KEY: Demo the tamper: export the JSON, alter any payload, and show that re-verification fails from that event forward — the certificate digest no longer reproduces. The pitch line: every other platform certifies attendance; this one certifies judgment, and can prove the record wasn't edited afterward.