Tools / Threat modeling / Eval

Diagram threat-model eval report

Twelve checks taken from the Trail of Bits Kubernetes assessment (2019), IETF RFC 6819, and the CNCF Financial User Group Kubernetes threat model. Documents read 22 August 2026. Current prompt pack v3.0.

No model has produced a v3 P-report JSON for the ten gold diagrams. The 47/48 and 44/48 values below are the retained v2.0 comparison. Version 3 adds method denominators that those twelve checks do not measure, so it remains unscored against that comparison.

Scores

Each check is 0 to 4. Twelve checks, 48 points. 4 matches the gold report. 3 is present and usable. 2 is thin. 1 is hinted. 0 is absent. Letter cut: A from 42, A- from 36, B from 29, C from 18, D from 8.

Artifact Points Grade What was scored
Prompt series v3.0 Not rescored Open Requires model runs, method-denominator checks, and SME review on the ten fixtures.
Prompt series v2.0 47 / 48 A What Track A asks the model to write.
Specified JSON v2.0 44 / 48 A What schema.json can hold. Reviewer is a human field, so this column cannot reach 48.
Gold rag-app inventory 8 / 48 D Q1 inventory only. Auspex would not use a manual threat list as gold; these files follow that limit.
Zeroshot stub (rag-app) 9 / 48 D Intentional miss for the scoring scripts. T1 is a catalog line, referent missing, no control point.

The three gold reports

RFC 6819 states attack assumptions and existing features before any threat. Trail of Bits pins a version and tests trust-boundary claims. CNCF writes a mitigation and a validation snippet on each leaf. All three are documents a reviewer can read without a parser.

Per-check scores

Columns: prompts (what the chain asks), schema (what the JSON can hold), gold inventory, zeroshot stub.

Check Prompts Schema Gold inv. Stub How this pack treats it
Named system, version, and scope 4322 P-norm asks for source_id, version, and commit, or unknown. Gold inventory has no version pin.
Attacker capabilities stated 4411 P-adv writes assumptions and positions per zone. Threats must cite attacker_position.
Architecture and trust boundaries 4443 P-norm, P-diag, and P-sol. Gold rag-app has four zones. Stub drops model-api and llm_subset.
In-scope and out-of-scope named 4310 P-scope. Schema stores the arrays. Gold inventory has no scope lists.
Existing security features listed first 4400 P-controls lists only features shown on the diagram, states shown coverage, and records expected controls not shown at a named referent. Empty is allowed only when none_drawn is true.
Each threat names a component or flow 4400 diagram_referent must exist in inventory. Stub T1 uses referent missing.
Preconditions or attacker position 4300 P-stride, P-phantom, and P-dedup require attacker_position from P-adv. Schema field is optional so old fixtures still validate.
Action at a named control point 4401 P-act limits control_point to a component, store, flow, or trust boundary. Mitigate and eliminate also require validation (test, log, or fail_condition).
What the design does not claim to stop 4401 P-adv writes claim_boundary.does_not_claim and box. P-report repeats that list.
Authors, method, and review 3301 P-report sets method and date. Reviewer stays empty for a human. Filling it in the prompt would be a fake sign-off.
A reviewer can read it without a parser 4400 P-report writes a full markdown projection of the matrix (grouped threat tables, every threat/position/control id). Export prompts serialize that stored string; they do not summarize it.
A way to check a mitigation landed 4400 Mitigate and eliminate require action.validation. P-qa checks actions_have_validation.

Verdict on the prompt series

Pack v3.0 requires a review profile and a traditional applicability decision. It uses typed STRIDE, conditional abuse and operational passes, PHANTOM-B, AI-to-traditional path coverage, pinned source manifests, and evidence-backed importance. Track B audits SRF layers. Track C joins vertical obligations. P-report remains the sole author of report.markdown. There is still no gold threat list.

StepTaken fromWhat it writes
P-adv RFC 6819 section 2.2; CNCF scenarios Assumptions, positions per zone, and claim_boundary.
P-controls RFC 6819 section 3; Trail of Bits isolation discussion Authn, TLS, filters, or isolation already drawn, or none_drawn.
P-act validation CNCF leaf tests; Trail of Bits retestable findings test, log, or fail_condition on the named control_point.
P-report All three gold reports Readable projection of the completed matrix: metadata, claim boundary, attacker positions, architecture, controls, review order, one table row per threat grouped by referent, coverage counts, and an empty reviewer. Export prompts emit that stored string.

Machine scores in the repository

Inventory precision and recall, format Jaccard, typed STRIDE, PHANTOM-B, composition paths, source provenance, CVE applicability, importance, SRF layer coverage, workflow fixtures, and schema checks run from eval/threat-model/ without calling a model. Those scripts do not award the letter grades on this page. SME sheets in sme/ are still empty, so closure stays false. Schema: eval/threat-model/schema.json. Prompts: /tools/prompts/threat-model/.

Vendor demo PDFs, fictional STRIDE samples, and Shostack's 2020 PCI reverse-engineering draft were not used. The PCI draft analyzes a standard, not a deployed system.