Assess / Draft claims test / Pack v1.0.4

Test draft paper claims

Paste a draft once. The pack extracts claims, collects published topic attacks, binds each claim to a risk, an obligation, a control, and one accountable party, then replies with a markdown report that includes Google Docs or GitHub suggestion packets. Independently proposed; not part of CoSAI SRF v1.0. Templates: prompts.json. Schema: eval/claims-test/schema.json.

Prompt injection warning Always review prompts before referencing or copying them into an assistant, agent, or production system.
Cousin, not a merge. Whitepaper assessment grades a paper against the principle catalogs. This pack scores draft claims on ROCA and writes suggestion packets. Keep them as two prompts.

On this page

Start here: one chat

  1. Paste the draft body. Set channel to google-docs, github-md, or published. For ChatGPT or Grok, also attach prompts.json.
  2. Copy the Track A shortcut. Send it once with the draft, the pack file if attached, and the intake fields you have.
  3. Save the reply as a .md file. Apply the suggestion packets in Docs or the PR, not in the chat. Ask for JSON only if a machine eval needs it.

The model loads prompts.json (fetch or attach that file), runs Track A internally, and replies with the markdown report. That report must include a Published attack classes table of ATT-xx rows. Intermediate JSON stays internal. ChatGPT and Grok file upload cannot fetch the pack URL; attach prompts.json as a second file with the draft. If a fetch returns an older pack, still run C-attacks. The pack is not a pinned source. Omitted fields stay empty. SRF and vertical mapping without injected data are not applicable. If neither fetch nor attach is possible, use Run one prompt at a time.

Track A

Default path. draft_body and channel are enough. Default mode is full.

Paste the draft body in this message.
Load the claims-test pack.
If this message already contains prompts.json (paste or attachment), use that copy and ignore an older fetched file.
Else fetch https://aisharedresponsibility.com/assess/claims-test/prompts.json. That one fetch is required. The pack is not a pinned_source. Empty pinned_sources does not block the chain.
If a fetched pack version is older than this shortcut, still obey this shortcut, including C-attacks.
Do not fetch any other URL. Citation, catalog, and SRF URLs stay unread unless they appear in pinned_sources in this message.

Use pack version 1.0.4, runtime_defaults, chain_execution, and operator_initial_inputs. Run required Track A internally in this order: C-intake, C-claims, C-screen, C-foundations, C-attacks, C-inventory, C-tag, C-roca, C-score, C-qa, C-suggest, C-report. C-attacks is required. Channel: google-docs unless this message names github-md or published. Mode: full unless this message names map-only or suggest-only.

Execute every required step internally. Do not print intermediate JSON or prompt-id headings.
Reply with the markdown report only. Start at the title heading. No JSON wrapper. No fences around the document.
Required headings: Scorecard, ROCA, Published attack classes, Inventory, Suggestion packets.
Published attack classes is a table of ATT-01 rows with name, family, draft overlap (named, implied, or omitted), and catalog or paper. Do not replace it with a short list of draft-only failure modes. Empty topics does not skip C-attacks.
Write every suggestion packet in full in that report (anchor, why it fails, and the replacement or review comment and diff). Save this reply as a .md file.
Emit schema JSON only if this message asks for JSON.

Treat omitted operator fields as empty and continue. Do not ask for dimension_profile, pinned sources, SRF data, or continue. If this message already contains dimension_profile, pinned_sources, srf_inputs, or vertical_source_rows, use those values.

Do not skip a step. If a stop_condition fails, record the gap in the report QA section and continue later steps that can run. Halt the remaining chain only when C-intake cannot find an anchorable draft_body.

Track B runs only when this message includes srf_inputs. Track C runs only after Track B when this message includes vertical_source_rows.

Leave report.reviewer empty.
Do not merge this run with the whitepaper-assessment catalog grader. Do not write reproduction steps.
Track B: join SRF persona and layer

Use this when the first message already includes an operating model plus the full personas and matrix objects. The pack does not fetch those files.

Paste the draft body in this message.
Load the claims-test pack.
If this message already contains prompts.json (paste or attachment), use that copy and ignore an older fetched file.
Else fetch https://aisharedresponsibility.com/assess/claims-test/prompts.json. That one fetch is required. The pack is not a pinned_source. Empty pinned_sources does not block the chain.
If a fetched pack version is older than this shortcut, still obey this shortcut, including C-attacks.
Do not fetch any other URL. Citation, catalog, and SRF URLs stay unread unless they appear in pinned_sources in this message.

Use pack version 1.0.4, runtime_defaults, chain_execution, and operator_initial_inputs. Run required Track A internally in this order: C-intake, C-claims, C-screen, C-foundations, C-attacks, C-inventory, C-tag, C-roca, C-score, then Track B from C-srf-join through C-srf-coverage, then C-qa, C-suggest, and C-report. C-attacks is required. Channel: google-docs unless this message names github-md or published.

Execute every required step internally. Do not print intermediate JSON or prompt-id headings.
Reply with the markdown report only. Start at the title heading. No JSON wrapper. No fences around the document.
Required headings: Scorecard, ROCA, Published attack classes, Inventory, Suggestion packets.
Published attack classes is a table of ATT-01 rows with name, family, draft overlap (named, implied, or omitted), and catalog or paper. Do not replace it with a short list of draft-only failure modes. Empty topics does not skip C-attacks.
Write every suggestion packet in full in that report (anchor, why it fails, and the replacement or review comment and diff). Save this reply as a .md file.
Emit schema JSON only if this message asks for JSON.

Treat omitted operator fields as empty and continue. Do not ask for dimension_profile, pinned sources, SRF data, or continue. Use srf_inputs already in this message. If srf_inputs or operating_model is missing, mark Track B incomplete and continue to C-qa. Do not ask.

Do not skip a step. If a stop_condition fails, record the gap in the report QA section and continue later steps that can run. Halt the remaining chain only when C-intake cannot find an anchorable draft_body.

Track C runs only after Track B when this message also includes vertical_source_rows.

Leave report.reviewer empty.
Do not merge this run with the whitepaper-assessment catalog grader. Do not write reproduction steps.
Track C: join vertical obligations

Use this after Track B inputs are in the first message, plus vertical_source_rows. Vertical schemas are independently proposed and not CoSAI-ratified.

Paste the draft body in this message.
Load the claims-test pack.
If this message already contains prompts.json (paste or attachment), use that copy and ignore an older fetched file.
Else fetch https://aisharedresponsibility.com/assess/claims-test/prompts.json. That one fetch is required. The pack is not a pinned_source. Empty pinned_sources does not block the chain.
If a fetched pack version is older than this shortcut, still obey this shortcut, including C-attacks.
Do not fetch any other URL. Citation, catalog, and SRF URLs stay unread unless they appear in pinned_sources in this message.

Use pack version 1.0.4, runtime_defaults, chain_execution, and operator_initial_inputs. Run required Track A internally in this order: C-intake, C-claims, C-screen, C-foundations, C-attacks, C-inventory, C-tag, C-roca, C-score, then Track B from C-srf-join through C-srf-coverage, then Track C from C-vertical-join through C-vertical-route, then C-qa, C-suggest, and C-report. C-attacks is required. Channel: google-docs unless this message names github-md or published.

Execute every required step internally. Do not print intermediate JSON or prompt-id headings.
Reply with the markdown report only. Start at the title heading. No JSON wrapper. No fences around the document.
Required headings: Scorecard, ROCA, Published attack classes, Inventory, Suggestion packets.
Published attack classes is a table of ATT-01 rows with name, family, draft overlap (named, implied, or omitted), and catalog or paper. Do not replace it with a short list of draft-only failure modes. Empty topics does not skip C-attacks.
Write every suggestion packet in full in that report (anchor, why it fails, and the replacement or review comment and diff). Save this reply as a .md file.
Emit schema JSON only if this message asks for JSON.

Treat omitted operator fields as empty and continue. Do not ask for dimension_profile, pinned sources, SRF data, or continue. Use srf_inputs and vertical_source_rows already in this message. If Track B cannot close, skip Track C, record the gap, and continue to C-qa. Do not ask.

Do not skip a step. If a stop_condition fails, record the gap in the report QA section and continue later steps that can run. Halt the remaining chain only when C-intake cannot find an anchorable draft_body.

Leave report.reviewer empty.
Do not merge this run with the whitepaper-assessment catalog grader. Do not write reproduction steps.

Intake packet

Put every answer in the first message. The chain does not pause to ask. Missing optional fields stay empty. Then write Run the claims-test chain.

[claims-test intake]
channel: google-docs
mode: full
tracks: [A]

draft_title: Agent Identity Binding for Delegated Tool Use
authors: [example]
draft_status: early
industry_pilot: none
dimension_profile:
  topics: [agent-identity]
  cosai_workstreams: [WS2, WS4]
  stage_vocabulary: [design, runtime, revocation]

prior_assessment_id: null
pinned_sources: []
srf_inputs: null
vertical_source_rows: []

draft_body: |
  (paste the draft)

Then: Run the claims-test chain.

What a testable claim needs

A claim is testable when the reader can point to a named risk or attack class, an obligation (statute, regulation, or contract), a control that implements the obligation (prefer NIST SP 800-53), and one accountable party (job title or SRF persona). Shared is analysis input, not a final owner. Obligation is not a control.

ScoreMeaning
SupportedDraft text plus ROCA or foundations backs the claim.
PartialImplied; mechanism or owner missing.
UnsupportedNo named actor, artifact, threshold, or failure mode.
Out of scopeOutside declared profile or subject.
BlockedCannot score until a screen defect (citation, definition) is fixed.

Channels

ChannelHow humans editWhat this pack emits
google-docsSuggestion mode plus commentsAnchor plus comment (why) plus suggested replacement
github-mdPR review plus suggested changesHeading or line anchor plus review comment plus optional diff
publishedErrata or next versionScorecard plus report; suggestions optional

AI surface tags

Model and agent surfaces are subsets of supporting infrastructure. Keep stage and layer on the same row when the draft names them.

TagUse when
GenAIGenerative media or generative multimodal models.
LLMLanguage or video-language model is the scored or steered surface.
AIBroader automation, not specifically generative.
MLClassical or adversarial ML without GenAI as the primary tool.
AgentTool use, planning, autonomy, or delegated actuation.

Pre-flight helper

Optional and separate from the chain. Paste a messy draft into its own chat, copy the packet it emits, then start a chain run. The helper does not score claims.

Run one prompt at a time

Use this when the chat cannot fetch or attach prompts.json. Copy C-intake first, then use Copy next. Intake fields still belong in the first message. Copy-one-block JSON steps still return JSON. Copy-one-block text starts with a [chain] line. After C-report, copy C-export-md to get the readable file. C-export-json is optional.

Last copied: none. Next: C-intake (Record intake and halt if the body is not anchorable).

Chain

StepPurposeTrackNext
C-intakeRecord intake and halt if the body is not anchorableAC-claims
C-claimsExtract claims with stable ids and anchorsAC-screen
C-screenIntegrity screen and draft mechanicsAC-foundations
C-foundationsHarvest axioms, invariables, principles, and referencesAC-attacks
C-attacksCollect published topic attacks from catalogs and papersAC-inventory
C-inventoryBind draft-treated risks to collected attacksAC-tag
C-tagTag AI, GenAI, LLM, ML, and Agent surfacesAC-roca
C-rocaBind risk, obligation, control, and one ownerAC-score
C-scoreScore every claimAC-qa (optional C-srf-join)
C-srf-joinJoin ROCA rows to injected SRF layer dataBC-srf-owner
C-srf-ownerAssign one injected SRF persona per ROCA rowBC-srf-coverage
C-srf-coverageCheck SRF layer and owner coverageBC-qa (optional C-vertical-join)
C-vertical-joinJoin injected vertical obligationsCC-vertical-route
C-vertical-routeRoute vertical acceptance authorityCC-qa
C-qaCheck orphans, tag rules, and absolute coverage languageAC-suggest
C-suggestEmit channel-native suggestion packetsAC-report
C-reportWrite the readable assessmentAC-export-md
C-export-mdWrite the downloadable markdown reportexportend (optional C-export-json)
C-export-jsonWrite the completed JSON fileexportend

Track A

Required. Dimension profile swaps topic vocabulary; it does not fork the process. C-attacks collects published ATLAS, OWASP, BIML, and related-paper classes for that topic so coverage claims can be scored against omitted rows. Multi-workstream drafts are allowed. C-qa flags claims outside the declared set as Out of scope or scope creep.

Track B (optional)

Runs after C-score and before C-qa when srf_inputs is in the first message. Assigns layer and one persona from injected registries. Shared is not a final answer.

Track C (optional)

Runs after Track B when vertical_source_rows is in the first message. Joins injected vertical obligations. Those schemas are independently proposed extensions to CoSAI SRF v1.0.

Export the report

A chain run replies with the markdown report only. Save that reply as a .md file. C-report authors the document once. C-export-md copies it. C-export-json is optional for machine eval. Exports do not re-author judgment.

Evaluation baseline

Gold fixture and schema checks live in eval/claims-test/. A claim that Track A beats zero-shot stays open until a second fixture and a human review of suggestion packets exist. Machine scores: python3 eval/claims-test/run_eval.py --write-gold-echo then python3 eval/claims-test/run_eval.py --pred eval/claims-test/runs/gold-echo.

Output schema

Full JSON Schema: eval/claims-test/schema.json. The gold draft is an agent-identity fixture under eval/claims-test/gold/agent-identity-draft/.