Assess

Whitepaper assessment changelog

Prompt versions 2.2, 2.1, and 2.0. What each version added or demoted, and which older reports stay valid. The live prompt is v2.2.

What changed in 2.2

Eleven steps became nine. Intake now runs before any catalog fetch. Novel stays a map result the user can challenge, and is no longer a form for adding entries to an admin-managed catalog. An assessment produced under 2.1 stays valid; it fetched full catalogs, emitted every Matched row, and packaged Novel rows for catalog addition.

ChangeReason
Intake before catalog fetch A truncated PDF already had 217 ids in context, which is how a partial read manufactures coverage gaps. Stop first. Fetch slim catalogs only if the read is whole.
Prompt fetches slim catalogs Full JSON carries src tables the mapping pass does not need. Slim files keep id, statement, category, related or section, and gap_index. Full files remain for a src check.
Extract, map, and classify merged Three numbered steps rewrote the same table twice. One pass per section, at most 20 statements, and most sections have fewer. The cap was being read as a quota.
Matched rows collapsed A clean match is a count plus catalog ids. Full table rows stay for Extended, Novel, Contradicted, and Unverified, where the Note does work.
Novel is a map result, not catalog ingest Users cannot edit the OSCAL source, the synthesis, or principle-candidates.json. The Note still names the closest ids so a miss can be challenged. OpenCRE is no longer a user obligation.
Citation audit bounded Fifteen or more entries run a pinned reference audit tool. Fewer than 15 resolve against one registry of record. A second registry runs only on a miss or a title mismatch. Four registries per row was work most assistants skipped or faked.
Pinning contradiction is conditional "Name a paper that violates the pinning principles" read as a required finding. The check now fires only if this paper argues for pinning and recommends an unpinned install.
P1 recurrence defined "Tier above" from P1 had nowhere to go. Three or more P2s of the same shape become one P1. P1 stays P1 and is reported once.
Coverage limited to declared scope Unused categories outside the paper's topic are out of scope, not a gap list. A classic axiom the advice would violate still gets named.
llms.txt fetched only on 404 Step 5 already named the live paths. Prefetching the inventory cost tokens on every regulatory check and added nothing when those paths resolved.
map-only switch A 4-page blog and an 80-page draft should not pay the same citation-registry bill. Named in the prompt header so reports stay comparable.
Draft mechanics does not repeat step 1 Production artifacts already have a home in the submission screen. Step 6 is for remaining editorial, link, structure, and machine-readability defects.

What changed in 2.1

Step 1 grew three checks. All of them replace judging prose by eye with counting something a reader can re-count. An assessment produced under 2.0 stays valid; it simply screened less of the submission.

ChangeReason
Claim support check added to step 1 Counts unsourced statistics, recommendations naming no version or setting a reader could act on, and experience claims with no environment or timeframe. A claim nobody can check is unusable whoever wrote it.
Currency check added to step 1 Names the newest dated item in the paper next to its publication date. A gap over a year is a fact about the paper rather than an inference about its author.
Production artifact check added to step 1 Markup surviving into a PDF, placeholder text, scaffolding addressed to a requester, and a figure carrying two values in two sections are all defects with a location and a fix.
Citation check now reads the shape of the bibliography A list where every entry is canonical and heavily cited, with nothing recent or specific to the subject, is worth reporting next to the resolver output.
Disclosure bullet narrowed further 2.0 barred inferring generation from prose rhythm. 2.1 extends that to vocabulary, punctuation habits, and sentence uniformity, and says plainly that no volume of stylistic signal reaches an integrity verdict.
Two false positives added to step 9 A standards document written at principle level on purpose reads as unactionable. A survey cites old work because that is its subject.

What changed in 2.0

The eight steps of v1.0 became eleven. Three are new. The rest are edits in place, and most of them are demotions. An assessment produced under v1.0 stays valid on its principle mapping; its step 0 findings are the ones to re-read, since v1.0 permitted intent-based conclusions that v2.0 caps at an observation.

ChangeReason
Governing rule added above step 0 Every severity ceiling below follows from it, so it has to be stated before the steps rather than implied by them.
New step 0, intake A partial read produces a confident and wrong coverage-gap list. Bad extraction manufactures findings rather than hiding them.
Disclosure bullet rewritten to observable facts v1.0 asked the reviewer to identify undisclosed sponsorship and undisclosed bot generation. Both require asserting something that by construction is not in evidence.
Prose-rhythm signal demoted to P4 House style guides, standards templates, single authors, and non-native English writers all produce uniform rhythm.
Citation check now resolves identifiers against registries Crossref, arXiv, DataCite, and Semantic Scholar answer the question that reading a reference list cannot.
Extraction chunked at roughly 20 statements per section A single pass over a long paper drops material without saying so.
Unverified status added A failed catalog fetch previously turned unchecked statements into Novel findings.
New step 9, one priority scale plus evidence ceilings v1.0 separated blocking from advisory inside draft mechanics only, so an integrity finding and a typo shared a report with no severity spine.
Recurrence escalation added One occurrence is a slip. A repeated defect is a process problem and reads differently.
Known false positive list added Without one, the same misclassifications recur on every run with nothing to catch them.
Tier column added to the findings table Lets a reader sort by severity instead of reading every row.
Candidate queue check removed from the report step The queue has not moved since it was created: four entries from one review date, none incorporated into a catalog, every source unverified. Sending a reviewer there as a deduplication check does nothing until it moves again. Novel candidates are still reported, now carrying the near-matches that were rejected and why.

Back to the prompt