Deep audit: what would make this worth adopting?
September 24, 2026 · Protocol 0.1-rc.4-candidate.3 · Author-defined candidate
The honest conclusion
There is a plausible, useful contribution here: helping people distinguish a persuasive AI-built artifact from evidence that justifies the next engineering investment. There is not yet evidence that this protocol changes HCAI practice, improves defect detection, saves money, or deserves the standing of an established standard. Better writing and passing software tests cannot establish those claims.
The promising unit is a bounded decision with inspectable evidence, not an all-purpose AI readiness score. The protocol should complement existing accessibility, human-AI design, security and risk-management practices—not rename or replace them. This is a design judgment from this audit, not an empirical finding or a novelty determination.
What established work teaches us
Exa returned 20 search results across four angles: accessibility standards, AI risk governance, human-AI interaction, and software assurance. Duplicate URLs/DOI versions and superseded drafts were not counted as independent support. The primary sources below ground the comparison; this was a targeted comparison, not an exhaustive systematic review.
| Reference | Relevant lesson | Consequence for this candidate |
|---|---|---|
| WCAG 2.2 | Technology-independent requirements are separated from supporting explanations and techniques. Its status also reflects a consensus process, not only document structure. | Keep stable criterion IDs, a defined scope and clear requirements. Do not borrow W3C status, A/AA/AAA labels, or certification language. |
| W3C guidance on test rules | A passing partial check does not establish every aspect of a success criterion. | State the automated-check boundary in every result; require human evidence-quality review. |
| NIST AI RMF 1.0 Core | Risk management involves context, diverse perspectives, and continuing governance across the lifecycle; its functions are not a simple ordered checklist. | Screen affected people and consequences. Keep this narrower engineering decision distinct from lifecycle risk management. |
| Guidelines for Human-AI Interaction, CHI 2019 | The original work reports multiple evaluation rounds, including practitioner application. Guidance and demonstrated applicability are different accomplishments. | Test whether people can interpret and apply this protocol; do not equate correspondence with comparable evaluation. |
| NIST SSDF AI community profile, SP 800-218A | The AI-specific profile supplements the SSDF and explicitly defines its scope. | Route security concerns to appropriate engineering review. This profile does not replace secure development or certify generated code. |
The W3C also distinguishes functional conformance checks from usability testing and recommends involving people with disabilities in usability evaluation. The protocol's own reading experience needs that work too. HTML is the primary reading alternative to the untagged PDFs; no WCAG conformance claim is made. W3C conformance guidance
Problems reproduced and corrections applied
| Audit finding in candidate.2 | Candidate.3 correction | What remains human or untested |
|---|---|---|
| Changing an acceptance criterion preserved an old passing validation. | Bind validation to both artifact and requirement/context fingerprints. Missing fingerprints block; mismatches require revision. | A new hash is not a new test. A person must inspect actual execution and test adequacy. |
| A proposed state could name an unrelated destination without affecting the decision. | Require resolvable transitions, entry reachability and possible exit paths. | Transition guards, bounded retries, runtime termination and full exception coverage. |
| An empty action-boundary list could pass. | Require explicit nonempty authority limits, including draft-only/manual cases. | Whether implementation actually enforces those limits. |
| Risk depended entirely on six self-selected labels. | Consequential-context flags set deterministic minimum depth; unknown cannot mean no. | Context answers and risk rationale still require competent judgment. Floors are provisional. |
| Human-centered concerns could be omitted while the owner’s efficiency story remained complete. | Add HCAI-1.4: affected roles, access/use, privacy/security, unequal effects and human agency, linked to requirements. | This is coverage screening, not assurance that every harm is discovered or mitigated. |
| A structural pass could be read as evidence-quality assurance. | Require a recorded human quality review and distinguish rule checks, supplied judgment and unverified authenticity in outputs. | The software cannot authenticate that the reviewer performed the inspection. |
| Routed pilot records could omit skipped gates and the reason for stopping. | Reconcile routing, NOT_EVALUATED gates and all follow-up reasons. | Whether the pilot occurred and what participants actually experienced. |
The first three findings were reproduced as PROCEED_TO_ENGINEERING before correction. Regression tests now exercise these and the added boundaries. The examples remain synthetic; this is evidence about implementation behavior only.
Does this make the user's work simpler?
The participant receives one question at a time and leaves with one next action. The facilitator maintains the detailed evidence record. A shared note can support several checks; the method does not require a new document for every field. References, criteria and schemas are available on demand.
The conversation worksheet makes this division explicit. It is not a shortcut around evidence. If evidence must be created from scratch, the short session stops and identifies the discovery work. Preparation, session and capture time remain visible.
There is a real tradeoff: stronger evidence requirements may make QUICK-6 less attainable. The <=15-minute target is conditional on an available evidence packet and a low-risk case. Do not advertise an end-to-end 15-minute implementation assessment. A pilot should test whether even this narrower target is useful and realistic. If not, revise the process or the target rather than hiding preparation time.
A falsifiable contribution, not a fame claim
The hypothesis worth testing is: for a defined review task, this protocol helps reviewers notice consequential missing/contradictory evidence and choose a justified next step, without imposing disproportionate burden or blocking sound proposals unnecessarily.
Evidence should be gathered in distinct stages:
- Bounded formative use. Observe a consenting external advisor/end user with a real workflow. Record preparation, elapsed time, stops, evidence difficulty, independent explanation of the result and next action. A declined or incomplete session is informative; do not remove it from the denominator. One eligible run is a release prerequisite, not proof of usability or effectiveness.
- Reproducibility and disagreement. Have reviewers independently assess the same versioned packets before seeing one another's results. Include adequate, defective and incomplete cases, not only obvious failures. Retain original judgments and reasons, report gate-level agreement and denominators, then adjudicate disputes with an appropriate independent domain reviewer. Do not call the author's preferred answer ground truth without a defensible reference process.
- Comparative effectiveness. Prespecify a suitable ordinary-review comparator, assignment/order, primary outcome, meaningful effect, burden tradeoff and analysis. Measure consequential omissions, false alarms/unnecessary stops, and appropriateness of the engineering decision—not just satisfaction. Choose sample size from the actual design and uncertainty needs. Keep the proposed visual-fidelity experiment separate unless an approved design explicitly studies this intervention.
- Transfer and maintenance. Replicate in additional workflows, sectors, accessibility needs and reviewer experience levels. Record adaptations and where the method fails. An English-language owner-led pilot does not establish universal usability. Retest implementations whenever normative rules change.
No stage above is reported as completed. No sample sizes, effects, adoption totals or endorsements are invented. Registration, participant protections and institutional requirements must be resolved before a controlled study; this roadmap is not research approval.
Adoption needs governance, not only downloads
Keep public, versioned requirements and examples; a no-install path; interoperable records; clear licenses; correction procedures; and a visible record of disputed interpretations. Allow domain profiles to add checks without weakening mandatory gates. Distinguish a request, a reported problem, a reproduced defect, software verification, actual use and empirical evidence. Do not turn downloads, email replies or interest into adoption metrics.
Read claims and governance, the update log, and the current verification record. These support scrutiny, not a claim to worldwide-standard status.