The two CHAI documents are unusually well-matched to this portfolio — not by subject matter, but because they are built on the same epistemic discipline the portfolio was founded on. The Risk Categorization Tool is a gate before interpretation: a solution is classified before it is allowed downstream. And the Tool makes "N/A" and "Need More Info" first-class response tokens alongside Low/Medium/High — the completed CHAI example even scores a modifier N/A with the rationale "Abstain b/c don't have this use case information." That is precisely this portfolio's "I cannot determine this," now carried by a multi-institution consensus body (Duke, Mayo, Stanford, Sutter, Providence, Mount Sinai, MSK, and others).
The practical value is therefore twofold: (1) adopt CHAI's vocabulary and evaluation structure so the portfolio speaks the language governance committees and vendors increasingly expect, and (2) use the portfolio to demonstrate the CHAI discipline instantiated at the clinical-reasoning layer, where most tools only assert it at the policy layer.
CHAI and this portfolio gate at different altitudes. Read together they form one continuous chain of earned authority — from a system's right to be deployed at all, down to a single recommendation's right to be spoken.
Pre-deployment. Scores 9 Life & Patient Safety + 10 Technology & Data modifiers as Low / Medium / High (or N/A / Need More Info). Any single High triggers a rigorous risk assessment. Sets the level of oversight, monitoring, and mitigation the system must carry.
Run-time, per decision. The nine gates set which of six authority tiers (T0 Silent → T5 Actionable) an answer has earned. Failing a gate caps the tier; it never forces a false answer. The honest exit — "I cannot determine this, and here is what would change that" — is always available.
Which CHAI Responsible-AI T&E chapter supplies the consensus evaluation rubric for each portfolio component, and the headline metrics to adopt.
| Portfolio component | Primary CHAI T&E chapter(s) | Headline methods/metrics to adopt |
|---|---|---|
| 01 · CKD Care Orchestration | §9 Clinical Decision Support; §10 EHR Information Retrieval (external-data reconciliation) | CDS usefulness/efficacy + fairness metrics; retrieval accuracy/attribution for the External Data Reconciliation Agent |
| 02 · Hypertension Framework | §9 Clinical Decision Support | Usefulness/efficacy + counterfactual fairness on next-best-action outputs |
| 04 · Complex Multi-Disease (CMDF) / fluid triage | §9 Clinical Decision Support | Safety & reliability under complexity; per-module benchmark endpoints |
| 05 · Hospitalization Packets | §2 Note Summarization; §10 EHR IR (evaluation inputs) | Ground-truth answer keys = the reference set for DocLens-style scoring |
| 06 · BRIDGE Extraction Heuristic | §10 EHR Information Retrieval; §2 Note Summarization | DocLens (completeness / conciseness / attribution); accuracy/sensitivity/specificity/PPV/NPV vs. ground truth |
| 08 · Synthetic Lab & Document Corpus | §2 Note Summarization; §10 EHR IR | DocLens/ACUEval fact-extraction vs. GROUNDTRUTH.json; de-ID recall against the PHI index |
| 09 · Aberrant / Counterfactual Labs | §9/§10 Safety & Reliability (data-integrity axis) | Implausibility-detection recall; runtime data-integrity monitoring benchmark |
| 07 · CRRF (nine-gate architecture) | Cross-cutting — the RAIG principles + the Risk Categorization Tool itself | The governance layer: maps to Risk Tool modifiers and to the RAIG lifecycle stages (see Tables 2–3) |
| Patients · Synthetic cohort | Fairness & Bias Management (all chapters) | Counterfactual sentiment/quality parity + Population Sensitivity testing on a deliberately diverse cohort |
Where a CHAI construct already has a working analog in the portfolio. These are the citations that let the portfolio claim CHAI-conformance today, and CHAI-as-external-warrant for its design choices.
| CHAI construct | Portfolio instantiation | Note |
|---|---|---|
| "N/A" & "Need More Info" as first-class responses | CRRF T0 (Silent) + the abstention principle; schema-validated "I cannot determine this" | Direct match — adopt CHAI's exact tokens in output schemas |
| Risk categorization → level of rigor | CRRF six-tier authority ladder (T0→T5); gate failures cap the tier | Risk sets deployment rigor; tier sets per-output rigor |
| DocLens / ACUEval ground-truth fact comparison | GROUNDTRUTH.json answer keys across 05, 06, 08 | The evaluation harness CHAI recommends is already built |
| PDQI-9 documentation quality | Narrative Quality Rules (_reference/); Narrative Synthesis Agent | Map narrative-quality rubric onto PDQI-9 items |
| Counterfactual fairness / Population Sensitivity | Diverse synthetic cohort (18 patients, varied race/ethnicity/SES) | Cohort is a ready-made counterfactual test bed |
| Monitoring / incident detection / drift | CRRF Gate 9 "Learn under governance" (versioned, auditable) | Principle present; runtime tooling is a known gap (see worked example) |
| Data-quality & fitness-for-use dimensions | CRRF Gate 3 "Check what we actually have"; data-confidence model | Grounded in Kahn / Weiskopf-Weng / SUITABILITY / METRIC (07 evidence base) |
| Detection & traceability of AI influence | CRRF Gate 8 "Show the work" — support, counter-evidence, provenance, reversal conditions | Traceability is a design default, not an add-on |
The join table. For each CRRF gate, the CHAI risk modifier(s) whose failure mode that gate is designed to catch at run-time. This is what lets a health system point to a specific gate as the mitigation control when a CHAI modifier scores Medium or High.
| CRRF gate | CHAI modifier(s) it interrogates | Domain |
|---|---|---|
| 1 · Understand the question (declare the unit) | Use Context & Complexity; Decision Autonomy | Life & Patient Safety |
| 2 · Know what data is needed | Sufficiency & Representativeness; Accuracy of Data | Technology & Data |
| 3 · Check what we actually have | Accuracy · Completeness · Veracity · Data Transparency · Sufficiency; Use of Sensitive Data | Technology & Data |
| 4 · Interpret before using | Accuracy of Data; Consequences of Failure | Tech & Data / LPS |
| 5 · Judge sufficiency (abstain if not) | Need More Info / N/A pathway; Consequences of Failure | Life & Patient Safety |
| 6 · See the whole picture (cross-stratum coherence) | Population Sensitivity / Disparity; Breadth of Potential Harm | Life & Patient Safety |
| 7 · Earn the authority level | Decision Autonomy; Distance From Patient; React Time | Life & Patient Safety |
| 8 · Show the work | AI Detection & Traceability; Data Transparency | Technology & Data |
| 9 · Learn under governance | AI Monitoring / Incident Detection; Lifecycle Management; Model Security; Cross-System Propagation | Technology & Data |
For any framework or app entering pre-deployment review: (1) complete the CHAI Risk Categorization Tool using the reusable template — this is the system's governance "birth certificate"; (2) for every modifier scored Medium or High, name the CRRF gate (Table 3) that mitigates it at run-time, and document any residual gap honestly; (3) structure the framework's validation report around the CHAI T&E chapter identified in Table 1, using the ground-truth harness already built. The worked example (CKD Care Orchestration Framework) shows all three steps end-to-end.