Phase 1 (Weeks 1–2) — Test Set Design and Metric Selection
The team works with the clinical sponsor to define the success criteria for the AI feature, derives test cases from de-identified historical encounters (typically 200–1,000 cases for an initial harness), and selects the metrics that match the clinical use case. Sensitivity-weighted versus specificity-weighted? Calibration curve thresholds? Fairness slicing across which patient demographics? The selection is documented and signed off by the clinical sponsor before any code is written.
End of Phase 1: a written eval methodology document with test set definition, metric selection rationale, and acceptance thresholds. This is the document that goes into hospital security review and (if applicable) FDA submission.


































