The Workbench · Craft

GMLP guides practice; it doesn't clear a device

A submission that cites Good Machine Learning Practice as though it were a standard a device could be certified against has misread what the document is for. FDA, Health Canada, and the UK's MHRA jointly published ten Good Machine Learning Practice guiding principles in October 2021, aimed at how a machine-learning-enabled medical device should be developed across its whole life cycle — from the data it's trained on through what happens after it ships. None of the three agencies treats GMLP as a checklist a device can satisfy for clearance, and none has built a conformity assessment around it. It's a description of good practice, meant to be read alongside whatever regulatory pathway and whatever standards already apply to a given device, not a substitute for either.

Ten principles, published jointly, binding nowhere

The joint publication is itself part of the point. FDA, Health Canada, and MHRA didn't each issue a separate document and call it convergence after the fact — they published one set of ten guiding principles together, signaling that the three agencies expect roughly the same things from a machine-learning-enabled device's development process regardless of which one eventually reviews the submission. None of the three has turned GMLP into a rule, a recognized consensus standard, or a required deliverable. It sits closer to where FDA's least-burdensome principle sits — a stated expectation that shapes how a reviewer reads a file, without itself being the provision a device is measured against.

What the ten principles actually ask for

The principles run from data through deployment. One asks that training and test datasets stay independent of each other, so a model's reported performance isn't inflated by data it has already seen. Another asks that clinical study participants and training and test datasets represent the intended patient population across relevant characteristics, not just whatever population happened to be convenient to collect. Another asks that testing demonstrate performance under clinically relevant conditions, generating evidence independent of the training data rather than treating training performance as a stand-in for real-world performance. And the last principle asks that a deployed model stay monitored after release, with re-training risks like performance drift actively managed rather than assumed away. None of these is phrased as a pass/fail gate; each is phrased as a practice a development program should follow.

Not a substitute for the standards it echoes

A team that treats GMLP adherence as the deliverable still owes the actual regulatory record a submission requires. IEC 62304's safety classification still has to be assigned and documented; a risk management file built to ISO 14971 still has to name each hazard and its control. GMLP's principles describe the posture a development program should take toward its data and its models — representative datasets, independent test sets, human-AI team performance rather than model performance alone — but they don't specify the artifact format a reviewer expects to see, and they don't relieve a sponsor of producing the software life-cycle documentation those other frameworks already require.

Where GMLP and PCCP meet without overlapping

A Predetermined Change Control Plan is a specific submission mechanism: it pre-clears a defined space of future modifications to an already-authorized device, reviewed and accepted by FDA as part of that device's own file. GMLP isn't a submission at all, and following it doesn't pre-clear anything on its own. The two meet at exactly one point — a PCCP that defines how a model may retrain or update has to describe verification methods and performance boundaries, and a development program that already practices GMLP's principles on data independence and deployed-model monitoring usually has the internal records that make writing a credible PCCP possible. A program with no GMLP practice behind it can still write a PCCP, but it's building the monitoring and re-training discipline from scratch instead of documenting discipline that already exists.

Where this meets the file

A machine-learning development file that cites “GMLP-aligned” as a label, without showing which of the ten principles the specific project actually followed and how, has recorded an aspiration rather than evidence. What the file should carry instead is each principle mapped to a concrete internal artifact — the dataset representativeness analysis, the train/test split protocol, the post-deployment monitoring plan — kept separate from and prior to whatever PCCP or 510(k) submission eventually cites them. A GMLP self-assessment worksheet built around the ten principles' own terms is previewed in the launch catalog. If your program documents this differently, the shelf takes that correction directly.

The Regulatory Toolkit launches soon — a free shelf of source-mapped templates, checklists and browser-only tools for regulatory teams. Get one email when it opens, or contribute a template.

All Workbench notes