Editor's note: This is the first of a two-part series on foundation model applications in the medical device industry.The second storywill be published on Tuesday.

A growing number of medical device companies are promoting "foundation models" — a type of artificial intelligence that can be adapted to multiple tasks. However, questions remain about the technology's use in the medical device industry.

Over the past year, GE HealthCarehas promoted an MRI research foundation modelPhilips announced plansto build an MRI foundation model with Nvidia; at last year's Radiological Society of North America (RSNA) annual meeting, multiple abstracts focused on how toevaluate and improve foundation models. The U.S. Food and Drug Administration (FDA) has also updated itsAI medical device database, saying it is exploring ways to identify and label products that incorporate foundation models.

However, experts say the definition of foundation models remains unclear, and it is difficult to determine whether currently available tools are truly helping radiologists and patients.

What is a foundation model?

Magdalini Paschali, a postdoctoral scholar in Stanford University's Department of Radiology, noted that foundation models have several key characteristics: they are trained on large datasets, often with much of the data unlabeled; they can handle multiple data types, such as images, text, medical history, and genomics; and they can address various tasks, such as detecting diseases not seen during training.

Professional photo of Magdalini Paschali
Magdalini Paschali is a postdoctoral scholar in Stanford University's Department of Radiology.
Image courtesy of Stanford University

Earlier this year, Paschalipublished a paperin RSNA's journal Radiology, attempting to define the technology more clearly.

In practice, "almost anything can be defined as a foundation model," said Akshay Chaudhari, assistant professor of radiology and biomedical data science at Stanford.

Chaudhari said the term "foundation model" wasfirst coined at Stanford in 2021. In medicine, one of the earliest versions wasGoogle's Med-PaLM, unveiled in late 2022— a large language model designed to answer medical questions.

Chaudhari said foundation models began receiving more attention at RSNA starting in 2023.

Traditional deep learning models used in radiology, such as those that detect pneumonia, typically focus on specific health conditions and rely on labeled data. Nina Kottler, associate chief medical officer for clinical AI at Radiology Partners, gave an example: radiologists need to review images, circle instances of pneumonia, or annotate them in text reports.

Chaudhari noted that foundation models are trained on millions of images rather than thousands, making it unrealistic to require labeled data.

Are foundation models more accurate?

Some medical device developers claim foundation models are more accurate than narrow AI models. For example, Aidoc, which makes radiology triage software, says foundation models can speed up the development ofmore accurate AI tools

Experts say a device's accuracy depends on how it is built. Stanford's Paschali said a foundation model used "out of the box," without any specialization or additional training, may perform worse than a specific AI tool; but once exposed to examples and context, it may perform better.


"The best thing we can do is look at the summary statements in the FDA clearance documents. At least in the market, we haven't really seen the benefits of these foundation models yet."

Akshay Chaudhari

Assistant professor of radiology and biomedical data science at Stanford University


Kottler said that because foundation models are trained on vast amounts of data, they may perform better at detecting rare events, such as brain aneurysms.

"When only a few people have a certain disease, finding it is like looking for a needle in a haystack," Kottler said. "You need a very accurate model to do that."

Professional photo of Nina Kottler
Nina Kottler is associate chief medical officer for clinical AI at Radiology Partners.
Image courtesy of Radiology Partners

Another area where foundation models excel is accelerating the development of other AI models. Kottler said building a traditional narrow AI model might take six months to clean data, label data, and train; different iterations of a foundation model can be done in weeks.

However, in practice, whether these advantages have translated into real benefits for patients or care teams is hard to know. Stanford's Chaudhari said that for FDA-cleared foundation models, little public information supports companies' claims.

"The best thing we can do is look at the summary statements in the FDA clearance documents," he said. "At least in the market, we haven't really seen the benefits of these foundation models yet."

Evaluating foundation models

Currently, FDA-authorized foundation models are all designed to solve specific tasks. For example,Aidoc's rib fracture triage toolis built on the company's foundation model. Aidoc received 510(k) clearance using an older version of its rib fracture triage tool as a predicate device. The 510(k) process requires manufacturers to demonstrate that their device is substantially equivalent to a predicate device that can be legally sold in the U.S.

Chaudhari said there are currently no guidelines for broader models that combine language with images or video.

Some hospitals have established systems to evaluate AI models, but these systems are not comprehensive. For example, hospitals identify a need — such as a model to detect pneumonia in X-ray images or a model to draft X-ray reports.

"Then they collect 1,000 images with known labels from their own institution and hold a competition to see which vendor performs best on that dataset," Chaudhari said. "It's rudimentary, but it's really the best we have right now because it's the only way to assess whether local performance truly fits the specific task."

Professional photo of Akshay Chaudhari
Akshay Chaudhari is assistant professor of radiology and biomedical data science at Stanford University.
Image courtesy of Stanford University

For large academic medical centers with data science teams, this may be feasible, but "many hospitals will just deploy the model and then get data of questionable quality from it," Chaudhari added.

Paschali said that to test a foundation model's accuracy, one should first define metrics and tasks based on what the model claims to do. Hospitals should also test the model's performance across different patient subgroups and different scanner types. Finally, hospitals should "stress test" the model to identify potential issues that may arise, such as with extremely rare diseases.

"That's why it's very important to work closely with radiologists, because they can help us design stress tests in a very thorough way," Paschali said. "Because they've seen so many cases, they know which tasks are difficult even for them."

Promisingly, foundation models trained on large datasets from different states, hospital types, and imaging equipment may require less evaluation before deployment.

"I don't think we've reached that point yet, at least there's no evidence in the world," Chaudhari said. "But that's the allure these foundation models can offer."

Another goal is to free up radiologists' time amid apersistent shortage of radiologistsin the U.S. and growing imaging volumes.

"If we ask them to verify outputs and review everything one by one," Chaudhari said, "does that really fulfill the promise of these models?"