The Potential and Ethical Challenges of Generative AI in Healthcare: A Summary of Expert Perspectives from the HIMSS Conference
Generative AI is accelerating its entry into the healthcare sector, but its rapid deployment and lagging regulation have sparked ethical controversies. Experts at the HIMSS conference discussed the potential of models like GPT-4 in scenarios such as clinical documentation and medical record summarization, while warning of risks including the "black box" problem, hallucination phenomena, algorithmic bias, and ambiguous accountability. Companies such as Microsoft, Epic, Google, Nuance, and Suki disclosed their latest progress and unanimously agreed that AI should be positioned as an assistive tool rather than a replacement for doctors.

Artificial intelligence is already widely used in healthcare, but generative AI is seen as a significant and risky step forward due to its rapid deployment and lack of regulatory oversight. Industry experts issued this warning at the HIMSS conference in Chicago this week.
As a result, a host of thorny ethical questions have emerged around GPT-4—OpenAI's latest large language model and the technology behind the paid version of the popular ChatGPT chatbot. Nevertheless, proponents believe generative AI has enormous potential to improve U.S. healthcare delivery.
Microsoft, one of the largest companies in the field, has been working with GPT-4 developer OpenAI over the past eight months to understand the technology's potential impact on healthcare. Peter Lee, corporate vice president of Microsoft Health, said during a Tuesday keynote panel that generative AI has already shown early success in several use cases, including simplifying the interpretation of Explanation of Benefits (EOB) notices and drafting prior authorization request forms.
Lee noted that doctors can prompt the AI to assist in interpreting difficult new cases or to rewrite patient conversations in a standardized format.
Microsoft recently announced plans to embed generative AI into Epic's clinical software. Epic is the largest provider of electronic health records (EHR) to U.S. hospitals.
Last week, Microsoft and Epic took the lead in rolling out GPT-integrated EHR workflows at select sites, used to automatically draft replies to patient messages. The two companies have also brought generative AI tools to Epic's hospital database, allowing non-specialists to ask the AI general questions directly without needing data scientists to query specific data.
Seth Howard, vice president of research and development at Epic, said in an interview that Epic is investing significant R&D resources in generative AI.
The healthcare IT company is also researching the use of generative AI to summarize patient medical histories and to translate patient-facing materials across different languages and reading levels to improve health literacy.
"This is a large category of work," Howard said. "We will release many use cases over the next year."
Google is also opening its proprietary large language model, Med-PaLM 2, to a select group of customers to explore use cases. Med-PaLM is specifically trained on medical data and can sift through and understand vast amounts of medical information. However, according to Google's AI team, the model still has room for improvement in the complexity of queries it can handle and in achieving product excellence.
Aashima Gupta, head of Google Cloud, said in an interview that Med-PaLM will not be used in patient-facing scenarios. Instead, hospitals could use the AI to analyze data to assist in diagnosing complex diseases, fill out medical records, or serve as a "concierge" service for patient portals.
Experts point out that one notable use case for generative AI is streamlining medical documentation. Doctors can spend up to six hours a day writing notes in the EHR, which reduces time with patients and can contribute to burnout.
Leading clinical documentation companies are integrating generative AI technology and claim it improves the accuracy and speed of their products.
Nuance, a Microsoft-owned documentation company, said last month it has integrated GPT-4 into its clinical note-taking software. This summer, providers using Nuance's existing documentation products can apply to test the beta version of the application, DAX Express.
Diana Nole, general manager of Nuance, said that in mid-June, 10 to 15 customers (totaling 300 to 500 physicians) will test the product in a private preview, followed by a full commercial release in early fall.
Suki, a documentation company that works with Google, launched its generative AI-powered "Gen 2" voice assistant earlier this month.
Punit Soni, CEO of Suki, told Healthcare Dive that Suki can now contextually generate clinical notes from conversations and auto-populate them. Soni said that closer to the end of the year, doctors will also be able to ask the AI questions and give commands, such as charting a patient's A1C levels over the past three months.
According to the CEO, Gen 2 is currently being piloted with a "very small" number of customers. Suki plans a full rollout later this year.
"The interest is incredible," Soni said. "If someone could help you do all your notes—wouldn't that be great?"
Heavy risks
Admittedly, generative AI makes mistakes, raising thorny ethical questions around accuracy, fairness, and accountability. This means the medical community needs to continue to play a significant role in deciding whether, when, and how to use such AI, HIMSS experts said.
One problem with AI is the "black box" issue: if you cannot clearly see how the AI reached a conclusion, how can the result be trusted?
It is easy to be misled into thinking this type of technology can reason, but users must remember that generative AI does not deliberate—it simply predicts the most plausible next set of words, said medical ethicist Reid Blackman.
"Maybe we can accept effective magic. That's one option. The other option is, if I'm going to do a cancer diagnosis, I have to fully understand why you're making that diagnosis... The problem is GPT doesn't give reasons. It gives things that look like reasons," Blackman said during a Tuesday panel.
Generative AI also has accuracy issues. Large language models are known for "hallucinations," providing factually incorrect or irrelevant answers. OpenAI's website states that GPT-4 "still has many known limitations that we are working to address, such as social biases, hallucinations, and adversarial prompts."
Bias is another key concern for AI, especially in healthcare. Experts say that if algorithms are trained on biased data, or if their results are applied in a biased way, they will reflect and perpetuate those biases.
Accountability also remains a concern. Because the field of generative AI is so new, it is unclear which stakeholder—the owner? the user? the developer?—should be responsible for any negative consequences arising from its use, said Kay Firth-Butterfield, CEO of the Centre for Trustworthy Technology.
"If something goes wrong, who do you sue? Is there anyone to sue?" Firth-Butterfield asked during a Tuesday keynote panel.
More than 27,000 stakeholders have signed an open letter calling on AI labs to immediately pause training of AI systems more powerful than GPT-4 for at least six months. The letter argues that the industry should use this pause to develop shared safety protocols and work with policymakers to establish AI governance systems.
"You should not be exploring or engaging with AI without understanding responsible AI issues," said Firth-Butterfield, who signed the letter. "We are rushing into the future without stepping back and designing it for ourselves."
'Mysterious,' 'awe-inspiring,' 'frustrating'
AI companies argue that many problems can be mitigated through careful human oversight. Experts say organizations should clearly understand the datasets used to train models and the exact circumstances in which they are appropriate.
Microsoft's Lee said that questions about generative AI are certainly important, especially regarding accountability, but the "black box" problem may not be as troubling as some suggest.
"We have never been able to prove it doesn't have reasoning capabilities," Lee said. "This is a mysterious, awe-inspiring, and frustrating research barrier that we have never been able to prove. As a scientist, I have to admit that from a certain stage of development starting with GPT-4, this black box problem may not actually truly exist."
Epic's Howard said there are also techniques to ensure generative AI produces reliable outputs—a field known as "prompt engineering," where scientists study how to ask the AI questions in ways that yield credible results.
Suki and Nuance both declined to disclose the accuracy rates of their generative AI-based documentation tools, but said accuracy is high enough that they are confident deploying them in real-world settings.
Soni said Suki may publish Gen 2's accuracy rates within a few months. Nole said Nuance has been working on AI for years.
"The accuracy levels we already have are advanced. We know how it works, we know how it behaves. Now, we're just using GPT to push it even further," Nole said.
As AI advances, some models have outperformed doctors on diagnostic tasks, raising concerns that machines might one day replace clinicians. HIMSS experts agreed that even with generative AI, that day is far off.
Alan Karthikesalingam, head of Google Health AI research, said in an interview that Med-PaLM's ability to pass medical exams "in no way represents medical knowledge or competence."
"There are a lot of substantive safety, bias, and ethical issues that need to be handled very carefully," Karthikesalingam said. But replacement? "I think overall, this hopefully should not be a huge concern for our doctors. These systems are very clearly positioned—they are tools."
Suki and Nuance emphasize that the notes generated by their AI allow clinicians to edit and approve what is ultimately written into the EHR. AI executives say generative AI is a tool for doctors, not a replacement.
Still, "anyone who predicts what this era will bring is a fool," Soni said. "What I can tell you is that for the foreseeable future, I see AI becoming a very powerful and serious assistant that makes our lives better."