I keep wanting to do away with the conversation around whether AI will replace humans in medicine. Unfortunately, it’s the topic that keeps coming back, like a terminator. The latest iteration is notable for the prominence of its authors and the publishing journal. In the Journal of the American Medical Association (JAMA), authors Ezekiel Emanuel, Abe Baker-Butler, and Neal and Vinod Khosla state their opinion (linked from the Khosla Ventures website here) that autonomous AI will exceed AI-aided physicians in giving the best medical care.
And, I’ve said before, they are asking the wrong question. Framing the problem as a human-machine matchup may appeal to a certain kind of industrial-revolution era storytelling trope, and is certainly great clickbait, but is this what we should really be caring about? When it comes to healthcare, in the long run, what are we most concerned with—providing the best health care, or fighting about who gets to deliver that healthcare? What I’ve argued is that instead of staging a cage match, we should be asking ourselves how can we provide the best care, with the tools we have available and the tools we are building. It’s a subtle shift, but one that avoids the need for staging a fair fight between humans and machines.
OK, putting aside for the moment the fact that they are asking the wrong question, let’s see if they are even answering their—perhaps misguided—question in a reasonable way.
They give three possibilities of what might provide the best healthcare—AI alone, humans alone, or AI/human hybrids. They do this by breaking down the process of medicine into five cognitive medical tasks, and then showing that AI alone is probably better at each of these.
I have two fundamental challenges to this.
- Is this really what makes up medicine?
The question they avoid is whether or not these five tasks actually capture the practice of medicine in any meaningful way. I’m not arguing that these tasks—eliciting medically relevant information, establishing a differential diagnosis, specifying diagnostic testing, prescribing guideline-concordant treatment, and managing chronic diseases—are not part of medicine (of course they are), but do they really capture the core tasks used in medicine? Or, are they just the things that people have actually tested and are convenient to use as comparisons? Managing chronic diseases is the most striking outlier—not all doctors (like radiologists, emergency medicine physicians, most surgeons, and others) manage chronic diseases. So why is that a core cognitive task, according to the authors? I think it’s that, instead of trying to define the core tasks of practicing medicine, they just chose some basic tasks that have been recently studied. But tasks don’t appear in papers because they represent the most important activities in healthcare; they appear because their outcomes are fairly easily measured, and because—and this is key—AI has a fighting chance at being better. No one proposes a study where the outcome is obvious: Can ChatGPT beat a human doctor at a physical exam? They also admit themselves that there is a bias against publishing papers when AI is unsuccessful. So if you base your concept of medicine on what tasks have been studied, you will choose the most easily measured tasks, you will choose tasks biased toward AI having a good chance of beating humans, and you might exclude tasks where AI failed.
Cherry picked elements like this aren’t representative of medicine as a whole. Unfortunately, in real healthcare, we can’t ignore the tasks that are hard to measure. We have to do them all.
- Are these really the only three choices?
Can we really only conclude that AI by itself, humans by themselves, or AI/human hybrids are the best healthcare providers? Are those the only choices? How about the idea that some tasks are better performed by AI, some by humans, and others by hybrids? This is what I talk about in my previously linked piece, and another one, where I agree that we should break medicine into tasks, but not—as they do—to reassemble them and pretend the combination stands for the entire practice of medicine, but to find ways to incorporate AI into healthcare for the sake of better patient care. I would love to not have to measure a lesion, calculate its volume, and compare it over multiple exams. I welcome AI in summarizing the thousands of pages of notes in the medical record of each patient. But if I have terminal cancer, I still want a human to help through the struggle to decide if I should try for one more round of chemotherapy, or focus on being comfortable. It’s a false trichotomy and an oversimplification to even pretend that we can reduce the entire practice of healthcare, with all its variations and details, into three possible versions. Even in our most automated industries, such as manufacturing, we understand that humans and machines have different roles to play throughout the enterprise.
I have additional problems with their reasoning.
The first is that in their review of the literature, they criticize the studies that show humans or human/AI hybrids outperforming AI alone on the basis that the AI used was not the most modern, or else not optimized for the task at hand. The same criticism is notably lacking in assessing the optimization of factors that might disprove their point. I have also made the point that we need to learn how to help people work well with AI. Design matters; a potentially valuable tool might fail simply because of a poorly designed interface. If AI-alone testing didn’t use the most modern version of AI, we have to assume that the AI/human hybrids are also not necessarily optimized. It’s easy to see that AI models get better over time—the models have version names and measurable metrics. But human-AI interfaces are harder to measure, and I don’t think we’ve done nearly as much work to optimize those. To their credit, the authors do admit this in their caveats section, but then they dismiss it with a handwave, writing “as long as a hybrid is human controlled, workflow design improvements seem unlikely to eliminate the human-introduced error present in hybrids and absent in AI alone.” This would be a more defensible statement if humans could only interact with computers via veto power, as they imply. Better-designed products with new forms of human/computer interface will almost certainly have better outcomes, and you can’t ignore that because it doesn’t support your argument.
Another point allows me to quote from an unlikely source: Geoffrey Hinton, the famous AI scientist who predicted in 2016 that radiologists would not be needed in 5 years, and began the AI-will-replace-radiologists furor that continues today. He has since taken ownership of that incorrect prediction, and he stated that one thing he forgot was that the problem was elastic. He specifically explained that as radiology gets more efficient, the demand increases, which is a very narrow economic approach to the issue. Rather than disagree with him again, I’d like to take his point in a more general way. “Healthcare” is not a bounded problem with a clear solution. As we improve our methods, our goals change. So saying that AI will be better at modern healthcare than humans, even if true, doesn’t account for the fact that the addition of AI to healthcare will change what medicine is, and that the future of healthcare may—and I believe certainly—still require humans in some capacity. But they won’t be doing all the tasks they do today. And thank goodness for that.
I do, by the way, have a history of disagreeing with Ezekiel Emanuel when it comes to AI and healthcare. Back in 2016 at the American College of Radiology meeting, I was in the room when he predicted the end of the need for radiologists, and he stated that the thing that convinced him was when AI beat a human champion at chess, and then another at Go. Even at the time, not yet immersed in AI, that seemed wrong to me. Chess and Go are bounded domains, with clear rules, limited moves, and easily measurable success metrics. Radiology is so different. It would be like watching a mathematician solve a problem and then assuming they must be a great pasta chef. The skill is admittedly amazing, but it’s not reasonable to generalize the ability, and then to apply it to a specific domain, like radiology. I spent much of that meeting explaining to colleagues why I thought he was wrong. And yet, here Emanuel goes again, bringing up chess. AI’s dominance over humans in chess is an interesting anecdote, and worth thinking about, but far from real evidence for future dominance in healthcare.
Lastly, I didn’t want to, as many other detractors have done, make this an ad hominem attack, questioning Emanuel or the two Khoslas’ conflicts of interest because of their potential financial gain from healthcare AI-related ventures: the Khoslas have direct financial stakes in AI health ventures, with Vinod an investor in OpenAI and Curai, and Neal serving as Curai’s CEO, while Emanuel discloses his own advisory and consulting ties to a roster of health-tech companies. Baker-Butler, his research assistant, shares co–first authorship without any disclosed interests of his own. But their questionable framing and reasoning do raise the question of why they wrote this opinion piece in the first place. You have to wonder if one of the goals of such an essay is not necessarily to give a rigorous answer, but to keep AI in the headlines. After all, most AI companies today are valued more on their potential for future impact than what they actually deliver now. The hype is a big part of the product, and even my commentary here may represent the success of the true purpose of their opinion piece: to create headlines for AI. Engagement and clicks are what matter.
The reason that I don’t like asking if AI will replace humans is that these questions imply that we are passive observers of history as it flows before us. What it ignores is our agency—the fact that AI is just a technology, a tool, and we can be part of shaping it into one that serves our best futures and our noblest intentions. That’s why I am constantly asking myself: how can we use AI to do a better job at helping patients? If all of us—especially those of us in healthcare and healthcare technology—can focus on that, instead of head-to-head battles, our patients will be much better off.