Abdulnour, R. E., Gin, B., & Boscardin, C. K. (2025).
New England Journal of Medicine, 393(8), 786–797.
Human–computer interactions have been occurring for decades, but recent technological developments in medical artificial intelligence (AI) have resulted in more effective and potentially more dangerous interactions. Although the hype around AI resonates with previous technological revolutions, such as the development of the Internet and the electronic health record, the appearance of large language models (LLMs) seems different. LLMs can simulate knowledge generation and clinical reasoning with humanlike fluency, which gives them the appearance of agency and independent information processing. Therefore, AI has the capacity to fundamentally alter medical learning and practice. As in other professions, the use of AI in medical training could result in professionals who are highly efficient yet less capable of independent problem solving and critical evaluation than their pre-AI counterparts.
Such a challenge presents educational opportunities and risks. AI can enhance simulation-based learning, knowledge recall, and just-in-time feedback and can be used for cognitive off-loading of rote tasks. With cognitive off-loading, learners rely on AI to reduce the load on their working memory, a strategy that facilitates mental engagement with more-demanding tasks. However, off-loading of complex tasks, such as clinical reasoning and decision making, can potentially lead to automation bias (overreliance on automated systems and risk of error), “deskilling” (loss of previously acquired skills), “never-skilling” (failure to develop essential competencies), and “mis-skilling” (reinforcement of incorrect behavior due to AI errors or bias). These risks are especially troubling because LLMs operate as unpredictable black boxes; they generate probabilistic responses with low reasoning transparency, which limits assessment of their reliability. For example, in one study, more than a third of advanced medical students missed erroneous LLM answers to clinical scenarios.
The article is paywalled, unfortunately.
Here is my brief summary:
This article addresses the challenge of supervising medical trainees who use artificial intelligence (AI) tools, particularly large language models (LLMs), in clinical reasoning. It notes that while AI can enhance learning through simulation, knowledge recall, and cognitive off-loading of rote tasks, off-loading complex tasks like clinical reasoning risks automation bias, "deskilling," "never-skilling," and "mis-skilling." LLMs are unpredictable black boxes with low reasoning transparency, making their reliability hard to assess.
The authors argue that critical thinking is foundational to adaptive practice in the age of AI. They propose the DEFT-AI framework (Diagnosis, Evidence, Feedback, Teaching, and recommendation for AI engagement) as a structured approach for educators to promote critical thinking during learner-AI interactions. The framework guides educators to probe the learner's clinical reasoning and AI use, evaluate supporting and opposing evidence, foster self-reflection, provide targeted teaching, and recommend safe AI engagement.
The article also describes two human-AI collaboration behaviors: centaur (strategic division of tasks, with human judgment leading) and cyborg (tight intertwining of user and AI throughout a task). Adaptive AI practice requires shifting between these modes based on task complexity and risk. The authors emphasize promoting AI literacy through evidence-based evaluation of AI tools and outputs, effective prompt engineering, and a "verify and trust" paradigm, concluding that AI interactions are here to stay and must be met with critical thinking and structured educational strategies.








