[1]
[2]
artificial intelligence,faculty development (simulation educator or simulation technician),generative AI,health professions education,and simulation
Article Type: brief-report
Simulation based healthcare education (SBHE) has progressed from a primarily technical training modality to a broad educational and systems-based strategy supporting clinical reasoning, teamwork, communication, professional identity, quality improvement and patient safety. The rapid emergence of artificial intelligence (AI), particularly generative AI and multimodal analytics, now presents another major transition in SBHE. AI, on the one hand, can accelerate scenario development, support facilitation, analyse performance and extend access to simulation, while on the other hand, it introduces concerns regarding accuracy, bias, transparency, privacy, authorship and overreliance. The central challenge is, therefore, not whether simulation should adopt AI, but how it should do so without weakening the human, relational and ethical foundations of experiential learning.
The nine contributions in this Short Reports on Simulation Innovations Supplement illustrate this tension between technological opportunity and human-centred practice. Three papers address the expanding role of AI and data-enabled methods. Levi and colleagues describe an evidence-anchored, timestamped multimodal annotation workflow for debriefing behaviourally anchored skills – Planning and Preparing, Authority and Assertiveness and Prioritising [1]. Power explores generative AI support for faculty during live simulation [2]. Xing and colleagues examine the use of generative AI to detect false references during a conference abstract review [3]. These innovations extend AI beyond automated content generation, positioning it as a supportive tool for facilitation, feedback, quality assurance and scholarly practice. AI-based applications in healthcare simulation may reduce workload and improve consistency; however, they should augment, rather than replace, educators’ professional judgement. Debriefing, for example, depends on more than identifying observable behaviours. Skilled facilitators interpret context, learner intent, team dynamics and emotional response while adapting inquiry to psychological safety and educational objectives. A timestamped analytic workflow may strengthen recall and provide a shared evidentiary basis for discussion, yet the educational meaning of an event still requires human interpretation. Similarly, AI support during live simulation may assist faculty with scenario progression, prompts or anticipated responses, but responsibility for learner welfare, curricular alignment and real-time adaptation remains with the facilitator. The same principle applies to scholarly review: AI may flag suspicious or unverifiable citations, but editorial decisions require transparent verification and accountable human oversight.
The remaining papers remind us that simulation innovation is not synonymous with digital technology. Scott and colleagues explore medical students’ experiences of paediatric simulated patients in ‘The wriggly child’ [4]. This contribution highlights the complexity of learning with children, where unpredictability, communication and emotional authenticity may be more educationally significant than technical fidelity. Another interesting aspect of this volume are the two courtroom-based simulation reports extending experiential learning into medicolegal and professional domains. Ritchie and colleagues describe a high-fidelity Coroners’ Court simulation for foundation doctors [5], while Allam and colleagues examine interdisciplinary learning and psychological safety within forensic psychiatry courtroom simulation [6]. These approaches enable learners to rehearse testimony, professional accountability, ethical reasoning and communication within environments that are difficult to access safely through routine clinical exposure. Te Arii and colleagues’ work with Aboriginal simulated participants in nursing education further broadens the meaning of fidelity [7]. Cultural safety cannot be reduced to adding demographic details to a scenario. It requires attention to power, history, identity, communication and the lived experience of communities. Involving Aboriginal simulated participants creates opportunities for co-designed learning that challenges assumptions and supports culturally safer care. This is particularly relevant in the era of AI because algorithmically generated scenarios may reproduce dominant norms unless diverse communities participate in design, validation and evaluation. Human-centred simulation must therefore remain culturally responsive as well as technologically responsible.
Faculty development is another unifying theme. Mossenson and colleagues present ‘Take Home Messages’, a card game designed for simulation faculty development [8], while Dick and colleagues apply cognitive apprenticeship to the teaching of surgical briefings [9]. Both contributions demonstrate that innovation can arise from thoughtful pedagogy rather than expensive infrastructure. Gamification can make faculty learning more accessible and participatory, while cognitive apprenticeship makes expert reasoning visible through modelling, coaching, scaffolding and reflection. These approaches are especially important as faculty roles evolve. Educators now require not only expertise in scenario design and debriefing, but also competence in evaluating AI outputs, recognising bias, protecting data and deciding when technology adds genuine educational value. Large language models can generate data quickly but have a potential for social and clinical biases. In SBHE, these risks could affect case authenticity, clinical accuracy, assessment validity and the representation of patients and communities. Responsible implementation therefore requires clear governance, faculty AI literacy, disclosure of AI use, independent verification of clinical content, data protection and continuing evaluation. International guidance on AI in health and education similarly emphasises human oversight, equity, transparency and accountability [10,11]. Efficiency should not become a substitute for educational rigour.
In this supplement, psychological safety remains a critical thread. It is explicitly foregrounded in the forensic psychiatry courtroom simulation and is implicit in work involving children, cultural safety, debriefing and faculty-supported AI. Psychological safety does not mean removing challenges. It means creating conditions in which learners can engage with uncertainty, error and emotionally demanding material without humiliation or avoidable harm [12]. AI tools should be assessed against this standard. Automated feedback that is decontextualised, overly authoritative or insensitive to learner readiness may undermine reflection. Conversely, carefully designed tools may help faculty identify patterns, support structured inquiry and personalise learning. The effect depends on implementation rather than technology alone. AI may help smaller or resource-constrained programmes generate adaptable scenarios, support novice faculty and analyse complex performance data. Generative systems may facilitate rapid prototyping of cases for rare events, interprofessional learning and systems testing. However, access to AI infrastructure, language models and technical support is uneven. Without deliberate attention to equity, AI could widen rather than reduce disparities in simulation education. Programmes should therefore evaluate cost, accessibility, language, cultural validity and local relevance alongside technical performance.
This supplement demonstrates that the future of healthcare simulation will be neither exclusively technological nor resistant to technology. It will be shaped by the integration of AI with established principles of experiential learning, skilled facilitation, cultural responsiveness, psychological safety and reflective practice. The nine innovations presented here collectively show a field extending its reach: from clinical performance to legal accountability, from faculty development to research integrity and from technical fidelity to cultural and emotional authenticity. Therefore, advancing simulation in the era of AI requires disciplined optimism. We should explore emerging capabilities but resist novelty for its own sake. We should use AI to strengthen educators, not marginalise them; to broaden access, not deepen inequity; and to improve evidence-informed reflection, not automate human judgement. The success of this new era will not be measured by how much AI is incorporated into simulation, but by whether its use produces safer learning, better-prepared professionals, more equitable systems and improved patient care.
None declared.
None declared.
None declared.
None declared.
None declared.
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
12.
[1]
[2]
[3]
anaesthesia,Anaesthetists' Non-Technical Skills (ANTS),ELAN,evidence-anchored feedback,medical education,multimodal annotation,non-technical skills,simulation,and simulation debriefing
Reliable assessment of non-technical skills (NTS) in medical simulation remains time-consuming and subjective. Anaesthetists’ Non-Technical Skills (ANTS) [1,2] provides structured behavioural categories, but routine use often yields scores and brief comments rather than a reusable record of the behavioural episodes underpinning each judgement. Timing, modality and interactional context are therefore lost. Prior video-supported simulation work has shown how operating theatre interaction can be reviewed through video and translated into teamwork/debriefing design [3]. Our contribution is to operationalise that logic as an ANTS-derived, debrief-ready annotation workflow. This gap matters for AI-assisted review, where experts want timestamped quotes and note that text alone misses contextual and non-verbal information [4]. We therefore used ELAN [5], a tool for fine-grained, timestamped multimodal annotation, to link selected NTS judgements to brief, reviewable evidence episodes.
Our framework is an ANTS-derived proof-of-concept workflow. Following senior physician review of five bradycardia simulations, we focused on three ANTS elements: Planning and Preparing, Authority and Assertiveness and Prioritising. Analysis was limited to the first 5 minutes, when information gathering and early deterioration management predominated.
The codebook comprised 21 positive and negative markers grouped by elements (Table 1). Markers retained the ANTS element structure and adapted global rating anchors. Three experienced physician examiners refined them during calibration as episode-level labels for visible behaviours in the reviewed clips. Markers were labelled by inference demand to distinguish directly observable behaviours from markers requiring greater expert judgement and attention. The ELAN schema (Table 2) centred on evidence episodes: the shortest time-bounded behavioural span judged sufficient to justify one or more markers, including necessary interactional context. Boundaries began when the relevant utterance/action sequence emerged and ended when that sequence or clinically meaningful interaction completed.

| ID | +/− | Name | Definition |
|---|---|---|---|
| Planning and preparing | |||
| PLN_POS_01 | + | States initial plan | Presents plan and next steps to team (shared mental model) |
| PLN_POS_02 | + | Role assignment | Explicitly assigns tasks/roles by name, gesture, eye contact |
| PLN_POS_03‡ | + | Anticipates next steps | Proactively thinks ahead about upcoming steps or resources |
| PLN_POS_04‡ | + | Updates plan | Updates plan with new information; communicates changes |
| PLN_POS_05‡ | + | Confirms understanding | Verifies team comprehension; invites questions/suggestions |
| PLN_NEG_01 | − | No plan | Acts without presenting plan/priorities; reactive approach |
| PLN_NEG_02‡ | − | Unclear delegation | Roles/tasks unclear or contradictory; team confusion |
| PLN_NEG_03† | − | Fails to update plan | Continues inappropriate approach despite situational change |
| Authority and assertiveness | |||
| AUT_POS_01 | + | Takes leadership | Establishes leadership role appropriately |
| AUT_POS_02 | + | Clear directives | Clear, concise instructions (who/what/when) |
| AUT_POS_03‡ | + | Closed-loop communication | Requests confirmation or read-back of instructions, ensuring task understanding |
| AUT_POS_04† | + | Assertive and respectful | Firm tone without aggression |
| AUT_POS_05‡ | + | Seeks help | Calls for senior/consultation/aid at appropriate time |
| AUT_NEG_01† | − | Hesitation/indecisiveness | Avoids decisions when needed |
| AUT_NEG_02 | − | Unclear orders | Confusing/contradictory instructions causing errors |
| AUT_NEG_03 | − | No closed loop | Tasks given without verifying receipt/completion |
| Prioritising | |||
| PRI_POS_01‡ | + | Life threats first | Addresses ABC/immediate dangers before secondary tasks |
| PRI_POS_02 | + | Delegates tasks | Delegates less urgent tasks; focuses on critical issues |
| PRI_POS_03† | + | Reassesses | Checks response/reassesses; updates priorities |
| PRI_NEG_01† | − | Fixation on low priority | Dwells on non-urgent details; delays critical steps |
| PRI_NEG_02‡ | − | Task overload | Too much in parallel; poor sequencing without clear prioritisation |
| Event scale levels (‘severity’) | Mild | Minor performance impact | |
| Moderate | Relevant teaching point | ||
| Significant | Important teaching point | ||
| Critical | Critical safety/leadership issue. | ||
| Element ratings | 1 | Unsafe/ineffective, requires immediate improvement | |
| 2 | Inconsistent, requires guidance, skill present but unreliable | ||
| 3 | Meets minimum standard, minor improvements needed | ||
| 4 | Effective and consistent; clear strengths; exemplary execution | ||
| 0 | Not assessable from available segment | ||
Note: Twenty-one behavioural markers (positive + and negative −) are grouped into three ANTS elements spanning two categories: Planning and Preparing, Authority and Assertiveness, and Prioritising, and are applied to time-bounded evidence episodes in ELAN. For each episode, annotators assigned the relevant marker(s) and the original pilot event scale field (‘severity’) reflecting the significance of the episode for debriefing, taking account of the clinical stakes of the scenario moment, whether arising from scenario demands, examinee actions or their interaction. Positive or negative handling was captured separately by marker valence. After all episodes in the analysed segment were annotated, each element received a global rating on a 0–4 scale (1 = unsafe/ineffective; 2 = inconsistent; 3 = meets minimum standard; 4 = effective/consistent; 0 = unassessable). Global rating anchors were adapted from the ANTS handbook [2]. The event scale is original to this protocol and was not derived from ANTS. † = high inference demand; ‡ = moderate inference demand. Unmarked markers are directly observable verbal or behavioural acts.
Abbreviations: ABC, airway-breathing-circulation; NTS, non-technical skills.

| Tier | ELAN type | Content |
|---|---|---|
| Leader/Nurse | Time-alignable | Pre-populated diarised transcript with timestamps |
| EVID_Event | Time-alignable | Parent evidence episode (timestamped span; may support one or more ANTS non-technical skills elements) |
| EVID_Planning | Symbolic_Subdivision + CV | Planning and preparing marker instances linked to parent event |
| EVID_Authority | Symbolic_Subdivision + CV | Authority and assertiveness marker instances linked to parent event |
| EVID_Prioritisation | Symbolic_Subdivision + CV | Prioritising marker instances linked to parent event |
| EVID_modality | Symbolic_Association + CV | Behavioural modality for the event (Verbal/Physical/Both) |
| EVID_severity | Symbolic_Association + CV | Original pilot single event-scale field (‘severity’) (Mild/Moderate/Significant/Critical) |
| EVID_QA_transcript | Symbolic_Association + CV | Transcript quality (Valid/Content error/Speaker error) |
| EVID_note | Symbolic_Association | Optional contextual notes (free text) |
| Planning_Rating | Controlled vocabulary | Global element rating (0 = unable to assess; 1–4) |
| Authority_Rating | Controlled vocabulary | Global element rating (0 = unable to assess; 1–4) |
| Prioritisation_Rating | Controlled vocabulary | Global element rating (0 = unable to assess; 1–4) |
Note: Indented tiers are child tiers of EVID_Event. The annotation file (.eaf) contains diarised, time-alignable transcript tiers (Leader/Nurse) and a parent evidence-episode tier (EVID_Event) marking the start and end of each behavioural episode used as the basis for scoring. EVID_Event marks a potentially multi-speaker behavioural episode over the diarised transcript; contextual meaning is carried primarily by the time-aligned transcript span, with additional detail recorded in EVID_note when needed. Element-specific marker tiers (EVID_Planning, EVID_Authority, EVID_Prioritisation) store marker instances linked to the parent episode as Symbolic_Subdivision annotations, allowing one episode to support multiple markers within the same ANTS element. Per-episode attributes are captured via Symbolic_Association tiers for modality (EVID_modality), the original pilot single event-scale field (EVID_severity), transcript quality (EVID_QA_transcript) and optional contextual notes (EVID_note). In the pilot, EVID_severity was intended to capture the significance of an episode for debriefing while taking account of the clinical stakes of the scenario moment, whether arising from scenario demands, examinee actions or their interaction. Global element ratings (Planning_Rating, Authority_Rating, Prioritisation_Rating) are recorded as controlled-vocabulary entries (0 = not assessable; 1–4).
Abbreviations: CV, controlled vocabulary; QA, quality assurance.
Each episode recorded modality, rating, interactional context and a piloted event scale intended to capture debriefing significance and clinical stakes. For example, the same 01:33–01:38 utterance (‘We need to start resuscitation. Let’s call for help’) was coded as positive Authority (‘seeks help’) but negative Planning (‘no plan’), preserving a mixed signal that a single score would obscure.
Annotation followed a three-step workflow: (1) identify evidence episode; (2) assign one or more element-linked markers and modality; (3) assign a global element rating (1–4, or 0 if unassessable).
We piloted the protocol on 19 resident-led bradycardia simulations. After calibration, two senior physician examiners received ELAN training and a one-page guide. Each independently annotated 14 study clips; nine were double annotated for agreement analysis. As a stress-test pilot, we report descriptive metrics without CIs.
Feasibility measures were annotation time, number of evidence episodes and annotated coverage. Agreement measures were global element ratings (quadratic-weighted κ), shared marker detection within a clip (symmetric F1) and 15-second per-marker temporal alignment (symmetric F1). Event-scale use was examined with Spearman correlations between clip-level mean event scores and element ratings.
Annotation required approximately 15–20 minutes per clip. Annotators identified a mean of 10.2 episodes per clip (SD 4.7), covering 82.5 seconds. Agreement was higher for Planning and Preparing (κw = 0.48) and Authority and Assertiveness (0.46) than for Prioritising (0.19). Raters were within ±1 point for 24/27 paired ratings, suggesting similar performance-band placement despite disagreement on exact scores.
Marker agreement varied by inference demand. Explicit markers performed best, including Seeks help (F1 = 0.92) and Takes leadership (0.86). High-inference markers, including Reassesses and Assertive and respectful, showed no overlap. In this small pilot, high-inference markers were less reliable, likely reflecting examiner thresholds and the difficulty of interpreting context and tone.
Requiring a 15-second alignment reduced the macro-average F1 by 0.21, indicating disagreement in boundary placement. Together, these patterns support current use for evidence-anchored debriefing and faculty review, where disagreement can guide discussion about the noticed evidence and its interpretation, but not formal assessment until marker rules, episode-boundary conventions and calibration are strengthened.
The event scale exposed a design problem. It combined clinical stakes with debriefing significance, although these may diverge: a clinically serious moment can be a less useful debriefing target, and a lower-stakes interaction may reveal important NTS behaviour. This ambiguity may explain why higher event scores aligned with higher ratings for one rater in Prioritising (ρ = +0.893), but with lower ratings for the other rater in Planning and Preparing (ρ = −0.718).
Next, we will expand the annotated dataset, tighten the high-inference labelling rules, split the event scale into clinical-stakes and debriefing-significance fields, use a third examiner for adjudication and test new scenarios. The framework is currently best suited to faculty debrief support and secondary analysis. Formal assessment or automated feedback applications require further calibration.
None declared.
None declared.
None declared.
The study followed the Declaration of Helsinki and its subsequent amendments. Participants provided written informed consent for participation and for research analysis. The Rambam Health Care Campus Institutional Review Board approved the study (0482-20-RMB).
None declared.
ASR and lexical normalisation were performed using faster-whisper (ivrit-ai/whisper-large-v3-ct2) and Google Gemini (gemini-3.1-pro-preview), respectively. Only de-identified transcript were analysed, in compliance with IRB approval. No audio or video left the institution, and all outputs were checked by the authors.
1.
2.
3.
4.
5.
[1]
artificial intelligence,computer-based/web-based/mobile technologies,computer science,debriefing,faculty development,large language models,simulation,and technology platforms
Delivering high-fidelity clinical simulation places significant cognitive demands on faculty. During live scenarios, facilitators must maintain clinical accuracy, manage scenario progression, support learners and prepare for an effective debrief, often while responding to unanticipated learner actions. These demands are particularly acute in emergency-based simulations for newly qualified doctors, where time-critical decision-making, escalation of care and structured handover are core learning objectives [1]. These competing educational, clinical and operational demands may contribute to increased faculty cognitive load during live simulation delivery, particularly in high-acuity scenarios.
Generative artificial intelligence (AI) has gained attention in healthcare education; however, most applications reported to date are learner-facing (e.g. tutoring) or used after sessions for content generation [2]. There is limited practical reporting of faculty-facing AI tools used during live simulation delivery. We identified a need for a low-friction, just-in-time faculty support tool that could assist with rapid access to credible clinical information and provide structured debrief scaffolds without disrupting scenario flow.
This short report describes the development, deployment and feasibility/usability pilot of a generative AI-supported faculty application designed as an educational innovation for live emergency simulation scenarios involving newly qualified doctors.
We developed and deployed a faculty-facing web application designed to provide cognitive support and facilitation augmentation during simulation delivery and debriefing in near real time. The web-based system uses a large language model (LLM) accessed via the OpenAI API (Application Programming Interface) to generate faculty-facing clinical prompts and debriefing scaffolds. The tool was designed as a bounded decision-support aid for simulation faculty rather than an autonomous clinical system. The LLM is not used to make clinical decisions, but to support simulation faculty with structured information retrieval and facilitation cues during live scenarios and debriefing. The tool was piloted during emergency medicine courses for newly qualified doctors. Scenarios focused on competencies including acute care management, escalation of care, structured communication and clinical deterioration recognition. Approximately 30 emergency simulation scenarios were conducted across six simulation days involving members of the simulation delivery team, including four medical faculty, three nurse facilitators and two simulation technical specialists with clinical backgrounds. Scenarios were 15–20 minutes in duration and involved learners working in pairs with a nurse faculty role-player embedded in the team. Additional participants acted as responding team members during cardiac arrest or clinical deterioration. Learners could escalate care and hand over to senior support (faculty) by telephone at any point.
The interface includes two modes:
1. Live support mode (during scenario delivery).
Faculty can enter a condition or clinical query to receive a concise, faculty-facing response intended to support facilitation. Outputs prioritise immediate actions, escalation steps, and, when relevant, drug treatments and doses. To support safe use and mitigate the risk of inappropriate clinical guidance, the tool implements guardrails to prioritise credible sources which are condition-specific responses grounded to a curated allowlisted set of emergency medicine reference pages (EMED). Explicit source citation and verification prompts are included in every response. If a condition is not recognised within the allowlist, the tool redirects the user to general EMED resources and avoids presenting definitive dosing or algorithmic instructions. Faculty remain the final decision-makers at all times. Outputs are framed for educational use only.
2) Debrief mode (post-scenario).
Faculty can paste a short scenario summary and debrief focus to generate a structured debrief scaffold aligned with the PEARLS framework [3]. The scaffold provides an opening script to establish psychological safety, suggested reaction and description questions, advocacy-inquiry prompts for analysis, targeted teaching points, take-home messages and a time-compressed ‘5-minute debrief’ option. This feature was designed to support newer faculty in organising and maintaining a consistent debrief structure.
The tool was iteratively refined during live delivery across multiple simulation days. Faculty feedback informed rapid adjustments to prompts, guardrails and interface elements between sessions, including refinement of response formatting to provide shorter, more facilitation-focused outputs during live scenario delivery, consistent with iterative design-based approaches in simulation education [4].
The application is used by faculty during emergency-based simulation to provide just-in-time clinical and facilitation support (Live support mode) and structured PEARLS-aligned debrief scaffolds (Debrief mode). Safety features include grounding to allowlisted trusted clinical references, human-in-the-loop decision-making and optional image generation. The tool is designed to support faculty cognition without disrupting scenario flow (Figure 1).


Conceptual model of faculty-facing generative AI support during live simulation delivery and debriefing
This innovation was evaluated as a feasibility and usability pilot rather than a research study. No learner outcome data were collected. Evaluation focused on faculty perceptions of usefulness, feasibility of use during live simulation and perceived impact on cognitive load and scenario flow.
Feedback was gathered through structured informal discussions during and after simulation sessions involving members of the simulation delivery team using the tool in real time. Feedback was synthesised using a rapid thematic approach to identify recurring usability, workflow and facilitation-support themes, with particular attention to the experiences of newer and more junior faculty members.
Most faculty reported that the application was feasible to use during live simulation without disrupting scenario delivery. The live support mode was most frequently used to confirm clinical guidance, check drug dosing and support escalation and handover prompts when responding to learner questions. Several participants reported that the Debrief mode supported more confident and structured debriefing, particularly for newer facilitators, by providing a clear scaffold and example phrasing consistent with PEARLS [3]. Faculty also reported perceived benefit in reducing cognitive burden during rapidly evolving scenarios. For example, the tool was used to rapidly generate escalation prompts and structured debrief questions during cardiac arrest and clinical deterioration scenarios.
The rapid iteration process allowed refinement of outputs in response to observed faculty needs across repeated runs of similar emergency scenarios. Faculty described the tool as useful and promising, with perceived benefit greatest for less experienced facilitators managing complex or rapidly evolving scenarios. More senior faculty reported selective use, primarily for confirmation and debrief structuring. No adverse effects on scenario flow were reported during the pilot period.
Future work will expand the tool to additional courses and clinical topics and will incorporate more formal evaluation, including structured usability measures and assessment of faculty cognitive load and confidence. Future work will also explore integration into broader simulation faculty development and programme delivery workflows. Further work is also needed on governance for safe use of generative AI during live simulation, including approach to source management, auditability of outputs and faculty development on appropriate use. The approach described here may be relevant to other high-acuity simulation settings where faculty must simultaneously manage clinical, educational and operational demands [2,5].
None declared.
None declared.
None declared.
This work was conducted as an educational innovation and quality improvement activity. It did not involve human subjects research and did not require institutional review board or research ethics committee approval. No identifiable data were collected.
None declared.
1.
2.
3.
4.
5.
[1]
[2]
[3]
[4]
AI-use disclosure,generative AI,simulated participants,simulated patients or standardized patients,and simulation scholarship
Artificial intelligence (AI) has made it possible to produce academic scholarship at an unprecedented speed. However, academic integrity is at stake given the bias automation and hallucination associated with any AI products [1]. Without involving people capable of critical thinking, the ‘human in the loop’ [2] approach, we risk degrading the quality of scholarship.
In health care/human simulation, a subcommittee from the Association of SP Educators (ASPE) Conference Committee reviews all abstract submissions (and peer reviewers’ comments) to recommend inclusion in the conference programme. The subcommittee was concerned that a few references were false because they referred to non-existent journals or titles. The subcommittee suspected the references were AI-generated. The research team conducted a series of post hoc analyses to evaluate the effectiveness of GenAI in detecting reference inaccuracy in the 2026 ASPE conference submissions. Based on the results from this study, we aim to create AI guardrails by ethically incorporating GenAI techniques in future abstract reviews and disseminating those guardrail strategies within the human simulation community.
This is a retrospective, observational study. We gathered all de-identified abstracts, references and peer reviewers’ comments submitted to the 2026 ASPE annual conference. Using GenAI, we scanned all reviewers’ comments with a keyword search prompt (Figure 1). The suspicious references identified by GenAI were then checked by the research team for inclusion or exclusion (‘human in the loop’). We further tested whether GenAI was able to detect additional fake AI-made references when embedded into the total list of references.


Flowchart for detecting problematic references from abstract reviewers’ comments
The research team wrote an AI prompt to examine all 484 reviewers’ comments in the 160 submitted abstracts. We ran the prompt in two GenAI platforms (ChatGPT Edu and Google Gemini 1.5 Pro) to create a list of comments that flagged potential reference issues. Next, two research team members examined the list, retaining entries that only pertained to problematic references. We defined a problematic reference as one ‘with no matching results for title nor authors’, ‘journal information incorrect’ and/or ‘journal doesn’t exist’. Then, the team crosschecked the suspicious references with Google Scholar and library databases (Figure 1). A total of 13 problematic references in five abstracts were found (Table 1).

| Abstract # | Example* | Common patterns of problematic references |
|---|---|---|
| Abstract 1 (Presentation) | Klaber RE, Watson J, Fraser A. Exploring SP perceptions of bias in simulation-based education. Simul Healthc. 2022;17(2):123–129 | No matching results for title nor authors; Journal title exists |
| Abstract 2 (Presentation) | Aldriwesh, M. A. (2022). Undergraduate-level teaching and learning approaches for interprofessional education in health professions: a systematic review. BMC Medical Education, 22(13). doi:10.1186/s12909-021-03073-0 | Matching results for title and journal; author info. missing |
| Abstract 3 (Presentation) | Copy-Paste/Data Integrity Flaws: Johnson, C. R., & Patel, D. S., The Hidden Dangers of Copy-Paste and Templating in Medical Records: A Failure of Attention to Detail. Journal of Clinical Documentation, 18(4), 210-218 (2019). | No matching results for title nor author; Journal listed does not exist. |
| Abstract 4 (Research/Poster) | Okuda, Y., Bryson, E. O., DeMaria, S., Jacobson, L., Quinones, J., Shen, B., & Levine, A. I. (2009). The impact of simulation-based training in medical education: A review. Mount Sinai Journal of Medicine: A Journal of Translational and Personalized Medicine, 76(4), 330–343. | Matching for title; author info was wrong; journal exists |
| Abstract 5 (Workshop) | Deering, S., Rosen, M., Ludi, V., Munroe, C., Pocrnich, A., Laky, C., & Napolitano, P. (2012). Speed mentoring: An innovative method to facilitate professional networking and mentoring among women in academic medicine. Journal of Graduate Medical Education, 4(2), 237–238. https://doi.org/10.4300/JGME-D-11-00183.1 | No matching for title no author; journal exists; doi does not exist |
| Series ID | AI-made references** | GenAI detection output |
| 1 | Benner, P., & Sutphen, M. (2022). Simulated empathy: Virtual patient encounters in advanced nursing education. Journal of Clinical Simulation Practice, 18(3), 145–159. | Suspicious; Authors are real but journal likely does not exist; article not traceable |
| 2 | Issenberg, S. B., & McGaghie, W. C. (2021). Augmented cognition in high-fidelity simulation training environments. International Journal of Medical Simulation, 9(2), 77–92. | Suspicious; Real authors but journal title not recognised |
| 3 | Gaba, D. M. (2020). Crisis resource management in hybrid human-digital simulation systems. Advances in Healthcare Simulation, 12(4), 201–218. | Suspicious; Well-known author but journal appears non-existent |
| 4 | Jeffries, P. R., & Rogers, K. J. (2023). Immersive simulation design for interprofessional collaboration outcomes. Nursing Simulation Review Quarterly, 27(1), 33–48. | Suspicious; Jeffries is real; journal likely fabricated |
| 5 | Cook, D. A., & Hatala, R. (2019). Adaptive simulation feedback systems for competency-based education. Journal of Medical Education Technology, 14(2), 88–104. | Suspicious; Real authors; incorrect or unknown journal |
| 6 | Tranvik, L. Q., & Moreno, F. J. (2024). Holographic mannequins in rural trauma training ecosystems. Global Journal of Synthetic Medicine, 6(1), 12–29. | Highly suspicious; Journal not known; topic sounds speculative |
| 7 | Patelson, R. K., Nguyen, T. P., & Olufemi, A. (2022). Neuroadaptive simulation loops for emergency response learning. Journal of Experimental Health Systems, 11(3), 201–219. | Highly suspicious; Unknown authors; journal not recognised |
| 8 | Dubois, C. L., & Henriksen, S. A. (2021). Emotion-responsive avatars in pediatric care simulations. International Review of Virtual Care, 5(4), 310–327. | Suspicious; Journal not recognised in field |
| 9 | Kowalski, M. Z., & Ibrahim, H. D. (2020). Biofeedback-integrated simulation for surgical precision training. Annals of Future Clinical Methods, 8(2), 55–70. | Highly suspicious; Journal likely fabricated |
| 10 | Alvarez, J. M., & Chen, Y. R. (2023). Quantum-driven scenario modeling in next-generation healthcare simulation. Journal of Theoretical Medical Engineering, 3(1), 1–18. | Highly suspicious; Buzzword-heavy title; journal not credible |
*The examples shown in the ‘Example’ column were original references that abstract authors submitted and that our research team identified as problematic references. One abstract may contain more than one problematic reference.
**AI-made references were those references generated using ChatGPT Edu.
We were also interested in GenAI’s capacity to recognise false references within the submitted abstracts. To test this, we prompted GenAI to create ten false references (Table 1) and inserted the ten AI-made references randomly into the complete references list. To observe whether GenAI could identify the problematic references, we ran a different prompt (Figure 1, Note 3) attempt to automate the scanning process for all references.
Our initial analysis showed that peer reviewers provided robust comments for identifying problematic references. GenAI extracted 124 reviewer comments with potentially problematic references. After the ‘human in the loop’ analysis, the final number of abstracts with problematic references was minimal (n = 5, 3% of all submissions; Table 1).
GenAI had difficulty distinguishing between real and false references. To evaluate GenAI’s detection of AI-generated problematic references, ten false AI-made references were modelled after the errors found in the abstract database. When those ten were embedded in the complete list of references, the automated AI scan identified only four as problematic, even though it had recognised all ten in a separate run of a fake-only reference list. The research team concluded that total automation of AI reference scanning may be quite limited in its effectiveness at this time.
One major finding from our study confirmed the quality of work by our reviewers, demonstrating their scholarly acumen and expertise in SP methodology [3]. There was a high level of accuracy among peer and research team reviewers in identifying problematic references. Overall, the prevalence of false references in the ASPE abstract submissions was minimal (below 5%), which is reassuring for submission quality.
Like other medical education researchers, we must recognise that our submission process has a gap in AI-use disclosure [4]. The ASPE Conference Abstract subcommittee will recommend the following AI guardrails for future submissions. Authors will be explicitly asked about AI use during abstract writing. Specifically, if they used AI to generate references, authors will be asked whether they have ensured that all references are accurate and formatted appropriately. Having this attestation during the submission process gives subcommittee reviewers an additional data point to guide the GenAI ‘human in the loop’ approach, furthering our aim of quality in the abstract submission and review process.
The research team would like to thank the Association of SP Educators (ASPE) Executive Committee for reviewing the manuscript. The views expressed in the manuscript are those of the authors and do not necessarily reflect the views of ASPE.
KX: Conceptualisation, Methodology, Data curation, Data analysis, Validation, Visualisation, Writing and Supervision.
KP: Conceptualisation, Methodology, Data curation, Data analysis, Validation and Writing.
KH: Conceptualisation, Methodology, Validation and Writing.
None declared.
None declared.
This project is not subject to human research ethics or institutional review board review.
The authors have no conflicts of interest to declare.
1.
2.
3.
4.
[1]
clinical or other education,medical education,paediatric care,simulated patients or standardized patients,and simulation modality
Simulation-based education (SBE) is well established in undergraduate health professions training, including in specialties like paediatrics. However, most paediatric simulations rely on manikins or adult simulated patients, limiting students’ opportunities to experience clinical and communication skills with real children and families prior to clinical placements [1]. At the University of Dundee, medical students complete a pre-clinical paediatrics block in second year and a clinical placement in fourth year. Student feedback consistently highlights a paucity of face-to-face interaction with young children in earlier years and feelings of under-preparedness for paediatric examination during clinical placements, a challenge reflected in the wider literature [2].
Studies of children and adolescent simulated patients (CASPs) predominantly focus on adolescent populations, with limited literature describing the involvement of younger children (0–12 years) in undergraduate education [3]. Given the influence of developmental stage on clinical examination, communication and rapport-building, there is a need for educational innovation that allows students to practise paediatric clinical skills with children across a range of developmental stages in a supportive environment and to augment their clinical paediatric placements.
We designed a simulation-based teaching session, inviting younger CASPs (aged 0–12 years) with their primary carers to engage with our undergraduate medical students in rehearsing paediatric clinical skills across developmental stages.
A pilot session was delivered in March 2025 for 20 year 3 medical students, who were recruited by poster. The session consisted of a structured pre-brief, which outlined the session’s aims, learning outcomes and included a focused revision of paediatric examination techniques and developmental milestones.
Students then rotated in pairs through five simulated clinical scenarios, each facilitated by a clinical tutor and involving a CASP and their parent or carer. Scenarios were aligned with core undergraduate paediatric curriculum topics and included cardiovascular and respiratory examinations, communication skills, and gross and fine motor developmental assessments. Communication scenarios focused on rapport-building using child-led conversation, without requiring CASPs to memorise scripts. Following the scenarios, a faculty-led debrief was conducted. Scenarios were selected to help facilitate a small pilot study but would be extended to include more clinical skills and communication scenarios in future teaching sessions.
CASPs were recruited from within the local paediatric department, via a staff email distribution list. Parents provided informed consent on behalf of their children. Children were matched with scenarios based on age and developmental stage. More CASPs were recruited than needed to minimise pressure and allow withdrawal at any point [4]. Ethical approval was granted by the University of Dundee.
This small-scale study invited students to participate in the pilot session and a follow-up focus group to explore their experiences of interacting with younger CASPs. To guide future curriculum developments, we captured post-event data and analysed transcripts using Braun and Clarke’s reflexive thematic analysis [5]. This qualitative approach was chosen to explore perceived educational value, confidence, realism and application of prior learning.
Three key themes were identified with representative quotes from students.
Gaining confidence: students reported reduced anxiety and increased confidence when interacting with children, describing the experience as preparation for future clinical placements.
•‘Takes away quite a lot of that scariness’.
Realistic exposure to developmental stages: students valued the realism of interacting with children across different developmental stages in a low-stakes educational environment, highlighting the opportunity to adapt examination and communication skills to children who were not always compliant or cooperative.
•‘Learning to cope with uncooperative kids is quite useful’.
Building on prior learning: students described the simulations as enabling them to apply previously learned theoretical knowledge of developmental milestones and paediatric examinations to real interactions, helping them understand how clinical skills differ between adults and children.
•‘I know what to do with an adult, but I didn’t know what parts [of the examination] to change in a child’?
Overall, students perceived the innovation as a valuable, authentic learning experience that supported their preparedness for paediatric clinical placements.
This pilot demonstrates the educational value of incorporating younger CASPs into undergraduate paediatric simulation. It also gave us insight into the local feasibility of incorporating CASPs into our curriculum. Future work should address sustainable CASP recruitment, with planned collaboration with local schools and early-years education settings, aligning participation with educational curricula to minimise disruption to school and caring schedules. Future teaching session design should align with children’s developmental age. Performing clinical skills in paediatrics requires adaptation to the child’s developmental stage; therefore, ensuring these are considered when designing teaching sessions is very important.
Further research should explore the experiences of CASPs and parents/carers, especially considering the psychological safety of CASPs and the benefits to patient volunteers. The longitudinal impact of this innovation on clinical performance during paediatric clinical placements and student confidence would be valuable to explore. This innovation may be transferable to other institutions seeking to enhance paediatric preparedness through developmentally appropriate simulations with CASPs.
None declared.
None declared.
None declared.
None declared.
None declared.
1.
2.
3.
4.
5.
[1]
educational design,innovation,medical education,realism,fidelity,authenticity,and simulation
The authors describe a novel, long-term simulation encounter with a patient journey culminating in an immersive, high-fidelity coroners’ court attendance. Survey data highlights that UK foundation doctors feel underprepared for medico-legal encounters [1]. Whilst there have been a range of studies to address this [2], only one appears to utilise an actual courtroom setting [3]. However, to the authors’ knowledge, this is the first description of an in situ courtroom simulation using the learners’ own medical documentation from a simulated case over 12 months prior.
This innovative programme was comprised of three sequential components: (i) a ‘busy day on-call’ simulation (BDOC) for new starter Foundation Year 1 (FY1) doctors, (ii) a complaints session and (iii) a coroners’ court inquest. The coroners’ inquest occurred over 12 months after BDOC for the same cohort following entry to Foundation Year 2 (FY2).
During their shadowing week FY1 doctors were divided into groups in separate classrooms to cover simulated ‘wards’. Each group had a task list for completion during the 2-hour session, requiring teamwork and prioritisation. At intervals doctors were individually given a simulation scenario of a patient having an acute asthma exacerbation using a high-fidelity manikin. All appropriate documentation from the doctors was labelled and stored.
During a teaching session 8 months later, the cohort were given a fictitious written complaint from the spouse of the simulated patient who had subsequently died following admission to intensive care. With a copy of their BDOC notes and generic guidance on responding to complaints, each doctor formulated their own response.
Four months later, the now FY2 cohort were sent a formal coroner’s letter drafted by NCIC solicitors, requiring a written statement and summons for an inquest hearing. The letter pack also included their original written notes, the complaint letter and their response. The trust solicitor, an experienced former coroner, selected three statements deemed to have significant learning potential.
Thirteen months following BDOC, the cohort attended a high-fidelity coroners’ court simulation in the former Carlisle Citadel Crown Courtroom (Figure 1), with the trust solicitor and assistant assuming coroner roles. The introductory presentation established the process of the session and learning objectives, with emphasis on psychological safety and the importance and acceptability of making mistakes. Three doctors were individually called to the witness stand and questioned ‘under oath’ by the ‘coroner’. Debriefing was conducted by experienced simulation faculty, with discussion and feedback from the solicitor.


Foundation doctor giving oral evidence under oath during coroners’ court simulation in the former Carlisle Citadel Crown Courtroom
The sequential nature of the simulation created a realistic psychological context for the learners. The faculty acknowledged that the doctors would have already demonstrated significant improvement in their clinical knowledge as well as written record keeping since their first week as an FY1. The aim was not for them to feel judged based on their performance, but to provide a realistic experience of attending a coroners’ court.
Thirteen months between BDOC and the court simulation formed a realistic timeframe for a patient death to a coroners’ inquest. Accurately, the doctors had an imperfect memory of the patient encounter and were reliant on their written notes for some details. Utilising a real superannuated courtroom and a solicitor with coroner’s experience led to a high-fidelity environment and proceedings. The cohort showed a range of initiative; many wrote a letter rather than a statement. Some learners strayed from establishing facts into providing opinion. Learning points were discussed afterwards along with the coroner’s recommendations for providing a statement. Pre- and post-session questionnaires assessed attendees’ confidence in coroners’ proceedings. The same questionnaire was used for follow-up 6 months later.
The 26 FY2 doctors in attendance provided written feedback rating their confidence on a four point scale before and after the session; not confident, satisfactory, confident and extremely confident. Reported confidence in understanding courtroom proceedings (Mdn 1, 2.5; Z = 4.01, p < 0.001, r = 0.85, n = 24), providing a written statement (Mdn 1, 2; Z = 3.82, p < 0.001, r = 0.88, n = 23) and answering coroners’ questions (Mdn 1, 2; Z = 4, p < 0.001, r = 0.92, n = 23) all increased significantly using Wilcoxon ranked-sum testing (median before, after; z-score, p-value, effect size). Sixteen respondents to the 6-month follow-up demonstrated maintained higher confidence ratings in all domains, compared to pre-session (Figure 2), shown to be significant using Kruskal–Wallis tests (data not shown).


Reported confidence ratings at the coroners’ court simulation ‘Before’ opening talks, ‘After’ the session close (n = 26) and at 6-month follow-up (n = 16) using a four-point scale (% frequency). (A) Confidence in providing written statements for legal proceedings. (B) Confidence in providing oral evidence in a coroners’ court
All FY2 doctors stated, at the end of the session and at 6 months, that the simulation session was useful (n = 25 and n = 16 respectively). FY2 doctors were asked to rate their agreement to the statement ‘I will make changes to my written documentation…’ following the session on a five-point scale. Twenty responses either agreed or strongly agreed (Median 4, IQR 1; n = 25) and responses were not significantly different at 6 months’ follow-up (U = 186, n1 = 25, n2 = 16, p = 0.354) using Mann–Whitney U testing.
Our aim to facilitate deeper learning using innovative methods was reflected in the feedback, with the majority of doctors confirming increased confidence and changed documentation practices, sustained at 6-month follow-up. Further research in this area will be important to determine the real-world impact of this type of simulation.
The authors wish to acknowledge Professor Matt Phillips, Honorary Clinical Professor Genitourinary Medicine (ULan), Deputy Medical Director NCIC NHS FT, for the original concept and his integral role in the planning and design of this novel simulation programme. The authors also acknowledge Ward Hadaway LLP for their role in implementing the courtroom simulation. Especial thanks to Mr Neil Smart, Barrister, for acting as Coroner and providing expert advice during debriefings.
None declared.
None declared.
None declared.
None declared.
None declared.
1.
2.
3.
[1]
[2]
[3]
[4]
expert witness,interprofessional collaborative practice,medical education,medico-legal,mental health,psychiatry,and simulation
Providing medico-legal evidence is core to Forensic Psychiatry, while Scottish Forensic Psychiatry residents gain experience in preparing written court reports, opportunities to provide oral evidence in court are infrequent. This training gap can negatively influence their performance as an expert witnesses and the legal process [1]. Simulation offers a safe, feedback-rich environment to build these skills [2] but face two distinct design challenges:
1.Resource barriers: the high financial and logistical costs of sourcing experienced faculty [2].
2The psychological safety paradox: maintaining an authentic adversarial challenge while ensuring a safe learning environment that prevents anxiety [3].
To address these intersecting challenges, we developed an interdisciplinary, high-fidelity courtroom simulation workshop.
To address the training gap, we designed a high-fidelity simulation workshop. Using fictional psychiatric court reports featuring divergent opinions helped ensure conceptual fidelity. Physical fidelity was enhanced by arranging large lecture rooms to mimic courtrooms and casting the audience as jurors. Video recording of performances were provided for self-reflection.
To overcome resource barriers, we pooled resources with local law schools and the Faculty of Advocates. This guaranteed a genuine experience that improved psychological fidelity. We also hoped that this interdisciplinary partnership would enrich the discussion in the debrief process.
To navigate the psychological safety paradox, we deliberately balanced the high fidelity of the simulation with psychological safety measures (Table 1), ensuring adversarial pressure remained educationally constructive.

| Safety measure | Description |
|---|---|
| Balanced opinions | Divergent clinical opinions are used to facilitate realistic cross-examination, the most challenging aspect of courtroom testimony. The authors also ensure that disagreements remain within the bounds of sound clinical reasoning. This approach helps improve conceptual and psychological fidelity of simulation. |
| Preparing appropriate expert witness reports | Fictional reports are utilised to prevent coaching expert witnesses on unresolved cases, which is illegal in Scotland. Initially, authors prepared these reports to reduce participant burden and provide psychological distance from opinions. However, following resident feedback favouring personal involvement and stylistic preference, participants now draft their own reports. Authors carefully review these submissions to ensure balanced arguments, as well as maintain conceptual and psychological fidelity. |
| The choice of legal participants | In choosing the legal participants in the simulations, the authors collaborate with law colleagues to nominate Advocates and Advocates in Training who prioritised joint learning over competitive participation. |
| Witness protection boundaries | Expert witnesses are explicitly prohibited from engaging in personal attacks and are given the autonomous option to pause proceedings for structural breaks. |
| Constrained adversarial roles | Counsel (Procurator Fiscal/Defence) is formally instructed to avoid aggressive, hostile tactics or personal degradation. |
| Judicial intervention | The Sheriff is explicitly tasked with assisting experts with courtroom procedures and must actively intervene if cross-examination becomes inappropriate. |
| Familiar peer roles | More senior peers serve as neutral figures (compared to the adversarial law colleagues) facilitating the workshop and keeping time. |
| Reflective co-facilitation | Ensuring joint medical and law faculty facilitation to deliver informed educational debriefing rather than a purely performance-based critique. |
| Structured debriefing using the PEARLS framework (Promoting Excellence and Reflective Learning in Simulation) | Utilising the PEARLS structured debriefing framework, grounded in ‘unconditional positive regard’ fosters psychological safety while offering a flexible, blended approach of facilitated self-discovery and directive teaching. This allows educators to tailor feedback by highlighting excellence or addressing specific performance gaps [4]. |
| Opting for self-assessments | To reduce participation anxiety, the authors choose self-assessment questionnaires over formal evaluations by educators or simulated jurors. Consequently, video recordings are shared exclusively with the respective participants to facilitate private personal reflection. |
We iterated the workshop format three times based on participant feedback. The refined model involves two mock criminal trials preceded by specialist tutorials covering key clinical and legal concepts that are chosen based on the feedback of previous participants. In each 40-minute trial, legal professionals play the roles of Sheriff, Procurator Fiscal and Defence, while two or four residents act as competing expert witnesses. The authors facilitate the session and lead a concluding 20-minute structured debrief.
We aimed to measure the workshop’s impact on residents’ self-assessed confidence as expert witnesses, and to capture their qualitative experience. Pre- and post-workshop self-assessment feedback forms were developed after a literature review of effective expert witness domains, ensuring they reflected skills directly relevant to courtroom practice. These included quantitative 5-point Likert-scale questions, analysed via paired samples t-tests (to detect within-subject change across matched pre- and post-intervention scores). Moreover, open-ended questions helped gather qualitative data about the attendees’ experience, capturing dimensions that quantitative scales alone could not reflect.
We delivered this workshop three times. Five Forensic Psychiatry residents attended once, four residents attended twice and seven residents attended thrice. Medical students and core psychiatry residents attended as well. Table 2 shows that mean score changes after the first workshop were positive and statistically significant across all quantitative domains. However, these scores were relatively and progressively lower in the latter two workshops. This trend across the three workshops corresponded with progressively higher reported baseline scores, suggesting a practice effect with the residents retaining skills and confidence over 18 months [5].

| Domain | Workshop 1 (N = 15) | Workshop 2 (N = 9) eight residents attended one workshop before | Workshop 3 (N = 10) three residents attended one workshop before seven residents attended two workshops before | ||||||
|---|---|---|---|---|---|---|---|---|---|
| δM | t (df) | p | δM | t (df) | p | δM | t (df) | p | |
| Preparation (reviewing prepared reports to clarify their clinical opinion and prepare for being questioned about diverging matters) | 1.4 | 5.1 | 0.0002 | 0.7 | 2.8 | 0.02 | 0.7 | 3.3 | 0.01 |
| Addressing the Sheriff (appropriate court etiquette and judicial communication) | 1.3 | 3.8 | 0.002 | 1.6 | 4.6 | 0.002 | 1.1 | 3.5 | 0.007 |
| Affirming/taking the oath (executing the formal ritual to testify, serving as the psychological transition into the high-fidelity witness role) | 1.5 | 4.2 | 0.0009 | 0.9 | 4.4 | 0.002 | 1 | 4.7 | 0.001 |
| Stating qualifications (presenting credentials under scrutiny to establish professional credibility) | 1.3 | 4.5 | 0.0005 | 0.8 | 5.3 | 0.0007 | 0.7 | 4.6 | 0.001 |
| Direct-examination (responding to questions from the instructing party) | 1.4 | 5.1 | 0.0002 | 1.1 | 4.3 | 0.003 | 0.5 | 3 | 0.015 |
| Cross-examination (responding to challenging questions from the opposing party) | 1.6 | 6.3 | 0.00002 | 1.2 | 5.5 | 0.0006 | 0.5 | 2.2 | 0.052 |
| Communication (clarity and effectiveness of verbal delivery) | 1.2 | 6.9 | 0.000007 | 0.8 | 3.5 | 0.008 | 0.7 | 3.3 | 0.01 |
| Giving evidence generally (overall performance as an expert witness) | 1.1 | 5.2 | 0.0001 | 1 | 4.2 | 0.003 | 0.3 | 1.2 | 0.279 |
Across the three workshops, the reported score improvements were greater among those who participated in the simulated trials, those with less courtroom experience and those at earlier stages of their residency.
Pre-workshop, the most frequently described challenge by both residents and law attendees was reciprocal intimidation. In addition, the residents commonly reported fear of public embarrassment and limited prior courtroom experience. These themes highlight the inherent tension between psychological threat and safety, which the authors sought to navigate.
The complete unfamiliarity with the process and the prospect of being publicly undermined (Forensic Psychiatry Resident)
The knowledge gap between professionals undermines witness control and flow of evidence (Law Attendee)
The most valued aspect of the workshops, by both residents and law attendees, was the integration of legal and clinical perspectives in debriefing. Moreover, residents appreciated practising direct/cross-examination, and the blended didactic-simulation format. These findings suggest that the interdisciplinary model was mutually enriching.
It was very informative with real lawyers (Forensic Psychiatry Resident)
Discussing participants’ performance and general points with lawyers (Forensic Psychiatry Resident)
Opportunity for dialogue between experts and council (Law Attendee)
While the specialist tutorials received unanimously positive feedback for addressing knowledge gaps, the widely varied suggestions for future topics highlight the need for a standardised psychiatric expert witness curriculum.
The simulation workshops are now integrated into the Scottish Forensic Psychiatry Residents Teaching Programme. To ensure longitudinal sustainability, the authors were recruited at different stages of their residency.
Further growth may include a standardised expert witness curriculum, diverse themes (e.g. civil trials or fatal accident inquiries) and widening the target audience (e.g. non-forensic psychiatry residents and other healthcare professionals).
Ultimately, the authors aim to share this psychologically safe interdisciplinary model to facilitate reproduction by other educators.
The authors wish to acknowledge law colleagues who helped deliver the workshops:
1.Edinburgh Napier University Law School, Edinburgh.
2.University of Strathclyde Law School, Glasgow.
3.Faculty of Advocates, Scotland.
None declared.
None declared.
None declared.
The NHS Health Research Authority decision-making tool has indicated that ethical approval is not required for this work as it is not research but a service evaluation/improvement/development project.
The authors have no conflicts of interest to declare.
1.
2.
3.
4.
5.
[1]
cultural safety,Indigenous health,innovation,nursing education,and simulated participants
In the southwest of Western Australia, on the unceded lands of the Noongar peoples, hundreds of registered nurses graduate from Murdoch University each year. Aboriginal and Torres Strait Islander (hereafter respectfully referred to as Aboriginal) health content is mandated by the Australian Nursing and Midwifery Accreditation Council (ANMAC) for all entry-to-practice nursing programmes. Despite this, students report uncertainty, anxiety and limited engagement with Aboriginal people, highlighting ongoing challenges in translating cultural safety knowledge into clinical practice. A persistent educational gap remains in nursing education: students struggle to translate cultural safety theory into culturally safe communication and practice in real clinical interactions with Aboriginal people.
Previous iterations of the unit relied primarily on didactic teaching approaches, which are misaligned with Aboriginal ways of knowing, being and learning that emphasise relational, experiential and community-centred approaches [1].
Simulation-based education provides an immersive alternative for experiential learning. In cultural safety education, simulations co-designed with Aboriginal people promote cultural authenticity and safety [2]. Evidence demonstrates that simulated cultural communication scenarios developed in partnership with Indigenous actors enhance rapport-building, cultural humility and learner confidence in cross-cultural clinical encounters [3]. In contrast, simulations lacking authentic Indigenous voices risk reinforcing stereotypes or superficial representations of culture [4]. In response, a culturally grounded, Noongar-led Simulated Participant (SP) simulation was developed to support nursing students in translating cultural safety learning into authentic clinical interactions with Aboriginal people.
A simulation was co-designed with Aboriginal SPs for undergraduate nursing students to support the application of cultural safety knowledge within a realistic and psychologically safe learning environment. The aim was to strengthen students’ capacity for culturally safe communication, rapport-building and engagement in culturally sensitive clinical interactions. The design intent of the simulations emphasised relational learning and culturally safe practice, enabling students to engage authentically with scenarios that reflected the experiences and cultural contexts of Aboriginal people through a strengths-based lens. This simulation was conceptually grounded in cultural safety theory and experiential learning theory, informed by Aboriginal ways of knowing, being and doing, and emphasising relational, reflective and immersive learning.
The scenario depicted a Noongar boy presenting to a clinic with his family with symptoms of acute otitis media. A First Nations Nursing Advisory Panel reviewed the scenario and provided cultural governance, strengthening cultural nuance, relational dynamics and language.
Collaboration with Aboriginal actors aligns with international evidence showing that Indigenous actor involvement enhances authenticity and cultural realism in simulation, enabling students to practise culturally safe communication without risk of harm [4].
The simulation was designed to support psychological safety and reflective learning. Participation in active roles was voluntary, and students were provided with full access to the scenario one week in advance to minimise uncertainty and avoid surprise. The clinical focus aligned with previously taught content on ear health, and students were given preparation time in class prior to the simulation. Learners participated in small groups of three and were able to pause or withdraw at any time using the phrase ‘this is not a simulation’. Throughout the activity, facilitators observed via live video stream and could intervene as needed to support student safety and learning.
Sessions involved two groups of three active participants, with the remaining students observing via live video stream. Observers remained in the classroom with a second facilitator and were guided to actively observe both the SPs and student nurses. Observation was structured through three reflective prompts: identifying what was done well, moments that felt awkward or uncomfortable and considering what they might do differently in a similar situation. This structured observation supported engagement, critical reflection and vicarious learning while reducing performance-related pressure for students not in active roles.
A structured pre-brief prepared students for the sensitive nature of the learning and normalised uncertainty and mistakes. Yarning, an Aboriginal conversational approach grounded in storytelling and dialogue [2], was incorporated into the debrief to support culturally safe reflection.
Aboriginal SPs played an active role in both scenario delivery and reflective dialogue during the debrief, drawing on lived experience to provide real-time feedback, share perspectives on navigating healthcare as Aboriginal people and support learners’ understanding of cultural power dynamics and respectful inquiry [3].
A post-simulation survey (n = 202) evaluated the learning experience of active participants and observers, with a mean satisfaction rating of 8.79 out of 10. The post-simulation survey was developed using a locally designed evaluation tool. Data were analysed using descriptive statistics for Likert-scale items, while open-ended responses were reviewed and grouped into common themes to summarise participant feedback. The survey explored perceived confidence in culturally safe communication, understanding of cultural safety principles, emotional responses and relevance to clinical practice. This approach is commonly used in Indigenous cultural simulation research and utilises learner reflection and self-report to assess awareness, rapport-building and cultural humility [3]. Overall, the reported outcomes indicate increased learner confidence, strengthened culturally safe communication skills and the development of cultural humility through reflection on relational practice.
Students described the simulation as challenging yet valuable. Many students reported initial discomfort performing in front of peers; however, this discomfort was recognised as important preparation for respectful engagement with Aboriginal people in practice. Both active participants and observers reported meaningful learning, indicating that vicarious engagement through observation can be effective in culturally sensitive simulations, consistent with the literature [5].
Learner reflections indicated increased awareness of cultural safety as a relational and reflective practice rather than a checklist of communication behaviours. As one student noted, ‘I entered this unit thinking being culturally safe meant being kind and compassionate to all backgrounds. I have been blown away by how much I have actually delved deep into self-reflection and realised that it’s not about treating everyone “equally” but acknowledging and respecting the differences and adapting or adjusting my approach as required’.
One class declined active participation due to fear of ‘getting it wrong’, reflecting the emotional weight of cultural safety learning. This aligns with findings by Marr et al., who note that cultural simulations may evoke anxiety and fear of causing offence, underscoring the importance of supportive facilitation and psychologically safe learning environments [3].
The SPs reported positive experiences and observed high levels of student engagement during post-simulation discussion. The yarning-based debrief supported a culturally safe space for questions about greeting protocols, family structures, connection to Country and relational dynamics, reinforcing the value of learning directly from Aboriginal people and co-constructed, lived-experience-led teaching [1]. Some students requested additional cultural simulations, recognising their potential role in bridging cultural theory and nursing practice.
Future developments include partnering with Aboriginal healthcare teams to deliver simulations for undergraduate and postgraduate nurses. Although opportunities within hospital-based education are limited, this approach has potential for transfer to other settings. Planned enhancements include improved student preparation, increased opportunities for cultural reflection and additional community-based scenarios. Ongoing collaboration with Aboriginal community members will remain central, strengthening partnerships and supporting culturally responsive health professional education. Sustainability of this initiative is supported through ongoing partnership with Aboriginal community members and advisory input, ensuring cultural governance, continuity of relationships and accountability in future simulation development. While this simulation was designed within a Noongar context, the co-design principles and use of culturally led SPs are scalable to other Indigenous and culturally diverse contexts, as well as to other health disciplines. This innovation demonstrates the role of culturally authentic simulation in bridging cultural safety theory and clinical practice, supporting nursing students to deliver safe, respectful, relationship-centred care for Aboriginal people.
The authors would like to thank Della Rae Morrison, a proud Bibulman Noongar woman who was instrumental to the success of this simulation, along with her colleague Marlanie Haerewa, a proud Nyikina and Ngāti Porou woman.
None declared.
No funding was received for the writing of this article.
None declared.
No ethics approval was required for this article.
No conflict of interest has been declared by the authors.
1.
2.
3.
4.
5.
[1]
[2]
[3]
[4]
debriefing,education,faculty development (simulation educator or simulation technician),gaming,and simulation
High-quality debriefing in healthcare simulation is essential to maximise participants’ reflection and to foster experiential learning [1]. Effective debriefers use a structured, yet flexible, approach that demonstrates artistry, promotes learner-centredness and grounds debriefing in psychological safety [2]. The journey for developing debriefing skills can be augmented by formal training and opportunities for reflection on practice. As leads of the not-for-profit company Vital Anaesthesia Simulation Training Ltd (VAST) [3], we have noticed challenges for novice facilitators in integrating various debriefing techniques into practice without a supportive structure and opportunities for practice.
The VAST Facilitator Course introduces core simulation theory and helps develop skills for design, delivery and debriefing of simulated scenarios. For many, it is the starting point of their simulation faculty development journey. The course is followed by opportunities for mentorship, self-reflection and ongoing practice. To augment the pedagogical approaches already embedded in the VAST Facilitator Course, we were motivated to design and implement a serious game that could offer rapid-cycle deliberate practice [4] and reflection on the core techniques used in debriefing. Serious games, when thoughtfully applied, can promote active learning, enhancing knowledge and skill development [5].
Our innovation is Take Home Messages, a card game that replicates elements of the analysis phase of debriefing, the core component of many debriefing frameworks, including that used in VAST [3]. Game design was grounded in theoretical frameworks relevant to simulation: experiential learning, cognitive load management, deliberate practice and reflective learning [6]. Take Home Messages is designed to scaffold skill development of a structured debriefing framework, to promote best-practice debriefing techniques [7], such as the advocacy–inquiry method [8] and to guide novice facilitator reflection. There are short clinical vignettes that require players to make observations and ask advocacy–inquiry questions. Subsequently, players use a variety of techniques to progress the debriefing through expanded exploration of ideas, finishing with application of learning to practice: the ‘take home message’.
The game Take Home Messages is played with custom-made cards designed to scaffold learning while still challenging participants to apply core debriefing techniques (Figure 1). Designed for four to eight players, Take Home Messages is played in a competitive and engaging environment where players earn points as they progress their debriefs to the ultimate step of applying learning to practice. There are opportunities to deploy wildcards that either enhance one’s debriefing, gaining points for the player or hinder a debriefing, decreasing a competitor’s points. Reflection on performance occurs in-action during gameplay, with overall reflection on learning from gameplay occurring at the conclusion of the game. Take Home Messages is introduced and played in a 1-hour session within the VAST Facilitator Course after course participants have already had interactive discussions on core debriefing theory, conversational techniques and VAST’s debriefing framework.


This figure shows a selection of Take Home Messages cards. The game scaffolds debriefing from initial advocacy–inquiry questions on an observation (top left) to exploration of specific concepts and application of learning into future practice (bottom right). The debriefing, and thus points scored, can be modified by ‘debriefing destroyers’ (red cards) and ‘debriefing defenders’ (green cards). Off-topic discussions are time wasters of no value
Take Home Messages was beta-tested during the VAST SIMposium in Nairobi, Kenya in 2024 (Figure 2). This international healthcare simulation faculty development event involved representatives from 21 countries who all played Take Home Messages and provided feedback in facilitated discussions during the session and anonymously in post-course evaluations. The game has undergone further iterative design through its application in faculty development initiatives in Australia, Bolivia, Canada, Fiji, Guatemala, India, Mongolia and Peru. Design changes included rationalisation of cards, simplification of gameplay and introducing the game in phases: first to build familiarity with rules, next introducing simulated debriefing and finally adding time pressure to simulate requirements for rapid thinking during debriefing.


A group of simulation facilitators playing Take Home Messages at the VAST SIMposium in Nairobi, Kenya
Feedback following gameplay has been consistently positive across this wide range of contexts, garnered through facilitated discussions during the course and anonymous post-course evaluations. In our informal thematic reflection on this feedback, we have found that playing Take Home Messages during simulation faculty development courses is engaging, builds confidence and provides an opportunity for skill integration through rapid and repeated practice of core debriefing elements. For example, a VAST Facilitator Course participant’s reflection on gameplay was: ‘The debriefing card game was interesting and [I] enjoyed it a lot. We could witness the transition in us from learner to facilitator’.
We are incorporating play of Take Home Messages into all VAST Facilitator Courses. In addition, we are developing a multi-centre randomised controlled trial to formally evaluate the impact of gameplay on development of simulation facilitation skills. This will be a mixed-method study, with quantitative exploration of how Take Home Messages influences the quality of advocacy–inquiry questions, according to the Advocacy-Inquiry Rubric [9] and qualitative focus groups that investigate the perceptions of participants and facilitators on how gameplay influences debriefing skill development. If you are interested in playing for yourself, cards can be purchased (here), with all proceeds going directly to support VAST’s activities in low-resource settings. The game is inherently adaptable across contexts through the use of ‘house rules’ a common board game design practice. It allows facilitators to modify gameplay to suit local needs. For example, in Mongolia it was adapted to a team-based format. Take Home Messages has been translated into French and Spanish for wider dissemination in VAST’s courses and beyond. We hope you find Take Home Messages both fun and educational. We would love to hear your feedback on the game.
Take Home Messages was designed by Drs Adam I. Mossenson, Patricia Livingston and Rodrigo Rubio-Martinez. We gratefully acknowledge Matthew Clarke, who created the artwork for the cards, and the many people around the world who have been involved in testing gameplay and providing feedback.
None declared.
None declared.
None declared.
None declared.
Drs Adam I. Mossenson, Patricia Livingston and Rodrigo Rubio-Martinez are the creators of Take Home Messages. Drs Adam I. Mossenson and Patricia Livingston are directors of Vital Anaesthesia Simulation Training (VAST) Ltd, which is a registered charity in Australia. Their work on VAST is entirely voluntary and non-remunerated. Proceeds from sales of Take Home Messages go entirely to VAST Ltd. Drs Adam I. Mossenson, Patricia Livingston and Rodrigo Rubio-Martinez have no financial benefit from sales of Take Home Messages.
1.
2.
3.
4.
5.
6.
7.
8.
9.
[1]
[2]
innovation,medical student preparation,operating theatre,surgery,and theory
The perioperative surgical briefing (PSB) is a vital practice for maintaining patient safety in the operating theatre. For medical students, the PSB can help solidify patient (e.g. comorbidities) and operative (e.g. perceived complexity) features fundamental to surgical curricula [1]. However, the PSB is frequently time pressured and information exchanges occur rapidly, potentially limiting opportunities for student learning. Priming students to the purpose and structure of the PSB could support meaningful engagement with this critical step in the patient journey, reinforcing core principles of perioperative decision-making and maximising intended learning [2].
We identified a gap in PSB preparation within our student curriculum, consistent with wider calls for more integrated preparation for operating-theatre learning [3]. To address this, we developed a simulated PSB task, based on cognitive apprenticeship theory, designed to enhance PSB understanding before students encounter the real environment. Cognitive apprenticeship theory was chosen as the underpinning theory given its focus on transforming implicit thought processes into explicit ones [4].
A focused simulated PSB task was embedded into an introductory session for final-year medical students during their general surgical rotation. Students had no prior formal PSB teaching, and exposure was limited to observation during prior surgical rotations. Groups of up to 12 students were provided with an exemplar operating list, reflective of cases they would see during their rotation (e.g. diabetic patient listed for inguinal hernia repair). The embedded task required students to perform a mock PSB, identifying relevant features and potential mitigation strategies. Students were divided into groups of four, each allocated one patient to present, to encourage active participation from all.
Guided by the methods dimension of cognitive apprenticeship theory [4], the task was structured as follows (Figure 1):
1.Modelling – Students were shown examples of relevant features discussed at PSB (e.g. management of an anticoagulated patient) as part of the task orientation.
2.Coaching – The same surgically trained facilitator, with experience and training in simulation debriefing, was available to guide each group discussion and answer queries throughout.
3.Scaffolding – A framework for structuring the PSB presentation was provided, consisting of (i) surgical, (ii) anaesthetic, (iii) operative and (iv) equipment features.
4.Articulation – Students presented their patient to the wider group, verbalising their findings and reasoning for mitigation strategies.
5.Reflection – Facilitated discussion explored PSB presentations to reinforce key learning points and clarify uncertainty.


An overview of the simulated PSB task structure, based on cognitive apprenticeship theory. For each phase in the theory, the corresponding phase of the PSB simulation is provided to demonstrate the student journey through the task
Although limited by the absence of the full multidisciplinary team, the task was designed to mirror real PSB workflow through the case mix used and commonly adopted structure in our setting.
The immediate impact of the simulated PSB task was evaluated using a structured feedback tool developed by the authors. In addition to feedback on enjoyment of the session, students provided pre- and post-task quantitative data on their understanding of the PSB. This outcome measure was purposefully broad so that students could base feedback on their learning priority (e.g. knowledge of PSB, preparedness to attend). A 7-point Likert scale was used as this balance validity of assessment with excessive options [5]. Free-text comments were available for students to provide additional feedback on the simulation’s effect on their operating theatre preparation. Ethical approval was granted by the Edinburgh Medical School Medical Education Research Ethics Committee (ref: 2024/35).
Over two academic years, 98 students participated. Paired-samples analysis using a Student’s t-test demonstrated a greater reported understanding of the PSB after participating in the task (mean difference 2.48, 95% CI 2.13–2.83, p < 0.001). Free-text feedback revealed that students valued the opportunity to explore features influencing the structure of an operating list and the PSB as a platform for discussing them:
Handy to consider how to prioritise patient cases by urgency and complexity.
Recognising anticoagulated and diabetic patients during PSB was also noted by students and was frequently the focus of micro-teaching (e.g. pharmacology of anticoagulants). Together, these findings suggest that the task supported gains in preparedness and clinical reasoning around perioperative risk. Feedback further highlighted the value of introducing this teaching earlier in the curriculum; all students had previously participated in PSB events through prior surgical-based rotations:
I think this would be best pitched at year 4 students – we’ve now been in a lot of operating theatres.
Informal feedback from surgical faculty suggested that students engaged more actively with the PSB process during the rotation, although this observation was not systematically captured and would benefit from a structured evaluation.
The simulated task enhanced student understanding of the PSB and its role in patient safety. Framing the activity using cognitive apprenticeship principles supported engagement, articulation of reasoning and targeted micro-teaching around clinically important considerations. Incorporating this short simulation into an introductory session also reduced logistical resources and ensured consistent delivery across rotations. The simulation model could also be modified for different PSB settings (e.g. greater focus on equipment in orthopaedic PSB) to integrate teaching of this essential practice across the surgical curriculum.
Our next focus is on examining the translational effects in the operating theatre environment using structured observations of student participation during live PSBs (e.g. Medical Students Non-Technical Skills), comparing students who have and have not completed the simulation. By demonstrating positive effects on longitudinal student engagement with PSBs, this low-resource intervention could be rolled out across our surgical curriculum to help bridge the gap between observational exposure and meaningful participation in the operating theatre.
None declared.
LD and KH are supported by fellowships funded by the NHS Lothian Medical Education Directorate.
None declared.
Ethical approval was granted by the Edinburgh Medical School Medical Education Research Ethics Committee (ref: 2024/35).
None declared.
1.
2.
3.
4.
5.