A Safety Verification Framework for AI-Driven XR Clinical Simulators

Authors: Dasa, D., Board, M., Rolfe, U., Dolby, T., Tang, W.

Editors: Ni, H., Cafolla, D.

Conference: International Conference on AI in Healthcare

Dates: 26/08/2026

Journal: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)

Publication Date: 12/08/2026

Volume: 16875

Pages: 17-29

Publisher: Springer Nature

eISSN: 1611-3349

ISSN: 0302-9743

DOI: 10.1007/978-3-032-35387-0_2

Abstract:

AI-driven XR healthcare simulation platforms raise questions of module correctness, safety-constraint enforcement, and bounded failure behaviour that end-user evaluation studies do not address. This paper reports a technical verification study for XR3DQuest, a modular AI-driven XR simulator; emergency anaphylaxis management serves as the verification test scenario (safety here concerns simulator-internal state integrity, not patient safety). The central contribution is a structured verification framework that decomposes the system into eight independently testable subsystem families, tested across both Unity and Unreal engines with the aim of establishing cross-engine equivalence, identifying both consistent behaviour and attributable divergence across implementations. A 280-assertion harness achieved 80.0% on the Unity build (five families at 100%); debrief content coverage identified the primary residual failure surface (68.8%), demonstrating the framework’s ability to localise unresolved weaknesses to a specific subsystem rather than treating the simulator as a monolithic black box. A 240-case adversarial safety test achieved 84.6%; compound safe and unsafe command parsing exposed the dominant structural failure mode (37%), illustrating how the framework isolates adversarial failure patterns not visible in aggregate end-user evaluation. Cross-engine results are broadly comparable (Unreal 92.1%); session-level verification confirmed 100% multi-turn state chaining. Six experienced practitioners rated the simulator; ratings were directionally consistent with harness findings, with error interception identified as the most fragile safety dimension. The study demonstrates that a modular AI-driven XR simulator can be decomposed and assessed as an auditable engineered platform, providing a governance baseline for AI-driven XR health simulation.

Source: Manual