Modular VR / AR Surgical Teaching System and Method Based on Local Large Model

By using a modular VR/AR surgical teaching system driven by a local large model, combined with lightweight hardware and multimodal assessment, the system achieves adaptive generation and secure inference of teaching content, solves the dynamic adaptability and privacy issues of existing systems, provides high-fidelity multi-sensory interaction and accurate assessment, and improves the effectiveness of surgical training.

CN122090684APending Publication Date: 2026-05-26TONGJI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TONGJI UNIV
Filing Date
2026-03-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing VR/AR surgical teaching systems lack dynamic adaptability, cannot adjust training difficulty according to the student's level, have risks of network latency and privacy leakage, lack comprehensive collection and analysis of multimodal data, cannot provide accurate feedback, and are bulky with unrealistic force feedback.

Method used

A modular VR/AR surgical teaching system driven by a local large model is adopted. It combines lightweight hardware with multimodal fusion evaluation algorithms to achieve adaptive content generation and secure reasoning. The system deploys a large model through edge computing, integrates multi-dimensional perception units and multi-sensory feedback, and performs real-time error correction and quantitative evaluation.

Benefits of technology

It achieves automated and dynamic adaptation of teaching content, reduces customization costs, ensures the relevance and security of teaching, provides high-fidelity multi-sensory interaction and accurate quantitative assessment, solves the latency and privacy issues of traditional systems, and improves training effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090684A_ABST
    Figure CN122090684A_ABST
Patent Text Reader

Abstract

This invention discloses a modular VR / AR surgical teaching system and method based on a local large-scale model. The local large-scale model training module is deployed in a local computing environment, storing structured medical knowledge and dynamically generating teaching scripts and assessment question banks. A lightweight VR / AR hardware group collects multi-dimensional user operation data and provides multi-sensory feedback. A full-process training control module, with a built-in physical simulation engine and evaluation logic, constructs a virtual surgical scene, simulates tissue mechanical changes in real time, compares user operations with standard surgical procedure models and generates error correction guidance, and generates a quantitative ability assessment report based on the entire process data. A multi-scene data mapping interface adopts a modular architecture based on a standardized intermediate communication protocol, achieving low-coupling logical isolation and dynamic integration between functional modules through a unified data exchange format and abstract interface specifications. This application adopts an edge-side deployment strategy to protect patient privacy and ensure real-time force feedback and tissue deformation rendering in the VR / AR scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical education technology, and in particular to a modular VR / AR surgical teaching system and method based on a local large model. Background Technology

[0002] Virtual reality (VR) and augmented reality (AR) technologies have been gradually applied to surgical teaching, aiming to solve the problems of high cost and limited resources in traditional cadaver dissection and animal experiments. However, existing VR / AR surgical teaching systems still have the following major shortcomings in their technical implementation: 1. Content generation relies on manual pre-setting, lacking dynamic adaptability. The surgical scenarios and operation steps in existing teaching systems are mostly pre-modeled and fixed in the software by developers. When faced with rare cases, anatomical variations, or new surgical procedures, it is necessary to re-model 3D and create animations, which is time-consuming and costly. Due to the lack of an automated content generation mechanism driven by medical knowledge, the system is unable to dynamically adjust the training difficulty or generate targeted error correction guidance in real time according to the trainee's operation level; 2. In achieving complex physical simulation and intelligent assessment, most existing systems place the core reasoning tasks on cloud servers. This architecture not only leads to operation delays and screen lag in environments with fluctuating network conditions, damaging the immersion and hand-eye coordination of surgical training; at the same time, uploading real patient case data to the cloud for processing poses a risk of patient privacy leakage. 3. Existing assessment indicators are mainly limited to geometric parameters such as instrument movement trajectory and operation time, lacking comprehensive collection and analysis of multimodal data such as operator force control, eye-tracking attention allocation, and decision-making logic. Due to the lack of a deep comparison algorithm based on expert gold standard path, the system cannot provide interpretable quantitative assessment of the compliance and current risks of key steps, making it difficult to provide accurate feedback; 4. Mainstream teaching equipment is bulky, and force feedback devices mostly use traditional motor drives, resulting in problems such as large response delays and lack of high-frequency vibration simulation, making it difficult to realistically reproduce the subtle tactile changes during soft tissue cutting and suturing, thus affecting the training effect of fine motor skills.

[0003] Therefore, developing a surgical teaching system that can deploy large models locally to achieve adaptive content generation and secure inference, and combine lightweight hardware with multimodal fusion evaluation algorithms, is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] To address the aforementioned problems in existing technologies, a modular VR / AR surgical teaching system based on a local large-scale model is provided, which is driven by a local large-scale model and is designed for multi-specialty surgical teaching.

[0005] Another objective of this application is to provide a modular VR / AR surgical teaching method based on a local large model.

[0006] To solve the above problems, the present invention adopts the following technical solution: a modular VR / AR surgical teaching system based on a local large model, including a local large model training module, which is deployed in a local computing environment, configured to store structured medical knowledge, and dynamically generate teaching scripts, practice tasks or assessment question banks based on requirements;

[0007] The lightweight VR / AR hardware kit is configured to collect multi-dimensional operational data of users in virtual or augmented display environments and provide multi-sensory feedback based on system instructions.

[0008] The full-process training control module communicates with the local large-scale model training module and the lightweight VR / AR hardware group. It has a built-in physical simulation engine and evaluation logic, and is configured to: construct a virtual surgical scene, respond to user operations in real time and simulate tissue mechanical changes; compare user operation data with standard surgical model during training and generate error correction guidance in real time; and generate a quantitative capability evaluation report based on the data of the entire training process.

[0009] It also includes a multi-scenario data mapping interface, adopts a modular architecture based on a standardized intermediate communication protocol, and achieves low-coupling logical isolation and dynamic integration between functional modules by defining a unified data exchange format and abstract interface specifications.

[0010] Furthermore, the local large model training module, which is deployed on an edge server or private cloud environment and communicates with the VR / AR interactive terminal, is configured to: perform knowledge distillation and domain adaptation on the general medical large model to construct a structured knowledge base containing a medical knowledge graph, a surgical procedure flowchart, and a complication rule base, and dynamically generate teaching scripts, practice tasks, and assessment question banks according to the department, surgical procedure, and student level.

[0011] Furthermore, the lightweight VR / AR hardware assembly includes a multimodal interaction terminal and a glasses terminal composed of a lightweight skeleton structure, as well as a tactile interaction peripheral connected in communication with it; the glasses terminal and the tactile interaction peripheral are respectively integrated with an eye-tracking sensor, an IMU sensor, a hand posture and force sensor, and a vibration and thermal feedback actuator; the hardware assembly is configured to: synchronously collect multimodal motion data of the user's eyes, head, and hands through each sensor and upload it to the system, and at the same time drive the vibration and thermal feedback actuator to generate corresponding tactile stimuli according to the instructions issued by the system, so as to provide feedback on force and temperature sensation;

[0012] Furthermore, the full-process training control module includes a basic medical knowledge base for storing medical theoretical data and standard surgical procedures; a surgical skills simulator configured to provide a virtual surgical operation environment; a chapter practice module and an intelligent assessment module for executing phased skills training tasks and comprehensive ability assessment tests; and a simulation and evaluation processing unit, which is communicatively connected to the basic medical knowledge base, the surgical skills simulator, the chapter practice module, and the intelligent assessment module, and is configured to: run simulation calculations based on physical and biomechanical models, simulate tissue deformation and mechanical feedback in real time, execute real-time error correction logic during training, compare user operation data with the standard surgical procedures and generate correction instructions; run a multi-dimensional evaluation algorithm based on the user's performance data in the chapter practice module and the intelligent assessment module to generate a quantitative evaluation report; and encapsulate the quantitative evaluation report into trainee ability profile data and output it to an externally connected local large-scale model training module.

[0013] Furthermore, the multi-scenario data mapping interface includes a general instruction parsing unit, a domain action mapping rule library, and a terminal adaptation unit. The general instruction parsing unit parses general instructions output by the local large-scale model training module or the full-process training control module into standard instruction tuples containing action type, target object, target area, execution parameters, and feedback requirements. The domain action mapping rule library converts these standard instruction tuples into specific domain actions based on department category, surgical procedure, instrument type, and risk level. For example, "displaying the hepatic hilum vascular system" is mapped to corresponding anatomical model retrieval, hepatic hilum region highlighting, and viewpoint positioning actions; "correcting clamping position" is mapped to safety threshold verification, instrument trajectory correction prompts, and tactile alarm actions. The terminal adaptation unit then performs parameter recalibration and protocol encapsulation of the specific domain actions based on the display resolution, sensor capabilities, actuator types, and communication protocols of different VR / AR devices, thereby enabling unified invocation and differentiated execution of the same general instruction across different terminals, departments, and surgical scenarios.

[0014] Furthermore, the multi-scenario mapping interface is used to define a unified device abstraction layer and a standardized content package format, enabling the same training engine to adapt to different terminals and switch to training tasks of different departments, procedures, or difficulty levels by changing the content package.

[0015] The present invention also provides a method for implementing the modular VR / AR surgical teaching system based on a local large model, the method comprising the steps of:

[0016] Step 1: Data Access and Governance. Access multi-source heterogeneous medical data, including impact data, case data, and standardized data from intraoperative video datasets. De-identify and organize sensitive information in the multi-source heterogeneous data, and construct a unified training data catalog and metadata index.

[0017] Step 2: Knowledge graph construction. Based on the training data directory and metadata index, feature information of anatomical structures, surgical instruments, operation steps, risk points and complication management procedures is extracted and mapped to the graph structure to construct a structured knowledge base containing unit anatomy knowledge graphs and surgical procedure graphs.

[0018] Step 3: Local large model adaptation. Load the general medical large model at the edge settlement node, and use lightweight fine-tuning technology in combination with the structured knowledge base to perform domain adaptation and generate a localized specialty surgery teaching model.

[0019] Step 4: Using the localized specialty surgery teaching model, teaching content is automatically generated based on the preset teaching template and the structured knowledge base, and packaged into a standardized training content package.

[0020] Step 5: Simulation scene loading and execution steps. The VR / AR training terminal loads the standardized training content package to generate a virtual surgical scene. During the user's operation, the biomechanical and physiological parameter simulation engine calculates tissue deformation, bleeding effect and thermal diffusion effect in real time.

[0021] Step Six: Real-time Interaction and Error Correction. Synchronously collect user operation data in the virtual surgical scenario. The operation data includes instrument position data, hand strength dataset, and gaze attention area data. Compare the operation data with standard surgical procedures in real time. When a deviation is detected, generate real-time error correction guidance. After training, generate error distribution analysis data and action alignment report.

[0022] Step 7: In response to the assessment instructions, load standard case scenarios or random emergency scenarios, quantify and score user performance based on the multimodal fusion assessment model, and generate assessment results that include capability profile analysis and interpretability debriefing analysis reports.

[0023] Furthermore, it also includes step eight, closed-loop feedback and adaptive recommendation steps, which extract the user's weak link characteristics based on the evaluation results, update the user learning profile, and drive the localized specialty surgery teaching model to dynamically adjust the recommendation measurement and difficulty level of subsequent teaching content.

[0024] Furthermore, the lightweight fine-tuning technology specifically includes knowledge distillation compression, LoRA, or Adapter fine-tuning methods; and in multi-institutional collaboration scenarios, a federated learning mechanism is used to aggregate the gradients of each institution to update the model parameters.

[0025] Furthermore, the standardized training content package described in step four is encapsulated in a predefined .medtrain or compatible format, and contains serialized scene configuration files and resource index tables.

[0026] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0027] 1. The system described in this application automates and dynamically adapts the generation of teaching content, significantly reducing customization costs and improving teaching relevance. By constructing a structured knowledge graph containing anatomical structures, surgical procedures, and complication rules, and combining knowledge distillation and lightweight fine-tuning techniques from a local large-scale model, this system can automatically parse the teaching logic from standardized SOPs and case data. Based on the real-time ability profile of trainees (such as weaknesses and operational precision deviations), it dynamically generates personalized practice levels, complication trigger scripts, and assessment question banks. Furthermore, the system described in this application can autonomously iterate and update the surgical content package based on the hospital's unique cases, significantly shortening the development cycle of new courses and achieving sustainable evolution of teaching content.

[0028] 2. The system described in this application adopts an edge-side deployment strategy, offloading large-scale model inference, biomechanical simulation calculations, and multimodal data processing to local servers or private cloud environments. Original case data (such as DICOM images and intraoperative videos) are only anonymized and indexed within the hospital's local area network, without needing to be uploaded to the public cloud. This not only completely avoids the legal risks of patient privacy leaks but also eliminates latency and jitter caused by wide area network transmission, ensuring real-time force feedback and tissue deformation rendering in VR / AR scenarios, achieving dual protection of data sovereignty and immersive experience.

[0029] 3. The system described in this application integrates a multi-dimensional sensing unit with eye tracking, hand posture, and force sensors, along with a diffractive waveguide optical module and a vibration / thermal feedback actuator, to construct a high-fidelity multi-sensory interactive channel. It can not only simulate visual tissue deformation and bleeding effects, but also calculate the interaction forces between instruments and tissues in real time based on a biomechanical model, and accurately reproduce cutting resistance, suture tension, and the sensation of thermal damage from electrocautery through tactile peripherals. This comprehensive feedback mechanism, integrating vision, touch, and temperature, effectively solves the problem of "lack of tactile feedback" in traditional training, allowing the muscle memory formed in the virtual environment to be more smoothly transferred to the real operating table.

[0030] 4. The system described in this application utilizes a multimodal fusion evaluation algorithm to simultaneously collect and analyze trainees' "instrument posture trajectory, force curve, gaze focus hotspot, and decision response time" in four dimensions. It can not only output a quantitative comprehensive score, but also generate a visualized error distribution heatmap and a key motion unit alignment report, accurately identifying whether trainees have deficiencies in anatomical cognition, operational procedures, or emergency decision-making. This provides instructors with clear intervention criteria, transforming teaching feedback from vague experience-based judgments to precise data-driven approaches.

[0031] 5. The system described in this application establishes logical isolation between the underlying hardware (VR / AR glasses and haptic gloves from different brands) and the upper-layer applications (different departments and surgical procedures) by defining a unified Device Abstraction Layer (UDAL) and a standardized .medtrain content package protocol. The same training engine and content package can be adapted to heterogeneous terminals without the need for repeated development for different hardware. At the same time, new surgical procedures can be quickly deployed by simply replacing the content package without reconstructing the system kernel, achieving a highly cohesive and loosely coupled modular architecture. This not only supports rapid replication and promotion across departments and hospitals, but also reserves flexible interfaces for future access to new sensors or expansion to new surgical specialties, significantly reducing the long-term operation and maintenance and expansion costs of the system. Attached Figure Description

[0032] Figure 1 This is a hardware architecture diagram of the system described in this application;

[0033] Figure 2 This is a diagram of the system software modules described in this application;

[0034] Figure 3 This is a flowchart of the system algorithm described in this application;

[0035] Figure 4 This is a flowchart of the system workflow described in this application;

[0036] Figure 5 This is a schematic diagram of the VR glasses for the system described in this application;

[0037] Figure 6 This is a schematic diagram of the multifunctional glove of the system described in this application;

[0038] Among them, 1. Vibration unit, 2. Main board, 3. Eye-tracking camera, 4. Waveguide glasses, 5. Optical module, 6. Drive and acquisition circuit box, 7. Shape memory alloy wire, 8. Finger piezoresistive array. Detailed Implementation

[0039] The technical solution of the present invention will be further described clearly and in detail below with reference to the embodiments and accompanying drawings.

[0040] like Figure 1-3As shown, this application provides a modular VR / AR surgical teaching system based on a local large model. By integrating a multi-dimensional perception unit with eye tracking, hand posture and force sensors, and in conjunction with a diffractive waveguide optical module 5 and a vibration / thermal feedback actuator, a high-fidelity multi-sensory interactive channel is constructed. This effectively solves the problem of "lack of tactile feedback" in traditional training, enabling the muscle memory formed by trainees in the virtual environment to be transferred more smoothly to the real operating table. The system includes a local large model training module, which is deployed in a local computing environment and configured to store structured medical knowledge and dynamically generate teaching scripts, practice tasks or assessment question banks based on needs.

[0041] The lightweight VR / AR hardware kit is configured to collect multi-dimensional operational data of users in virtual or augmented display environments and provide multi-sensory feedback based on system instructions.

[0042] The full-process training control module communicates with the local large-scale model training module and the lightweight VR / AR hardware group. It has a built-in physical simulation engine and evaluation logic, and is configured to: construct a virtual surgical scene, respond to user operations in real time and simulate tissue mechanical changes; compare user operation data with standard surgical model during training and generate error correction guidance in real time; and generate a quantitative capability evaluation report based on the data of the entire training process.

[0043] It also includes a multi-scenario data mapping interface, adopts a modular architecture based on a standardized intermediate communication protocol, and achieves low-coupling logical isolation and dynamic integration between functional modules by defining a unified data exchange format and abstract interface specifications.

[0044] After the system starts, the local large model training module first loads the pre-set structured medical knowledge graph and desensitized case data. It analyzes the current teaching needs through retrieval enhancement generation technology, dynamically instantiates teaching scripts and task sequences containing anatomical structural parameters, pathophysiological characteristics and surgical operation logic, and transmits the generated standardized scenario configuration files to the full-process training control module through the multi-scenario data mapping interface.

[0045] After receiving the configuration, the full-process training control module calls the built-in physical simulation engine to initialize the virtual surgical environment. Based on the patient's specific image data, it calculates tissue biomechanical models and fluid dynamic parameters. Simultaneously, it sends rendering commands and feedback strategies to the lightweight VR / AR hardware group through a multi-scene data mapping interface. The lightweight VR / AR hardware group integrates eye-tracking, hand posture, and force sensors to collect the user's eye movement trajectory, six-DOF hand pose, and force data in real time. This data is then encapsulated through a unified device abstraction layer protocol and transmitted back to the full-process training control module. The full-process training control module performs spatiotemporal alignment and deviation calculation between the received real-time operation data stream and the standard surgical procedure model. When it detects that the operation trajectory, force threshold, or gaze focus deviates from the preset safety range, it immediately triggers error correction logic, generates correction commands, and sends them through the multi-scene data mapping interface. The system is sent to the lightweight VR / AR hardware group to drive the diffractive waveguide optical module 5 to project visual guidance signs and control the vibration / thermal feedback actuator to output corresponding tactile resistance or temperature change signals to simulate the interactive feel of real tissue. Throughout the training process, the full-process training control module continuously records the user's eye movement, hand posture, and force sensor data to form a multimodal operation log, decision response time, and complication handling process. After training, the local large model training module combines preset evaluation dimensions to perform multi-dimensional feature extraction and weighted analysis on the log data, generate a capability evaluation report, and feed the evaluation results back to the local knowledge base to optimize the generation strategy of subsequent teaching scripts. This achieves a fully automated technical process from dynamic content generation, real-time data acquisition and multi-sensory feedback simulation based on eye tracking, hand posture, and force sensors to closed-loop quantitative evaluation.

[0046] In some cases, general-purpose large-scale models lack a refined sense of anatomical space and rigorous surgical logic, and have a large number of parameters, making them difficult to run in real time on local hardware. The local large-scale model training module described in this application, deployed on an edge server or private cloud environment and connected to a VR / AR interactive terminal, is configured to: perform knowledge distillation and domain adaptation on the general-purpose medical large-scale model to construct a structured knowledge base containing a medical knowledge graph, surgical procedure flowchart, and complication rule base; and dynamically generate teaching scripts, practice tasks, and assessment question banks based on department, surgical procedure, and student level. Through the lightweight model after knowledge distillation, combined with an edge computing architecture, network transmission latency and inference time are significantly reduced. By introducing a structured knowledge graph and rule base as constraints, each surgical step and each complication triggering logic reviewed by the model has clear graph nodes or rule bases, ensuring the medical rigor of the teaching content.

[0047] For example, the implementation process of the adaptive training system for laparoscopic cholecystectomy in general surgery based on a private cloud deployment is as follows: Four edge servers with NVIDIA A800 GPUs are deployed in the private cloud environment of a tertiary hospital's data center. The servers run the Llama-3-Medical-8B lightweight medical model, which has undergone knowledge distillation and quantification. The model is pre-loaded with two hundred anonymized laparoscopic cholecystectomy surgery videos, corresponding DICOM image data, and the latest clinical guidelines from the hospital within the past five years. The model uses natural language processing technology to extract anatomical entities and constructs a medical knowledge graph containing nodes such as the gallbladder triangle, cystic artery, and common bile duct and their three-dimensional spatial adjacency relationships. The model also sorts out the temporal logic of standard steps such as exposure, separation of adhesions, clamping of arteries, cutting, and removal to construct a surgical procedure flowchart. Furthermore, it defines the triggering conditions for misclamping of the common bile duct, such as clamping position less than two millimeters away from the common bile duct and failure to identify variant structures, to construct a complication rule base, forming a structured local knowledge base. When a junior resident initiates a training request for laparoscopic cholecystectomy in general surgery via a VR / AR interactive terminal, the local large-scale model training module parses the department tag, surgical procedure type, and junior level profile of the resident in the request. Based on retrieval enhancement generation technology, it retrieves relevant anatomical nodes and operational rules from the structured knowledge base, dynamically synthesizes a personalized teaching script that includes high-level variation parameters of the cystic artery, the pathological state of mild edema of the gallbladder wall, and random intraoperative bleeding events, generates a practice sequence and assessment question bank containing standard paths and branch tasks, and transmits the generated standardized scene description file to the full-process training control module via an encrypted channel on the hospital's intranet. After parsing the scene file, the full-process training control module calls the physical simulation engine to instantiate a virtual surgical environment, loads a patient-specific biomechanical model according to the script parameters, and distributes rendering resources and interaction logic to the lightweight VR / AR hardware group through a multi-scene data mapping interface. The eye-tracking integrated into the hardware group... The system uses hand posture and force sensors to collect and transmit trainee operation data in real time. The control module compares the real-time data stream with the standard path of the surgical procedure flowchart in time and space. When the system detects that the trainee's instrument is close to the common bile duct safety threshold or triggers the complication rule base conditions, it immediately generates a visual guidance command to drive the diffraction waveguide optical module 5 to highlight the safety area and controls the vibration / thermal feedback actuator to output resistance to simulate tissue adhesion for real-time error correction. After training, the control module summarizes the full-cycle multimodal operation log. The local large model training module combines the completion of key nodes in the medical knowledge graph and the risk avoidance performance of the complication rule base to conduct a multi-dimensional quantitative assessment of the trainee and generate a capability report. At the same time, the assessment results are fed back to the local knowledge base as new training samples to iteratively optimize the subsequent content generation strategy. This achieves fully automated operation from professional knowledge structuring and dynamic personalized generation of teaching content to real-time closed-loop feedback, while ensuring that the data does not leave the hospital.

[0048] In real surgical environments, tissue cutting is accompanied by frictional heat, electrocautery is accompanied by temperature rise, and different pathological tissues have specific hardness and resistance. Therefore, these key physiological signals cannot be reproduced by vision or a single vibration alone. The lightweight VR / AR hardware group provided by the system described in this application includes a multimodal interactive terminal and a glasses terminal composed of a lightweight skeleton structure, as well as a tactile interactive peripheral device that is communicatively connected to it. Figure 5 The VR glasses shown include a vibration unit 1, a motherboard 2, an eye-tracking camera 3, a waveguide lens 4, and an optical module 5. Their working principle is as follows: the optical module 5 collimates and shapes the virtual surgical scene image generated by the motherboard 2, outputting parallel light incident on the waveguide lens 4; the waveguide lens 4, based on a diffraction grating structure, couples, performs total internal reflection transmission, and decouples the incident light for output, achieving lightweight, large field-of-view near-eye display; the eye-tracking camera 3 acquires infrared reflection images of the user's eyeballs at a high frame rate, and the motherboard 2's built-in image processing algorithm calculates the pupil center coordinates, gaze vector, and blink state in real time for dynamic rendering optimization and operation intent recognition; the motherboard 2, as the core control unit, integrates a multi-core processor and internal... The system stores and communicates with the operating system, spatial positioning algorithm, gesture recognition model, and tactile feedback scheduler. It synchronously processes data streams from the eye-tracking camera 3, IMU, and external gloves, and sends PWM or analog voltage signals to the vibration unit 1 based on the risk level and operation deviation instructions output by the surgical simulation engine. The vibration unit 1 is attached to the side wing of the frame or the headband contact area and performs mechanical vibrations of a specific frequency and amplitude according to the received control parameters. It is used to provide auxiliary tactile feedback such as spatial orientation prompts, error operation warnings, or task completion confirmations in non-hand areas, thus forming a closed head-mounted display and sensing system that integrates visual presentation, eye-tracking interaction, computational control, and local tactile reminders.

[0049] like Figure 6As shown, the fingertip piezoresistive array 8 consists of multiple flexible piezoresistive sensor units distributed in the contact areas of the index, middle, and ring fingers. It is used to collect real-time data on the normal pressure distribution and dynamic changes applied by the user when holding surgical instruments. The signals are digitized by the analog-to-digital converter module within the drive and acquisition circuit box 6 and then transmitted to the MCU. The shape memory alloy wire 7 is embedded in the dorsal region of the palm and finger joints, serving as an active deformation feedback actuator. After the MCU calculates the target deformation based on the tissue stiffness, cutting resistance, or slippage risk level output by the simulation engine, a precise current is applied through the drive circuit to induce controllable thermo-contraction, thereby applying a reverse constraint force to the user's hand to simulate real tissue resistance or instrument rebound. Simultaneously... When a vibrator integrated into the fingertip or palm detects excessive speed, deviation from the trajectory, or accidental contact with a dangerous area, the MCU triggers a high-frequency pulse vibration to provide an instantaneous tactile warning. The workflow is as follows: the sensor collects hand posture and force data, transmits the data to the MCU for spatiotemporal feature modeling and operation intention recognition, combines the physical parameters of the virtual environment to generate a tactile feedback strategy, and drives the circuit to control the deformation of the alloy wire or the start and stop of the vibrator, realizing a closed-loop interaction from sensory input to force / vibration dual-mode tactile output. The drive and acquisition circuit box 6 serves as the local processing hub, undertaking signal conditioning, protocol encapsulation, low-power management, and wireless communication with the VR host, ensuring that the glove has independent operation capability and low-latency response characteristics.

[0050] Existing surgical teaching systems cannot fully demonstrate the viscoelasticity, anisotropy, and hemodynamic characteristics of real human tissues. Therefore, the system described in this application provides a full-process training control module, including a basic medical knowledge base for storing medical theoretical data and standard surgical procedures; a surgical skills simulator configured to provide a virtual surgical operation environment; a chapter practice module and an intelligent assessment module for executing phased skills training tasks and comprehensive ability assessment tests; and a simulation and evaluation processing unit, which is communicatively connected to the basic medical knowledge base, surgical skills simulator, chapter practice module, and intelligent assessment module, and configured to: run simulation calculations based on physical and biomechanical models, simulate tissue deformation and mechanical feedback in real time, execute real-time error correction logic during training, compare user operation data with the standard surgical procedures and generate correction instructions; run a multi-dimensional evaluation algorithm based on the user's performance data in the chapter practice module and intelligent assessment module to generate a quantitative evaluation report; and encapsulate the quantitative evaluation report into student ability profile data and output it to an externally connected local large model training module.

[0051] For example, during the training of the "Gallbladder Triangle Anatomy" chapter in laparoscopic cholecystectomy, the full-process training control module first calls the standard surgical procedure specifications stored in the basic medical knowledge base, extracts theoretical data such as the definition of the safe separation plane normal vector, the maximum allowable traction force threshold, and the standard operating path curve, and loads them into the surgical skills simulator. Subsequently, the surgical skills simulator performs real-time simulation calculations on the trainee's virtual operation based on the built-in physical and biomechanical models. When it detects that the traction force applied by the trainee exceeds the preset threshold or the instrument trajectory deviates from the safe plane, it immediately executes real-time error correction logic, applying a countermeasure by dynamically adjusting the damping coefficient of the virtual environment. The system generates resistance, renders visual warning zones, and generates correction instructions to force operations back to standard specifications. After training, the intelligent assessment module automatically summarizes the full-cycle operation data and runs a multi-dimensional evaluation algorithm to calculate quantitative indicators from the dimensions of efficiency, quality, and standardization, generating a quantitative evaluation report. This report is then packaged into student ability profile data containing student ability weakness feature tags and directly output to the externally connected local large model training module to drive the large model to dynamically adjust the pathological scene weights and difficulty coefficients of subsequent training scripts based on the profile features, thereby achieving a complete data closed loop from single skill training execution to adaptive optimization of global teaching strategies.

[0052] In some embodiments, the multi-scenario mapping interface is used to define a unified device abstraction layer and a standardized content package format, so that the same training engine can be adapted to different terminals, and can be switched to different departments, different procedures or different difficulty levels of training tasks by changing the content package.

[0053] The system described in this application is applied to the following examples.

[0054] Case 1: Teaching of Laparoscopic Cholecystectomy in General Surgery

[0055] In the basic anatomy learning phase, trainees wear lightweight VR glasses to access a 3D model of the liver. A voice command, "Show the hepatic hilum vascular system," triggers a dynamic anatomical dissection of the local large model. An AR overlay of a portal vein variation probability heatmap provides a basic understanding of the human liver's structure. During pathological simulation training, a cholecystitis adhesion model is loaded. A biomechanical engine generates fibrotic tissue characteristics based on the patient's CT data, and real-time tissue tearing effects are rendered when a virtual dissection hook applies tension. Trainees then begin practicing key operations. AR projects a gold guide line into the surgical field, dynamically correcting the placement angle of the vascular clips. A haptic glove simulates the stepped resistance of titanium clip closure, while a spatial positioning system records instrument trajectories and generates a hotspot cluster map of accidental common bile duct contact. Trainees practice resection within this simulated environment. In the assessment phase, the system first injects a virtual patient heart rate drop event, tracking the trainee's decision chain between pausing or continuing the resection. Then, a standard path engine compares the curvature differences in instrument movement trajectories, finally outputting an eight-dimensional capability matrix, including indicators such as stress response speed and 3D spatial accuracy, to accurately assess the trainee's operational level.

[0056] Case Study 2: Teaching Total Knee Arthroplasty

[0057] During the bone structure learning phase, trainees wear lightweight VR glasses to access a 3D model of the knee joint. Voice commands such as "mark areas of high-wearing cartilage" trigger dynamic marking of femoral condyle cartilage defects on a large local model. An AR-overlay of bone density distribution heatmaps provides a visual understanding of bone degeneration characteristics. In biomechanical training, an osteoporotic joint model is loaded, and a finite element analysis engine generates load-bearing characteristics based on the patient's bone density data. A virtual bone saw renders real-time bone fragment splatter effects during osteotomy. Trainees then begin key osteotomy practice. AR projects a golden-angle reference plane into the surgical field, dynamically correcting the saw blade's cutting trajectory. Haptic gloves simulate vibration feedback under different bone hardness levels, while a motion tracking system records instrument amplitude and generates a pressure distribution cloud map of excessively deep osteotomies. Trainees complete prosthesis implantation procedures based on biomechanical guidance. During the assessment phase, the system dynamically triggers bone cement embolism complications, tracking the trainee's decision-making logic regarding revision or thrombolysis. The biomechanical engine then calculates the probability of pulmonary embolism, finally outputting a surgical accuracy-risk correlation surface graph, including dimensions such as prosthesis fit and complication predictability, enabling a quantitative assessment of orthopedic techniques.

[0058] Case Study 3: Teaching Coronary Artery Bypass Grafting

[0059] In the cardiovascular pathology learning phase, trainees wear lightweight VR glasses to access a coronary artery model. Voice commands such as "Locate the plaque rupture risk area" trigger dynamic highlighting of vulnerable plaques in the left anterior descending artery on a large local model. AR overlays hemodynamic simulations of the flow field, enabling precise understanding of vascular lesion mechanisms. During micromanipulation training, a myocardial ischemia pathology model is loaded, and a computational fluid dynamics engine generates blood flow parameters based on angiography data. A virtual needle holder provides real-time warnings of needle distance deviation during suturing. Trainees then practice vascular anastomosis. AR projects suture grid lines with 0.2 mm precision, dynamically calibrating the needle insertion angle. Tactile gloves simulate the stepped resistance of vascular wall penetration, while an optical tracking system records the needle tip trajectory, generating a spatial clustering map of posterior wall suturing errors. Trainees complete vascular bridge construction under sub-millimeter guidance. In the assessment phase, the system simulates a ventricular fibrillation crisis during surgery, tracking the compliance of the trainee's defibrillation procedure. A timeline engine then compares the response time to the gold standard, finally outputting an emergency decision tree report highlighting key capability gaps such as instrument operation precision and the completeness of physiological parameter monitoring.

[0060] Case Study 4: War Trauma First Aid Teaching

[0061] During the critical injury assessment training phase, trainees wear lightweight VR glasses to access a blast injury torso model. A voice command, "Mark the liver injury risk area," triggers a local large-scale model to dynamically mark the subcostal impact injury area. An AR simulation of internal bleeding rates is overlaid, quickly establishing a recognition of injury severity levels. In emergency care training, a multi-organ rupture pathology model is loaded, and a physiological parameter engine generates vital sign curves based on a trauma database. Real-time feedback on vascular pressure thresholds is provided when virtual compression bandages are applied. Trainees then practice first aid hemostasis. AR projects graded compression guidance areas onto the wound, dynamically adjusting the position and intensity of force application. Tactile gloves simulate the transmission of arterial pulsation, while a distributed sensing system records team collaboration latency and generates a heat map of resource allocation conflicts. Trainees complete trauma management under dynamic life monitoring. In the assessment phase, the system introduces secondary shock complications, tracks trainees' priority decisions regarding fluid resuscitation and surgical hemostasis, and then the team collaboration engine analyzes role coordination effectiveness. Finally, it outputs a battlefield injury management capability matrix, including core battlefield medicine indicators such as wound control timeliness, multi-threaded task management, and crisis prediction.

[0062] like Figure 4 As shown, the present invention also provides a method for implementing the modular VR / AR surgical teaching system based on a local large model, the method comprising the following steps:

[0063] Step 1: Data Access and Governance. Access multi-source heterogeneous medical data, including impact data, case data, and standardized data from intraoperative video datasets. De-identify and organize sensitive information in the multi-source heterogeneous data, and construct a unified training data catalog and metadata index.

[0064] Step 2: Knowledge graph construction. Based on the training data directory and metadata index, feature information of anatomical structures, surgical instruments, operation steps, risk points and complication management procedures is extracted and mapped to the graph structure to construct a structured knowledge base containing unit anatomy knowledge graphs and surgical procedure graphs.

[0065] Step 3: Local Large Model Adaptation. A general medical large model is loaded at the edge settlement node. Lightweight fine-tuning techniques, combined with the structured knowledge base, are used for domain adaptation to generate a localized specialty surgical teaching model. These lightweight fine-tuning techniques specifically include knowledge distillation compression, LoRA, or Adapter fine-tuning. Furthermore, in multi-institutional collaboration scenarios, a federated learning mechanism is used to aggregate gradients from each institution to update model parameters.

[0066] Step 4: Using the localized specialty surgery teaching model, teaching content is automatically generated based on the preset teaching template and the structured knowledge base, and packaged into a standardized training content package; the standardized training content package is packaged in a predefined .medtrain or compatible format, and contains a serialized scenario configuration file and a resource index table.

[0067] Step 5: Simulation scene loading and execution steps. The VR / AR training terminal loads the standardized training content package to generate a virtual surgical scene. During the user's operation, the biomechanical and physiological parameter simulation engine calculates tissue deformation, bleeding effect and thermal diffusion effect in real time.

[0068] Step Six: Real-time Interaction and Error Correction. Synchronously collect user operation data in the virtual surgical scenario. The operation data includes instrument position data, hand strength dataset, and gaze attention area data. Compare the operation data with standard surgical procedures in real time. When a deviation is detected, generate real-time error correction guidance. After training, generate error distribution analysis data and action alignment report.

[0069] Step 7: In response to the assessment instructions, load standard case scenarios or random emergency scenarios, quantify and score user performance based on the multimodal fusion assessment model, and generate assessment results that include capability profile analysis and interpretability debriefing analysis reports.

[0070] Step 8: Closed-loop feedback and adaptive recommendation step. Based on the evaluation results, extract the user's weak link characteristics, update the user learning profile, and drive the localized specialty surgery teaching model to dynamically adjust the recommended measurement and difficulty level of subsequent teaching content.

[0071] Finally, it should be pointed out that the above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A modular VR / AR surgical teaching system based on a local large model, characterized in that, It includes a local large model training module, which is deployed in a local computing environment, configured to store structured medical knowledge, and dynamically generates teaching scripts, practice tasks or assessment question banks based on requirements; The lightweight VR / AR hardware kit is configured to collect multi-dimensional operational data of users in virtual or augmented display environments and provide multi-sensory feedback based on system instructions. The full-process training control module communicates with the local large model training module and the lightweight VR / AR hardware group. It has a built-in physical simulation engine and evaluation logic, and is configured to: construct a virtual surgical scene, respond to user operations in real time and simulate tissue mechanical changes, compare user operation data with standard surgical model during training, and generate error correction guidance in real time. A quantitative capability assessment report is generated based on data from the entire training process. It also includes a multi-scenario data mapping interface, adopts a modular architecture based on a standardized intermediate communication protocol, and achieves low-coupling logical isolation and dynamic integration between functional modules by defining a unified data exchange format and abstract interface specifications.

2. The modular VR / AR surgical teaching system based on a local large model according to claim 1, characterized in that, The local large model training module, which is deployed on an edge server or private cloud environment and communicates with VR / AR interactive terminals, is configured to: perform knowledge distillation and domain adaptation on a general medical large model to construct a structured knowledge base containing a medical knowledge graph, a surgical procedure flowchart, and a complication rule base; and dynamically generate teaching scripts, practice tasks, and assessment question banks according to the department, surgical procedure, and student level.

3. The modular VR / AR surgical teaching system based on a local large model according to claim 1, characterized in that, The lightweight VR / AR hardware assembly includes a multimodal interaction terminal and a glasses terminal consisting of a lightweight skeleton structure, as well as a tactile interaction peripheral connected to them. The glasses terminal and the tactile interaction peripheral are respectively equipped with an eye-tracking sensor, an IMU sensor, a hand posture and force sensor, and a vibration and thermal feedback actuator. The hardware assembly is configured to: synchronously collect multimodal motion data of the user's eyes, head, and hands through each sensor and upload it to the system; at the same time, drive the vibration and thermal feedback actuator to generate corresponding tactile stimuli according to the instructions issued by the system, so as to provide feedback on force and temperature.

4. The modular VR / AR surgical teaching system based on a local large model according to claim 1, characterized in that, The full-process training control module includes a basic medical knowledge base, which is used to store medical theoretical data and standard surgical procedures. Surgical skills simulator, configured to provide a virtual surgical operating environment; The chapter exercise module and the intelligent assessment module are used to perform phased skills training tasks and comprehensive ability assessment tests. The simulation and evaluation processing unit is connected to the basic medical knowledge base, surgical skills simulator, chapter exercise module and intelligent assessment module respectively, and is configured to: run simulation calculations based on physical and biomechanical models, simulate tissue deformation and mechanical feedback in real time, execute real-time error correction logic during training, compare user operation data with the standard surgical procedure specifications and generate correction instructions. Based on the user's performance data in the Demon Refining Module and Intelligent Assessment Module of the chapter, a multi-dimensional evaluation algorithm is run to generate a quantitative evaluation report; the quantitative evaluation report is encapsulated as student ability profile data and output to the externally connected local large model training module.

5. The modular VR / AR surgical teaching system based on a local large model according to claim 1, characterized in that, The multi-scenario data mapping interface includes a general instruction parsing unit, a domain action mapping rule base, and a terminal adaptation unit. The general instruction parsing unit parses upper-level training control instructions into standard instruction tuples containing action type, target object, target area, execution parameters, and feedback requirements according to preset semantic slots. The domain action mapping rule base converts the standard instruction tuples into corresponding specific domain action sets based on department type, surgical procedure nodes, instrument category, and risk level. These specific domain action sets include at least scene loading actions, instrument-driven actions, line-of-sight guidance actions, force feedback or thermal feedback actions, and evaluation and acquisition actions. The terminal adaptation unit is used to combine the display capabilities, sensor configurations, and feedback actuator parameters of different VR / AR terminals to perform parameter recalibration and protocol encapsulation on the specific domain action set, so as to drive the corresponding terminal to execute training tasks. The conversion logic includes: mapping general instructions such as "display, annotation, and guidance" to anatomical structure highlighting, operation path overlay, or risk area visualization actions; mapping general instructions such as "grasp, cut, separate, and suture" to corresponding instrument model invocation, motion constraint loading, and tissue interaction feedback actions; and mapping general instructions such as "alarm, error correction, and scoring" to vibration or thermal feedback, voice or graphic prompts, and evaluation data recording actions.

6. A method for implementing the modular VR / AR surgical teaching system based on a local large model as described in claim 1, characterized in that, The method includes the following steps: Step 1: Data Access and Governance. Access multi-source heterogeneous medical data, including impact data, case data, and standardized data from intraoperative video datasets. De-identify and organize sensitive information in the multi-source heterogeneous data, and construct a unified training data catalog and metadata index. Step 2: Knowledge graph construction. Based on the training data directory and metadata index, feature information of anatomical structures, surgical instruments, operation steps, risk points and complication management procedures is extracted and mapped to the graph structure to construct a structured knowledge base containing unit anatomy knowledge graphs and surgical procedure graphs. Step 3: Local large model adaptation. Load the general medical large model at the edge settlement node, and use lightweight fine-tuning technology in combination with the structured knowledge base to perform domain adaptation and generate a localized specialty surgery teaching model. Step 4: Training content generation package. Using the localized specialty surgery teaching model, teaching content is automatically generated based on the preset teaching template and the structured knowledge base, and packaged into a standardized training content package. Step 5: Simulation scene loading and execution steps. The VR / AR training terminal loads the standardized training content package to generate a virtual surgical scene. During the user's operation, the biomechanical and physiological parameter simulation engine calculates tissue deformation, bleeding effect and thermal diffusion effect in real time. Step Six: Real-time Interaction and Error Correction. Synchronously collect user operation data in the virtual surgical scenario. The operation data includes instrument position data, hand strength dataset, and gaze attention area data. Compare the operation data with standard surgical procedures in real time. When a deviation is detected, generate real-time error correction guidance. After training, generate error distribution analysis data and action alignment report. Step 7: Intelligent Assessment and Review. In response to assessment instructions, standard case scenarios or random emergency scenarios are loaded. Based on a multimodal fusion assessment model, the user's performance is quantitatively scored, and an assessment result including a capability profile analysis and an interpretable review analysis report is generated.

7. The method of the modular VR / AR surgical teaching system based on a local large model according to claim 6, characterized in that, It also includes step eight, closed-loop feedback and adaptive recommendation, which extracts the user's weak link characteristics based on the evaluation results, updates the user learning profile, and drives the localized specialty surgery teaching model to dynamically adjust the recommendation measurement and difficulty level of subsequent teaching content.

8. The method of the modular VR / AR surgical teaching system based on a local large model according to claim 6, characterized in that, The lightweight fine-tuning techniques specifically include knowledge distillation compression, LoRA, or Adapter fine-tuning methods; and in multi-institutional collaboration scenarios, a federated learning mechanism is used to aggregate the gradients of each institution to update the model parameters.

9. The method of the modular VR / AR surgical teaching system based on a local large model according to claim 6, characterized in that, The standardized training content package described in step four is encapsulated in a predefined .medtrain or compatible format and contains serialized scene configuration files and resource index tables.