Medical training device combining AR and VR

By combining modules such as 3D image acquisition, motion capture, central processing unit, AR interaction and VR equipment, the efficient application of AR and VR technologies in medical training devices is achieved, solving the problems of insufficient realism, poor interactivity and safety hazards in existing technologies, providing personalized and flexible training solutions, and supporting unified assessment and model iteration across institutions.

CN120808664APending Publication Date: 2025-10-17SHENZHEN CITY BAOAN DISTRICT MATERNAL & CHILD HEALTH HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511078964.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies fail to effectively combine AR and VR technologies, and are unable to provide medical training devices with high realism, strong interactivity and precise guidance. There are hardware delays, interface interaction burdens and privacy and security risks, and there is a lack of unified cross-institutional assessment and continuous model iteration mechanisms.

Method used

It adopts a combination of 3D image acquisition module, motion capture module, central processing unit, AR interaction module, VR equipment, storage module and display module, and integrates deep convolutional neural network with long short-term memory network to achieve multimodal data synchronization and real-time correction, support wireless transmission and cloud update, and provide immersive training experience and personalized solutions.

Benefits of technology

It improves the realism and sense of presence of training, increases the flexibility and content richness of training, realizes unified assessment and continuous model iteration across institutions, solves hardware latency and privacy and security issues, and enhances the quantifiability and efficiency of training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808664A_ABST
    Figure CN120808664A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical apparatus and instruments, and discloses an AR and VR combined medical training device, which comprises a 3D image acquisition module; the action capture module is used for capturing operation actions of trainees in real time and transmitting action data to the central processing unit; the central processing unit outputs guidance information pre-stored in the operation process according to the operation process; the AR interaction module is used for acquiring an augmented reality or virtual reality model generated by the central processing unit in real time; the VR equipment receives the virtual scene data and related information sent by the central processing unit and is used for helping students to be familiar with the operation process and environment before training and providing more comprehensive immersive experience for the students during training; and the display module is used for displaying the three-dimensional image of the operation target object, the action analysis data of the trainee, the guidance information and the correction information in real time. An augmented reality or virtual reality model containing a real scene and virtual guidance information is generated by using a central processing unit.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to but is not limited to the technical field of medical devices, and particularly relates to a medical training device combining AR and VR. BACKGROUND

[0002] In the field of medical training, traditional training methods have many shortcomings. For example, using artificial prostheses for virtual surgery training has poor realism and immediacy, and cannot allow trainees to actually experience various details and conditions in surgery, so the practicability is low. At the same time, there are differences in viewing angles between trainees and instructors during training, and the number of instructors is limited, which makes it difficult for trainees to obtain accurate surgical guidance in time and easily forms incorrect surgical habits, making it difficult to correct later. In addition, most traditional training methods rely on fixed venues and equipment, limiting the flexibility of time and space for training.

[0003] With the development of technology, Virtual Reality (VR) technology and Augmented Reality (AR) technology have gradually matured. VR technology can create a virtual world, and users can fully immerse themselves in it through VR devices, obtaining a 3D stereoscopic immersive experience; AR technology superimposes virtual information onto the real world, allowing users to see both virtual information and the real world, greatly enhancing the immediacy and realism. However, there are still few devices that effectively combine AR and VR technology for medical training, which cannot fully meet the needs of medical training for high realism, strong interactivity, and precise guidance.

[0004] The closest prior art can be referred to Chinese patent CN106251752A "AR and VR combined medical training system". This patent also proposes: using a 3D image acquisition module to obtain three-dimensional images of surgical target objects and instruments, an action capture module to capture the surgical actions of the operator in real time, a central processor to generate correction and guidance information and build an AR / VR model, and an AR interaction module to superimpose and display information to training personnel to improve training realism and reduce costs.

[0005] This prior art mainly stays at the level of module stacking and function description, does not give the synchronization registration strategy of multi-modal data (images / actions) on the time axis, the fusion details of deep convolutional neural networks and long short-term memory networks, and does not solve the key problems of low-latency rendering in AR presentation, spatial registration drift, and standard action quantitative evaluation in industrialization. The disclosure only generally describes "improving realism and shortening the cycle", lacks objective quantitative indicators and cloud case updating mechanism, and is difficult to support unified assessment across institutions and continuous model iteration. In addition, industry reports also point out that existing VR / AR medical training often has hardware latency, interface interaction burden, and privacy security risks, which are not technically solved in this patent. SUMMARY

[0006] In view of the problems existing in the prior art, the application provides a medical training device combining AR and VR.

[0007] The application is implemented as a medical training device combining AR and VR, which comprises:

[0008] a 3D image acquisition module, which acquires three-dimensional images of a surgical target object and surgical instruments used by a training personnel in real time during a surgery process and transmits the acquired images to a central processor;

[0009] an action capture module, which is connected with the central processor and is used for capturing surgical actions of the training personnel in real time and transmitting action data to the central processor;

[0010] the central processor, which receives data transmitted by the 3D image acquisition module and the action capture module, outputs pre-stored guidance information according to a surgical process, compares surgical actions of the training personnel with pre-set standard surgical actions in real time and gives correction information, and models three-dimensional images acquired by the 3D image acquisition module, the guidance information and the correction information to generate an augmented reality or virtual reality model; the stored standard surgical data is formulated with reference to a large number of expert surgical cases;

[0011] an AR interaction module, which is connected with the central processor, acquires the augmented reality or virtual reality model generated by the central processor in real time, and displays the guidance information and the correction information in a field of view of the training personnel;

[0012] a VR device, which is connected with the central processor, receives virtual scene data and related information sent by the central processor, and is used for helping trainees to be familiar with a surgical process and environment before training and providing more comprehensive immersive experience for the trainees during training;

[0013] a storage module, which is connected with the central processor and is used for storing surgical cases, standard surgical data, medical images, guidance information and training materials;

[0014] a display module, which is electrically connected with the central processor and is used for displaying three-dimensional images of the surgical target object, action analysis data of the training personnel, the guidance information and the correction information in real time.

[0015] Further, the storage module has a data updating function and can update surgical cases, medical images and other materials regularly to ensure timeliness and accuracy of training content.

[0016] Further, the display module can adopt a touch display screen to facilitate operation and marking of a guide personnel during monitoring and review.

[0017] Further, the AR interaction module is an AR glasses using a liquid crystal on silicon (LCOS) chip, a silicon-based organic light-emitting diode chip or a micro light-emitting diode chip, and the implementation mode is a video see-through or an optical see-through.

[0018] Further, the motion capture module uses a time-of-flight (TOF) sensor to capture the surgical motion from a different angle than the 3D image acquisition module.

[0019] In combination with the above technical solutions and the technical problems solved, the technical solutions to be protected by the present application have the following advantages and positive effects:

[0020] 1. Improved realism and presence: The present application captures three-dimensional images of real surgical target objects and surgical instruments through the 3D image acquisition module, combines motion data obtained from multiple angles by the motion capture module, and generates an augmented reality or virtual reality model containing real scenes and virtual guidance information using a central processing unit. During the training process, the trainee can experience the surgical environment firsthand, as if he or she were performing a real surgical operation, greatly improving the realism and presence of the training, and making the training results more closely resemble actual surgical conditions.

[0021] 2. Improved training flexibility: The VR device of the present application allows trainees to familiarize themselves with virtual scenes before training, without being limited by time and space. The data transmission module supports wireless transmission, making the layout of the entire training device more flexible, and allowing for free adjustment according to different training needs and site conditions, providing a more convenient and efficient way for medical training.

[0022] 3. Rich training content and methods: The large number of surgical cases and training materials stored in the storage module provide a rich variety of materials for training. Individualized training programs can be developed for different types of surgery, trainee levels and training goals, and a variety of training activities can be carried out. At the same time, in combination with the immersive experience of the VR device and the real-time information superimposition of the AR interaction module, a variety of training methods are provided for trainees, meeting the learning habits and needs of different trainees, helping to improve the trainees' interest in learning and participation, and promoting the trainees' better mastery of medical knowledge and skills. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is a schematic diagram of the structure of a medical training device combining AR and VR;

[0024] Figure 2 is a structural diagram of the AR interaction module;

[0025] Figure 3 is a schematic diagram of the central processing unit circuit;

[0026] Figure 4 is a structural diagram of the VR device;

[0027] In the figure: 1, 3D image acquisition module; 2, motion capture module; 3, central processor; 4, AR interaction module; 5, VR device; 6, storage module; 7, display module. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with examples. It should be understood that the specific examples described herein are only used to explain the present application and not to limit the present application.

[0029] The existing AR or VR single mode medical training system mostly stays at the level of "watching-mimicking": the spatial registration drift, the time delay accumulation and the subjectivity of action evaluation make it difficult to enter the standardized teaching pipeline of the training base and the surgical robot manufacturer. In view of this industry pain point, the present application takes the bimodal synchronization of three-dimensional image acquisition and motion capture as the entrance, and through the multi-modal feature fusion in the central processor, the "vision-action-feedback" closed loop is solidified into measurable, replayable and transferable data flow, which opens up the path from individual training to group standard construction.

[0030] The three-dimensional image acquisition module acquires the point cloud of the surgical area and the instrument posture by using a multi-view ToF / structured light array, and the original point cloud is compressed into sparse key frames by voxel filtering and ICP registration; the motion capture module outputs the skeleton joint sequence and the instrument six-degree-of-freedom trajectory in parallel. The central processor performs Kalman filtering cascade on the two-way data under a unified timestamp, corrects the instantaneous frame loss caused by occlusion, and completes the spatial synchronization of rigid and non-rigid targets according to the anatomical coordinate system of the surgical scene.

[0031] The core of the working mechanism is the series-parallel fusion of deep convolutional neural network and long short-term memory network: DCNN extracts the spatial texture features of instrument trajectory and tissue surface deformation, and LSTM captures the time dependence and rhythm pattern of surgical action; the fusion layer adopts an attention gate mechanism to emphasize the feature contribution degree of key operation stages (such as penetration, separation, and suture). The action embedding vector generated thereby is compared with the pre-stored standard case library for metric learning, and real-time correction information and risk reminders are generated.

[0032] The augmented reality end is not simply superimposed with subtitles, but is self-adaptively arranged with guide arrows, virtual tangent lines and safety boundary surfaces in the field of view according to the surgical stage and operation deviation. The LCOS or MicroLED optical see-through display projects information through optical waveguide coupling, and the system completes rendering and superposition within an end-to-end delay of less than 20 ms, ensuring that the operator obtains an auxiliary view without occluding key anatomical structures in the real scene; at the same time, the VR device provides a full virtual surgical environment when offline, allowing the trainee to replay his own operation trajectory and superimpose the thermal map error distribution from different perspectives.

[0033] The storage module manages the expert cases, standard action sequences, and training logs in a hierarchical index, synchronizes with the cloud server in both directions at regular intervals, and runs clustering analysis and baseline update services on the cloud side. When new procedures or instruments enter the clinic, the platform can automatically issue the latest parameter set and case video during the night batch processing period, ensuring the timeliness and consistency of the training content. The central processor re-quantizes the model weights after receiving the updates to prevent local inference from drifting.

[0034] In engineering implementation, the system uses distributed middleware to decouple the acquisition, inference, and rendering threads, uses zero-copy queues and GPU direct DMA to reduce transmission delay. The watchdog monitors the heartbeats of each sensor node, and automatically switches to the prediction interpolation mode when it goes offline, ensuring uninterrupted training. Through the above mechanisms, the device completes the closed-loop design from accurate data acquisition, objective action evaluation to multi-end real-time feedback in industrial scenarios, significantly improving the quantifiable degree and promotion efficiency of medical training.

[0035] As shown in Figures 1-4 , a medical training device combining AR and VR, the device comprises:

[0036] a 3D image acquisition module 1, which acquires real-time three-dimensional images of the surgical target object and the surgical instruments used by the training personnel during the operation, and transmits the acquired images to the central processor;

[0037] an action capture module 2 connected with the central processor 3, which is used to capture the surgical actions of the training personnel in real time and transmit the action data to the central processor;

[0038] a central processor 3, which receives the data transmitted by the 3D image acquisition module 1 and the action capture module 2, outputs the pre-stored guidance information according to the surgical process, compares the surgical actions of the training personnel with the pre-set standard surgical actions in real time and gives correction information, and models the three-dimensional images acquired by the 3D image acquisition module 2, the guidance information, and the correction information to generate an augmented reality or virtual reality model; the stored standard surgical data is formulated based on a large number of expert surgical cases;

[0039] an AR interaction module 4 connected with the central processor 3, which acquires the augmented reality or virtual reality model generated by the central processor 3 in real time, and displays the guidance information and the correction information in the field of view of the training personnel;

[0040] a VR device 5 connected with the central processor 3, which receives the virtual scene data and related information sent by the central processor 3, and is used to help the trainees familiarize themselves with the surgical process and environment before training and provide them with a more comprehensive immersive experience during training;

[0041] The storage module 6 is connected with the central processor 3, and is used for storing surgical cases, standard surgical data, medical images, guidance information and training materials;

[0042] The display module 7 is electrically connected with the central processor 3, and is used for displaying the three-dimensional image of the surgical target object, the action analysis data of the training personnel, the guidance information and the correction information in real time.

[0043] The storage module 6 has a data updating function, and can periodically update the surgical cases, medical images and the like, so as to ensure the timeliness and accuracy of the training content.

[0044] The display module 7 can adopt a touch display screen, so that the guide personnel can operate and mark when monitoring and reviewing.

[0045] The AR interaction module 4 is AR glasses adopting a liquid crystal on silicon (LCOS) chip, a silicon-based organic light-emitting diode chip or a micro light-emitting diode chip, and has a video see-through or optical see-through implementation mode.

[0046] The action capture module 2 adopts a time-of-flight (TOF) sensor to capture surgical actions from different angles with the 3D image acquisition module.

[0047] The three-dimensional image acquisition and action capture part adopts a multi-camera ToF array combined with structured light projection, and a reference camera group is arranged above and on both sides of a surgical area to complete unified coordinate system calibration in cooperation with a calibration plate; a lightweight reflective marker point is attached to an instrument to enhance the stability of feature extraction. After the original point cloud is subjected to voxel filtering and outlier elimination, it enters an iterative closest point registration process to obtain continuous key frames; the action capture subsystem synchronously acquires a quaternion sequence of a skeleton joint and a six-degree-of-freedom trajectory of an instrument, and a Kalman filter and a sliding window interpolation module built in a central processor correct occlusion and frame loss, and finally generate a time-aligned spatial action data stream.

[0048] The algorithm layer deploys a deep convolutional network extractor on a GPU to extract spatial texture features and tissue surface deformation features corresponding to the instrument trajectory, a long short-term memory network processes the time correlation of the action sequence, and the outputs of the two are fused through an attention gate to generate an action embedding vector. The embedding vector is measured and matched with a standard surgical action library in the storage module, a deviation threshold is calculated in real time, and correction information is generated. The central processor packages correction arrows, virtual safety boundary surfaces and key step prompts into rendering instructions, which are superimposed into the visual field of the training personnel through the light waveguide display of the AR interaction module, and the complete scene stream is sent to a VR device for offline review and multi-view playback.

[0049] The data and content management adopts a double-layer structure of local storage plus cloud synchronization: locally indexed by time stamp and procedure label, recording original point cloud, skeleton sequence, feature parameters and score results; the cloud periodically receives incremental data, runs clustering analysis and baseline update services, generates new weight files and case resources and pushes them to the terminal. The overall software architecture of the system decouples the collection, reasoning and rendering threads with a message bus, uses zero-copy shared memory queues and GPU direct DMA to reduce end-to-end delay to within twenty milliseconds; the hardware layer monitors the heartbeat of each sensing node with an independent watchdog, immediately switches to prediction interpolation mode in case of disconnection, ensuring continuous availability of the training process.

[0050] Embodiment 1: Laparoscopic surgery training system

[0051] In this embodiment, the three-dimensional image acquisition module uses a binocular stereo endoscope to acquire the surgical field and instrument pose under the laparoscope in real time; the motion capture module installs a set of inertial measurement units (IMUs) on the wrists and fingers of the trainees, for high-precision capture of the operator's surgical action trajectory; the central processor integrates a convolutional neural network and a long short-term memory network fusion algorithm model to perform online comparison and analysis on the labeled expert laparoscopic operation video. The AR interaction module is a head-mounted video see-through AR glasses that can superimpose the paths and force control information that need to be cut and sutured in the surgical field in the trainee's field of view; the VR device provides preoperative abdominal anatomical structure and operation process roaming functions.

[0052] After the training is completed, the system stores the three-dimensional trajectory, force feedback curve and correction information of the operation process to the storage module, and generates a review report on the touch display screen; the teacher can gradually replay the trainee's operation actions in the review report and add annotations and text comments at key steps; at the same time, the storage module can periodically obtain the latest laparoscopic surgery video and case library from the cloud and automatically update the local database to ensure the cutting-edge and diversity of the training content.

[0053] Embodiment 2: Orthopedic arthroscopic surgery training system

[0054] This embodiment uses a structured light depth camera as a three-dimensional image acquisition module to obtain stereo images of soft tissues and bones under the arthroscope of the trainee; the motion capture module is arranged on both sides of the operating table, and a time-of-flight (TOF) sensor array is used to capture the instrument holding and joint manipulation actions of the trainee from multiple angles; the central processor integrates a spatio-temporal convolution network enhanced with an attention mechanism for real-time comparison of the differences between the trainee and the standard expert model at key suture and cutting steps, and automatically generates feedback suggestions. The AR interaction module is an optical see-through AR glasses that can superimpose bone dissection level identifiers and repair area outlines in the field of view; the VR device loads a complete knee joint or shoulder joint model to provide free roaming and multi-scene simulation in a three-dimensional space.

[0055] The system is also configured with a personalized teaching mode: a "difficult point practice" mode can be called on the display module, focusing on training the actions of the trainees that repeatedly make mistakes at specific steps; the storage module supports video and action data tagging management, and continuously optimizes the action recognition accuracy by regularly updating the algorithm library; after the training is completed, an ability assessment report for the trainee can be generated, including reaction time, trajectory smoothness, force control stability, and other dimensions, so that the teacher and the trainee can jointly develop a subsequent practice plan.

[0056] Embodiment 3: Remote collaborative medical training system

[0057] This embodiment adds a multi-user collaboration function on the basis of the above: the three-dimensional image acquisition module and the motion capture module are respectively deployed in the training center and the remote branch, and the data stream is transmitted to the central processor through high-speed network; the central processor supports cloud deployment, and can simultaneously access the data streams of multiple trainees and instructors, compare the surgical actions of different trainees, and superimpose real-time guidance annotations on the instructor's AR terminal. The AR interaction module and the VR device both support network synchronization, and the instructor can "walk to" the trainee through the virtual operating room scene, view the trainee's surgical field, and demonstrate on site.

[0058] The system storage module has a multi-tenant management function, which can perform permission grading on the surgical case library of each institution; the display module supports picture-in-picture and multi-window display, and the trainee can watch the instructor's demonstration or the operation of classmates in the same group; at the same time, the system regularly synchronizes the excellent training records of each institution to the central database to build a standard surgical action data set across institutions, and continuously improve the algorithm model and training content.

[0059] Embodiment 4: Emergency trauma surgery simulation training system

[0060] In this embodiment, the three-dimensional image acquisition module uses a mobile optical tracking device to construct real-time images and environmental models for irregular flowing scenes in the trauma emergency environment (such as an outdoor emergency tent); the motion capture module uses a wearable inertial sensor network to capture the body position and hand movements of the rescuer during the rescue process. The central processor integrates a scene enhancement module based on a generative adversarial network (GAN), which can generate realistic blood flow, tissue tearing, and other emergency trauma special effects in AR and VR environments in real time, and superimpose the correct first aid operation path and timing prompts in the trainee's field of view.

[0061] During the training process, the display module can be switched to a "countdown" mode to simulate the pressure of the golden time of first aid; the storage module packages and archives data such as actions, time nodes, and changes in vital signs of the whole process, providing complete data for subsequent clinical conversion training and accident review; the system regularly obtains the latest trauma treatment guidelines and cases from authoritative medical first aid alliances to ensure that the training content keeps up with the progress of first aid standards and technologies.

[0062] Embodiment 5: Robot-assisted surgery training system

[0063] This embodiment integrates a surgical robot port in the central processor, obtains the robot end view and end effector pose through a 3D image acquisition module; the motion capture module synchronizes the student's hand operation with the robot teaching rod input signal; the central processor maps the student's action to the robot control command, and compares the actual execution path of the robot with the standard path in real time to generate a correction instruction. The AR interaction module identifies the robot motion trajectory deviation and force feedback overload area in the student's glasses; the VR device simulates the entire operating table environment, including an anesthesia machine, a monitor, and a team collaboration process.

[0064] The system storage module records the joint angle, end force control data, and image synchronous flow of the robot platform, and provides an API interface for third-party algorithm plug-ins to access; the display module can call the "safe zone demonstration" function to present the robot motion reachable space and the operating range of surgical instruments in three dimensions on the touch screen, helping the student master the boundaries of robot operation. The system updates the robot firmware version and the latest robot surgery case library regularly to ensure the consistency of the training and clinical use environment.

[0065] It should be noted that the embodiments of the present application can be realized by hardware, software or a combination of software and hardware. The hardware part can be realized by using special logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or a specially designed hardware. Those skilled in the art can understand that the above-mentioned devices and methods can be realized by using computer executable instructions and / or included in processor control code, such as provided on a carrier medium, such as a magnetic disk, CD or DVD-ROM, a programmable memory, such as a read-only memory (firmware), or a data carrier, such as an optical or electronic signal carrier. The devices of the present application and their modules can be realized by hardware circuits, such as very large scale integrated circuits or gate arrays, semiconductors, such as logic chips, transistors, or programmable hardware devices, such as field programmable gate arrays, programmable logic devices, etc. They can also be realized by software executed by various types of processors, or by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0066] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any modification, equivalent replacement and improvement within the technical range disclosed by the present application and within the spirit and principle of the present application should be covered within the protection scope of the present application.

Claims

1. A medical training device combining AR and VR, characterized in that: include: A three-dimensional image acquisition module, used to acquire three-dimensional images of surgical target objects and surgical instruments used by trainees in real time and output image data; Motion capture module, used to capture the trainer's surgical movements in real time and output motion data; The central processing unit is connected to the 3D image acquisition module and the motion capture module, and is used to receive the image data and motion data and perform the following operations: Compare the trainee's surgical movements with pre-stored standard surgical movements in real time and generate correction information; constructing an augmented reality model and a virtual reality model based on the image data, motion data, and correction information; Analyze surgical action characteristics based on deep convolutional neural network and long short-term memory network fusion algorithm; An AR interaction module, connected to the central processor, is used to overlay guidance and correction information in the trainer's field of view; VR equipment, connected to the central processing unit, is used to present the surgical process and environment in a virtual scene; A storage module, connected to the central processing unit, for storing expert surgical case data, standard surgical data and training materials; The display module is electrically connected to the central processing unit and is used to display three-dimensional images, motion analysis results and correction information in real time.

2. The device according to claim 1, characterized in that The storage module has a timed update function, which is used to regularly obtain and store new surgical case data and medical images from the cloud server to ensure the timeliness and accuracy of the training content.

3. The device according to claim 1, characterized in that The display module is a capacitive touch screen display, which is used to support training personnel in marking and reviewing guidance information and correction information.

4. The device according to claim 1, characterized in that The AR interaction module uses a silicon-based liquid crystal LCOS chip or a silicon-based organic light-emitting device or a micro-light-emitting diode chip to realize augmented reality glasses with video see-through or optical see-through.

5. The device according to claim 1, characterized in that The motion capture module uses a time-of-flight sensor to capture the trainee's surgical movements from multiple angles.

6. A medical training method based on the combination of AR and VR, characterized in that: The steps include: Acquiring a three-dimensional image of the surgical target object and surgical instruments using a three-dimensional image acquisition module; Use the motion capture module to obtain the trainees' real-time surgical motion data; In the central processing unit, a surgical action recognition algorithm that integrates a convolutional neural network and a long short-term memory network is used to extract and compare features of the action data; Generate correction information based on the comparison results and build an augmented reality model together with pre-stored guidance information; Output the augmented reality model to the trainer's field of view through the AR interactive module, and output the virtual scene through the VR device; Store and annotate training process data for subsequent review and statistical analysis.

7. The method according to claim 6, wherein: The surgical action recognition algorithm includes: Extract spatiotemporal features based on 3D convolutional networks; Perform temporal correlation analysis on action sequences based on long short-term memory networks; Weighting key steps based on attention mechanism; The network parameters are jointly optimized within the central processing unit to improve the accuracy of the correction information.

8. A medical training software system combining AR and VR, characterized by: include: An image preprocessing module, used for denoising and reconstructing the original image output by the 3D image acquisition module; Motion analysis module, which is used to execute the deep learning-based motion recognition algorithm and output feature vectors; The model generation module is used to fuse the feature vector, guidance information and correction information to construct a 3D visualization model; The user interaction module is used to conduct two-way data interaction with trainers through AR interaction interface and VR interface.

9. A three-dimensional image reconstruction model for AR and VR medical training, characterized in that: The model is based on a three-dimensional convolutional network for transfer learning. It inputs multi-view image data and outputs a high-precision stereo mapping of surgical instruments and tissue structures.

10. A pre-trained surgical action dataset for AR and VR medical training, characterized by: Contains multiple annotated expert surgical cases. The data format includes 3D coordinate sequences and corresponding action category labels, which are used for model fine-tuning when training the above-mentioned surgical action recognition algorithm.

Citation Information

Patent Citations

  • Augmented reality (AR) and virtual reality (VR) combined medical training system

    CN106251752A