Virtual reality-based emergency medical rescue drill teaching system

By combining virtual reality technology and intelligent algorithms, trainees' movements are monitored in real time and aligned with standard movement sequences, solving the problems of limited resources and insufficient interaction in traditional medical rescue training, and achieving efficient teaching results and error correction.

CN121053839BActive Publication Date: 2026-02-17SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511576045.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-17
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Traditional medical assistance training resources are scarce, and trainees have difficulty obtaining real-time interaction and personalized guidance, making it difficult to correct incorrect operating habits.

Method used

The training system employs a virtual reality-based emergency medical rescue drill. Through intelligent algorithms and data acquisition modules, it monitors trainees' actions in real time and aligns them with standard action sequences, providing immediate feedback and quantitative assessment.

Benefits of technology

Standardized training was achieved, which improved teaching effectiveness, allowed for timely correction of erroneous operations, and prevented the formation of bad habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053839B_ABST
    Figure CN121053839B_ABST
Patent Text Reader

Abstract

The application discloses a virtual reality-based emergency medical rescue drill teaching system, comprising a computer terminal, a data acquisition module, a plurality of sensor nodes for collecting physical quantity data, a VR device including a head-mounted display device and a simulated training dummy; wherein the computer terminal is configured to extract student static posture features and standard action dynamic trajectory features; align the student action sequence with the standard action sequence in time and output a determination result; trigger state transition or generate a quantitative evaluation report based on the determination result. Thus, through the VR device, the data acquisition module, and the intelligent algorithm corresponding to the functions, the input standard demonstration video is converted into a standard action sequence, so that in the teaching process, the interaction process between the VR and the student can realize standardized training of medical rescue, thereby improving the teaching effect, and in the student's self-practice process, errors can be corrected in time, so that the student does not develop incorrect operation habits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical rescue technology, and in particular to an emergency medical rescue drill and teaching system based on virtual reality. Background Technology

[0002] Traditional medical assistance training models have significant limitations in the technical instruction component. Previously, the teaching of various medical assistance techniques typically relied on a single method: simply showing demonstration videos.

[0003] While demonstration videos can showcase the operational procedures and key points of medical assistance to some extent, this teaching method has many shortcomings. From a resource allocation perspective, traditional training often faces resource constraints. Medical assistance training requires professional venues, equipment, and experienced instructors, but in reality, sufficient resources are often unavailable for a large number of trainees. Complete real-time action comparison teaching requires a sufficient number of professional instructors who can simultaneously observe and accurately judge each trainee's actions, providing timely and targeted feedback and guidance. However, limited by factors such as instructor resources and venue size, such teaching conditions are difficult to meet.

[0004] When relying solely on demonstration videos for instruction, learners can only passively watch and find it difficult to truly understand the operational details and techniques. Due to the lack of real-time interaction and personalized guidance, errors made during self-practice cannot be corrected promptly, easily leading to the formation of incorrect operating habits. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, one objective of this application is to propose a virtual reality-based emergency medical rescue drill and teaching system capable of accurately providing standardized training for trainees.

[0006] The virtual reality-based emergency medical rescue drill and teaching system according to an embodiment of this application includes: a computer terminal equipped with an intelligent algorithm; a data acquisition module including multiple sensor nodes deployed on the trainee's body for collecting physical quantity data; and a VR device including a head-mounted display device and a simulated training dummy. The computer terminal is configured to: extract the trainee's static posture features and the dynamic trajectory features of standard movements; align the trainee's movement sequence with the standard movement sequence in time and output a judgment result; and trigger a state transition or generate a quantitative evaluation report based on the judgment result.

[0007] Its effects are as follows: through VR devices, data acquisition modules, and intelligent algorithms that work together to achieve corresponding functions, the input standard demonstration video is transformed into a standard action sequence. This allows for standardized training in medical rescue through VR interaction with students during the teaching process, thereby improving teaching effectiveness. Furthermore, it allows for timely correction of errors during students' self-practice, preventing the development of incorrect operating habits.

[0008] Furthermore, the data acquisition module includes: a pressure sensor deployed on the palm, a torque sensor on the torso strap, and accelerometers arranged at key body nodes, including 15 standard medical articular points and 3 medical operation nodes.

[0009] Furthermore, the computer terminal includes: a static pose feature extraction module equipped with a ResNet50 model and a dynamic trajectory feature extraction module equipped with a 3D-CNN model; the static pose features extracted by the ResNet50 model and the dynamic trajectory features extracted by the 3D-CNN model are finally output as a joint coordinate sequence representing the standard action.

[0010] Furthermore, the ResNet50 model is pre-trained with medical pose data and configured to: locate the two-dimensional pixel coordinates of 15 standard medical joints and 3 medical operation nodes; generate high-dimensional feature vectors describing the relative positions, angles and body contours of the nodes, forming a single-frame pose snapshot.

[0011] Furthermore, the 3D-CNN model is configured as follows: receiving a sequence of continuous frame features extracted by ResNet50, processing the spatial dimension (width, height) and temporal dimension simultaneously through a three-dimensional convolutional kernel; learning the spatiotemporal patterns of displacement, velocity, acceleration, and overall motion trajectory of joints in the continuous frame sequence; and decoding the fused spatiotemporal features into a three-dimensional spatial coordinate sequence of 18 nodes through a fully connected layer to form the joint coordinate sequence of the standard motion.

[0012] Furthermore, the computer terminal also includes a determination module configured to: align the trainee joint coordinate sequence with the standard joint coordinate sequence in time using a dynamic time warping algorithm (DTW) under the constraint of the maximum time offset; and calculate the Euclidean norm difference between the trainee joint coordinates and the standard coordinates within a specific operation time window as the attitude error.

[0013] Furthermore, the determination module presets preset thresholds associated with the medical operation state, including: a first threshold set for the posture error in the first medical operation state and a second threshold set for the posture error in the second medical operation state; when the posture error of all core joints in the current state is continuously lower than the corresponding first threshold or second threshold, a state transition is triggered.

[0014] Furthermore, the computer terminal also includes a state transition module, configured to: when the state transition is triggered, trigger the VR device to switch to the next medical operation state in the virtual environment; guide the student to perform the new operation through visual prompts, audio instructions or task list updates.

[0015] Furthermore, the computer terminal also includes an evaluation module configured to: extract key operating parameters and assign weight factors; generate a quantitative evaluation report and a three-dimensional action playback of the erroneous steps.

[0016] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0017] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0018] Figure 1 This is a modular schematic diagram of a virtual reality-based emergency medical rescue drill and teaching system according to some embodiments of this application. Detailed Implementation

[0019] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0020] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0021] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout.

[0022] according to Figure 1This application discloses an emergency medical rescue drill and teaching system based on virtual reality, which includes a computer terminal equipped with intelligent algorithms, a data acquisition module, and VR equipment.

[0023] The data acquisition module includes pressure sensors deployed in the trainee's palms, torque sensors on the torso straps, and accelerometers positioned at key body nodes (such as the sternal angle and L5 lumbar vertebrae) to acquire physical data such as pressure, torque, and acceleration. For example, it includes 15 IMU sensor nodes with a sampling rate of 100Hz distributed at standard medical joints of the human body, such as the head, C7 cervical vertebra, T7 thoracic vertebra, both shoulders, both elbows, both wrists, L5 lumbar vertebrae, both hips, both knees, and both ankles. In some preferred examples, dedicated nodes for medical procedures are added, including the sternal angle node for detecting compression depth, the metacarpal eminence node for monitoring bandaging force, and the submental point node for assessing airway opening angle. The feedback device includes tactile gloves with 0-50N pressure sensors integrated into the fingertips and a torso strap equipped with a 6-axis force and torque sensor at the lumbar spine.

[0024] VR devices, including head-mounted displays (HMDs) with OLED screens featuring a 90Hz refresh rate and a 110° field of view; and in subsequent examples, mannequins are placed in the room where the VR devices are set up for simulation training.

[0025] The computer terminal includes a static posture feature extraction module and a dynamic trajectory feature extraction module. The static posture feature extraction module is configured to receive the dataset input from the data acquisition module, extract the static posture features of the trainee, and identify and locate the two-dimensional pixel coordinate information of multiple preset body joints. The dynamic trajectory feature extraction module is configured to receive video stream data from an input standard instructional video, analyze the continuous video frame sequence, and generate a standard joint coordinate sequence representing the standard movement. Based on the static posture feature extraction module and the dynamic trajectory feature extraction module, the joint coordinate sequence generated by the trainee in real time (generated by fusing IMU sensor node data) is aligned with the standard joint coordinate sequence on the time axis, and the judgment result is output.

[0026] More specifically, the static pose feature extraction module includes a residual neural network (ResNet50) model, and the dynamic trajectory feature extraction module includes a three-dimensional convolutional neural network (3D-CNN) model; the static pose features extracted by the ResNet50 model and the dynamic trajectory features extracted by the 3D-CNN model are finally output as a sequence of joint coordinates representing the standard movement.

[0027] For example, for the 15 standard medical joints (such as head, C7, T7, both shoulders, both elbows, both wrists, L5, both hips, both knees, and both ankles) and 3 medical operation nodes (sternal angle, metacarpal prominence, and submental point) in this embodiment, ResNet50 is pre-trained and fine-tuned with medical pose data to locate the two-dimensional pixel coordinates of these nodes in the current frame image and form a high-dimensional feature vector describing their relative position, angle, and body contour; the output of this level provides a detailed pose "snapshot" for each frame.

[0028] 3D-CNNs include three-dimensional convolutional kernels (width, height, and temporal depth), enabling them to simultaneously process spatial and temporal information from multiple consecutive frames in a video clip (e.g., a sequence of frames within a small time window). They receive a sequence of consecutive frame features (or its reduced-dimensional representation) extracted by ResNet50 as input, and through operations such as 3D convolutional layers and pooling layers, learn patterns in the displacement, velocity, acceleration of joints between frames, and the overall motion trajectory. For example, in the action of "chest compressions," 3D-CNNs can learn the regular depth change trajectory of the sternal angle node within the compression cycle, the coordination of elbow flexion and extension, and the dynamic features of trunk stability (L5 node) during force application. This layer ensures the temporal continuity of the action. The output layer of the ResNet50 network outputs the fused high-dimensional spatiotemporal features, which are then decoded and mapped to a joint coordinate sequence through fully connected layers, thus providing each frame of the input video clip with, for example, a vector containing the coordinates (e.g., X, Y, Z) of all 18 nodes in 3D space.

[0029] In this way, the continuous motion trajectories of all key nodes in space form the data foundation for the subsequent medical operation state machine (FSM) to perform state recognition, action comparison, and real-time feedback. Thus, the entire network learns a complex mapping from raw video pixels to precise joint coordinate sequences through end-to-end training.

[0030] In some embodiments, the computer terminal further includes a determination module that determines whether the posture error (i.e., the norm difference between the student's joint coordinates and the teacher's corresponding coordinates) within a specific allowed time window of operation exceeds a preset threshold.

[0031] Specifically, the decision-making module uses Dynamic Time Warping (DTW) to align the trainee's real-time joint coordinate sequence (generated by fusing IMU sensor node data) with the instructor's standard joint coordinate sequence on the time axis, focusing on the state corresponding to the current operation (e.g., "chest compressions"). The module calculates the posture error between the trainee's joint coordinate sequence and the instructor's standard sequence within a specific time window allowed for the operation (the window length is set according to the standard duration of the specific medical operation, such as the time required for a complete compression cycle or an effective airway opening action). This posture error is quantified as the norm difference (e.g., Euclidean distance) between the trainee's joint coordinates and the instructor's corresponding coordinates, and is calculated for key joint points relevant to the current state (e.g., in the "chest compressions" state, focusing on monitoring the sternal angle, elbow joints, and L5 lumbar vertebra). A posture error less than a preset threshold is a key condition for triggering a state transition.

[0032] The preset threshold includes a first threshold and a second threshold.

[0033] For example, in the first medical operation state (such as chest compressions), the first threshold corresponding to the core joint point (such as the sternal angle node) that directly reflects the compression depth is set relatively strictly. It is typically required that the posture error (i.e., the norm difference between the trainee's joint coordinates and the instructor's corresponding coordinates) be maintained within the millimeter-level error range to meet the stringent requirements of clinical operation standards for compression depth. In contrast, in the second medical operation state (such as environmental assessment), the second threshold corresponding to the core joint point (such as the head node) is set relatively leniently; this second threshold is greater than the first threshold. The system determines whether the trainee's operation in the current medical operation state meets the standard based on the following condition: within the allowed time window for the specific operation, the posture error of all monitored core joint points in that state must continuously be lower than their respective preset thresholds (for example, lower than the first threshold for the first medical operation state, and lower than the second threshold for the second medical operation state), thereby confirming that the operation meets the standard requirements.

[0034] In some examples, the computer terminal also includes a state transition module that triggers a state transition when the value is less than a preset threshold.

[0035] Specifically, the process of triggering a state transition includes:

[0036] First, based on the current joint coordinate sequence, a state machine (FSM) identifies the trainee's state (e.g., "chest compressions"). This is achieved through an action classifier (e.g., an LSTM-based model) that analyzes sequence patterns to match predefined states.

[0037] Then, for the current state, the determination module calculates the posture error within the time window and compares it with the core node threshold for that state. The core node is defined by the state: for example, the core nodes for chest compressions are the sternal angle, both elbows, and L5; for an open airway, the core nodes are the submental point and C7 cervical vertebra.

[0038] If the error of all core nodes remains below the threshold, a state transition signal is triggered. The transition signal includes the next state identifier (e.g., transitioning from "chest compressions" to "artificial respiration") and the transition confidence (calculated based on the error level; a confidence level higher than 0.9 is required to trigger the transition).

[0039] Upon triggering, the VR device immediately updates the virtual scene: for example, it highlights the next operating tool (such as an airbag mask) on the headset, plays the audio command "Start artificial respiration," and updates the task list. The transfer latency is controlled within 50ms.

[0040] Furthermore, if the error exceeds a threshold, the system will not trigger a transfer, but will instead provide real-time feedback (such as visual cues of the error) and allow the learner to retry. After multiple consecutive failures (e.g., 3 times), the system can automatically demonstrate the correct action or adjust the difficulty.

[0041] After triggering the state transition, the VR device will guide the trainee into the next medical procedure in the virtual environment (such as switching to the "artificial respiration" state). The virtual reality system will clearly inform the trainee of the next action to be performed through visual cues (such as highlighting the next operation target), audio instructions, or updating the task list.

[0042] Therefore, the system performs real-time motion synchronization and correction, collects trainee IMU data, and generates trainee joint coordinate sequences through sensor fusion technology. A Dynamic Time Warping (DTW) algorithm is used to align the trainee's motion sequence with the teacher's standard sequence. This algorithm finds the warped path with the minimum cumulative distance while satisfying medical operation-specific constraints (such as a maximum time offset limit of 1.5 seconds).

[0043] In some embodiments, the computer terminal further includes an evaluation module, which extracts key operating parameters and performs quantitative evaluation.

[0044] Taking cardiopulmonary resuscitation (CPR) training as an example, the ResNet50 model extracts static pose features from keyframes of the video to accurately locate the two-dimensional pixel coordinates of 18 joints. For example, during the preparation phase of compressions (timestamp t=3.2 seconds), it outputs the spatial coordinates of the sternal angle node (x=12.3cm, y=5.8cm, z=0cm), the flexion angles of both elbow joints (left elbow 162°±2°, right elbow 160°±2°), and the pitch angle of the trunk L5 node (5°±1°). At the same time, the 3D-CNN model analyzes the dynamic trajectory features of the continuous frame sequence, capturing the displacement of the sternal angle node (maximum depression depth 5.3cm), metacarpal pressure sensor data (peak pressure 42.7N), and changes in lumbar torque (peak 32.1N·m) within the chest compression cycle (t=5.0 to 5.8 seconds), forming a spatiotemporal parameter sequence of standard movements.

[0045] During actual practice, trainees align their action sequences with the standard sequence using a Dynamic Time Warping (DTW) algorithm. When a timing logic error is detected (such as skipping the airway opening and directly performing chest compressions), the system immediately freezes the scene and issues a voice prompt: "Error! Missing airway opening step." In compression depth monitoring, displacement values ​​are calculated in real-time through double integration of the z-axis data from the sternal angle accelerometer. If three consecutive compressions are insufficient (e.g., 4.5cm is below the 5.0cm lower limit), the haptic glove fingertips will trigger a 5Hz vibration feedback, and the head-mounted display will simultaneously show a red warning text: "Depth Deviation -0.5cm." A protective mechanism is also activated: when the lumbar torque sensor reading consistently exceeds the 35 N·m safety threshold (e.g., 38.6 N·m), a red warning light illuminates on the trunk straps, and the virtual compression resistance is dynamically reduced by 20% to prevent injury.

[0046] For example, the Dynamic Time Warping (DTW) algorithm is used to align the student's real-time generated joint coordinate sequence (denoted as sequence A, length M) with the teacher's standard joint coordinate sequence (denoted as sequence B, length N). Both sequence A and sequence B contain three-dimensional coordinate data of multiple joint points (e.g., 18 nodes) at a sampling rate of 100Hz. Therefore, the number of points in each sequence may be different, i.e., M≠N, but DTW can minimize the cumulative distance by finding the optimal warping path.

[0047] Let sequence A be A =( a 1, a 2,..., aM ), where each point ai It is a high-dimensional vector representing the coordinates of all relevant nodes at time point i. For example, the X, Y, and Z coordinates of 18 nodes are concatenated to form a 54-dimensional vector. Similarly, sequence B is... B =( b 1, b2,..., bN Therefore, the goal of the dynamic time warping algorithm is to find a warped path. P =( p 1, p 2,..., pK ), of which each pk =( ik , jk ) represents point i in sequence A. k Point j of sequence B k Alignment, and path P satisfies:

[0048] Boundary conditions are i 1=1, j 1=1 and i K = M , j K = N .

[0049] Monotonicity is i k +1≥ i k and j k+1 ≥ j k .

[0050] Continuity is i k+1 - i k ≤1 and j k+1 - j k ≤1, meaning one-step alignment is allowed.

[0051] Then, the cumulative distance matrix D is constructed, including: first, calculating the local distance matrix d, where each element... This is the norm difference. Then, the cumulative distance matrix D(i,j) represents the minimum cumulative distance from (1,1) to (i,j), recursively defined as:

[0052] ;

[0053] Considering the real-time requirements of medical procedures, the system imposes a maximum time offset constraint, such as 1.5 seconds, meaning that for any alignment point (i,j), | i - j |≤ T max ×SR, where SR is the sampling rate. In this example... T max=1.5 seconds, sampling rate 100Hz, therefore the maximum index offset is 150 points, which can be achieved by limiting the range of cumulative distance calculation, that is, only calculating the value satisfying | i - j | ≤ 150 for D(i,j), thereby reducing computational complexity and ensuring the rationality of alignment.

[0054] Then, the minimum cumulative path is found by backtracking from D(M,N) to D(1,1). Finally, the aligned sequences A' and B' have the same length L (L≈max(M,N)), where each point corresponds to the optimal alignment.

[0055] Furthermore, to meet the real-time feedback requirements of VR systems, i.e., latency below 100ms, the DTW algorithm in this embodiment adopts a sliding window strategy. The system divides the long-time sequence into overlapping subsequences, i.e., the window length W corresponds to a specific operation time window, such as 2 seconds, or 200 points, and performs DTW alignment independently within each window. Simultaneously, multi-threaded processing is used, including: one thread responsible for data acquisition and preprocessing, and another thread performing DTW calculations.

[0056] This allows for GPU-accelerated matrix operations, such as parallel computation of the distance matrix d and cumulative matrix D using the CUDA library, reducing computation time from O(MN) to approximately O(M) through constraint windows. Multiple variants of the pre-computed standard sequence B (e.g., different speed versions) can be implemented to speed up alignment.

[0057] In a further embodiment, after alignment, the determination module calculates the posture error within a specific operation time window. The specific operation time window is dynamically defined according to the standard duration of the medical operation. For example, for chest compressions, a complete compression cycle is approximately 0.6 seconds (corresponding to 60 points), so the window length W is set to 60 points; for opening the airway, the window may be longer (e.g., 1.5 seconds, 150 points) to cover the entire sequence of actions.

[0058] The attitude error is defined as the Euclidean norm difference between the trainee's joint coordinates and the standard coordinates, but it is calculated by aggregation for multiple joints. Let the aligned trainee sequence be... A ′=( a 1′, a 2′,..., aL ′), the standard sequence is B ′=( b 1′, b 2′,..., bL For each time point t within the time window, calculate the error for each joint:

[0059] ;

[0060] Where k represents the key index (e.g., k=1 to 18). and These are the three-dimensional coordinates (X, Y, Z) of the student and the standard sequence at the k-th joint point at time t.

[0061] Then, for the entire time window, the aggregated error yields the pose error E within the window. The aggregation method employs a weighted average to emphasize key joints:

[0062] ;

[0063] in w k This is a weighting factor. Higher weights (e.g., 0.3) are assigned to core points (such as the sternal angle during chest compressions), while lower weights (e.g., 0.05) are assigned to non-core points. The total weights are normalized to 1. The weighting is based on the importance of the medical procedure, t. window Let t = start be the time window, and t = start be the start time. For a specific time point t, the Euclidean distance in three-dimensional space between the kth joint of the student and the kth joint of the standard teacher is given, and W represents the number of time points contained in a specific operation time window.

[0064] For example, during chest compressions, the sternal angle node (compression depth) has the highest weight (0.3), followed by the double elbow node (postural stability) (0.2 each), the L5 node (trunk posture) has a weight of 0.2, and the remaining nodes share the remaining weights. The weights were determined through clinical expert evaluation and machine learning optimization (e.g., using reinforcement learning to adjust and maximize training effectiveness).

[0065] Furthermore, error calculation also considers the continuity requirements over time. The system not only checks the average error within the window but also monitors the variance and peak value of the error. For example, it requires that the error of each core joint point be below a threshold for at least 95% of the time points within the window, and that the maximum error not exceed 1.5 times the threshold, to avoid misjudgments caused by transient errors. For critical operations such as compression depth, integral error is also used: the cumulative difference of the Z-axis displacement of the sternal angle node within the window is calculated to ensure that the overall depth meets the standard.

[0066] According to any of the aforementioned embodiments, a quantitative evaluation report is generated after training, comprehensively calculating the pressing depth attainment rate (weight factor 0.30), operation timing correctness (weight factor 0.50), and joint coordination (weight factor 0.20). For example, in a certain training session, the trainee's average pressing depth was 5.2cm (standard value 5.5cm, allowable deviation ±0.5cm), and the depth score was 0.3×[1-|5.2-5.5| / 0.5]=0.27; due to one state transition error (out of a total of 5 steps), the timing score was 0.5×(1-1 / 5)=0.40; the correlation coefficient of elbow and wrist joint angular velocity reached 0.87 (passing threshold 0.80), and the coordination score was 0.20, resulting in a final total score of 87. The system also generates a 3D motion replay of the erroneous steps to assist trainees in targeted improvement.

[0067] Based on the above-analyzed data, the system constructs a medical operation state machine (FSM), defining the state sequence as environmental assessment to open airway to artificial respiration to chest compressions. State transitions must meet strict conditions: for example, when transitioning from environmental assessment to open airway, the error between the trainee's hand positions and the standard coordinates must be less than a threshold of 0.25; the transition from open airway to artificial respiration requires the angle between the subchinus point and earlobe to remain stable within the range of 45°±5° for at least 1.0 second.

[0068] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system.

[0069] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0070] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0071] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A virtual reality-based emergency medical rescue drill teaching system, characterized by, The computer terminal comprises a static posture feature extraction module with a residual neural network model and a dynamic trajectory feature extraction module with a three-dimensional convolutional neural network model; The data acquisition module comprises a plurality of sensor nodes arranged on the body of the student for collecting physical quantity data; The VR device comprises a head-mounted display device and a simulated training dummy; The computer terminal is configured to: extract the static posture features of the student and the dynamic trajectory features of the standard action, the static posture features extracted by the residual neural network model, and the dynamic trajectory features extracted by the three-dimensional convolutional neural network model, and finally output a joint coordinate sequence representing the standard action; The determination module comprises: aligning the joint coordinate sequence of the student and the standard joint coordinate sequence of the teacher in time under the constraint of the maximum time offset through a dynamic time warping algorithm; calculating the Euclidean norm difference between the joint coordinates of the student and the standard coordinates within the operation time window as the posture error; comprising: The length of the joint coordinate sequence A is M, and the length of the standard joint coordinate sequence B of the teacher is N; The sequences A and B both contain three-dimensional coordinate data of a plurality of joint nodes, and a sampling rate RS is set; aM The sequence A =( a 1, a 2,..., ai ), where each point bN It is a high-dimensional vector representing the coordinates of all relevant nodes at time point i, the sequence. B =( b 1, b 2,..., pK ), where each point bi is a high-dimensional vector representing the coordinates of all teachers' standard key points at time point i; The goal of the dynamic time warping algorithm is to find a warping path P =( p 1, p 2,..., pk ) where each ik =( jk , Then, the construction of the cumulative distance matrix D is performed, comprising: ) represents an alignment of point i k of sequence A with point j k of sequence B, and the path P satisfies: Boundary conditions are i 1 = 1, j 1 = 1 and i K = M , j K = N; Monotonicity is i k +1≥ i k and j k+1 ≥ j k ; continuity is i k+1 - i k ≤ 1 and j k+1 - j k ≤ 1 ; finding the minimum cumulative path from D(M, N) to D(1, 1) through backtracking, and finding the best alignment between each point in the aligned sequences A' and B' with the same length L; First, a local distance matrix d is computed, where each element is the norm difference Then, the accumulated distance matrix D(i,j) represents the minimum accumulated distance from (1,1) to (i,j) and is recursively defined as: ; The maximum time offset constraint is applied, such that for any alignment point (i,j), the following must be satisfied: i - j ∣ T max × SR, where SR is the sampling rate; aL After alignment, the determination module calculates the pose error within the operation time window, including: aggregated calculation for multiple joint nodes, wherein the aligned student sequence is A ′=( a 1′, a 2′,..., bL ′), the standard sequence is B ′=( b 1′, b 2′,..., Then, for the entire time window, the posture error E within the window is obtained by aggregating the error, and the aggregation method adopts weighted average: ′); for each time point t within the time window, the error of each joint node is calculated: ; where k denotes the joint index, and are the three-dimensional coordinates (X, Y, Z) of the kth joint of the student and standard sequences, respectively, at time point t. aligning the student action sequence and the standard action sequence in time and outputting the determination result; ; wherein w k is a weight factor, the weights are set for core nodes and non-core nodes, the core nodes have higher weights than the non-core nodes, and the total weights are normalized to 1, t window is a time window, t=start is a start time, is the Euclidean distance in three-dimensional space between the kth node of the student and the kth node of the standard teacher node at a certain time point t, and W represents the number of time points included in the operation time window. triggering state transition or generating a quantitative evaluation report based on the determination result, wherein the triggering state transition comprises: a preset threshold value associated with the medical operation state is preset, and when the posture error of all core joint nodes in the current state continuously falls below the corresponding preset threshold value, the state transition is triggered. The data acquisition module comprises:

2. The virtual reality-based emergency medical rescue drill teaching system according to claim 1, characterized in that, a pressure sensor arranged on the palm and a torque sensor arranged on the torso belt; accelerometers arranged at key nodes of the body, including a plurality of standard medical joint nodes and a plurality of medical operation nodes. The residual neural network model is pre-trained with medical posture data and is configured to:

3. The virtual reality-based emergency medical rescue drill teaching system according to claim 2, wherein locate the two-dimensional pixel coordinates of 15 standard medical joint nodes and 3 medical operation nodes; and generate a high-dimensional feature vector describing the relative position, angle, and body contour of the nodes to form a single-frame posture snapshot. The three-dimensional convolutional neural network model is configured to:

4. The virtual reality-based emergency medical rescue drill teaching system according to claim 3, characterized in that, receive the continuous frame feature sequence extracted by the residual neural network, and simultaneously process the spatial dimension and the time dimension through a three-dimensional convolution kernel; learn the displacement, velocity, acceleration of the joint nodes and the spatiotemporal pattern of the overall action trajectory in the continuous frame sequence; decode the fused spatiotemporal features into a three-dimensional spatial coordinate sequence of 18 nodes through a fully connected layer to form the joint coordinate sequence of the standard action. The preset threshold values associated with the medical operation state of the determination module comprise:

5. The virtual reality-based emergency medical rescue drill teaching system according to claim 4, characterized in that, a first threshold value set for the posture error in the first medical operation state and a second threshold value set for the posture error in the second medical operation state; ​ When the pose error of all core nodes of the current state continuously falls below the corresponding first threshold or second threshold, a state transition is triggered; The process of triggering the state transition includes: Based on the current joint coordinate sequence, the state machine identifies the state in which the student is located; For the current state, the determination module calculates the pose error within the time window and compares it with the core node threshold of the state; If the error of all core nodes continuously falls below the threshold, a state transition signal is triggered, which includes the next state identifier and the transition confidence; After triggering, the VR device immediately updates the virtual scene.

6. The virtual reality-based emergency medical rescue drill teaching system according to claim 5, wherein The computer terminal further includes a state transition module configured to: When the state transition is triggered, the VR device switches to the next medical operation state in the virtual environment; Through visual prompts, audio instructions or task list updates, the student is guided to perform new operations.

7. The virtual reality-based emergency medical rescue drill teaching system according to claim 1, wherein The computer terminal further includes an evaluation module configured to: Extract key operation parameters and assign weight factors; Generate a quantitative evaluation report and a three-dimensional action playback of the error steps.

Citation Information

Patent Citations

  • Teaching method and mobile terminal

    CN107833283A

  • Glove-based acupuncture training method, system and platform and storage medium

    CN110826835A

  • Handmade ceramic teaching system applying visual sensor and use method of handmade ceramic teaching system

    CN120340323A