Cardiopulmonary resuscitation training system based on multi-mode artificial intelligence combined with virtual reality technology
Through a training system with multimodal artificial intelligence combined with virtual reality technology, the attitude and mechanical data of CPR operations are collected and integrated in real time, and virtual first aid scenarios are dynamically generated and multi-sensory feedback is provided. This solves the problems of inaccurate posture evaluation and unindividualized feedback in the existing system, and efficient and standardized CPR skills training is achieved.
Patent Information
- Application Number
- CN202510635082.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-24
AI Technical Summary
The existing virtual reality CPR training system lacks real-time feedback mechanism, cannot accurately identify operator posture deviations and provide personalized correction guidance, and data alignment and collaborative processing are difficult to achieve real-time and accurate operation quality evaluation and targeted feedback.
The training system using multimodal artificial intelligence combined with virtual reality technology is adopted. Through the multimodal perception module, the multi-modal perception module collects posture and mechanical data in real time, the multi-source data collaborative module performs data fusion and evaluation, the virtual scene construction module dynamically generates virtual first aid scenarios, and the decision module recognizes posture deviations and generates multi-sensory feedback instructions to achieve standardized, intelligent and diversified cardiopulmonary resuscitation skills training.
It significantly improved the pertinence, standardization and training effect of CPR training, and solved the problems of inaccurate posture evaluation, unindividual feedback, single scenarios and serious resource dependence in traditional training.
Smart Images

Figure CN120199131A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and particularly to a cardiopulmonary resuscitation training system based on multimodal artificial intelligence combined with virtual reality technology. Background Art
[0002] Cardiopulmonary resuscitation (CPR), as a crucial first-aid measure to save the lives of patients with cardiac arrest, its effective implementation is directly related to the survival rate and neurological function prognosis of patients. Traditional cardiopulmonary resuscitation training mainly relies on mannequin training devices, and trainees learn compression techniques under professional guidance. With the development of technology, a variety of real-time feedback devices have emerged on the market. These devices usually monitor key indicators such as compression depth, frequency, and chest wall recoil through sternum sensors, and provide simple warning feedback when the operation does not meet the standards.
[0003] Virtual reality (VR) technology creates a three-dimensional virtual environment to simulate diverse first-aid scenarios, providing multi-sensory simulations such as vision, hearing, and touch for training, creating an immersive experience, and making the training closer to the actual situation. Virtual technology has advantages such as customizable scenarios, low resource consumption, and unified training standards, and has been initially applied in the field of medical education. However, most existing virtual reality CPR training systems only focus on scenario simulation and process demonstration, lack effective integration with real-time feedback mechanisms, and cannot accurately identify the posture deviation of operators and provide personalized corrective guidance. In addition, data alignment and collaborative processing between the virtual environment and actual operations still face technical challenges, making it difficult to achieve real-time and accurate operation quality assessment and targeted feedback. Summary of the Invention
[0004] In view of this, the present invention proposes a cardiopulmonary resuscitation training system based on multimodal artificial intelligence combined with virtual reality technology. By real-time sensing and analyzing the posture and mechanical parameters of the training object, constructing a dynamically adjustable virtual first-aid scenario, establishing an accurate mapping relationship between the operation posture and compression efficacy, and generating multi-sensory personalized feedback instructions based on the quantitative evaluation results, it realizes standardized, intelligent, and scenario-diversified cardiopulmonary resuscitation skill training, and solves technical problems such as inaccurate posture assessment, non-personalized feedback, single scenario, and high resource dependence in traditional training.
[0005] The technical solution of the present invention is realized as follows:
[0006] The present invention provides a cardiopulmonary resuscitation training system based on multimodal artificial intelligence combined with virtual reality technology, including:
[0007] A multimodal perception module for collecting posture data, mechanical data, and environmental data of the training object when performing cardiopulmonary resuscitation operations;
[0008] A multi-source data collaboration module, which is used to perform cross-modal feature alignment and fusion on attitude data, mechanical data, and environmental data, generate an operation quality evaluation parameter set through spatio-temporal correlation modeling, and establish a mapping relationship between the operation attitude and the pressing efficiency;
[0009] A virtual scene construction module, which is used to generate a virtual first-aid environment containing variable-complexity obstacles, simulated patient physiological states, and environmental interference factors according to a preset first-aid scene template and the operation quality evaluation parameter set;
[0010] A decision-making module, which is used to identify operation attitude deviations and their quantitative impacts on the pressing efficiency based on the mapping relationship, and generate feedback control signals, including three-dimensional space trajectory guidance strategies, hierarchical tactile feedback signals, and acoustic prompt instructions;
[0011] An interactive closed-loop execution module, which is used to input the feedback control signals into the virtual first-aid environment and update the operation quality evaluation parameter set according to the operation data after the feedback response.
[0012] Preferably, the attitude data includes the three-dimensional coordinates of 25 skeletal key points of the upper limbs and the trunk, the body center of gravity distribution, the elbow joint angle, the arm straightness, the trunk verticality, the shoulder and scapula position parameters, and the posture stability index; the mechanical data includes the pressing depth, the pressing frequency, the thoracic cavity rebound rate, the pressing force, the pressing path offset, the pressing duration ratio, the pressing consistency index, and the arm force application angle; the environmental data includes the operation space geometric features, the obstacle distribution information, the environmental light intensity, the background noise level, the human activity information in the virtual scene, and the interference factor complexity quantification value.
[0013] Preferably, the multi-modal perception module includes:
[0014] A visual acquisition unit, including a depth camera and an RGB camera, which is used to acquire attitude data;
[0015] A mechanical detection unit, including a pressure sensor, an acceleration sensor, and an IMU, which is used to acquire mechanical data;
[0016] A biometric acquisition unit, including a blood oxygen saturation sensor and a pupil reflex monitoring module, which is used to obtain biometric parameters associated with the simulated patient physiological state;
[0017] An environmental perception unit, which is used to acquire environmental data.
[0018] Preferably, the multi-source data collaboration module includes:
[0019] A spatio-temporal alignment unit, which is used to align multi-modal sensor data to a unified spatio-temporal coordinate system through timestamp synchronization and spatial coordinate transformation, and generate a spatio-temporal synchronized data stream;
[0020] A feature extraction unit, which is used to extract pose features from pose data, extract pressing dynamics features from mechanical data through wavelet transform, and extract interference factor features through environmental perception data analysis, so as to obtain a multi-dimensional feature vector characterizing the operation quality;
[0021] A parameter set generation unit, which is used to analyze the multi-dimensional feature vector based on a sliding time window to form an operation quality evaluation parameter set, including a pose feature parameter set and a mechanical performance index set;
[0022] A mapping relationship modeling unit, which is used to construct a pose-force correlation model through a deep neural network, output the influence coefficient of the operation pose on the pressing efficiency, and establish a mapping relationship between the operation pose and the pressing efficiency.
[0023] Preferably, the calculation expression of the influence coefficient EI of the operation pose on the pressing efficiency is:
[0024]
[0025] In the formula, EI represents the influence coefficient of pressing efficiency; P k is the k-th pose feature parameter; W k is the dynamic weight of the pose feature parameter; n is the total number of pose feature parameters; Ω is the pressing depth in the mechanical performance index set; D(Ω) represents the compliance function of the pressing depth Ω with the target depth; λ is the pressing frequency in the mechanical performance index set; T(λ) represents the modulation factor of the pressing frequency λ; f(P k ) is the pose deviation function, and the calculation is as follows:
[0026]
[0027] Among them, S k is the reference standard value corresponding to P k ; σ k is the feature adaptation coefficient. When |P k −S k | is greater than the preset value, f(P k ) approaches 1, indicating that the pose deviation is serious; when |P k −S k | approaches 0, f(p k ) approaches 0, indicating that the pose meets the standard.
[0028] Preferably, both the compliance function D(Ω) and the modulation factor T(λ) adopt non-linear function forms, which are used to perform weighted correction on the pressing depth and the pressing frequency, where:
[0029]
[0030] In the formula, Ω ref and λ refThey are the target values of the compression depth and compression frequency respectively, and a and b are modulation coefficients.
[0031] Preferably, the virtual scene construction module includes:
[0032] A scene template library that stores three-dimensional components, texture maps, and physiological parameter reference values of four types of scenes, namely hospitals, homes, public places, and outdoor environments, using a hierarchical structured data model. Each type of scene contains N pluggable preset templates;
[0033] A dynamic scene configuration engine that, based on a procedural content generation algorithm and a constraint optimization model, analyzes the spatial trajectory offset in the operation quality evaluation parameter set through a sliding window to generate the obstacle layout and environmental interference factors that match the current training stage in real time;
[0034] A virtual character generator for constructing virtual patients with physiological state feedback and simulating bystanders and assistants in the first aid scene;
[0035] A scene parameter scheduler for adjusting the obstacle density, environmental interference intensity, and time pressure according to the attitude stability and compression effectiveness indicators in the operation quality evaluation parameter set to achieve optimal control of the scene difficulty.
[0036] Preferably, the scene parameter scheduler adjusts the scene difficulty by calculating the scene complexity score SC:
[0037] SC = α·log(1 + OD + η)·β PE ·γ 1―SR
[0038] In the formula, SC represents the scene complexity score; OD represents the obstacle density, that is, the number of obstacles in the unit space; η represents the environmental interference factor; PE represents the operation proficiency score, which is calculated from the operation quality evaluation parameter set; SR represents the historical success rate, that is, the proportion of the historical successful times of the training object in similar scenes; α, β, and γ are adjustment coefficients, and α > 0, 0 < β < 1, γ > 1.
[0039] Preferably, the decision-making module includes:
[0040] An attitude evaluation unit for comparing the skeletal key point data with the standard reference model through a spatio-temporal trajectory matching algorithm to identify at least one type of attitude defect such as shoulder forward tilt deviation, elbow flexion abnormality, and torso offset angle exceeding the standard, and calculating the contribution value of each defect to the compression efficiency influence index EI based on the attitude-force correlation model;
[0041] A biomechanical analysis unit for generating a comprehensive compression quality score based on the compression depth deviation value, frequency fluctuation coefficient, and thoracic cavity rebound lag time, in combination with the operation quality evaluation parameter set;
[0042] The strategy optimization unit adopts a reinforcement learning framework based on training stage segmentation, and optimizes the personalized error correction strategy by dynamically adjusting the weight of the pressing depth compliance rate, the frequency stability coefficient, and the proportion of the posture deviation correction amount;
[0043] The multi-modal feedback generation unit is used to generate three-dimensional trajectory guidance marks, graded tactile pulse sequences, and priority voice prompts based on the pressing efficiency impact index EI and the comprehensive score of pressing quality, where the generation priority of the content of the voice prompt is consistent with the contribution value ranking.
[0044] Preferably, the multi-modal perception module is communicatively connected to the multi-source data collaboration module, the multi-source data collaboration module is communicatively connected to the virtual scene construction module and the decision-making module, the decision-making module is communicatively connected to the interactive closed-loop execution module, and the interactive closed-loop execution module is communicatively connected to the virtual scene construction module to form a closed-loop system for data acquisition, processing, decision-making, feedback, and optimization.
[0045] The present invention has the following beneficial effects compared with the prior art:
[0046] (1) Through the integrated application of multi-modal perception technology and virtual reality technology, the present invention realizes the real-time, accurate perception, analysis, and evaluation of the cardiopulmonary resuscitation operation posture and mechanical parameters, establishes a quantitative mapping relationship between the operation posture and the pressing efficiency, and simultaneously dynamically generates a virtual first aid scene that matches the training difficulty and provides personalized multi-sensory feedback guidance, thereby significantly improving the pertinence, standardization degree, and training effect of cardiopulmonary resuscitation training. This system forms a complete closed loop of data acquisition, processing, decision-making, feedback, and optimization, expands CPR training from simple pressing quality monitoring to comprehensive operation behavior evaluation and guidance, and effectively solves technical problems such as inaccurate posture evaluation, non-personalized feedback, single training scene, and severe resource dependence in traditional training;
[0047] (2) The calculation model of the pressing efficiency impact coefficient EI established by the present invention realizes the accurate quantification of the degree to which different posture defects affect the quality of CPR by performing non-linear function mapping on multi-dimensional posture feature parameters and standard reference values, and combining the pressing depth compliance function and the frequency modulation factor. This model can not only accurately identify posture defects such as shoulder forward tilt deviation, elbow flexion abnormality, and excessive trunk offset angle, but also calculate the specific contribution value of each defect to the pressing efficiency, enabling the system to preferentially provide correction guidance for the posture problem with the greatest impact, thereby improving the pertinence and efficiency of training and effectively overcoming the limitation of the traditional technology that cannot quantify the relationship between posture and pressing effect;
[0048] (3) The virtual scene construction module of the present invention stores four types of scene templates using a hierarchical structured data model, and through a procedural content generation algorithm combined with the scene complexity scoring formula SC, dynamically adjusts the scene obstacle density, environmental interference intensity, and time pressure according to the trainer's operation proficiency, posture stability, and historical success rate, realizing the adaptive adjustment of training difficulty. This performance-based scene generation technology keeps the training content always in the optimal challenge range, avoids the monotony and adaptability problems brought by the fixed scene mode, and at the same time enhances the immersion and authenticity of the training;
[0049] (4) The decision-making module and the interactive closed-loop execution module of the present invention jointly construct a feedback system based on multi-sensory channels, and provide intuitive, real-time, and multi-dimensional corrective guidance for the trainer by combining three-dimensional trajectory guidance markers, hierarchical tactile pulse sequences, and priority voice prompts. The system optimizes the personalized error correction strategy using a reinforcement learning framework based on training stage segmentation, making the feedback content match the trainer's current skill level and cognitive characteristics, avoiding information overload, and improving the effectiveness and acceptance of the feedback. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 It is the system framework diagram of the present invention;
[0052] Figure 2 It is the technical implementation diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will combine the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0054] As Figure 1 shown, the present invention provides a cardiopulmonary resuscitation training system based on multi-modal artificial intelligence combined with virtual reality technology, including:
[0055] A multi-modal perception module for collecting the posture data, mechanical data, and environmental data of the training object when performing cardiopulmonary resuscitation operations;
[0056] A multi-source data collaboration module, which is used to perform cross-modal feature alignment and fusion on attitude data, mechanical data, and environmental data, generate an operation quality evaluation parameter set through spatio-temporal correlation modeling, and establish a mapping relationship between the operation attitude and the pressing efficiency;
[0057] A virtual scene construction module, which is used to generate a virtual first-aid environment containing variable-complexity obstacles, simulated patient physiological states, and environmental interference factors according to a preset first-aid scene template and the operation quality evaluation parameter set;
[0058] A decision-making module, which is used to identify the operation attitude deviation and its quantitative impact on the pressing efficiency based on the mapping relationship, and generate feedback control signals, including a three-dimensional space trajectory guidance strategy, a hierarchical tactile feedback signal, and an acoustic prompt instruction;
[0059] An interactive closed-loop execution module, which is used to input the feedback control signal into the virtual first-aid environment and update the operation quality evaluation parameter set according to the operation data after the feedback response.
[0060] Among them, the multi-modal perception module is communicatively connected to the multi-source data collaboration module, the multi-source data collaboration module is communicatively connected to the virtual scene construction module and the decision-making module, the decision-making module is communicatively connected to the interactive closed-loop execution module, and the interactive closed-loop execution module is communicatively connected to the virtual scene construction module, forming a closed-loop system for data acquisition, processing, decision-making, feedback, and optimization.
[0061] As Figure 2 shown, the technical process of the present invention is as follows: The system first collects visual, mechanical, biometric, and environmental data of the training object performing cardiopulmonary resuscitation operations through the multi-modal perception module; then the multi-source data collaboration module performs spatio-temporal alignment, feature extraction, and parameter set generation on these heterogeneous data, and establishes a mapping relationship between the attitude and the pressing efficiency; the virtual scene construction module dynamically generates a virtual first-aid environment that matches the training requirements based on the operation quality evaluation parameter set through a scene template library, a dynamic scene configuration engine, a virtual character generator, and a scene parameter scheduler; the decision-making module analyzes the quantitative impact of the operation attitude deviation on the pressing efficiency, and generates multi-sensory feedback control signals through attitude evaluation, biomechanical analysis, and strategy optimization; finally, the interactive closed-loop execution module inputs these control signals into the virtual first-aid environment, executes feedback through scene rendering, tactile feedback, and acoustic signal processing, and at the same time monitors the operation response and updates the operation quality evaluation parameter set to form a complete closed-loop training system.
[0062] Specifically, in an embodiment of the present invention, the pose data includes three-dimensional coordinates of 25 skeletal key points of the upper limb and the torso, body center of gravity distribution, elbow joint angle, arm straightness, torso verticality, shoulder and scapula position parameters, and pose stability index; the mechanical data includes pressing depth, pressing frequency, thoracic cavity rebound rate, pressing force, pressing path offset, pressing duration ratio, pressing consistency index, and arm force application angle; the environmental data includes operational space geometric features, obstacle distribution information, environmental light intensity, background noise level, human activity information in the virtual scene, and interference factor complexity quantization value.
[0063] In this embodiment, the multi-modal perception module includes:
[0064] A visual acquisition unit, including a depth camera and an RGB camera, for acquiring pose data;
[0065] A mechanical detection unit, including a pressure sensor, an acceleration sensor, and an IMU, for acquiring mechanical data;
[0066] A biometric acquisition unit, including a blood oxygen saturation sensor and a pupil reflex monitoring module, for obtaining biometric parameters associated with the physiological state of the simulated patient;
[0067] An environmental perception unit, for acquiring environmental data.
[0068] In a specific example, the multi-modal perception module adopts a distributed sensor architecture. The visual acquisition unit adopts a dual-view configuration combining a depth camera and an RGB camera. The depth camera adopts structured light or ToF technology and is installed at the front side of the training site at a certain angle with the ground to ensure complete coverage of the operator's whole body; the RGB camera is installed directly above the training site to provide a top-down planar image. This unit adopts a two-stream fusion architecture: First, the initial key point positions are extracted from the RGB image through the HRNet network, and at the same time, the depth camera data is used to enhance the accuracy of spatial position inference through point cloud projection and feature matching; then, time series filtering and Kalman prediction algorithms are applied to eliminate detection jitter and output stable 25-point skeletal model data, including the three-dimensional spatial coordinates of the shoulder, torso, elbow, wrist, and finger joints. The system also calculates key pose indicators such as body center of gravity distribution, elbow joint angle, arm straightness, torso verticality, and pose stability.
[0069] The mechanical detection unit integrates a pressure sensor, an acceleration sensor, and an Inertial Measurement Unit (IMU). Among them, the pressure sensor is embedded in the chest of the training model to collect data on the pressing depth, pressing force, and chest wall rebound rate; the acceleration sensor is installed at the wrist position of the trainer to collect the hand movement trajectory and acceleration changes; the IMU sensor integrates a three-axis accelerometer, a gyroscope, and a magnetometer and is installed on the back of the trainer's hand to measure the arm force application angle and the deviation of the pressing path.
[0070] The biometric acquisition unit includes a blood oxygen saturation sensor and a pupil reflex monitoring module, which are used to simulate and monitor the physiological state changes of virtual patients. These biometric parameters are associated with the virtual patient model, providing the trainer with approximate real patient state feedback and enhancing the realism and sense of urgency of the training.
[0071] The environment perception unit consists of an environment camera, a distance sensor array, and an audio sensor. The environment camera is a fish-eye camera, which is used to obtain the geometric features of the operation space and the distribution of obstacles; the distance sensor array is composed of multiple ultrasonic sensors, which are evenly distributed on the edge of the training space to monitor the space boundary and dynamic obstacles in real time; the audio sensor is used to collect the environmental noise level and acoustic features.
[0072] Specifically, in an embodiment of the present invention, the multi-source data collaboration module includes:
[0073] The spatio-temporal alignment unit is used to align multi-modal sensor data to a unified spatio-temporal coordinate system through timestamp synchronization and spatial coordinate transformation, generating a spatio-temporally synchronized data stream.
[0074] Specifically, the spatio-temporal alignment unit adopts a multi-level synchronization mechanism to process heterogeneous data from different sensors. Time alignment is achieved through a hardware synchronization scheme based on a shared high-precision clock source, combined with a timestamp interpolation correction algorithm at the software level, to align all sensor data to a unified time axis. Spatial alignment is completed in three steps: first, a transformation matrix between each sensor and the global coordinate system is established during the calibration phase; then, during operation, the spatial positions of different devices are transformed to the global coordinate system through rigid body transformation; finally, the transformation result is optimized through the Iterative Closest Point (ICP) algorithm to reduce the cumulative error.
[0075] The feature extraction unit is used to extract pose features from pose data, extract pressing dynamics features from mechanical data through wavelet transform, and extract interference factor features through environmental perception data analysis, obtaining a multi-dimensional feature vector characterizing the operation quality.
[0076] Specifically, the feature extraction unit adopts specialized algorithms for different types of data: for pose data, a spatial feature extraction network based on the PointNet++ architecture is used to extract spatial pose features from the point cloud formed by skeletal key points; a temporal convolutional network is used to extract dynamic features from the key point trajectories of consecutive frames; the two sets of features are fused through an attention mechanism to form a complete pose representation. For mechanical parameter data, wavelet transform is used for time-frequency analysis to decompose the original pressure and acceleration signals into different frequency components, and the energy distribution of the frequency band corresponding to the normal pressing frequency range is mainly analyzed to extract peak features and statistical characteristics. For environmental perception data, a multi-scale feature fusion method is used to extract environmental interference features, including spatial obstacle distribution, noise features, etc.
[0077] The parameter set generation unit is used to analyze the multi-dimensional feature vector based on a sliding time window to form an operation quality evaluation parameter set, including a pose feature parameter set and a mechanical performance index set.
[0078] Specifically, the parameter set generation unit analyzes the temporal variation of the feature vector based on a multi-scale sliding window mechanism. The system simultaneously maintains a short-term window, a medium-term window, and a long-term window, which are used to capture immediate action features, analyze pressing periodicity, and evaluate overall operation stability respectively. For the feature sequence within each window, statistics and temporal patterns are calculated to construct an operation quality evaluation parameter set including a pose feature parameter set and a mechanical performance index set. The pose feature parameter set includes the elbow angle stability index, the torso vertical deviation index, the shoulder alignment index, etc.; the mechanical performance index set includes key performance indicators such as pressing depth, pressing frequency, and chest wall rebound rate.
[0079] The mapping relationship modeling unit is used to construct a pose-force correlation model through a deep neural network, output the influence coefficient of the operation pose on the pressing efficiency, and establish the mapping relationship between the operation pose and the pressing efficiency.
[0080] Specifically, a deep neural network is used to construct a pose-force correlation model, which adopts a two-stream network architecture: one stream processes the pose feature parameter set, and the other stream processes the mechanical performance index set. The features of the two paths are fused through an attention mechanism to generate a comprehensive feature representation. The training of the model adopts a combination of supervised learning and self-supervised learning. The supervised part uses the data annotated by experts, and the loss function includes a weighted combination of the pressing depth prediction error, the pressing frequency prediction error, and the expert score prediction error; the self-supervised part learns the internal correlation between the pose and the force from the unannotated data through contrast learning to enhance the generalization ability of the model.
[0081] In this embodiment, the specific implementation of the posture-force correlation model adopts a sequence modeling architecture based on Transformer, which consists of an encoder and a decoder. The encoder is composed of 6 layers of Transformer structures, each layer contains 8 attention heads, and the hidden layer dimension is 512. The long-term dependence relationship of the posture feature sequence is captured through the self-attention mechanism. The decoder is composed of 4 layers of Transformer structures, which fuse the posture features and mechanical performance features through the cross-attention mechanism to generate the influence coefficient of the posture on the pressing efficiency. The model is trained with a batch size of 64, the initial learning rate is set to 0.0002 and the cosine annealing strategy is used, and the number of training epochs is 200. The model input is the posture feature parameters and mechanical performance indicators for 30 consecutive seconds (about 50-60 presses), and the output is the influence coefficient EI of the operating posture on the pressing efficiency.
[0082] The calculation expression of the influence coefficient EI of the operating posture on the pressing efficiency is as follows:
[0083]
[0084] where EI represents the influence coefficient of the pressing efficiency, which is used to quantify the comprehensive influence degree of different posture deviations on the pressing efficiency and provide a basis for the decision-making module to generate a targeted feedback strategy; P k is the k-th posture feature parameter, which is an element in the posture feature parameter set; W k is the dynamic weight of the posture feature parameter, which is adaptively adjusted according to the training progress and individual differences; n is the total number of posture feature parameters; Ω is the pressing depth in the mechanical performance index set; D(Ω) represents the compliance function of the pressing depth Ω with the target depth; λ is the pressing frequency in the mechanical performance index set; T(λ) represents the modulation factor of the pressing frequency λ; f(P k ) is the posture deviation function, and the calculation is as follows:
[0085]
[0086] where S k is the reference standard value corresponding to P k , σ k is the feature adaptation coefficient. When |P k −S k | is greater than the preset value, f(P k ) approaches 1, indicating that the posture deviation is serious; when |P k −S k | approaches 0, f(P k ) approaches 0, indicating that the posture meets the standard.
[0087] In this embodiment, both the compliance function D(Ω) and the modulation factor T(λ) adopt non-linear function forms and are used to perform weighted correction on the pressing depth and the pressing frequency, where:
[0088]
[0089] In the formula, Ω ref and λ ref are respectively the target values of the pressing depth and the pressing frequency, and a and b are modulation coefficients. When Ω is close to Ω ref , D(Ω) tends to be between 0.5 and 1, indicating that the pressing depth meets the standard; when the deviation between λ and λ ref is larger, T(λ) decays exponentially, indicating that the influence of the rhythm problem on the pressing efficiency increases significantly.
[0090] Specifically, W k is adaptively calculated through the attention mechanism to reflect the importance of different pose parameters in the current operation state. The calculation formula is:
[0091]
[0092] In the formula, q k is the query vector of the current operation state, V is the parameter importance knowledge base, and d is the vector dimension. The dynamic adjustment of the weight ensures that the system focuses on the pose defects that most affect the operation effect.
[0093] In this embodiment, the value range of the pressing efficiency influence coefficient EI is [0, 1], where 0 means that the operation pose does not affect the pressing efficiency at all, and 1 means that the operation pose seriously affects the pressing efficiency and results in ineffective pressing.
[0094] The mapping relationship modeling unit also implements a reverse mapping function, which inversely deduces the ideal pose parameter combination according to the target pressing efficiency, and provides the target value of pose correction for the decision-making module. This function is realized through the gradient optimization method, with the goal of minimizing the predicted EI value, searching for the optimal solution in the pose parameter space, and considering the ergonomic constraints to ensure that the recommended pose meets the physiological conditions of the operator.
[0095] The multi-source data collaboration module and the multi-modal perception module perform data exchange through a cache queue in a publish-subscribe mode to ensure low-latency data transmission. The processing results of the collaboration module are transmitted to the decision-making module and the virtual scene construction module through a standard interface, forming a complete link of data processing and feedback.
[0096] Specifically, in an embodiment of the present invention, the virtual scene construction module includes:
[0097] The scenario template library stores 3D components, texture maps, and physiological parameter reference values for four types of scenarios, namely hospitals, homes, public places, and outdoor environments, using a hierarchical structured data model. Each type of scenario contains N pluggable preset templates.
[0098] Specifically, a hierarchical structured data model is used to store four types of basic scenario templates for hospitals, homes, public places, and outdoor environments. Hospital scenarios include typical medical environments such as emergency rooms, wards, and operating rooms; home scenarios include common living spaces such as bedrooms, living rooms, and kitchens; public places include crowded areas such as shopping malls, subway stations, and restaurants; outdoor environments include open spaces such as parks, mountains, and beaches. Each type of scenario contains multiple pluggable preset templates, and the number of templates can be dynamically increased according to system expansion requirements. The scenario templates organize data in a three-layer structure: the basic layer stores the spatial geometric structure, obstacle prototypes, and texture maps; the interaction layer defines the operable objects, character positions, and interaction trigger areas; the parameter layer contains lighting parameters, sound configurations, and physiological parameter reference values.
[0099] The system uses a knowledge-graph-based scenario semantic representation, representing the relationships between scenario elements as connection edges in the knowledge graph, such as "first aid kit - located at - nurse station", "hospital bed - blocking - movement path", etc. This representation method enables the system to understand the semantic associations between scenario elements and maintain the logical rationality of the scenario during generation. The template components adopt an interface standardization design, enabling seamless combination of templates from different sources and greatly increasing the possibility of scenario variations.
[0100] The dynamic scenario configuration engine, based on the procedural content generation algorithm and the constraint optimization model, analyzes the spatial trajectory offset in the operation quality evaluation parameter set through a sliding window to generate the obstacle layout and environmental interference factors that match the current training stage in real time.
[0101] Specifically, based on the procedural content generation (PCG) algorithm and the constraint optimization model, it generates the obstacle layout and environmental interference factors that match the current training stage in real time. The engine first analyzes the spatial trajectory offset in the operation quality evaluation parameter set through a sliding window to identify the spatial activity patterns of the operator in different postures, and then dynamically adjusts the layout of key elements in the virtual environment according to the recognition results. The procedural content generation algorithm adopts a hierarchical structure: the global planning layer is responsible for determining the main activity areas and key paths; the local optimization layer processes the positions and attributes of specific obstacles; the detail generation layer adds environmental textures and auxiliary objects.
[0102] The constraint optimization model formulates the scenario generation problem as a multi-objective optimization problem, comprehensively considering the constraint conditions in three dimensions: space limitation, training objectives, and operation difficulty. The space limitation constraint ensures that the generated scenario conforms to the actual conditions of the physical space; the training objective constraint ensures that the scenario design can provide corresponding challenges for specific training focuses (such as posture maintenance, stable pressing frequency, etc.); the operation difficulty constraint adjusts the complexity of the challenges according to the current skill level of the trainer. The constraint optimization adopts a hybrid method combining genetic algorithm and simulated annealing to avoid falling into local optimal solutions while ensuring the convergence efficiency.
[0103] By continuously monitoring the changes in the key indicators in the operation quality evaluation parameter set, when significant changes in operation performance are detected (such as a continuous decline in the quality of multiple presses), the system triggers the scenario reconfiguration process, adjusts the distribution of obstacles or the intensity of interference factors in the current environment, so that the training environment always maintains an appropriate level of challenge, enhancing the pertinence and effectiveness of training.
[0104] A virtual character generator is used to construct virtual patients with physiological state feedback and simulate bystanders and assistants in the first aid scene.
[0105] Specifically, the virtual character generator is used to construct virtual patients with physiological state feedback and simulate bystanders and assistants in the first aid scene. The virtual patient model adopts a multi-level structure: the appearance layer realizes the visual performance of the patient, including facial expressions, skin color changes, and limb movements; the physiological layer simulates the internal physiological state, including key vital signs such as heart rate, blood oxygen saturation, and pupil reaction; the feedback layer associates the operation quality with the changes in physiological state to achieve the dynamic response of the patient's state to the operation effect.
[0106] The virtual character generator is also responsible for generating on-site bystanders and assistants. These characters adopt an artificial intelligence control system based on a behavior tree and can perform different actions according to the scene state and operation progress, such as assisting in moving obstacles, delivering first aid equipment, or providing verbal encouragement. The system dynamically adjusts the behavior patterns of these assistants according to the training objectives, providing more support in the primary training stage and possibly introducing certain interference in the advanced training stage to simulate the complex interpersonal interactions in the real first aid environment.
[0107] A scenario parameter scheduler is used to adjust the obstacle density, environmental interference intensity, and time pressure according to the posture stability and pressing effectiveness indicators in the operation quality evaluation parameter set to achieve the optimal regulation of the scenario difficulty.
[0108] Specifically, the scenario parameter scheduler calculates the complexity level of the current scenario based on the scenario complexity score SC formula and dynamically adjusts the scenario parameters based on this:
[0109] SC = α·log(1 + OD + η)·βPE ·γ 1―SR
[0110] In the formula, SC represents the scene complexity score; OD represents the obstacle density, that is, the number of obstacles in the unit space; η represents the environmental interference factor, which is obtained by weighted summation of the visual interference intensity, auditory interference intensity, action interference intensity and time pressure intensity; PE represents the operation proficiency score, which is calculated from the operation quality evaluation parameter set; SR represents the historical success rate, that is, the proportion of the historical successful times of the training object in the similar scenarios; α, β, and γ are adjustment coefficients, and α>0, 0<β<1, γ>1.
[0111] In this embodiment, the obstacle density OD represents the number of obstacles in the unit space, and its calculation needs to consider the volume, distribution and influence degree on the operation of the obstacles, and is obtained by the following method:
[0112]
[0113] In the formula, V i represents the volume of the i-th obstacle; w i represents the obstacle weight factor, which reflects the hindrance degree of the obstacle to the operation behavior; D i represents the distance influence factor between the obstacle and the operation core area; V represents the total volume of the effective operation space.
[0114] In the actual system, the obstacle recognition and volume calculation obtain data through the depth camera and distance sensor array of the environmental perception unit, generate the obstacle space model through the three-dimensional reconstruction algorithm, and then calculate the OD value by applying the above formula.
[0115] In this embodiment, the operation proficiency score PE reflects the current operation skill level of the training object, and is calculated by comprehensively analyzing the posture feature parameter set and mechanical performance index set in the operation quality evaluation parameter set:
[0116] PE = w p ×PS + w f ×PF
[0117] In the formula, PS is the posture stability score, which reflects the standard degree of the operation posture and is calculated according to the posture feature parameter set; PF is the mechanical performance score, which reflects the effectiveness of the pressing operation and is calculated according to the mechanical performance index set; w p and w f are weight coefficients, satisfying w p + w f = 1.
[0118] In this embodiment, the historical success rate SR represents the proportion of the historical successful times of the training object in the similar scenarios, and is calculated by time-weighted average to make the most recent training results have a higher weight:
[0119]
[0120] Wherein, R i is the success indication variable for the i-th training, with success being 1 and failure being 0; t i is the timestamp of the i-th training; t is the current timestamp; λ is the time decay coefficient, controlling the decay rate of the weight of historical data; the summation range is the past M training records, and M takes values from 10 to 20.
[0121] In this embodiment, the scheduling strategy of the scenario parameter scheduler is based on the "zone of proximal development" theory in educational psychology, aiming to provide a training environment with moderate challenges. When the SC value is too low (indicating that the current scenario is too simple), the system will increase the obstacle density and the intensity of environmental interference, or introduce time pressure factors; when the SC value is too high (indicating that the current scenario is too complex), the system will reduce the degree of interference and simplify the scenario layout to ensure that the training difficulty always remains near the ability boundary of the trainer, maximizing the learning effect.
[0122] The scenario parameter scheduler also implements a smooth transition mechanism for scenario difficulty. The system does not suddenly change the scenario parameters, but achieves smooth transition through progressive changes, avoiding abrupt changes in the training experience. This smooth transition is achieved through an interpolation algorithm, which dynamically adjusts the change rate according to the trainer's adaptation speed, ensuring the coherence and immersion of the training experience.
[0123] The units of the virtual scenario construction module cooperate through an event-driven message passing mechanism. When the operation quality evaluation parameter set is updated, the scenario parameter scheduler calculates a new scenario complexity score SC; the change in the SC value triggers the dynamic scenario configuration engine to regenerate or adjust scenario elements; the scenario change simultaneously notifies the virtual character generator to update the character state and behavior. The interaction between the virtual scenario construction module and other modules of the system is implemented through a standardized interface. It receives the operation quality evaluation parameter set from the multi-source data collaboration module as the basis for scenario generation and complexity adjustment; provides the current scenario state information to the decision-making module to support situation-aware decision-making; and receives the real-time operation feedback from the interaction closed-loop execution module for dynamically adjusting scenario elements.
[0124] Specifically, in an embodiment of the present invention, the decision-making module includes:
[0125] The posture evaluation unit is used to compare the skeletal key point data with the standard reference model through a spatio-temporal trajectory matching algorithm, identify at least one type of posture defect such as forward shoulder deviation, abnormal elbow flexion, and excessive trunk offset angle, and calculate the contribution value of each defect to the pressing efficiency impact index EI based on the posture-force correlation model.
[0126] Specifically, the posture evaluation unit compares the bone key point data collected in real time with the standard reference model through a spatio-temporal trajectory matching algorithm to identify the types of posture defects during cardiopulmonary resuscitation operations. This unit adopts a two-stream spatio-temporal neural network architecture to process spatial configuration information and time series information respectively. The spatial stream processes the human posture graph composed of bone key points through a graph convolutional network to extract joint position relationships and angle features; the temporal stream analyzes the action change features in consecutive frames through a temporal convolutional network to capture action coherence and stability. The feature maps of the two streams are integrated through a feature fusion layer to form a spatio-temporal unified posture representation.
[0127] In this embodiment, the spatio-temporal trajectory matching algorithm adopts the dynamic time warping algorithm, which introduces an elastic deformation tolerance parameter to allow adaptation to the speed differences of different operators while keeping the order of key actions unchanged. The matching process calculates a comprehensive matching score based on two dimensions: the geometric similarity and the temporal consistency of the key point trajectories:
[0128] M(T,R) = α·G s (T,R)+(1―α)·T s (T,R)
[0129] where T represents the trajectory sequence to be evaluated, R represents the standard reference trajectory, G s represents the geometric similarity function, T s represents the temporal consistency function, and α is a balance parameter. The geometric similarity is obtained by calculating the weighted sum of the Euclidean distances between corresponding key points; the temporal consistency is evaluated by analyzing the time distribution and relative order of key action points.
[0130] Based on the matching results, the posture evaluation unit can accurately identify three main types of posture defects: shoulder forward lean deviation (the forward movement distance of the shoulder key point beyond the vertical reference line), elbow flexion abnormality (the elbow joint angle is less than the recommended straight-arm angle), and excessive trunk deviation angle (the angle between the trunk midline and the vertical reference line exceeds the allowable range). For each identified defect, the system calculates the specific contribution value of its impact on the compression efficacy index EI based on the posture-force correlation model to form a sorted contribution list. This contribution value calculation uses a gradient-based sensitivity analysis method. By changing each posture parameter separately through the control variable method, the change amplitude of the EI value is observed to quantify the impact degree of each defect.
[0131] The posture evaluation unit also implements a posture stability evaluation function. By analyzing the variance of key point positions and angle fluctuations within a continuous time window, the posture maintenance stability index is calculated. This index reflects the operator's ability to maintain a standard posture during continuous compressions.
[0132] A biomechanical analysis unit for generating a comprehensive score of compression quality according to the deviation value of compression depth, the frequency fluctuation coefficient, and the thoracic cavity rebound lag time, in combination with an operation quality evaluation parameter set.
[0133] Specifically, the biomechanical analysis unit evaluates the quality of the compression operation from a biomechanical perspective, mainly based on three core indicators: the deviation value of compression depth, the frequency fluctuation coefficient, and the thoracic cavity rebound lag time. The deviation value of compression depth is obtained by calculating the deviation between the actual compression depth and the target depth range; the frequency fluctuation coefficient is quantified by analyzing the variability of the time intervals of consecutive compression cycles; and the thoracic cavity rebound lag time is calculated by measuring the time required for the thoracic cavity to return to its original position after the compression is released.
[0134] This unit uses a comprehensive scoring model to integrate multiple indicators and generate a comprehensive score of compression quality PQ:
[0135] PQ = w d ·F d (Δd) + w f ·F f (σ f ) + w r ·F r (t r ) + w e ·F e (E)
[0136] Wherein, F d , F f , F r and F e are the scoring functions for depth deviation, frequency fluctuation, rebound lag, and energy transfer efficiency respectively, Δd, σ f , t r and E are the corresponding measured values, and w d , w f , w r and w e are the weight coefficients of each index. Each scoring function uses a non-linear mapping to convert the original measured value into a standardized score. For example, the depth deviation scoring function F d adopts a Gaussian response curve:
[0137]
[0138] Where σ d is an adjustable tolerance parameter, reflecting the tolerance for depth deviation.
[0139] The biomechanical analysis unit introduces the evaluation of energy transfer efficiency. This indicator comprehensively evaluates the physical efficacy of the pressing operation by analyzing the consistency between the acting direction of the force and the displacement direction of the chest during pressing, and calculating the proportion of the force converted into the effective pressing depth. The energy transfer efficiency is highly correlated with the pressing posture, providing an important intermediate verification indicator for the posture-force correlation.
[0140] The strategy optimization unit adopts a reinforcement learning framework based on training stage segmentation, and optimizes the personalized error correction strategy by dynamically adjusting the weight of the pressing depth compliance rate, the frequency stability coefficient, and the proportion of the posture deviation correction amount.
[0141] Specifically, the strategy optimization unit adopts a reinforcement learning framework based on training stage segmentation to optimize the personalized error correction strategy for training objects with different training stages and individual characteristics. This unit divides the cardiopulmonary resuscitation training process into four stages: the basic posture establishment stage, the single element mastery stage, the comprehensive skill formation stage, and the stability improvement stage. Different optimization goals and constraint conditions are set for each stage to form a targeted strategy optimization space.
[0142] The reinforcement learning framework adopts a deep reinforcement learning model based on the Actor-Critic architecture. The state space includes the operation quality evaluation parameter set, the historical feedback effect data, and the current training stage information; the action space is defined as the parameter combination of the feedback strategy, including three key dimensions: the weight of the pressing depth compliance rate, the frequency stability coefficient, and the proportion of the posture deviation correction amount; the reward function is constructed based on the operation improvement effect after feedback, the slope of the learning curve, and the subjective acceptance of the training object.
[0143] The policy network is trained by the Advantage Actor-Critic algorithm (A2C) and adopts a hierarchical structure design: the first layer selects the macro policy based on the training stage, and the second layer adjusts the specific policy parameters according to the individual characteristics and current performance. To solve the sample efficiency problem in reinforcement learning, the system adopts a model-enhanced strategy optimization method, combines a pre-trained human biomechanical model, and trains the policy network by combining a small amount of actual interaction data with a large amount of simulated data.
[0144] The strategy optimization introduces an adaptive adjustment mechanism for error correction priorities. Instead of using a fixed-weight evaluation standard, the system dynamically adjusts the priorities of different error correction dimensions according to the cognitive characteristics and operation habits of the training object. For example, for visual learning type trainers, the system will increase the weight of posture guidance; for tactile sensitive type trainers, the proportion of force feedback will be enhanced. This adaptive optimization mechanism significantly improves the pertinence and effectiveness of the feedback.
[0145] The strategy update adopts a gradient update method based on temporal difference, and the update rule of the strategy parameter θ is:
[0146]
[0147] where π(a│s; θ) is the action probability distribution output by the policy network, A(s, a) is the advantage function, representing the degree of advantage of taking action a in state s relative to the average performance, and α is the learning rate. The advantage function is calculated through the temporal difference error:
[0148] A(s, a) = r + γ·V(s ′ ) - V(s)
[0149] where r is the immediate reward, V(s) is the state value function, γ is the discount factor, and s' is the subsequent state.
[0150] The multi-modal feedback generation unit is used to generate three-dimensional trajectory guidance markers, graded tactile pulse sequences, and priority voice prompts based on the pressing efficiency impact index EI and the comprehensive score of pressing quality, where the generation priority of the content of the voice prompt is consistent with the contribution value ranking.
[0151] Specifically, the multi-modal feedback generation unit converts the error correction policy output by the policy optimization unit into specific multi-sensory feedback control signals, including three feedback forms: three-dimensional trajectory guidance markers, graded tactile pulse sequences, and priority voice prompts. This unit adopts a feedback signal synthesizer architecture, including three subsystems: a visual channel processor, a tactile signal generator, and a voice feedback synthesizer.
[0152] The three-dimensional trajectory guidance markers are generated by the visual channel processor and guide the operator to adjust the posture through graphical elements (such as dynamic paths, pose assistance frameworks, and direction indicators) in the virtual reality environment. This system adopts a color-changing dynamic trajectory technology, and the trajectory color and brightness change with the operation compliance, providing continuous and intuitive visual feedback to the operator. The trajectory generation is based on the interpolation path between the ideal posture and the current posture, and a spline-based smoothing algorithm is used to ensure natural and smooth motion transitions.
[0153] The graded tactile pulse sequences are generated by the tactile signal generator and transmitted to the operator through a multi-point tactile feedback device. The system designs a tactile pattern library that encodes different information, including error prompt patterns, correction guidance patterns, and positive reinforcement patterns. The intensity, frequency, and rhythm of the tactile signals are dynamically adjusted according to the degree of deviation, forming a gradient tactile perception. For example, slight deviations are prompted by low-frequency short pulses, while severe deviations are warned by high-frequency continuous vibrations. The selection of the tactile part also has semantic relevance. For example, the tactile feedback for elbow flexion problems is located in the upper arm area.
[0154] The priority voice prompts are generated by a voice feedback synthesizer. The prompt priorities are determined by sorting according to the contribution values of posture defects to EI, ensuring that the problems most affecting the operation effect are corrected first. The content of the voice prompts adopts a three-level structure: the first-level prompt is a short instruction (such as "Straighten the elbow"); the second-level prompt includes the specific adjustment direction and amplitude (such as "Abduct the right elbow by 15 degrees"); the third-level prompt includes method guidance (such as "Lower the shoulders and keep the upper body vertical"). The voice generation system uses parametric speech synthesis technology, which can adjust the speech rate and pitch according to the urgency, enhancing the situational adaptability of the prompts.
[0155] The system establishes a feedback coordination mechanism across sensory channels to avoid information redundancy and channel interference. The coordination mechanism distributes feedback based on three dimensions: information priority, sensory load, and cognitive characteristics: critical posture problems are transmitted through both visual and voice channels; fine adjustments are mainly guided by touch and vision; immediate feedback is mainly through touch; and comprehensive evaluation is mainly through voice. This mechanism effectively prevents information overload and improves the acceptability and effectiveness of feedback.
[0156] Data exchange between the decision-making module and other modules is carried out through a standardized interface. It receives the operation quality evaluation parameter set and the pressing efficiency influence coefficient EI from the multi-source data collaboration module; outputs multi-sensory feedback control signals to the interactive closed-loop execution module; and at the same time receives the operation response data fed back by the interactive closed-loop execution module to form a complete decision-feedback-adjustment cycle.
[0157] The decision-making module adopts a hierarchical computing architecture, which assigns tasks with different time sensitivities to different processing levels to ensure the low-latency performance of critical feedback. Basic calculations such as posture evaluation and biomechanical analysis are executed at a high frequency in the real-time layer; strategy optimization is updated at a medium frequency in the middle layer; and the global optimization of the strategy model is carried out periodically in the background layer.
[0158] Specifically, in an embodiment of the present invention, the interactive closed-loop execution module is used to input the multi-sensory feedback control signal into the virtual first-aid environment and update the operation quality evaluation parameter set according to the operation data after the feedback response. This module adopts a multi-channel parallel processing architecture to achieve the synchronous output and status monitoring of visual, tactile, and auditory feedback, forming a complete feedback-adjustment-evaluation cycle.
[0159] In this embodiment, this module is composed of four functional units: a virtual reality scene rendering engine, a tactile feedback execution unit, an acoustic signal processing unit, and an operation response monitoring unit.
[0160] The virtual reality scene rendering engine receives the three-dimensional trajectory guidance marker data generated by the decision-making module and integrates it into the first-aid environment generated by the virtual scene construction module. Based on modern graphics rendering technology, the engine supports the real-time rendering of high-precision pose guidance elements, such as pose correction trajectory lines, ideal operation position markers, and dynamic guidance arrows. The system adopts a semi-transparent gradient visual design, enabling the guidance elements to provide clear and intuitive pose adjustment guidance for the operator without disturbing the overall visual effect of the scene.
[0161] The rendering of the trajectory guidance marker adopts a priority hierarchical technology, dividing the guidance elements into three levels: the key guidance layer, the auxiliary prompt layer, and the status feedback layer, and determining the display priority according to the EI contribution value and the criticality of the operation. The system intuitively expresses the severity of the pose deviation and the urgency of adjustment by controlling the color gradient, flicker frequency, and transparency change of the guidance elements. At the same time, it uses dynamic size scaling to strengthen visual attention guidance, prompting the operator to correct the most influential pose problems first.
[0162] The tactile feedback execution unit is responsible for receiving and converting the hierarchical tactile pulse sequence control signal, driving the multi-point tactile feedback device to generate precise tactile stimuli. The system uses a distributed tactile actuator array to cover key joint positions such as the operator's wrist, elbow, and shoulder, and can provide accurately positioned tactile feedback for different pose problems. The encoding of the tactile signal adopts a multi-dimensional parameter space, including vibration frequency, amplitude, duration, and rhythm pattern, to construct a rich tactile language.
[0163] The drive control of the tactile feedback adopts a model-based predictive control method. By establishing a tactile actuator response model, it predicts the relationship between the actual output and the target perception effect, realizing more precise control of the tactile experience.
[0164] The acoustic signal processing unit is responsible for converting the priority voice prompt into a clear and distinguishable acoustic signal. The system uses spatial positioning audio technology to make the voice prompt directional, thereby enhancing the spatial perception and intuitiveness of the sound feedback. For example, the prompt for the right elbow pose problem will be emitted from the right side to strengthen the spatial positioning awareness. The acoustic signal processing adopts a dynamic priority queue management mechanism, and reasonably arranges the playback order and interval of the voice prompts according to the operation process stage and the severity of the pose problem, avoiding the cognitive burden caused by overly dense prompts.
[0165] The operation response monitoring unit is responsible for collecting the response data of the operator to the feedback signal, evaluating the feedback effect, and updating the operation quality evaluation parameter set. The unit calculates the change rate of key parameters by setting time windows before and after the feedback, and quantitatively evaluates the influence effect of the feedback signal on the operation behavior. The response monitoring adopts a time series comparison algorithm:
[0166] ΔP = w1·(P t –Pt―1 ) + W2·(P t – P t―2 ) + W3·(P t – P t―3 )
[0167] Among them, P t represents the operation parameter value at the current moment, and P t―n represents the parameter value n time units ago. w1, w2, and w3 are time decay weights, satisfying w1 > w2 > w3 and ∑w = 1. This multi-period comparison method can capture both immediate changes and short-term trends, and comprehensively evaluate the feedback effect.
[0168] Response monitoring also includes feedback acceptance evaluation. By analyzing the response patterns and improvement effects of the operator to different types of feedback, a personalized feedback effect profile is established. The system identifies four typical response patterns: immediate response type (quickly adjusts to feedback), gradual adaptation type (requires multiple feedbacks to complete adjustment), selective acceptance type (sensitive only to specific channel feedback), and low sensitivity type (slow to respond to multiple feedbacks). For different response patterns, the system provides feedback effect evaluation data to the decision-making module to support the strategy optimization unit in adjusting the feedback strategy.
[0169] The data exchange between the interactive closed-loop execution module and other modules is based on an event-driven mechanism to achieve efficient cooperation. It receives three types of feedback control signals from the decision-making module; transmits real-time operation status data to the virtual scene construction module to support dynamic scene adjustment; collects feedback response data and updates the operation quality evaluation parameter set, and sends it back to the multi-source data collaboration module for further analysis. The module adopts a parallel computing architecture internally to ensure the synchronous execution and real-time response capabilities of multi-channel feedback.
[0170] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cardiopulmonary resuscitation training system based on multimodal artificial intelligence combined with virtual reality technology, characterized in that: include: A multimodal perception module is used to collect posture data, mechanical data, and environmental data of the trainees when they perform cardiopulmonary resuscitation operations; The multi-source data collaboration module is used to perform cross-modal feature alignment and fusion of posture data, mechanical data, and environmental data, generate an operation quality assessment parameter set through spatiotemporal correlation modeling, and establish a mapping relationship between operation posture and compression efficiency; A virtual scene construction module is used to generate a virtual emergency environment containing obstacles of variable complexity, simulated patient physiological states and environmental interference factors according to a preset emergency scene template and an operation quality assessment parameter set; A decision module, which is used to identify the operation posture deviation and its quantitative impact on the pressing efficiency based on the mapping relationship, and generate feedback control signals, including three-dimensional space trajectory guidance strategy, graded tactile feedback signals and acoustic prompt instructions; The interactive closed-loop execution module is used to input the feedback control signal into the virtual emergency environment and update the operation quality evaluation parameter set according to the operation data after the feedback response.
2. The system according to claim 1, characterized in that Posture data include the three-dimensional coordinates of 25 key skeletal points of the upper limbs and torso, body center of gravity distribution, elbow joint angle, arm extension, torso verticality, shoulder and scapula position parameters, and posture stability index; mechanical data include compression depth, compression frequency, chest rebound rate, compression force, compression path offset, compression duration ratio, compression consistency index and arm force angle; environmental data include operating space geometric characteristics, obstacle distribution information, ambient light intensity, background noise level, character activity information in the virtual scene, and quantified values of the complexity of interference factors.
3. The system according to claim 2, characterized in that The multimodal perception module includes: Visual acquisition unit, including depth camera and RGB camera, used to collect posture data; Mechanical detection unit, including pressure sensor, acceleration sensor and IMU, used to collect mechanical data; a biometric acquisition unit, including a blood oxygen saturation sensor and a pupil reflex monitoring module, for acquiring biological parameters associated with the physiological state of the simulated patient; Environmental perception unit, used to collect environmental data.
4. The system according to claim 2, characterized in that The multi-source data collaboration module includes: A spatiotemporal alignment unit, used to align multimodal sensor data to a unified spatiotemporal coordinate system through timestamp synchronization and spatial coordinate conversion, and generate a spatiotemporal synchronized data stream; A feature extraction unit is used to extract posture features from posture data, extract pressing dynamics features from mechanical data through wavelet transform, extract interference factor features through environmental perception data analysis, and obtain a multi-dimensional feature vector that characterizes the operation quality; A parameter set generating unit, used for analyzing the multi-dimensional feature vector based on the sliding time window to form an operation quality evaluation parameter set, including a posture feature parameter set and a mechanical performance index set; The mapping relationship modeling unit is used to construct a posture-force association model through a deep neural network, output the influence coefficient of the operation posture on the pressing efficiency, and establish a mapping relationship between the operation posture and the pressing efficiency.
5. The system according to claim 4, characterized in that The calculation expression of the influence coefficient EI of the operation posture on the pressing efficiency is: Where, EI represents the compression efficiency influence coefficient; P k is the kth posture feature parameter; W k is the dynamic weight of the posture feature parameter; n is the total number of posture feature parameters; Ω is the pressing depth in the mechanical performance index set; D(Ω) represents the conformity function between the pressing depth Ω and the target depth; λ is the pressing frequency in the mechanical performance index set; T(λ) represents the modulation factor of the pressing frequency λ; f(P k ) is the posture deviation function, which is calculated as follows: Among them, S k P k The corresponding reference standard value, σ k is the characteristic adaptation coefficient, when |P k ―S k |When it is greater than the preset value, f(P k ) approaches 1, indicating that the posture deviation is serious; when |P k ―S k | approaches 0, f(P k ) is close to 0, indicating that the posture meets the standard.
6. The system according to claim 5, characterized in that The conformity function D(Ω) and the modulation factor T(λ) are both in the form of nonlinear functions, which are used to perform weighted correction on the compression depth and compression frequency, where: In the formula, Ω ref and λ ref are the target values of compression depth and compression frequency respectively, and a and b are the modulation coefficients.
7. The system according to claim 1, characterized in that The virtual scene building modules include: The scene template library uses a hierarchical structured data model to store three-dimensional components, material maps, and physiological parameter benchmark values for four types of scenes: hospitals, homes, public places, and outdoor environments. Each type of scene contains N pluggable preset templates. Dynamic scene configuration engine, based on programmatic content generation algorithm and constraint optimization model, analyzes the spatial trajectory offset in the operation quality evaluation parameter set through sliding window, and generates obstacle layout and environmental interference elements matching the current training stage in real time; A virtual character generator for building virtual patients with physiological status feedback as well as bystanders and assistants in simulated emergency scenes; The scene parameter scheduler is used to adjust the obstacle density, environmental interference intensity and time pressure according to the posture stability and press effectiveness indicators in the operation quality evaluation parameter set to achieve optimal control of the scene difficulty.
8. The system according to claim 7, characterized in that The scene parameter scheduler adjusts the scene difficulty by calculating the scene complexity score SC: SC=α·log(1+OD+η)·β PE ·c 1―SR Where SC represents the scene complexity score; Od represents the obstacle density, that is, the number of obstacles in a unit space; η represents the environmental interference factor; PE represents the operation proficiency score, which is calculated by the operation quality evaluation parameter set; SR represents the historical success rate, that is, the proportion of historical success times of the training subjects in similar scenarios; α, β, and γ are adjustment coefficients, and α>0, 0<β<1, and γ>1.
9. The system according to claim 5, characterized in that The decision-making modules include: A posture evaluation unit is used to compare the skeleton key point data with the standard reference model through a spatiotemporal trajectory matching algorithm, identify at least one type of posture defect among shoulder forward tilt deviation, elbow flexion abnormality and excessive trunk deviation angle, and calculate the contribution value of each defect to the compression efficiency impact index EI based on the posture-force association model; A biomechanical analysis unit, which is used to generate a comprehensive compression quality score based on the compression depth deviation value, frequency fluctuation coefficient and chest rebound lag time in combination with an operation quality assessment parameter set; The strategy optimization unit uses a reinforcement learning framework based on segmented training phases to optimize the personalized error correction strategy by dynamically adjusting the weight of the compression depth compliance rate, the frequency stability coefficient, and the posture deviation correction ratio; The multimodal feedback generation unit is used to generate three-dimensional trajectory guidance marks, graded tactile pulse sequences and priority voice prompts based on the pressing effectiveness impact index EI and the comprehensive score of pressing quality, wherein the content generation priority of the voice prompt is consistent with the contribution value ranking.
10. The system according to claim 1, characterized in that The multimodal perception module is communicatively connected with the multi-source data collaboration module, the multi-source data collaboration module is communicatively connected with the virtual scene construction module and the decision-making module, the decision-making module is communicatively connected with the interactive closed-loop execution module, and the interactive closed-loop execution module is communicatively connected with the virtual scene construction module, forming a closed-loop system of data collection, processing, decision-making, feedback and optimization.
Citation Information
Cited By
VR teaching experience enhancement system and method
CN120523333A
Humanoid robot multi-mode dynamic jump test system, method, equipment and medium
CN120538863A
Training and examination robot system and equipment based on multi-modal large model and computer program product
CN120876176A
A training robot system, device and computer program product based on a multi-modal large model
CN120876176B
Physiological state multi-modal simulation method, device and equipment based on deep learning
CN120954739A