Virtual reality interactive training system and method based on multi-modal feedback

Through the distributed rendering architecture and the neural network of the space-time attention mechanism, combined with the three-level evaluation system and the multimodal feedback mechanism, the shortcomings of the virtual reality interaction training system in multimodal feedback, scene construction flexibility and data processing accuracy are solved, and efficient and flexible virtual training scenarios and natural human-computer interaction are achieved, improving the authenticity and effect of training.

CN119937798AActive Publication Date: 2025-05-06CHENGDU JINJIELI POLICE EQUIP CO LTD

Patent Information

Application Number
CN202510422381.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing virtual reality interactive training system has shortcomings in multimodal feedback mechanism, scene construction flexibility and data processing accuracy, and cannot achieve collaborative feedback from multiple senses, scene switching is inflexible, data accuracy is low, and it is difficult to meet the real-time requirements.

Method used

A distributed rendering architecture is used to build a multi-dimensional training scenario library, and multi-modal feedback and intelligent scene switching are achieved through LSTM-GRU hybrid neural network based on spatiotemporal attention mechanism extraction and weighted fusion of multi-source data. At the same time, a three-level evaluation system is designed to evaluate the user's operational deviation in real time and transmit feedback signals through the multimodal feedback channel.

Benefits of technology

It realizes flexible construction and real-time rendering of multi-dimensional training scenarios, improves data processing accuracy and rendering efficiency, supports natural human-computer interaction, improves the authenticity and effect of training, and realizes automatic teaching function through feedback signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937798A_ABST
    Figure CN119937798A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of interactive training, in particular to a virtual reality interactive training system and method based on multi-modal feedback. Comprising a multi-dimensional training scene construction unit, a multi-source data interaction unit and an operation deviation feedback unit. According to the method, the multi-dimensional training scene library and the two-channel semantic analysis engine are constructed through the distributed rendering architecture to interact and switch scenes, real-time rendering supporting large-scale scenes is achieved, the rendering efficiency and scene complexity are improved, meanwhile, the natural man-machine interaction capability is achieved, the operation deviation of a user is evaluated, and the user experience is improved. The intensity of the feedback signal is in direct proportion to the operation deviation, the deviation during user operation training is detected, the training result is evaluated according to the operation deviation, the user is reminded to conduct adjustment in time through the feedback signal, the automatic teaching function is achieved, meanwhile, the intensity of the deviation can be fed back through the intensity of the feedback signal, and the user experience is improved. And the user is further promoted to adjust.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of interactive training, and in particular to a virtual reality interactive training system and method based on multimodal feedback. Background Art

[0002] Virtual reality uses computer simulation to create a three-dimensional virtual world, providing users with simulations of vision and other senses, allowing users to feel as if they are in the real world and observe things in the three-dimensional space instantly and without restrictions. Virtual reality (VR) technology has developed rapidly in recent years and is widely used in education and training, medical rehabilitation, military simulation and other fields. However, the existing virtual reality interactive training system still has many deficiencies in terms of multimodal feedback mechanism, scene construction flexibility and data processing accuracy, such as: Existing systems usually only provide single-mode feedback (such as tactile feedback or visual feedback), and cannot achieve coordinated feedback of multiple senses such as touch, vision, and hearing. At the same time, virtual scene switching usually relies on preset trigger conditions, lacking flexibility and intelligence. Users cannot quickly switch training scenes according to actual needs, which limits the applicability of the system, resulting in unrealistic user experience and limited training effects. In virtual reality training, the user's behavioral data (such as motion trajectory, posture, operation timing, etc.) needs to be collected and processed in real time. Existing systems often use a single sensor or a simple data processing method, resulting in low data accuracy and low processing efficiency, which makes it difficult to meet real-time requirements. In view of this, we propose a virtual reality interactive training system and method based on multimodal feedback. Summary of the invention

[0003] The purpose of the present invention is to provide a virtual reality interactive training system and method based on multimodal feedback, so as to solve the problems raised in the above background technology, such as how to efficiently collect and process multi-source heterogeneous data, improve data processing accuracy, build flexible and multi-dimensional virtual training scenes, and support intelligent scene switching. At the same time, how to design a multimodal feedback mechanism to achieve comprehensive and accurate user feedback, and establish a three-level deviation evaluation system to evaluate user operation deviations in real time.

[0004] In order to solve the above technical problems, one of the purposes of the present invention is to provide a virtual reality interactive training method based on multimodal feedback, comprising the following steps: The scene is divided into multiple sub-scenes, which are rendered on different computing nodes respectively. The distributed rendering architecture builds a multi-dimensional training scene library. The LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weights the fusion of multi-source data, including posture and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the computer's real-time switching scene, and the scene is interactively switched through the dual-channel semantic parsing engine; Multi-source data is compared with standard data through a three-level evaluation system to evaluate the user's operation deviation, transmit feedback signals to the user through feedback channels and generate training results. The feedback channels include touch, vision and hearing, where the intensity of the feedback signal is proportional to the operation deviation.

[0005] Preferably, the distributed rendering architecture constructs a multi-dimensional training scene library, comprising the following steps: Define the geometric complexity index, traverse each geometric body in the scene, calculate its complexity index, and record the complexity of each area. At the same time, identify dynamically changing objects in the scene and use the space partitioning algorithm to dynamically divide the scene into multiple sub-scenes; Each sub-scene is assigned to a different computing node, and a new global timestamp is generated according to the rendering frame rate. The global timestamp is sent to each computing node. Each node performs rendering operations according to the assigned sub-scene and the current timestamp, and the rendering results of all sub-scenes are integrated into the main scene to generate a complete virtual environment.

[0006] Preferably, the space partitioning algorithm comprises the following steps: The user's perspective, including position and direction, is tracked in real time by a computer. The visual cone is calculated based on the user's perspective to determine the user's current focus area. Taking the focus area as a benchmark, the scene grid units within the threshold layer circle are extended outward, and the space is recursively divided into sub-areas. The distance weight of the sub-area close to the benchmark is used as a rendering priority indicator. Sub-areas with the same distance weight are rendered synchronously through timestamps, and the sub-areas are dynamically updated in real time according to changes in perspective.

[0007] Preferably, the LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weightedly fuses multi-source data, comprising the following steps: Collection of posture and motion data, collection of spatial position and depth information, and collection of muscle activity data; Different data are aligned in time through the hardware clock, and different weights are assigned according to the reliability and importance of the data. The LSTM unit designed to capture long-term dependencies and the GRU unit designed to process short-term dependencies form a hybrid network structure. Multi-source data are input into the LSTM and GRU units respectively, and the features of the multi-source data are extracted. The extracted features are weighted and fused according to the weights to generate a comprehensive feature vector.

[0008] Preferably, the dual channels include a voice channel and an action channel, and the scenes are interactively switched through the dual channel semantic parsing engine, which includes the following steps: The voice channel uses a BERT-LSTM hybrid model for semantic understanding. BERT extracts contextual semantic features, LSTM decodes temporal dependencies, outputs structured instructions, and sets a voice confidence score for validating voice input. If the voice confidence score is lower than the confidence threshold, an error signal is output; if the confidence score is not lower than the confidence threshold, a switching signal is output. The action channel is based on 3D-CNN's skeletal posture semantic analysis, inputs Kinect V4 skeletal data, captures spatial and temporal correlations through 3D convolution kernels, outputs gesture semantic encoding, and sets an action confidence score for judging the effectiveness of gesture recognition. If the action confidence score is lower than the confidence threshold, an error signal is output; if the confidence score is not lower than the confidence threshold, a switching signal is output.

[0009] Preferably, the parsing engine interactively switches scenes, including the following gestures: Posture 1: Receive the collected posture and motion data, spatial position and depth information, and muscle activity data, generate Kinect V4 skeleton data, trigger the action channel to output gesture semantic coding, and execute interactive switching scenes according to the switching signal; Posture 2: Receive real-time voice data, extract contextual semantic features, trigger the voice channel to output structured instructions, and execute interactive switching scenarios according to the switching signal; Posture 3. Perceive the error signal of the interactive switching scene of the action channel and the voice channel. If the number of error signals is higher than the error threshold, establish a voice-action command mapping table, weightedly fuse the confidence scores of the voice and action channels, output the matching degree of the dual-channel command and the semantic mapping matrix, and set the comprehensive confidence. If the comprehensive confidence exceeds the total threshold, the input is considered valid and the switching scene is triggered. Otherwise, a warning signal is issued.

[0010] Preferably, the step of evaluating the user's operation deviation comprises the following steps: Construct standard operation data as training targets, including standard posture, motion trajectory, and spatial position, and align the standard operation data with the user's real-time operation data in the time dimension; For each time step , calculate the comprehensive feature vector With standard operating data The deviation between , the calculation formula is as follows: ; in, is the data dimension, i is the number of training target bits, and according to the importance of the data, a weight is assigned to each data to generate a weighted deviation.

[0011] Preferably, the transmitting feedback signal comprises the following steps: The weighted deviation is mapped to the intensity range of the feedback signal using a normalization function, and the corresponding feedback signals are generated through the feedback channels, including the following postures: Haptic feedback: reflects deviations through vibration intensity and transmits vibration signals through tactile devices; Visual feedback: Deviations are indicated by color changes, and the color changes are displayed in the computer interface; Auditory feedback: Deviations are indicated by rhythmic changes, and the sound of rhythmic changes is played through a speaker; Training results are generated based on feedback signals based on a three-level evaluation system, including a primary evaluation that calculates deviations in real time and provides feedback, an intermediate evaluation that periodically evaluates training progress and effects, and an advanced evaluation that comprehensively analyzes long-term training data and provides personalized recommendations.

[0012] The second object of the present invention is to provide a virtual reality interactive training system based on multimodal feedback, including any one of the virtual reality interactive training methods based on multimodal feedback described above, including a multi-dimensional training scene construction unit, a multi-source data interaction unit and an operation deviation feedback unit; The multi-dimensional training scene construction unit is used to divide the scene into multiple sub-scenes, which are rendered on different computing nodes respectively, and the distributed rendering architecture constructs a multi-dimensional training scene library; The multi-source data interaction unit is used to extract features and weightedly fuse multi-source data based on the LSTM-GRU hybrid neural network of the spatiotemporal attention mechanism, the multi-source data including posture and motion data, spatial position and depth information, and muscle activity data, interactively input the multi-source data into the scene switched by the computer in real time, and interactively switch the scene through the dual-channel semantic parsing engine; The operation deviation feedback unit is used to compare multi-source data with standard data through a three-level evaluation system, evaluate the user's operation deviation, transmit feedback signals to the user through a feedback channel and generate training results, and make the strength of the feedback signal proportional to the operation deviation.

[0013] Compared with the prior art, the present invention has the following beneficial effects: Through the distributed rendering architecture, a multi-dimensional training scene library and a dual-channel semantic parsing engine are built to interactively switch scenes, realize real-time rendering that supports large-scale scenes, improve rendering efficiency and scene complexity, and at the same time achieve natural human-computer interaction capabilities and improve authenticity. During interactive training, the user's operation deviation is evaluated, and the feedback signal is transmitted to the user through the feedback channel to generate training results, so that the intensity of the feedback signal is proportional to the operation deviation. It not only detects the deviation of the user's operation training, evaluates the training results according to the operation deviation, but also reminds the user to make timely adjustments through the feedback signal, realizing the function of automatic teaching. At the same time, the intensity of the deviation can be fed back through the intensity of the feedback signal, further prompting the user to make adjustments. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is an overall flow chart of the embodiment; Figure 2 A schematic diagram of space division in an embodiment; Figure 3 It is a schematic diagram of the scene switching process of the analysis engine interaction of the embodiment. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0016] Examples, such as Figure 1-3 As shown, one of the purposes of the present invention is to provide a virtual reality interactive training method based on multimodal feedback, comprising the following steps: S1. Divide the scene into multiple sub-scenes and render them on different computing nodes. The distributed rendering architecture builds a multi-dimensional training scene library to support real-time rendering of large-scale scenes, improve rendering efficiency and scene complexity; Specifically, the distributed rendering architecture builds a multi-dimensional training scene library, including the following steps: Define geometric complexity indicators, such as the number of geometric bodies, the number of faces, the number of vertices, etc., traverse each geometric body in the scene, calculate its complexity indicators, and record the complexity of each area. At the same time, identify dynamically changing objects in the scene, such as moving objects and characters, and mark the area containing dynamic objects as dynamic areas for subsequent processing. Use the space partitioning algorithm to dynamically divide the scene into multiple sub-scenes; Each sub-scene is assigned to a different computing node, and a new global timestamp is generated according to the rendering frame rate. The global timestamp is sent to each computing node. Each node performs rendering operations according to the assigned sub-scene and the current timestamp, and the rendering results of all sub-scenes are integrated into the main scene to generate a complete virtual environment, improve rendering efficiency, support large-scale complex scenes, and enhance system scalability. Performance can be improved by adding nodes. In addition, using timestamp synchronization is conducive to ensuring the consistency of time of all nodes, improving user experience, and enhancing the immersiveness of virtual reality. As for the rendering method, the commonly used existing technology can be adopted, such as the Distributed Rendering Engine (DRE) and other frameworks to coordinate the rendering tasks of each node. The master node is responsible for task allocation, and each slave node performs rendering tasks.

[0017] In order to reduce the operation intensity and memory usage, the space partitioning algorithm includes the following steps: The user's perspective is tracked in real time by a computer, including position and direction. The visual cone is calculated according to the user's perspective to determine the user's current focus area. The scene grid units within the threshold layer circle are extended outward based on the focus area, and the space is recursively divided into sub-areas. The distance weight of the sub-area close to the benchmark is used as the rendering priority indicator. Sub-areas with the same distance weight are rendered synchronously through timestamps, and the sub-areas are dynamically updated in real time according to the change of perspective, such as Figure 2 As shown, after determining the user's focus area, multiple layers of grid units are extended outward, including the layer circle where sub-area a is located and the layer circle where sub-area b is located. Each layer circle has multiple sub-areas, which belong to the same distance weight. Among them, within the threshold layer circle range, the user's possible viewing angle range that can be adjusted within the guaranteed rendering time can be analyzed according to user habits, and the weight a corresponding to sub-area a>the weight b corresponding to sub-area b. When the user belongs to the focus area, multiple sub-areas a are rendered first, and the rendering results of each sub-scene are combined according to the timestamp to ensure overall synchronization. Then, sub-area b is rendered and dynamically adjusted in real time according to the user's perspective. This not only reduces the operating intensity caused by the simultaneous rendering of multiple scenes and saves resources, but also, according to the synchronous adjustment of the outward diffusion of the perspective, the sub-areas can be dynamically updated in real time as the perspective is updated, so that scenes that are not within the threshold layer circle range can be deleted first, avoiding more space occupation and improving the operating speed.

[0018] S2, LSTM-GRU hybrid neural network based on spatiotemporal attention mechanism extracts features and weights and fuses multi-source data, including posture and motion data, spatial position and depth information, and muscle activity data. Multi-source data is interactively input into the scene switched by the computer in real time, and the scene is interactively switched through the dual-channel semantic parsing engine to avoid a single interactive channel and lack of natural human-computer interaction capabilities, thereby improving authenticity; In the virtual reality interactive training system, the user's behavior data (such as posture, motion trajectory, spatial position, depth information, muscle activity, etc.) needs to be collected and processed in real time. However, the existing system usually uses a single sensor or a simple data processing method, resulting in low data accuracy and low processing efficiency, which is difficult to meet the real-time requirements. To address this problem, the features of multi-source data are extracted and weighted fused to improve the accuracy and efficiency of data processing. The LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weighted fuses multi-source data, including the following steps: Collection of posture and motion data: Install sensors (including inertial sensors, optical sensors, and electromagnetic sensors) on various joints of the user, such as wrists, elbows, knees, etc., to collect the user's motion data in real time, including acceleration, angular velocity, position, and posture. The rotation angle of the posture is represented by Euler angles, quaternions represent three-dimensional rotations to avoid the gimbal lock problem, and transformation matrices represent the transformation of the coordinate system. Collection of spatial position and depth information: The depth camera outputs the depth value of each pixel, forming a depth map that is converted into a three-dimensional point cloud to represent the user's position in space. Computer vision algorithms (such as OpenPose and HRNet) are used to estimate the user's posture from the depth map. It is also possible to consider separating the user from the background to improve the accuracy of posture estimation. The user's three-dimensional coordinates are mapped to the virtual scene to achieve position synchronization. Muscle activity data collection: electrodes are attached to the user's muscle surface, such as forearms and thighs, to collect muscle electrical signals, including the amplitude and frequency of the electromyographic signals. A bio-amplifier is used to amplify weak electromyographic signals, remove noise, such as power frequency interference and motion artifacts, and extract the time domain and frequency domain features of the electromyographic signals, such as the root mean square value (RMS) and average power frequency (APF). Different data are aligned in time through the hardware clock, and data from different sensors are aligned to the same timeline according to timestamps. For data with inconsistent timestamps, interpolation methods (such as linear interpolation) can be used to complete them to ensure the consistency of data in time and provide a unified time reference. Considering that some sensors may have data transmission delays or inconsistent sampling rates, synchronization algorithms (such as timestamp-based alignment algorithms) are used to adjust the data to eliminate the impact of delays and jitters. Different weights are given according to the reliability and importance of the data. For example, some sensors may be more reliable in specific environments, or some data sources may be more important in training. It can be pre-set or automatically determined by machine learning models). The LSTM unit designed to capture long-term dependencies and the GRU unit designed to process short-term dependencies form a hybrid network structure. Multi-source data are input into the LSTM and GRU units respectively to extract features of the multi-source data. For example, posture data may extract motion patterns, and depth data may extract spatial relationships. The extracted features are weighted and fused according to weights to generate a comprehensive feature vector, ensuring that data features with high importance occupy a larger weight. In summary, the comprehensiveness of feature extraction is improved. The entire system design supports real-time data processing and feedback, and is suitable for scenarios with high real-time requirements such as virtual reality training.

[0019] Furthermore, the dual channel includes a voice channel and an action channel, and the scenes are interactively switched through the dual channel semantic parsing engine, which includes the following steps: The voice channel uses the BERT-LSTM hybrid model for semantic understanding. BERT extracts contextual semantic features, LSTM decodes temporal dependencies, outputs structured instructions, and sets a voice confidence score for valid voice input. If the voice confidence score is lower than the confidence threshold, an error signal is output. If the confidence score is not lower than the confidence threshold, a switching signal is output. In the output layer of the model, the probability of each category is normalized using the Softmax function to obtain a probability distribution. The category with the highest probability is the prediction result of the model, and its corresponding probability value is the confidence score. For example, assuming that the probability distribution output by the model is [0.1, 0.3, 0.6], the confidence is 0.6, indicating that the model has a confidence of 60% for the third category. According to experimental and test data, a confidence threshold (such as 0.7 or 0.8) can be set. When the confidence of the model is lower than the threshold, the system considers that the instruction is incorrect and cannot switch scenes. When the error signal is output, the error correction mechanism (such as fuzzy matching) is enabled to try to correct the instruction. The action channel is based on 3D-CNN's skeletal posture semantic analysis, inputs Kinect V4 skeletal data, captures spatial and temporal correlations through 3D convolution kernels, and outputs gesture semantic encoding (such as "right hand swipe → scene switch"). The dynamic gesture library contains 8 types of scene control gestures and sets an action confidence score for judging the effectiveness of gesture recognition. If the action confidence score is lower than the confidence threshold, an error signal is output. If the confidence score is not lower than the confidence threshold, a switching signal is output.

[0020] It is worth noting that the parsing engine interactively switches scenes, including the following gestures: Posture 1: Receive the collected posture and motion data, spatial position and depth information, muscle activity data, generate Kinect V4 skeleton data, trigger the action channel to output gesture semantic coding, execute interactive switching scenes according to the switching signal, and realize that the action can switch scenes, which is convenient for real-time switching of behaviors according to user behaviors, making user training more realistic; Posture 2: Receive real-time voice data, extract contextual semantic features, trigger the voice channel to output structured instructions, execute interactive switching scenes according to the switching signal, realize voice switching scenes, and improve the convenience of switching; Posture 3: Sense the error signal of the interactive switching scene of the action channel and the voice channel. If the number of error signals is higher than the error threshold, establish a voice-action command mapping table (for example: voice "zoom in" ↔ gesture "both hands out"), perform weighted fusion on the confidence scores of the voice and action channels, output the matching degree of the dual-channel command and the semantic mapping matrix, and set the comprehensive confidence. If the comprehensive confidence exceeds the total threshold (such as 0.8), the input is considered valid and the scene switching is triggered. Otherwise, a warning signal is issued; In summary, not only can dual-channel control of switching scenes be achieved to improve the accuracy of scene switching, but when single-channel switching is inaccurate, timely verification through two channels at the same time can improve the accuracy of command recognition and reduce misjudgment.

[0021] S3. Compare multi-source data with standard data through a three-level evaluation system to evaluate the user's operation deviation, transmit feedback signals to the user through feedback channels and generate training results. The feedback channels include touch, vision and hearing. The strength of the feedback signal is proportional to the operation deviation. It not only detects the deviation of the user's operation training, evaluates the training results according to the operation deviation, but also reminds the user to make adjustments in time through feedback signals, realizing the function of automatic teaching. At the same time, the strength of the feedback signal can feedback the strength of the deviation, further prompting the user to make adjustments. Therefore, evaluating the user's operational deviation includes the following steps: Construct standard operation data as training targets, including standard posture, motion trajectory, and spatial position, and align the standard operation data with the user's real-time operation data in the time dimension; For each time step , calculate the comprehensive feature vector With standard operating data The deviation between , the calculation formula is as follows: ; in, is the data dimension (such as posture, position, muscle activity, etc.), i is the number of training target bits, and a weight is assigned to each data according to its importance to generate a weighted deviation. By assigning different weights to different data sources, the weighted deviation can more accurately reflect the criticality of user operations. For example, in a training scenario, posture data may be more important than spatial position data, so a higher weight is assigned to ensure that the evaluation results are closer to actual needs; Specifically, it also includes the use of cosine similarity algorithm to calculate the deviation between the user's actual joint angle and the standard angle for key joint angle deviation, the use of Hausdorff distance to calculate the deviation between the user's actual motion trajectory and the standard trajectory for judging the consistency of the motion trajectory, and the use of dynamic time warping (DTW) algorithm to calculate the deviation between the user's operation timing and the standard timing for judging the accuracy of the operation timing. The user's operation deviation is evaluated in real time, accurate feedback guidance is provided, and training efficiency is improved.

[0022] Based on the above, transmitting the feedback signal includes the following steps: The weighted deviation is mapped to the intensity range of the feedback signal using a normalization function, and the corresponding feedback signals are generated through the feedback channels, including the following postures: Tactile feedback: The deviation is reflected by the vibration intensity, and the vibration signal is transmitted through a tactile device (such as a vibration feedback glove), and the normalized deviation is mapped to the vibration intensity. For example, using a linear mapping, vibration intensity = maximum vibration intensity * deviation; Visual feedback: Display deviation through color change and display color change in computer interface, map deviation to color change, such as from green (small deviation) to red (large deviation), such as: color = interpolation (deviation, green, red); Auditory feedback: The deviation is prompted by rhythm changes, and the sound of rhythm changes is played through speakers or headphones to map the deviation to the rhythm changes. For example, a faster rhythm indicates an increase in deviation, such as: rhythm speed = basic speed + (maximum speed increment * deviation); Training results are generated based on the feedback signal based on a three-level evaluation system, including a primary evaluation that calculates deviations in real time and provides feedback, an intermediate evaluation that periodically evaluates training progress and effects, and an advanced evaluation that comprehensively analyzes long-term training data to provide personalized suggestions. Among them, the primary evaluation: real-time deviation calculation and feedback, to ensure that users can promptly discover deviations and make adjustments during operation, avoid error accumulation, and improve the real-time and effectiveness of training; the intermediate evaluation regularly analyzes the user's training data to evaluate the training progress and effect. For example, the user's training data is summarized and analyzed every week or month to determine whether the user has achieved the predetermined training goal. By analyzing the user's long-term deviation data, the user's operation improvement is evaluated. If the user's deviation persists The intermediate evaluation will trigger further feedback mechanisms, such as increasing the intensity of training or adjusting the training content. The advanced evaluation conducts a comprehensive analysis of the user's long-term training data to identify the user's operation mode, strengths and weaknesses. Based on the analysis results, the advanced evaluation provides users with personalized training suggestions. For example, if the user has a long-term deviation in a certain operation, the advanced evaluation will recommend that the user conduct targeted training or adjust the operation method. The three-level evaluation system covers all stages of the training process from real-time operation to long-term data analysis, ensuring that each stage has a corresponding evaluation and feedback mechanism. Through hierarchical evaluation, the system can use resources more efficiently, adjust training strategies in a timely manner, avoid resource waste, and improve user participation and training effects. In addition, the training effect is evaluated based on the user's operation deviation. The smaller the deviation, the better the training result. The feedback signal can be adjusted according to the importance of the data source, so that users can focus on the main issues more easily. For example, the vibration intensity of the tactile feedback can be proportional to the weighted value of the posture deviation, helping users to quickly adjust their movements and facilitating the automatic provision of teaching guidance. At the same time, the deviation of the secondary data source is prevented from causing excessive interference to the feedback signal, allowing users to focus on the main issues. For example, if the spatial position deviation is small, its impact on the feedback signal will also be reduced, helping users to adjust their operations more efficiently.

[0023] The second object of the present invention is to provide a virtual reality interactive training system based on multimodal feedback, including any one of the virtual reality interactive training methods based on multimodal feedback described above, including a multi-dimensional training scene construction unit, a multi-source data interaction unit and an operation deviation feedback unit; The multi-dimensional training scene construction unit is used to divide the scene into multiple sub-scenes, which are rendered on different computing nodes respectively. The distributed rendering architecture builds a multi-dimensional training scene library. The multi-source data interaction unit is used to extract features and weightedly fuse multi-source data based on the LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism. The multi-source data includes posture and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the scene switched by the computer in real time, and the scene is interactively switched through the dual-channel semantic parsing engine. The operation deviation feedback unit is used to compare multi-source data with standard data through a three-level evaluation system, evaluate the user's operation deviation, transmit feedback signals to the user through the feedback channel and generate training results, and make the strength of the feedback signal proportional to the operation deviation.

[0024] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only preferred examples of the present invention and are not intended to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A virtual reality interactive training method based on multimodal feedback, characterized in that: The steps include: The scene is divided into multiple sub-scenes, which are rendered on different computing nodes respectively. The distributed rendering architecture builds a multi-dimensional training scene library. The LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weights the fusion of multi-source data, including posture and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the computer's real-time switching scene, and the scene is interactively switched through the dual-channel semantic parsing engine; Multi-source data is compared with standard data through a three-level evaluation system to evaluate the user's operation deviation, transmit feedback signals to the user through feedback channels and generate training results. The feedback channels include touch, vision and hearing, where the intensity of the feedback signal is proportional to the operation deviation.

2. The virtual reality interactive training method based on multimodal feedback according to claim 1, characterized in that: The distributed rendering architecture constructs a multi-dimensional training scene library, including the following steps: Define the geometric complexity index, traverse each geometric body in the scene, calculate its complexity index, and record the complexity of each area. At the same time, identify dynamically changing objects in the scene and use the space partitioning algorithm to dynamically divide the scene into multiple sub-scenes; Each sub-scene is assigned to a different computing node, and a new global timestamp is generated according to the rendering frame rate. The global timestamp is sent to each computing node. Each node performs rendering operations according to the assigned sub-scene and the current timestamp, and the rendering results of all sub-scenes are integrated into the main scene to generate a complete virtual environment.

3. The virtual reality interactive training method based on multimodal feedback according to claim 2, characterized in that: The space partitioning algorithm comprises the following steps: The user's perspective, including position and direction, is tracked in real time by a computer. The visual cone is calculated based on the user's perspective to determine the user's current focus area. Taking the focus area as a benchmark, the scene grid units within the threshold layer circle are extended outward, and the space is recursively divided into sub-areas. The distance weight of the sub-area close to the benchmark is used as a rendering priority indicator. Sub-areas with the same distance weight are rendered synchronously through timestamps, and the sub-areas are dynamically updated in real time according to changes in perspective.

4. The virtual reality interactive training method based on multimodal feedback according to claim 1, characterized in that: The LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weightedly fuses multi-source data, including the following steps: Collection of posture and motion data, collection of spatial position and depth information, and collection of muscle activity data; Different data are aligned in time through the hardware clock, and different weights are assigned according to the reliability and importance of the data. The LSTM units designed to capture long-term dependencies and the GRU units designed to process short-term dependencies are composed of a hybrid network structure. Multi-source data are input into the LSTM and GRU units respectively, the features of the multi-source data are extracted, and weighted fusion is performed according to the weights to generate a comprehensive feature vector.

5. The virtual reality interactive training method based on multimodal feedback according to claim 4, characterized in that: The dual channels include a voice channel and an action channel, and scenes are interactively switched through a dual channel semantic parsing engine, which includes the following steps: The voice channel uses a BERT-LSTM hybrid model for semantic understanding. BERT extracts contextual semantic features, LSTM decodes temporal dependencies, outputs structured instructions, and sets a voice confidence score for validating voice input. If the voice confidence score is lower than the confidence threshold, an error signal is output; if the confidence score is not lower than the confidence threshold, a switching signal is output. The action channel is based on 3D-CNN's skeletal posture semantic analysis, inputs Kinect V4 skeletal data, captures spatial and temporal correlations through 3D convolution kernels, outputs gesture semantic encoding, and sets an action confidence score for judging the effectiveness of gesture recognition. If the action confidence score is lower than the confidence threshold, an error signal is output; if the confidence score is not lower than the confidence threshold, a switching signal is output.

6. The virtual reality interactive training method based on multimodal feedback according to claim 5, characterized in that: The parsing engine interactively switches scenes, including the following gestures: Posture 1: Receive the collected posture and motion data, spatial position and depth information, and muscle activity data, generate Kinect V4 skeleton data, trigger the action channel to output gesture semantic coding, and execute interactive switching scenes according to the switching signal; Posture 2: Receive real-time voice data, extract contextual semantic features, trigger the voice channel to output structured instructions, and execute interactive switching scenarios according to the switching signal; Posture 3. Perceive the error signal of the interactive switching scene of the action channel and the voice channel. If the number of error signals is higher than the error threshold, establish a voice-action command mapping table, weightedly fuse the confidence scores of the voice and action channels, output the matching degree of the dual-channel command and the semantic mapping matrix, and set the comprehensive confidence. If the comprehensive confidence exceeds the total threshold, the input is considered valid and the switching scene is triggered. Otherwise, a warning signal is issued.

7. The virtual reality interactive training method based on multimodal feedback according to claim 4, characterized in that: The step of evaluating the user's operation deviation comprises the following steps: Construct standard operation data as training targets, including standard posture, motion trajectory, and spatial position, and align the standard operation data with the user's real-time operation data in the time dimension; For each time step , calculate the comprehensive feature vector With standard operating data The deviation between , the calculation formula is as follows: ; in, is the data dimension, i is the number of training target bits, and according to the importance of the data, a weight is assigned to each data to generate a weighted deviation.

8. The virtual reality interactive training method based on multimodal feedback according to claim 7, characterized in that: The transmitting feedback signal comprises the following steps: The weighted deviation is mapped to the intensity range of the feedback signal using a normalization function, and the corresponding feedback signals are generated through the feedback channels, including the following postures: Haptic feedback: reflects deviations through vibration intensity and transmits vibration signals through tactile devices; Visual feedback: Deviations are indicated by color changes, and the color changes are displayed in the computer interface; Auditory feedback: Deviations are indicated by rhythmic changes, and the sound of rhythmic changes is played through a speaker; Training results are generated based on feedback signals based on a three-level evaluation system, including a primary evaluation that calculates deviations in real time and provides feedback, an intermediate evaluation that periodically evaluates training progress and effects, and an advanced evaluation that comprehensively analyzes long-term training data and provides personalized recommendations.

9. A virtual reality interactive training system based on multimodal feedback, used to implement the virtual reality interactive training method based on multimodal feedback according to any one of claims 1 to 8, characterized in that: It includes a multi-dimensional training scenario construction unit, a multi-source data interaction unit, and an operation deviation feedback unit; The multi-dimensional training scene construction unit is used to divide the scene into multiple sub-scenes, which are rendered on different computing nodes respectively, and the distributed rendering architecture constructs a multi-dimensional training scene library; The multi-source data interaction unit is used to extract features and weightedly fuse multi-source data based on the LSTM-GRU hybrid neural network of the spatiotemporal attention mechanism, the multi-source data including posture and motion data, spatial position and depth information, and muscle activity data, interactively input the multi-source data into the scene switched by the computer in real time, and interactively switch the scene through the dual-channel semantic parsing engine; The operation deviation feedback unit is used to compare multi-source data with standard data through a three-level evaluation system, evaluate the user's operation deviation, transmit feedback signals to the user through a feedback channel and generate training results, and make the strength of the feedback signal proportional to the operation deviation.

Citation Information

Patent Citations

  • Virtual reality scene interaction method and system of SaaS platform

    CN118295538A

  • AR (Augmented Reality) guide system for enhancing stone forest tourism experience

    CN118314306A

  • System for physical-virtual environment fusion

    US20200394409A1

  • System and method for implementing a vocal user interface by combining a speech to text system and a speech to intent system

    WO2017091883A1

  • Generating digital avatar

    WO2021125843A1

Cited By

  • Virtual reality interaction method and system applied to classical famous picture display

    CN120428866A

  • VR teaching experience enhancement system and method

    CN120523333A

  • Scene design method and system based on virtual reality

    CN120560505A

  • Immersive virtual teaching system based on multi-modal perception and dynamic evaluation

    CN120599153A

  • Application program management system and method driven by virtual engine

    CN120635279A