Virtual Reality Interactive Training System and Method Based on Multimodal Feedback
Through the distributed rendering architecture and the neural network of the space-time attention mechanism, combined with the three-level evaluation system, the shortcomings of the existing virtual reality interaction training system in multimodal feedback, scene switching and data processing accuracy are solved, and efficient, flexible and real virtual reality interaction training is achieved.
Patent Information
- Application Number
- CN202510422381.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing virtual reality interactive training system has shortcomings in multimodal feedback mechanism, scene construction flexibility and data processing accuracy, and cannot achieve collaborative feedback from multiple senses, scene switching is inflexible, data accuracy is low, and it is difficult to meet the requirements of real-time.
A distributed rendering architecture is used to build a multi-dimensional training scenario library, and multi-source data is extracted and weighted fusion through an LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism to realize multi-modal feedback and intelligent scene switching, and the user's operation deviation is evaluated through a three-level evaluation system to generate feedback signals.
It realizes comprehensive and accurate multimodal feedback, improves data processing accuracy and rendering efficiency, supports real-time rendering of large-scale complex scenes, improves the flexibility and authenticity of training scenes, and enhances user experience and training effects.
Smart Images

Figure CN119937798B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of interactive training, and more specifically, to a virtual reality interactive training system and method based on multi-modal feedback. Background Art
[0002] Virtual reality uses a computer to simulate and generate a virtual world in a three-dimensional space, providing users with simulations of senses such as vision, enabling users to feel as if they are on the scene and allowing them to observe things in the three-dimensional space immediately and without limitation. Virtual reality (VR) technology has developed rapidly in recent years and is widely used in fields such as education and training, medical rehabilitation, and military simulation. However, existing virtual reality interactive training systems still have many deficiencies in multi-modal feedback mechanisms, scene construction flexibility, and data processing accuracy. For example:
[0003] Existing systems usually only provide single-modal feedback (such as tactile feedback or visual feedback), and cannot achieve collaborative feedback of multiple senses such as touch, vision, and hearing. At the same time, virtual scene switching usually depends on preset trigger conditions, lacking flexibility and intelligence. Users cannot quickly switch training scenes according to actual needs, restricting the applicability of the system, resulting in an untrue user experience and limited training effects;
[0004] In virtual reality training, users' behavioral data (such as movement trajectories, postures, operation timings, etc.) need to be collected and processed in real time. Existing systems often use single sensors or simple data processing methods, resulting in low data accuracy and low processing efficiency, and it is difficult to meet real-time requirements. In view of this, we propose a virtual reality interactive training system and method based on multi-modal feedback. Summary of the Invention
[0005] The purpose of the present invention is to provide a virtual reality interactive training system and method based on multi-modal feedback, so as to solve the problems proposed in the above background art, such as how to efficiently collect and process multi-source heterogeneous data, improve data processing accuracy, construct flexible and multi-dimensional virtual training scenes, and support intelligent scene switching. At the same time, how to design a multi-modal feedback mechanism to achieve comprehensive and accurate user feedback, and establish a three-level deviation evaluation system to evaluate users' operation deviations in real time.
[0006] To solve the above technical problems, one of the purposes of the present invention is to provide a virtual reality interactive training method based on multi-modal feedback, including the following steps:
[0007] Divide the scene into multiple sub-scenes and render them on different computing nodes respectively, and a distributed rendering architecture constructs a multi-dimensional training scene library;
[0008] The LSTM-GRU hybrid neural network based on spatio-temporal attention mechanism extracts features and weighted fuses multi-source data, including pose and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the scenes switched by the computer in real time, and the scenes are interactively switched through a dual-channel semantic parsing engine;
[0009] The multi-source data is compared with the standard data through a three-level evaluation system to evaluate the user's operation deviation. A feedback signal is transmitted to the user through the feedback channel and a training result is generated. The feedback channel includes touch, vision, and hearing. Among them, the intensity of the feedback signal is proportional to the operation deviation.
[0010] Preferably, the distributed rendering architecture constructs a multi-dimensional training scene library, including the following steps:
[0011] Define the index of geometric complexity, traverse each geometric body in the scene, calculate its complexity index, and record the complexity of each area. At the same time, identify the dynamically changing objects in the scene, and use the space partitioning algorithm to dynamically divide the scene into multiple sub-scenes;
[0012] Each sub-scene is assigned to a different computing node. A new global timestamp is generated according to the rendering frame rate and sent to each computing node. Each node performs a rendering operation according to the assigned sub-scene and the current timestamp, and integrates the rendering results of all sub-scenes into the main scene to generate a complete virtual environment.
[0013] Preferably, the space partitioning algorithm includes the following steps:
[0014] The computer tracks the user's perspective in real time, including position and direction. Calculate the viewing frustum according to the user's perspective to determine the area currently focused by the user. Based on the focused area, extend the scene grid cells within the threshold layer circle outward, recursively divide the space into sub-regions, and use the distance weight of the sub-regions close to the reference as the rendering priority index. The sub-regions with the same distance weight are rendered synchronously through timestamps, and the sub-regions are dynamically updated in real time according to the perspective change.
[0015] Preferably, the LSTM-GRU hybrid neural network based on spatio-temporal attention mechanism extracts features and weighted fuses multi-source data, including the following steps:
[0016] Collection of pose and motion data, collection of spatial position and depth information, collection of muscle activity data;
[0017] Align different data in time through a hardware clock, assign different weights according to the reliability and importance of the data, design a hybrid network structure composed of LSTM units for capturing long-term dependencies and GRU units for processing short-term dependencies, input multi-source data into the LSTM and GRU units respectively, extract the features of the multi-source data, and perform weighted fusion on the extracted features according to the weights to generate a comprehensive feature vector.
[0018] Preferably, the dual channels include a voice channel and an action channel, and the scene is interactively switched through a dual-channel semantic parsing engine, which includes the following steps:
[0019] The voice channel uses a BERT-LSTM hybrid model for semantic understanding. BERT extracts context semantic features, LSTM decodes temporal dependencies, outputs structured instructions, and sets a voice confidence score for determining the validity of voice input. If the voice confidence score is lower than the confidence threshold, an error signal is output. If the confidence score is not lower than the confidence threshold, a switching signal is output;
[0020] The action channel is based on 3D-CNN bone pose semantic analysis. Input Kinect V4 bone data, capture spatial and temporal correlations through 3D convolutional kernels, output gesture semantic encodings, and set an action confidence score for judging the validity of gesture recognition. If the action confidence score is lower than the confidence threshold, an error signal is output. If the confidence score is not lower than the confidence threshold, a switching signal is output.
[0021] Preferably, the parsing engine interactively switches the scene, including the following postures:
[0022] Posture 1: Receive the collected posture and motion data, spatial position and depth information, muscle activity data, generate Kinect V4 bone data, trigger the action channel to output gesture semantic encodings, and execute the interactive switching scene according to the switching signal;
[0023] Posture 2: Receive real-time voice data, extract context semantic features, trigger the voice channel to output structured instructions, and execute the interactive switching scene according to the switching signal;
[0024] Posture 3: Sense the error signals of the interactive switching scene of the action channel and the voice channel. If the number of error signals is higher than the error threshold, establish a voice-action instruction mapping table, perform weighted fusion on the confidence scores of the voice and action channels, output the matching degree of the dual-channel instruction and the semantic mapping matrix, and set a comprehensive confidence. If the comprehensive confidence exceeds the total threshold, the input is considered valid and the switching scene is triggered. Otherwise, a warning signal is issued.
[0025] Preferably, the evaluation of the user's operation deviation includes the following steps:
[0026] Construct standard operation data as the training target, including standard postures, motion trajectories, and spatial positions, and align the standard operation data with the user's real-time operation data in the time dimension;
[0027] For each time step , calculate the comprehensive feature vector and the deviation from the standard operation data , and the calculation formula is as follows:
[0028] ;
[0029] wherein, is the data dimension, i is the number of digits of the training target, and weights are assigned to each type of data according to the importance of the data to generate a weighted deviation.
[0030] Preferably, the transmission feedback signal includes the following steps:
[0031] Use a normalization function to map the weighted deviation to the intensity range of the feedback signal, and generate corresponding feedback signals through the feedback channel respectively, including the following postures:
[0032] Tactile feedback: Reflect the deviation through the vibration intensity and transmit the vibration signal through the tactile device;
[0033] Visual feedback: Display the deviation through color changes and display the color changes in the computer interface;
[0034] Auditory feedback: Hint the deviation through rhythm changes and play the sound of rhythm changes through the speaker;
[0035] Generate training results based on the feedback signal according to a three-level evaluation system, including a primary evaluation that calculates the deviation in real time and provides feedback, a secondary evaluation that periodically evaluates the training progress and effect, and a high-level evaluation that comprehensively analyzes long-term training data and provides personalized suggestions.
[0036] The second object of the present invention is to provide a virtual reality interaction training system based on multi-modal feedback, including the virtual reality interaction training method based on multi-modal feedback described in any one of the above, including a multi-dimensional training scenario construction unit, a multi-source data interaction unit, and an operation deviation feedback unit;
[0037] The multi-dimensional training scenario construction unit is used to divide the scenario into multiple sub-scenarios, render them on different computing nodes respectively, and construct a multi-dimensional training scenario library with a distributed rendering architecture;
[0038] The multi-source data interaction unit is used to extract features and weighted-fuse multi-source data based on the LSTM-GRU hybrid neural network with spatio-temporal attention mechanism. The multi-source data includes pose and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the scenes switched by the computer in real time, and the scenes are interactively switched through the dual-channel semantic parsing engine;
[0039] The operation deviation feedback unit is used to compare the multi-source data with the standard data through a three-level evaluation system, evaluate the operation deviation of the user, transmit the feedback signal to the user through the feedback channel and generate a training result, and make the intensity of the feedback signal proportional to the operation deviation.
[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0041] By constructing a multi-dimensional training scene library through a distributed rendering architecture and interactively switching scenes through a dual-channel semantic parsing engine, it realizes real-time rendering supporting large-scale scenes, improves the rendering efficiency and scene complexity, and at the same time realizes natural human-computer interaction capabilities and improves authenticity. When performing interactive training, it evaluates the operation deviation of the user, transmits the feedback signal to the user through the feedback channel and generates a training result, and makes the intensity of the feedback signal proportional to the operation deviation. It not only realizes the detection of the deviation during the user's operation training, evaluates the training result according to the operation deviation, but also reminds the user to make timely adjustments through the feedback signal, realizing the function of automatic teaching. At the same time, the intensity of the feedback signal can reflect the intensity of the deviation, further prompting the user to make adjustments. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the overall flow block diagram of the embodiment;
[0043] Figure 2 It is the schematic diagram of space division of the embodiment;
[0044] Figure 3 It is the schematic diagram of the process of the parsing engine interactively switching scenes of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] Embodiment, as Figures 1-3 shown, one of the purposes of the present invention is to provide a virtual reality interaction training method based on multi-modal feedback, including the following steps:
[0047] S1. Divide the scene into multiple sub - scenes and render them on different computing nodes respectively. The distributed rendering architecture constructs a multi - dimensional training scene library to achieve real - time rendering of large - scale scenes, improve rendering efficiency and scene complexity;
[0048] Specifically, the steps for the distributed rendering architecture to construct a multi - dimensional training scene library are as follows:
[0049] Define metrics for geometric complexity, such as the number of geometric bodies, the number of faces, the number of vertices, etc. Traverse each geometric body in the scene, calculate its complexity metrics, and record the complexity of each region. At the same time, identify dynamically changing objects in the scene, such as moving objects, characters, etc., and mark the regions containing dynamic objects as dynamic regions for subsequent processing. Use a space - partitioning algorithm to dynamically divide the scene into multiple sub - scenes;
[0050] Allocate each sub - scene to different computing nodes, generate a new global timestamp according to the rendering frame rate, and send the global timestamp to each computing node. Each node performs a rendering operation based on the allocated sub - scene and the current timestamp, and integrates the rendering results of all sub - scenes into the main scene to generate a complete virtual environment, improve rendering efficiency, support large - scale complex scenes, enhance system scalability, and the performance can be improved by adding nodes. Also, using timestamp synchronization helps to ensure that the time of all nodes is consistent, improve the user experience, and enhance the immersion of virtual reality. As for the rendering method, it can adopt the commonly used existing technologies, such as frameworks like Distributed Rendering Engine (DRE), to coordinate the rendering tasks of each node. The master node is responsible for task allocation, and each slave node executes the rendering task.
[0051] To reduce the running intensity and memory occupation, specifically, the space - partitioning algorithm includes the following steps:
[0052] The computer tracks the user's perspective in real - time, including position and direction, calculates the viewing frustum based on the user's perspective, determines the user's current focus area. Taking the focus area as a reference, extend the scene grid cells within the threshold layer circle outward, recursively divide the space into sub - regions, use the distance weight of the sub - regions close to the reference as the rendering priority metric, render the sub - regions with the same distance weight through timestamp synchronization, and dynamically update the sub - regions in real - time according to the perspective change, such as Figure 2As shown in the figure, after determining the user's area of interest, extend multiple layers of grid cells outward, including the layer circles where sub-region a and sub-region b are located. Each layer circle has multiple sub-regions with the same distance weight. Within the threshold layer circle, the perspective range that the user may adjust within the guaranteed rendering time can be analyzed according to the user's habits. And the weight a corresponding to sub-region a > the weight b corresponding to sub-region b. When the user belongs to the area of interest, render multiple sub-regions a first, combine the rendering results of each sub-scene according to the time stamp to ensure overall synchronization, then render sub-region b, and dynamically adjust in real time according to the user's perspective. This not only reduces the running intensity caused by rendering multiple scenes simultaneously, saves resources, but also, according to the synchronous adjustment spreading outward with the perspective, can dynamically and real-time update the sub-regions as the perspective is updated, enabling scenes outside the threshold layer circle to be deleted first, avoiding occupying too much space and improving the running speed.
[0053] S2. Use the LSTM-GRU hybrid neural network based on the spatio-temporal attention mechanism to extract features and weighted fuse multi-source data. The multi-source data includes pose and motion data, spatial position and depth information, and muscle activity data. Input the multi-source data interactively into the scenes that are switched in real time by the computer, and switch the scenes through the dual-channel semantic parsing engine to avoid a single interaction channel and lack of natural human-computer interaction ability, thus improving authenticity.
[0054] In a virtual reality interactive training system, the user's behavior data (such as pose, motion trajectory, spatial position, depth information, muscle activity, etc.) needs to be collected and processed in real time. However, existing systems usually adopt a single sensor or simple data processing methods, resulting in low data accuracy and low processing efficiency, and it is difficult to meet the real-time requirements. To address this problem, by extracting the features of multi-source data and weighted fusing them, the LSTM-GRU hybrid neural network based on the spatio-temporal attention mechanism is used to extract features and weighted fuse multi-source data, including the following steps:
[0055] Collection of pose and motion data: Install sensors (including inertial sensors, optical sensors, and electromagnetic sensors) on each joint of the user, such as the wrist, elbow, knee, etc., to collect the user's motion data in real time, including acceleration, angular velocity, position, and pose. Represent the rotation angle of the pose by Euler angles, represent the three-dimensional rotation by quaternions to avoid the gimbal lock problem, and represent the coordinate system transformation by a transformation matrix.
[0056] Collection of spatial position and depth information: The depth camera outputs the depth value of each pixel, forms a depth map and converts it into a three-dimensional point cloud to represent the user's position in space. Computer vision algorithms (such as OpenPose, HRNet) are used to estimate the user's pose from the depth map. It is also possible to consider separating the user from the background to improve the accuracy of pose estimation. The three-dimensional coordinates of the user are mapped into the virtual scene to achieve position synchronization;
[0057] Collection of muscle activity data: Electrodes are attached to the user's muscle surface, such as the forearm, thigh, etc., to collect the electrical signals of the muscles, including the amplitude, frequency, etc. of the electromyogram signals. A bioamplifier is used to amplify the weak electromyogram signals and remove noises such as power frequency interference and motion artifacts. The time-domain and frequency-domain features of the electromyogram signals are extracted, such as the root mean square value (RMS), average power frequency (APF);
[0058] Align different data in time through a hardware clock, align the data of different sensors to the same timeline according to the timestamps. For data with inconsistent timestamps, interpolation methods (such as linear interpolation) can be used for completion to ensure the consistency of the data in time to provide a unified time reference. Considering that some sensors may have data transmission delays or inconsistent sampling rates, synchronization algorithms (such as timestamp-based alignment algorithms) are used to adjust the data to eliminate the influence of delays and jitters, and different weights are assigned according to the reliability and importance of the data. For example, some sensors may be more reliable in a specific environment, or some data sources are more important in training. A hybrid network structure composed of LSTM units designed to capture long-term dependencies and GRU units for processing short-term dependencies is designed. The multi-source data are respectively input into the LSTM and GRU units to extract the features of the multi-source data. For example, the pose data may extract the motion pattern, and the depth data may extract the spatial relationship, and the extracted features are weighted and fused according to the weights to generate a comprehensive feature vector to ensure that the data features with high importance occupy a larger weight. In summary, the comprehensiveness of feature extraction is improved. The entire system design supports real-time data processing and feedback and is applicable to scenarios with high real-time requirements such as virtual reality training.
[0059] Furthermore, the dual channels include a voice channel and an action channel, and the scene is interactively switched through a dual-channel semantic parsing engine, which includes the following steps:
[0060] The voice channel uses a BERT-LSTM hybrid model for semantic understanding. BERT extracts context semantic features, and LSTM decodes temporal dependencies, outputs structured instructions, and sets a voice confidence score for determining the validity of the voice input. If the voice confidence score is lower than the confidence threshold, an error signal is output. If the confidence score is not lower than the confidence threshold, a switching signal is output. Among them, in the output layer of the model, the Softmax function is used to normalize the probability of each category to obtain a probability distribution. The category with the highest probability is the prediction result of the model, and the corresponding probability value is the confidence score. For example, assuming the probability distribution output by the model is [0.1, 0.3, 0.6], the confidence is 0.6, indicating that the model's confidence in the third category is 60%. According to experimental and test data, a confidence threshold (such as 0.7 or 0.8) can be set. When the confidence of the model is lower than this threshold, the system considers the instruction incorrect and cannot switch scenarios. When an error signal is output, an error correction mechanism (such as fuzzy matching) is enabled to try to correct the instruction;
[0061] The action channel is based on 3D-CNN for skeletal pose semantic analysis. It inputs Kinect V4 skeletal data, captures spatial and temporal correlations through 3D convolutional kernels, and outputs gesture semantic encodings (such as "right hand swipe → scene switching"). Among them, the dynamic gesture library contains 8 types of scene control gestures, and sets an action confidence score for judging the effectiveness of gesture recognition. If the action confidence score is lower than the confidence threshold, an error signal is output. If the confidence score is not lower than the confidence threshold, a switching signal is output.
[0062] It should be noted that the parsing engine interacts to switch scenarios, including the following postures:
[0063] Posture 1: Receive the collected posture and motion data, spatial position and depth information, and muscle activity data, generate Kinect V4 skeletal data, trigger the action channel to output gesture semantic encodings, and execute the interactive scene switching according to the switching signal, so that the action can switch scenarios, which is convenient to switch behaviors in real time according to the user's behavior and makes the user training more realistic;
[0064] Posture 2: Receive real-time voice data, extract context semantic features, trigger the voice channel to output structured instructions, and execute the interactive scene switching according to the switching signal, so that the voice can switch scenarios and improve the convenience of switching;
[0065] Pose three, error signals for the interactive switching scenario between the perception action channel and the speech channel. If the number of error signals is higher than the error threshold, establish a voice-action instruction mapping table (e.g., voice "zoom in" ↔ gesture "expand both hands outward"), perform weighted fusion on the confidence scores of the voice and action channels, output the matching degree between the dual-channel instruction and the semantic mapping matrix, and set the comprehensive confidence. If the comprehensive confidence exceeds the total threshold (such as 0.8), it is considered that the input is valid and the switching scenario is triggered. Otherwise, a warning signal is issued;
[0066] In summary, it not only realizes the control of the switching scenario through the dual-channel, improves the accuracy of scene switching, but also when the single-channel switching is inaccurate, it verifies through the two channels simultaneously in a timely manner, improves the accuracy of instruction recognition, and reduces misjudgment.
[0067] S3. Compare the multi-source data with the standard data through a three-level evaluation system to evaluate the operation deviation of the user, transmit the feedback signal to the user through the feedback channel and generate the training result. The feedback channel includes touch, vision, and hearing. Among them, the intensity of the feedback signal is proportional to the operation deviation. It not only realizes the detection of the deviation during the user's operation training, evaluates the training result according to the operation deviation, but also reminds the user to make adjustments in a timely manner through the feedback signal, realizes the function of automatic teaching. At the same time, the intensity of the feedback signal can reflect the intensity of the deviation, further prompting the user to make adjustments;
[0068] Therefore, evaluating the operation deviation of the user includes the following steps:
[0069] Construct standard operation data as the training target, including standard postures, motion trajectories, and spatial positions, and align the standard operation data with the user's real-time operation data in the time dimension;
[0070] For each time step , calculate the comprehensive feature vector and the deviation between the standard operation data . The calculation formula is as follows:
[0071] ;
[0072] Among them, is the data dimension (such as posture, position, muscle activity, etc.), i is the number of digits of the training target, and weights are assigned to each type of data according to the importance of the data to generate the weighted deviation. By assigning different weights to different data sources, the weighted deviation can more accurately reflect the key points of the user's operation. For example, in the training scenario, the posture data may be more important than the spatial position data, so a higher weight is assigned to ensure that the evaluation result is closer to the actual needs;
[0073] Specifically, it also includes calculating the deviation between the user's actual joint angle and the standard angle using the cosine similarity algorithm for the key joint angle deviation, calculating the deviation between the user's actual motion trajectory and the standard trajectory using the Hausdorff distance for judging the consistency of the motion trajectory, calculating the deviation between the user's operation timing sequence and the standard timing sequence using the dynamic time warping (DTW) algorithm for judging the operation timing accuracy, evaluating the user's operation deviation in real time, providing accurate feedback guidance, and improving the training efficiency.
[0074] On this basis, the transmission of the feedback signal includes the following steps:
[0075] Using a normalization function to map the weighted deviation to the intensity range of the feedback signal, and generating corresponding feedback signals through the feedback channels respectively, including the following postures:
[0076] Tactile feedback: Reflect the deviation through the vibration intensity, and transmit the vibration signal through a tactile device (such as a vibration feedback glove), mapping the normalized deviation to the vibration intensity. For example, using linear mapping, vibration intensity = maximum vibration intensity * deviation;
[0077] Visual feedback: Display the deviation through color change, and display the color change in the computer interface, mapping the deviation to the color change. For example, from green (small deviation) to red (large deviation), such as: color = interpolation (deviation, green, red);
[0078] Auditory feedback: Prompt the deviation through the rhythm change, and play the sound of the rhythm change through a speaker or headphones, mapping the deviation to the rhythm change. For example, an accelerating rhythm indicates an increasing deviation, such as: rhythm speed = base speed + (maximum speed increment * deviation);
[0079] Training results are generated based on the feedback signal based on a three-level evaluation system, including a primary evaluation that calculates deviations in real time and provides feedback, an intermediate evaluation that periodically evaluates training progress and effects, and an advanced evaluation that comprehensively analyzes long-term training data to provide personalized suggestions. Among them, the primary evaluation: real-time deviation calculation and feedback, to ensure that users can promptly discover deviations and make adjustments during operation, avoid error accumulation, and improve the real-time and effectiveness of training; the intermediate evaluation regularly analyzes the user's training data to evaluate the training progress and effect. For example, the user's training data is summarized and analyzed every week or month to determine whether the user has achieved the predetermined training goal. By analyzing the user's long-term deviation data, the user's operation improvement is evaluated. If the user's deviation persists The intermediate evaluation will trigger further feedback mechanisms, such as increasing the intensity of training or adjusting the training content. The advanced evaluation conducts a comprehensive analysis of the user's long-term training data to identify the user's operation mode, strengths and weaknesses. Based on the analysis results, the advanced evaluation provides users with personalized training suggestions. For example, if the user has a long-term deviation in a certain operation, the advanced evaluation will recommend that the user conduct targeted training or adjust the operation method. The three-level evaluation system covers all stages of the training process from real-time operation to long-term data analysis, ensuring that each stage has a corresponding evaluation and feedback mechanism. Through hierarchical evaluation, the system can use resources more efficiently, adjust training strategies in a timely manner, avoid resource waste, and improve user participation and training effects.
[0080] In addition, the training effect is evaluated based on the user's operation deviation. The smaller the deviation, the better the training result. The feedback signal can be adjusted according to the importance of the data source, so that users can focus on the main issues more easily. For example, the vibration intensity of the tactile feedback can be proportional to the weighted value of the posture deviation, helping users to quickly adjust their movements and facilitating the automatic provision of teaching guidance. At the same time, the deviation of the secondary data source is prevented from causing excessive interference to the feedback signal, allowing users to focus on the main issues. For example, if the spatial position deviation is small, its impact on the feedback signal will also be reduced, helping users to adjust their operations more efficiently.
[0081] The second object of the present invention is to provide a virtual reality interactive training system based on multimodal feedback, including any one of the virtual reality interactive training methods based on multimodal feedback described above, including a multi-dimensional training scene construction unit, a multi-source data interaction unit and an operation deviation feedback unit;
[0082] The multi-dimensional training scene construction unit is used to divide the scene into multiple sub-scenes, which are rendered on different computing nodes respectively. The distributed rendering architecture builds a multi-dimensional training scene library.
[0083] The multi-source data interaction unit is used to extract features and weighted-fuse multi-source data based on the LSTM-GRU hybrid neural network with spatio-temporal attention mechanism. The multi-source data includes pose and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the scene that is switched in real time by the computer, and the scene is interactively switched through the dual-channel semantic parsing engine;
[0084] The operation deviation feedback unit is used to compare the multi-source data with the standard data through a three-level evaluation system, evaluate the operation deviation of the user, transmit the feedback signal to the user through the feedback channel and generate a training result, and make the intensity of the feedback signal proportional to the operation deviation.
[0085] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A virtual reality interactive training method based on multimodal feedback, characterized in that: The steps include: The scene is divided into multiple sub-scenes, which are rendered on different computing nodes. The distributed rendering architecture builds a multi-dimensional training scene library. Each sub-scene is assigned to a different computing node. A new global timestamp is generated based on the rendering frame rate and sent to each computing node. Each node performs rendering operations based on the assigned sub-scene and the current timestamp, and integrates the rendering results of all sub-scenes into the main scene to generate a complete virtual environment. The LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weights the fusion of multi-source data, including posture and motion data, spatial position and depth information, and muscle activity data. The multi-source data is interactively input into the scene switched by the computer in real time, and the scene is interactively switched through the dual-channel semantic parsing engine, which includes a voice channel and an action channel. Multi-source data are compared with standard data through a three-level evaluation system, and training results are generated based on the three-level evaluation system according to feedback signals, including a primary evaluation that calculates deviations in real time and provides feedback, an intermediate evaluation that periodically evaluates training progress and effects, and an advanced evaluation that comprehensively analyzes long-term training data and provides personalized suggestions. It evaluates the user's operation deviation, transmits feedback signals to the user through feedback channels, and generates training results. The feedback channels include touch, vision, and hearing, and the strength of the feedback signal is proportional to the operation deviation.
2. The virtual reality interactive training method based on multimodal feedback according to claim 1, characterized in that: The distributed rendering architecture constructs a multi-dimensional training scene library, including the following steps: Define the geometric complexity index, traverse each geometric body in the scene, calculate its complexity index, and record the complexity of each area. At the same time, identify dynamically changing objects in the scene and use the space partitioning algorithm to dynamically divide the scene into multiple sub-scenes; Each sub-scene is assigned to a different computing node, and a new global timestamp is generated according to the rendering frame rate. The global timestamp is sent to each computing node. Each node performs rendering operations according to the assigned sub-scene and the current timestamp, and the rendering results of all sub-scenes are integrated into the main scene to generate a complete virtual environment.
3. The virtual reality interactive training method based on multimodal feedback according to claim 2, characterized in that: The space partitioning algorithm comprises the following steps: The user's perspective, including position and direction, is tracked in real time by a computer. The visual cone is calculated based on the user's perspective to determine the user's current focus area. Taking the focus area as a benchmark, the scene grid units within the threshold layer circle are extended outward, and the space is recursively divided into sub-areas. The distance weight of the sub-area close to the benchmark is used as a rendering priority indicator. Sub-areas with the same distance weight are rendered synchronously through timestamps, and the sub-areas are dynamically updated in real time according to changes in perspective.
4. The virtual reality interactive training method based on multimodal feedback according to claim 1, characterized in that: The LSTM-GRU hybrid neural network based on the spatiotemporal attention mechanism extracts features and weightedly fuses multi-source data, including the following steps: Collection of posture and motion data, collection of spatial position and depth information, and collection of muscle activity data; Different data are aligned in time through the hardware clock, and different weights are assigned according to the reliability and importance of the data. The LSTM units designed to capture long-term dependencies and the GRU units designed to process short-term dependencies are composed of a hybrid network structure. Multi-source data are input into the LSTM and GRU units respectively, the features of the multi-source data are extracted, and weighted fusion is performed according to the weights to generate a comprehensive feature vector.
5. The virtual reality interactive training method based on multimodal feedback according to claim 4, characterized in that: The dual channels include a voice channel and an action channel, and scenes are interactively switched through a dual channel semantic parsing engine, which includes the following steps: The voice channel uses a BERT-LSTM hybrid model for semantic understanding. BERT extracts contextual semantic features, LSTM decodes temporal dependencies, outputs structured instructions, and sets a voice confidence score for validating voice input. If the voice confidence score is lower than the confidence threshold, an error signal is output; if the confidence score is not lower than the confidence threshold, a switching signal is output. The action channel is based on 3D-CNN's skeletal posture semantic analysis, inputs Kinect V4 skeletal data, captures spatial and temporal correlations through 3D convolution kernels, outputs gesture semantic encoding, and sets an action confidence score for judging the effectiveness of gesture recognition. If the action confidence score is lower than the confidence threshold, an error signal is output; if the confidence score is not lower than the confidence threshold, a switching signal is output.
6. The virtual reality interactive training method based on multimodal feedback according to claim 5, characterized in that: The parsing engine interactively switches scenes, including the following gestures: Posture 1: Receive the collected posture and motion data, spatial position and depth information, and muscle activity data, generate Kinect V4 skeleton data, trigger the action channel to output gesture semantic coding, and execute interactive switching scenes according to the switching signal; Posture 2: Receive real-time voice data, extract contextual semantic features, trigger the voice channel to output structured instructions, and execute interactive switching scenarios according to the switching signal; Posture 3. Perceive the error signal of the interactive switching scene of the action channel and the voice channel. If the number of error signals is higher than the error threshold, establish a voice-action command mapping table, weightedly fuse the confidence scores of the voice and action channels, output the matching degree of the dual-channel command and the semantic mapping matrix, and set the comprehensive confidence. If the comprehensive confidence exceeds the total threshold, the input is considered valid and the switching scene is triggered. Otherwise, a warning signal is issued.
7. The virtual reality interactive training method based on multimodal feedback according to claim 4, characterized in that: The step of evaluating the user's operation deviation comprises the following steps: Construct standard operation data as training targets, including standard posture, motion trajectory, and spatial position, and align the standard operation data with the user's real-time operation data in the time dimension; For each time step , calculate the comprehensive feature vector With standard operating data The deviation between , the calculation formula is as follows: ; in, is the data dimension, i is the number of training target bits, and according to the importance of the data, a weight is assigned to each data to generate a weighted deviation.
8. The virtual reality interactive training method based on multimodal feedback according to claim 7, characterized in that: The transmitting feedback signal comprises the following steps: The weighted deviation is mapped to the intensity range of the feedback signal using a normalization function, and the corresponding feedback signals are generated through the feedback channels, including the following postures: Haptic feedback: reflects deviations through vibration intensity and transmits vibration signals through tactile devices; Visual feedback: Deviations are indicated by color changes, and the color changes are displayed in the computer interface; Auditory feedback: Deviations are indicated by rhythmic changes, and the sound of rhythmic changes is played through a speaker; Training results are generated based on feedback signals based on a three-level evaluation system, including a primary evaluation that calculates deviations in real time and provides feedback, an intermediate evaluation that periodically evaluates training progress and effects, and an advanced evaluation that comprehensively analyzes long-term training data and provides personalized recommendations.
9. A virtual reality interactive training system based on multimodal feedback, comprising the virtual reality interactive training method based on multimodal feedback according to any one of claims 1 to 8, characterized in that: It includes a multi-dimensional training scenario construction unit, a multi-source data interaction unit, and an operation deviation feedback unit; The multi-dimensional training scene construction unit is used to divide the scene into multiple sub-scenes, which are rendered on different computing nodes respectively, and the distributed rendering architecture constructs a multi-dimensional training scene library; The multi-source data interaction unit is used to extract features and weightedly fuse multi-source data based on the LSTM-GRU hybrid neural network of the spatiotemporal attention mechanism, the multi-source data including posture and motion data, spatial position and depth information, and muscle activity data, interactively input the multi-source data into the scene switched by the computer in real time, and interactively switch the scene through the dual-channel semantic parsing engine; The operation deviation feedback unit is used to compare multi-source data with standard data through a three-level evaluation system, evaluate the user's operation deviation, transmit feedback signals to the user through a feedback channel and generate training results, and make the strength of the feedback signal proportional to the operation deviation.
Citation Information
Patent Citations
Virtual reality scene interaction method and system of SaaS platform
CN118295538A
AR (Augmented Reality) guide system for enhancing stone forest tourism experience
CN118314306A