A Virtual Reality-Based Interactive Control Method and System

By using multimodal data fusion and dynamic weighting mechanisms, combined with multi-channel feedback and environmental adaptation, the problem of non-standard actions in virtual reality interactive control has been solved, achieving efficient action recognition and immersive experience, and improving user interaction efficiency and comfort.

CN120704516BActive Publication Date: 2026-04-03CHANGCHUN BAIRUI INTERNATIONAL EXHIBITION GROUP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing virtual reality interactive control systems, non-standard target user movements lead to low recognition success rates and reduced interaction efficiency.

Method used

Employing a multimodal motion capture module, a temporal convolutional neural network model, and a dynamic weighted fusion unit, combined with a multi-channel feedback module and an environment adaptation module, the system monitors user limb data, joint angles, and muscle signals, and adjusts data fusion weights and scene parameters in real time to provide visual, tactile, and audio feedback, ensuring accurate motion recognition and environmental adaptability.

Benefits of technology

It significantly improves motion capture accuracy and environmental adaptability, enhances interaction efficiency, increases user immersion and comfort, reduces recognition failures caused by non-standard movements, and improves user training compliance and learning effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704516B_ABST
    Figure CN120704516B_ABST
Patent Text Reader

Abstract

This application discloses an interactive control method and system based on virtual reality, relating to the field of monitoring and analysis technology. It includes a multimodal motion capture module for capturing limb spatial coordinate data, joint rotation angle data, and muscle activation signals corresponding to the target user; a motion intent parsing module, including a temporal convolutional neural network model, comprising an input layer and an output layer. The input layer receives skeletal joint point sequences, angular velocity data, and electromyographic feature vectors from the multimodal motion capture module, while the output layer extracts spatiotemporal features through a three-layer dilated causal convolution and connects to a classifier to generate initial motion data corresponding to the target user; a correction module constructs a human joint angle constraint library, judges the initial motion data corresponding to the target user, and outputs correction commands; and an environment adaptation module adjusts the environment in the virtual scene. This application improves interaction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monitoring and analysis technology, and in particular to an interactive control method and system based on virtual reality. Background Technology

[0002] Virtual reality technology is an advanced, digital human-computer interface technology. Its characteristic is that the computer generates an artificial virtual environment, creating an artificial environment with visual perception as the main focus, but also including auditory and tactile perception. People can perceive the computer-generated virtual world through multiple sensory channels such as vision, hearing, touch, and acceleration. They can also interact with the virtual world in the most natural ways, such as movement, voice, and actions, thereby creating an immersive experience.

[0003] In existing virtual reality interactive control systems, there are frequent instances where the target user's actions are not properly executed, resulting in a decrease in the user's interaction efficiency and requiring improvement. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this application provides an interactive control method and system based on virtual reality.

[0005] In a first aspect, this application provides an interactive control system based on virtual reality, comprising:

[0006] The multimodal motion capture module is used to capture the spatial coordinates of the target user's limbs, joint rotation angles, and muscle activation signals through a monitoring device.

[0007] The motion intent parsing module includes a temporal convolutional neural network model and a dynamic weighted fusion unit. The temporal convolutional neural network model includes an input layer and an output layer. The input layer is used to receive the skeletal joint sequence, angular velocity data, and electromyographic feature vector from the multimodal motion capture module. The output layer is used to extract spatiotemporal features through a three-layer dilated causal convolution and connect it to a Softmax classifier to generate initial motion data corresponding to the target user. The dynamic weighted fusion unit is used to adjust the fusion weights of the data from each monitoring device in real time according to the environmental complexity parameter.

[0008] The correction module is used to build a human joint angle constraint library, judge the initial motion data corresponding to the target user, and output correction instructions based on the judgment results.

[0009] The multi-channel feedback module includes a visual guidance projection unit, a tactile vibration array, and a bone conduction audio unit. The visual guidance projection unit is used to generate a semi-transparent reference motion trajectory in a virtual scene. The tactile vibration array is used to place linear resonant actuators on the target user and generate differentiated vibration patterns based on the motion error vector. The bone conduction audio unit is used to select feedback content based on a preset voice command library and according to the error type.

[0010] Preferably, the multimodal motion capture module includes a binocular infrared depth camera array, an inertial measurement unit sensor group, and a surface electromyography sensor;

[0011] The binocular infrared depth camera array is used to acquire depth images of the virtual scene. Based on the depth images, the contour data corresponding to the target user is extracted, and the key joints corresponding to the target user are located using a joint point detection algorithm. The coordinates of the key joints are confirmed, and the detected joint coordinates are converted into the world coordinate system to generate a skeletal joint point sequence.

[0012] The inertial measurement unit sensor group is used to acquire real-time attitude data corresponding to each joint of the target user, and extract angular velocity data from the real-time attitude data;

[0013] The surface electromyography (EMG) sensor is used to collect muscle activation signals, extract features, and normalize the extracted features to identify the EMG feature vector.

[0014] Preferably, the process of adjusting the fusion weights of data from each monitoring device in real time based on environmental complexity parameters specifically includes:

[0015] Through formula Confirm the complexity of the environment ;

[0016] in, These represent the density of virtual objects and the variance of illumination variation in the virtual scene, respectively. This represents the movement speed corresponding to the target user. Represented as preset weighting coefficients;

[0017] Environmental complexity Compare with the preset environmental complexity threshold range;

[0018] If environmental complexity In When the weights are in between, the weights are allocated according to the preset first allocation strategy;

[0019] If environmental complexity In When the weights are in between, the weights are allocated according to the preset second allocation strategy;

[0020] If environmental complexity If the weights are not specified, then the weights will be allocated according to the preset third allocation strategy.

[0021] Preferably, the correction module specifically includes:

[0022] Within a preset time period, a human joint angle constraint library is established. The human joint angle constraint library stores data on the range of motion of each joint and sets kinematic chain verification rules.

[0023] Define action recognition confidence , ,in, This is represented by the corresponding number for each sensor. These are represented as the weights and recognition probabilities corresponding to each sensor, respectively.

[0024] Set dynamic threshold ,in, This represents the preset initial confidence level. This is represented as a preset correction factor;

[0025] When the confidence level is identified If the action is valid, then the action is determined to be valid.

[0026] Preferably, the correction module further includes:

[0027] If the first action recognition fails, the virtual object state remains unchanged, and prompts and guidance are provided through the multi-channel feedback module;

[0028] If the second action recognition fails, the joint angle constraint will be automatically relaxed.

[0029] If the third action recognition fails, the environment adaptation module is triggered to adjust the difficulty of the environment in the virtual scene.

[0030] Preferably, the environment adaptation module includes a parameter dynamic adjustment unit and a load assessment unit;

[0031] The parameter dynamic adjustment unit is used to modify the mass and friction coefficient of virtual objects in the virtual scene in real time.

[0032] The load assessment unit is used to determine the load assessment index using fixation point data and heart rate data.

[0033] Preferably, the process of modifying the mass and coefficient of friction of virtual objects in a virtual scene specifically includes:

[0034] Formulas for adjusting the mass and friction coefficient parameters of virtual objects in a virtual scene:

[0035] Quality of virtual objects: ;

[0036] Coefficient of friction of virtual objects: ;

[0037] in, This represents the number of consecutive failed action recognition attempts. This is expressed as a preset standard mass and coefficient of friction;

[0038] And when If this occurs, the emergency assistance mode is activated, which is used to automatically complete the remaining range of motion.

[0039] Preferably, the process of determining the workload assessment index using fixation point data and heart rate data specifically includes:

[0040] Within a preset time window, the target user's gaze point data and heart rate data are collected through a preset device.

[0041] The fixation point data and heart rate data corresponding to the target user are preprocessed, and the fixation point dispersion corresponding to the target user is calculated using the preprocessed fixation point data and heart rate data. Heart rate variability index ;

[0042] Through formula Identify the load assessment index corresponding to the target user. ,in, These represent the weighting coefficients corresponding to fixation point dispersion and heart rate variability, respectively.

[0043] Load assessment index corresponding to the target user Compared with the preset load assessment threshold Perform a comparison;

[0044] If the target user's corresponding load assessment index At that time, the scene elements in the virtual scene are dynamically simplified.

[0045] Secondly, this application provides an interactive control method based on virtual reality, comprising the following steps:

[0046] The monitoring device captures the target user's limb spatial coordinate data, joint rotation angle data, and muscle activation signals.

[0047] Action intent parsing includes a temporal convolutional neural network model and a dynamic weighted fusion unit. The temporal convolutional neural network model includes an input layer and an output layer. The input layer is used to receive the skeletal joint sequence, angular velocity data, and electromyographic feature vector from the multimodal motion capture module. The output layer is used to extract spatiotemporal features through a three-layer dilated causal convolution and connect to a Softmax classifier to generate initial action data corresponding to the target user. The dynamic weighted fusion unit is used to adjust the fusion weights of the data from each monitoring device in real time according to the environmental complexity parameter.

[0048] Construct a human joint angle constraint library, judge the initial action data corresponding to the target user, and output correction instructions based on the judgment results;

[0049] Multi-channel feedback includes a visually guided projection unit, a tactile vibration array, and a bone conduction audio unit. The visually guided projection unit is used to generate a semi-transparent reference motion trajectory in a virtual scene. The tactile vibration array is used to place linear resonant actuators on the target user and generate differentiated vibration patterns based on the motion error vector. The bone conduction audio unit is used to select feedback content based on a preset voice command library and according to the error type.

[0050] Adjust the environment in the virtual scene.

[0051] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to execute any of the above-described interactive control systems based on virtual reality.

[0052] In summary, this application includes at least one of the following beneficial technical effects:

[0053] 1. This application provides an interactive control system based on virtual reality. Through the fusion of multimodal data and dynamic weighting mechanism, it significantly improves the accuracy and environmental adaptability of motion capture. By using a temporal convolutional neural network to extract spatiotemporal features and combining them with a human joint constraint library, it achieves both accurate interpretation of action intent and physiological rationality. Multi-channel feedback constructs a three-dimensional correction system, enhancing the naturalness and guidance efficiency of human-computer interaction. The environment adaptation module ensures the dynamic matching between virtual scenes and real training needs, which not only guarantees the standardization of actions but also improves user training compliance and action learning effects through a highly immersive experience. This effectively reduces the occurrence of unsuccessful recognition due to non-standard actions of the target user, thereby effectively improving the interaction efficiency of the target user.

[0054] 2. By assessing the load evaluation index of target users in real time, it can accurately capture the physiological and behavioral state of users in the virtual reality environment. When the load evaluation index exceeds the preset threshold, it automatically simplifies scene elements, thereby effectively reducing the cognitive burden of target users. Moreover, dynamic adaptive optimization not only enhances the immersion and comfort of target users, but also significantly improves interaction efficiency. Attached Figure Description

[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 This is a schematic diagram of an interactive control system based on virtual reality, according to an embodiment of this application.

[0057] Figure 2 This is a flowchart of an interactive control method based on virtual reality, according to an embodiment of this application. Detailed Implementation

[0058] The following is in conjunction with the appendix Figure 1-2 This application will be described in further detail.

[0059] Example 1

[0060] This application discloses an interactive control system based on virtual reality.

[0061] Reference Figure 1 An interactive control system based on virtual reality, comprising:

[0062] The multimodal motion capture module is used to capture the spatial coordinates of the target user's limbs, joint rotation angles, and muscle activation signals through a monitoring device.

[0063] The motion intent parsing module includes a temporal convolutional neural network model and a dynamic weighted fusion unit. The temporal convolutional neural network model includes an input layer and an output layer. The input layer is used to receive the skeletal joint sequence, angular velocity data, and electromyographic feature vector from the multimodal motion capture module. The output layer is used to extract spatiotemporal features through a three-layer dilated causal convolution and connect it to a Softmax classifier to generate initial motion data corresponding to the target user. The dynamic weighted fusion unit is used to adjust the fusion weights of the data from each monitoring device in real time according to the environmental complexity parameter.

[0064] The correction module is used to build a human joint angle constraint library, judge the initial motion data corresponding to the target user, and output correction instructions based on the judgment results.

[0065] The multi-channel feedback module includes a visual guidance projection unit, a tactile vibration array, and a bone conduction audio unit. The visual guidance projection unit is used to generate a semi-transparent reference motion trajectory in a virtual scene. The tactile vibration array is used to place linear resonant actuators on the target user and generate differentiated vibration patterns based on the motion error vector. The bone conduction audio unit is used to select feedback content based on a preset voice command library and according to the error type.

[0066] The environment adaptation module is used to adjust the environment in the virtual scene.

[0067] Furthermore, the multimodal motion capture module includes a binocular infrared depth camera array, an inertial measurement unit sensor group, and a surface electromyography sensor;

[0068] The binocular infrared depth camera array is used to acquire depth images of the virtual scene. Based on the depth images, the contour data corresponding to the target user is extracted, and the key joints corresponding to the target user are located using a joint point detection algorithm. The coordinates of the key joints are confirmed, and the detected joint coordinates are converted into the world coordinate system to generate a skeletal joint point sequence.

[0069] The inertial measurement unit sensor group is used to acquire real-time attitude data corresponding to each joint of the target user, and extract angular velocity data from the real-time attitude data;

[0070] The surface electromyography (EMG) sensor is used to collect muscle activation signals, extract features, and normalize the extracted features to identify the EMG feature vector.

[0071] Specifically, the above data acquisition requires multi-sensor time synchronization to ensure data timestamp alignment and the establishment of a unified coordinate system. The target user's corresponding skeletal joint sequence, angular velocity data, and electromyographic feature vector are input into a temporal convolutional neural network model for multimodal feature fusion, thereby confirming the fused initial motion data.

[0072] The above technical solutions can accurately acquire skeletal joint sequence, angular velocity data, and electromyographic feature vectors, providing a reliable data foundation for subsequent motion recognition and interactive control.

[0073] It should be noted that the process of adjusting the fusion weights of data from each monitoring device in real time based on environmental complexity parameters specifically includes:

[0074] Through formula Confirm the complexity of the environment ;

[0075] in, These represent the density of virtual objects and the variance of illumination variation in the virtual scene, respectively. This represents the movement speed corresponding to the target user. Represented as preset weighting coefficients;

[0076] Environmental complexity Compare with the preset environmental complexity threshold range;

[0077] If environmental complexity In When the weights are in between, the weights are allocated according to the preset first allocation strategy;

[0078] If environmental complexity In When the weights are in between, the weights are allocated according to the preset second allocation strategy;

[0079] If environmental complexity If the weights are not specified, then the weights will be allocated according to the preset third allocation strategy.

[0080] Specifically, different allocation strategies result in different weighting coefficients for the data collected by the binocular infrared depth camera array, the inertial measurement unit sensor group, and the surface electromyography sensor, thus providing a reliable data foundation for the subsequent calculation of recognition confidence.

[0081] It should be noted that the correction module specifically includes:

[0082] Within a preset time period, a human joint angle constraint library is established. The human joint angle constraint library stores data on the range of motion of each joint and sets kinematic chain verification rules.

[0083] For example, the range of motion for shoulder abduction is 0-180°, and the range of motion for knee flexion is 0-135°; the kinematic chain verification rule for grasping action is a temporal constraint of [shoulder flexion > 30° → elbow flexion > 90° → wrist dorsiflexion > 20°], and allows for an error of 15 degrees in the intermediate links;

[0084] Define action recognition confidence , ,in, This is represented by the corresponding number for each sensor. These represent the weights and recognition probabilities corresponding to each sensor, respectively. The recognition probabilities can be obtained by fitting historical data.

[0085] Set dynamic threshold ,in, This represents the preset initial confidence level. This is represented as a preset correction factor;

[0086] When the confidence level is identified If the action is valid, then the action is determined to be valid.

[0087] Furthermore, the correction module specifically includes:

[0088] If the first action recognition fails, the virtual object state remains unchanged, and prompts and guidance are provided through the multi-channel feedback module;

[0089] If the second action recognition fails, the joint angle constraint will be automatically relaxed.

[0090] If the third action recognition fails, the environment adaptation module is triggered to adjust the difficulty of the environment in the virtual scene.

[0091] It should be noted that the environment adaptation module includes a parameter dynamic adjustment unit and a load assessment unit;

[0092] The parameter dynamic adjustment unit is used to modify the mass and friction coefficient of virtual objects in the virtual scene in real time.

[0093] The load assessment unit is used to determine the load assessment index using fixation point data and heart rate data.

[0094] Furthermore, the process of modifying the mass and coefficient of friction of virtual objects in a virtual scene specifically includes:

[0095] Formulas for adjusting the mass and friction coefficient parameters of virtual objects in a virtual scene:

[0096] Quality of virtual objects: ;

[0097] Coefficient of friction of virtual objects: ;

[0098] in, This represents the number of consecutive failed action recognition attempts. This is expressed as a preset standard mass and coefficient of friction;

[0099] And when If this occurs, the emergency assistance mode is activated, which is used to automatically complete the remaining range of motion.

[0100] Specifically, dynamically adjusting the mass and coefficient of friction of objects in a virtual environment can significantly enhance the realism and application value of simulations. First, by optimizing mass parameters, the motion inertia, collision feedback, and energy transfer effects of objects can be simulated more accurately. For example, in game development, increasing the mass of a metal box will result in a stronger impact force and displacement trajectory upon collision, while reducing the weight of a balloon will produce a lighter floating effect. Second, adjusting the coefficient of friction can effectively control the contact behavior between objects. For instance, in racing games, reducing the coefficient of friction between tires and ice can realistically recreate vehicle slippage, while increasing the friction of rubber materials can achieve a sudden stop. This parameter adjustment not only enhances the realism of the user experience but also has practical significance for industrial simulation—in robotic arm grasping operations, precise friction coefficient settings can predict the risk of workpiece slippage; in the construction field, combinations of friction parameters for components of different materials can simulate structural displacement during earthquakes. Simultaneously, dynamic parameter optimization can balance computational resources, improving system operating efficiency by reducing the physical calculation accuracy of non-critical objects. Overall, the flexible adjustment of mass and friction coefficient is the core technical means to build a highly realistic virtual environment, which can not only meet the needs of artistic expression, but also provide a reliable digital twin platform for professional fields such as scientific research experiments and product testing.

[0101] Furthermore, the process of confirming the workload assessment index using fixation point data and heart rate data specifically includes:

[0102] Within a preset time window, the target user's gaze point data and heart rate data are collected through a preset device.

[0103] The fixation point data and heart rate data corresponding to the target user are preprocessed, and the fixation point dispersion corresponding to the target user is calculated using the preprocessed fixation point data and heart rate data. Heart rate variability index ;

[0104] Through formula Identify the load assessment index corresponding to the target user. ,in, These represent the weighting coefficients corresponding to fixation point dispersion and heart rate variability, respectively.

[0105] Load assessment index corresponding to the target user Compared with the preset load assessment threshold Perform a comparison;

[0106] If the target user's corresponding load assessment index At that time, the scene elements in the virtual scene are dynamically simplified.

[0107] Specifically, by assessing the user's workload index in real time, the system can accurately capture the user's physiological and behavioral state in the virtual reality environment. When the workload index exceeds a preset threshold, the system automatically simplifies scene elements, such as reducing the number of non-critical objects, lowering the complexity of lighting and shadows, or extending the task time limit, thereby effectively reducing the user's cognitive burden. Dynamic adaptive optimization not only enhances the user's immersion and comfort but also significantly improves interaction efficiency. For example, in virtual driving training, when a user feels stressed due to complex scenes, the system automatically reduces the number of background vehicles and simplifies road textures, allowing the user to focus more on the core task and avoid misoperations or fatigue caused by information overload.

[0108] Example 2

[0109] This application also discloses an interactive control method based on virtual reality.

[0110] Reference Figure 2 A virtual reality-based interactive control method includes the following steps:

[0111] The monitoring device captures the target user's limb spatial coordinate data, joint rotation angle data, and muscle activation signals.

[0112] Action intent parsing includes a temporal convolutional neural network model and a dynamic weighted fusion unit. The temporal convolutional neural network model includes an input layer and an output layer. The input layer is used to receive the skeletal joint sequence, angular velocity data, and electromyographic feature vector from the multimodal motion capture module. The output layer is used to extract spatiotemporal features through a three-layer dilated causal convolution and connect to a Softmax classifier to generate initial action data corresponding to the target user. The dynamic weighted fusion unit is used to adjust the fusion weights of the data from each monitoring device in real time according to the environmental complexity parameter.

[0113] Construct a human joint angle constraint library, judge the initial action data corresponding to the target user, and output correction instructions based on the judgment results;

[0114] Multi-channel feedback includes a visually guided projection unit, a tactile vibration array, and a bone conduction audio unit. The visually guided projection unit is used to generate a semi-transparent reference motion trajectory in a virtual scene. The tactile vibration array is used to place linear resonant actuators on the target user and generate differentiated vibration patterns based on the motion error vector. The bone conduction audio unit is used to select feedback content based on a preset voice command library and according to the error type.

[0115] Adjust the environment in the virtual scene.

[0116] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention, they should all fall within the protection scope of the present invention.

[0117] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0118] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. An interactive control system based on virtual reality, characterized in that, include: The multimodal motion capture module is used to capture the spatial coordinate data of the target user's limbs, joint rotation angle data, and muscle activation signals through a monitoring device. The motion intent parsing module includes a temporal convolutional neural network model and a dynamic weighted fusion unit. The temporal convolutional neural network model includes an input layer and an output layer. The input layer is used to receive the skeletal joint sequence, angular velocity data, and electromyographic feature vector from the multimodal motion capture module. The output layer is used to extract spatiotemporal features through a three-layer dilated causal convolution and connect it to a Softmax classifier to generate initial motion data corresponding to the target user. The dynamic weighted fusion unit is used to adjust the fusion weights of the data from each monitoring device in real time according to the environmental complexity parameter. The process of adjusting the fusion weights of data from each monitoring device in real time based on environmental complexity parameters specifically includes: Through formula Confirm the complexity of the environment ; in, These represent the density of virtual objects and the variance of illumination variation in the virtual scene, respectively. This represents the movement speed corresponding to the target user. Represented as preset weighting coefficients; Environmental complexity Compare with the preset environmental complexity threshold range; If environmental complexity In When the weights are in between, the weights are allocated according to the preset first allocation strategy; If environmental complexity In When the weights are in between, the weights are allocated according to the preset second allocation strategy; If environmental complexity In this case, the weights are allocated according to the preset third allocation strategy; The correction module is used to build a human joint angle constraint library, judge the initial motion data corresponding to the target user, and output correction instructions based on the judgment results. The correction module specifically includes: Within a preset time period, a human joint angle constraint library is established. The human joint angle constraint library stores data on the range of motion of each joint and sets kinematic chain verification rules. Define action recognition confidence , ,in, This is represented by the corresponding number for each sensor. These are represented as the weights and recognition probabilities corresponding to each sensor, respectively. Set dynamic threshold ,in, This represents the preset initial confidence level. This is represented as a preset correction factor; When the confidence level is identified If so, the action is determined to be a valid action; The correction module further includes: If the first action recognition fails, the virtual object state remains unchanged, and prompts and guidance are provided through the multi-channel feedback module; If the second action recognition fails, the joint angle constraint will be automatically relaxed. If the third action recognition fails, the environment adaptation module is triggered to adjust the difficulty of the environment in the virtual scene. The multi-channel feedback module includes a visual guidance projection unit, a tactile vibration array, and a bone conduction audio unit. The visual guidance projection unit is used to generate a semi-transparent reference motion trajectory in a virtual scene. The tactile vibration array is used to place linear resonant actuators on the target user and generate differentiated vibration patterns based on the motion error vector. The bone conduction audio unit is used to select feedback content based on a preset voice command library and according to the error type. The environment adaptation module is used to adjust the environment in the virtual scene.

2. The interactive control system based on virtual reality according to claim 1, characterized in that, The multimodal motion capture module includes a binocular infrared depth camera array, an inertial measurement unit sensor group, and a surface electromyography sensor. The binocular infrared depth camera array is used to acquire depth images of the virtual scene. Based on the depth images, the contour data corresponding to the target user is extracted, and the key joints corresponding to the target user are located using a joint point detection algorithm. The coordinates of the key joints are confirmed, and the detected joint coordinates are converted into the world coordinate system to generate a skeletal joint point sequence. The inertial measurement unit sensor group is used to acquire real-time attitude data corresponding to each joint of the target user, and extract angular velocity data from the real-time attitude data; The surface electromyography (EMG) sensor is used to collect muscle activation signals, extract features, and normalize the extracted features to identify the EMG feature vector.

3. The interactive control system based on virtual reality according to claim 1, characterized in that, The environment adaptation module includes a parameter dynamic adjustment unit and a load assessment unit; The parameter dynamic adjustment unit is used to modify the mass and friction coefficient of virtual objects in the virtual scene in real time. The load assessment unit is used to determine the load assessment index using fixation point data and heart rate data.

4. The interactive control system based on virtual reality according to claim 3, characterized in that, The process of modifying the mass and coefficient of friction of virtual objects in a virtual scene includes: Formulas for adjusting the mass and friction coefficient parameters of virtual objects in a virtual scene: Quality of virtual objects: ; Coefficient of friction of virtual objects: ; in, This represents the number of consecutive failed action recognition attempts. This is expressed as a preset standard mass and coefficient of friction; And when If this occurs, the emergency assistance mode is activated, which is used to automatically complete the remaining range of motion.

5. The interactive control system based on virtual reality according to claim 3, characterized in that, The process of identifying the workload assessment index using fixation point data and heart rate data specifically includes: Within a preset time window, the target user's gaze point data and heart rate data are collected through a preset device. The fixation point data and heart rate data corresponding to the target user are preprocessed, and the fixation point dispersion corresponding to the target user is calculated using the preprocessed fixation point data and heart rate data. Heart rate variability index ; Through formula Identify the load assessment index corresponding to the target user. ,in, These represent the weighting coefficients corresponding to fixation point dispersion and heart rate variability, respectively. Load assessment index corresponding to the target user Compared with the preset load assessment threshold Perform a comparison; If the target user's corresponding load assessment index At that time, the scene elements in the virtual scene are dynamically simplified.

6. A virtual reality-based interactive control method, applied to the virtual reality-based interactive control system described in any one of claims 1-5, characterized in that, Includes the following steps: The monitoring device captures the target user's limb spatial coordinate data, joint rotation angle data, and muscle activation signals. Action intent parsing includes a temporal convolutional neural network model and a dynamic weighted fusion unit. The temporal convolutional neural network model includes an input layer and an output layer. The input layer is used to receive the skeletal joint sequence, angular velocity data, and electromyographic feature vector from the multimodal motion capture module. The output layer is used to extract spatiotemporal features through a three-layer dilated causal convolution and connect to a Softmax classifier to generate initial action data corresponding to the target user. The dynamic weighted fusion unit is used to adjust the fusion weights of the data from each monitoring device in real time according to the environmental complexity parameter. The process of adjusting the fusion weights of data from each monitoring device in real time based on environmental complexity parameters specifically includes: Through formula Confirm the complexity of the environment ; in, These represent the density of virtual objects and the variance of illumination variation in the virtual scene, respectively. This represents the movement speed corresponding to the target user. Represented as preset weighting coefficients; Environmental complexity Compare with the preset environmental complexity threshold range; If environmental complexity In When the weights are in between, the weights are allocated according to the preset first allocation strategy; If environmental complexity In When the weights are in between, the weights are allocated according to the preset second allocation strategy; If environmental complexity In this case, the weights are allocated according to the preset third allocation strategy; Construct a human joint angle constraint library, judge the initial action data corresponding to the target user, and output correction instructions based on the judgment results; Multi-channel feedback includes a visually guided projection unit, a tactile vibration array, and a bone conduction audio unit. The visually guided projection unit is used to generate a semi-transparent reference motion trajectory in a virtual scene. The tactile vibration array is used to place linear resonant actuators on the target user and generate differentiated vibration patterns based on the motion error vector. The bone conduction audio unit is used to select feedback content based on a preset voice command library and according to the error type. Adjust the environment in the virtual scene.

7. A computer-readable storage medium, characterized in that: The system stores instructions that, when executed on a computer, cause the computer to perform an interactive control system based on virtual reality as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Virtual reality interactive training system and method based on multi-modal feedback

    CN119937798A