LBE enhanced multimodal perception interactive VR display control method and system

By collecting and fusion of multimodal data in large spatial scenarios and dynamically adjusting the range of activities and device responses, the problems of unsmooth interaction and user security risks in the existing technology are solved, and a more efficient and safer VR interactive experience is achieved.

CN119536526BActive Publication Date: 2025-05-06NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510098853.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is difficult to realize real-time fusion of multimodal data, accurate mapping of virtual and physical space, and dynamic adaptation of device responses in large-space scenarios, resulting in poor interaction, poor user security risks and poor experience.

Method used

Multimodal perceptual data is collected through VR headset equipment and hand controllers, and the Kalman filtering algorithm is used to fuse position and action data, calculate the user's continuous motion trajectory and dynamically adjust the range of activities. At the same time, intelligent algorithms are used to synchronize multimodal data, feature extraction and weighted fusion, interactive state vectors are generated, users' real-time behavior and physiological feedback are analyzed, and device response parameters are dynamically adjusted.

Benefits of technology

Real-time and precise fusion of multimodal data, accurate matching of virtual and physical space, and dynamic adaptation of device responses have been achieved, improving interaction flexibility, security and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119536526B_ABST
    Figure CN119536526B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of VR display control technology, and in particular to an LBE enhanced multimodal perception interactive VR display control method and system. The present invention proposes the following scheme: a user's multimodal perception data is collected through a VR head display device and a hand controller, and the position data and motion data are fused using a Kalman filter algorithm to calculate the user's continuous motion trajectory, and the activity range is dynamically adjusted based on the continuous motion trajectory to generate a directional dynamic range. The multimodal data is synchronized, feature extracted, and weighted fused to generate an interaction state vector. The user's real-time behavior, action intention, and physiological feedback are analyzed according to the interaction state vector, the user's operation sensitivity and emotional state are judged, and the response parameters of the VR device are dynamically adjusted. The present invention effectively improves the real-time, flexibility, and immersiveness of large-space VR interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of VR display control technology, and in particular to an LBE-enhanced multi-modal perception interactive VR display control method and system. Background Art

[0002] In large-space scenes based on location-based entertainment (LBE), VR interaction provides users with an immersive experience. However, current technological development still faces challenges such as difficulty in multimodal data fusion, inaccurate mapping of virtual and physical spaces, insufficient interactive response sensitivity, and static activity range management. First, VR interaction requires the collection of multimodal data from users, including position data, motion data, voice data, and tactile feedback data. However, due to the inconsistency of the time axis, format, and feature dimensions of different modal data, multimodal fusion is complex and has poor real-time performance. Secondly, in large-space scenes, it is difficult for existing technologies to dynamically adjust the activity range of the virtual environment based on the user's real-time motion trajectory. The mismatch between the virtual space and the physical space may lead to unsmooth interaction and user safety risks. In addition, the response sensitivity of the device is usually fixed, and there is a lack of dynamic adaptation mechanism based on user behavior, emotions, and physiological feedback. Users may have a poor experience when they are nervous or tired. Finally, the traditional activity range management method is based on a static boundary model, which is difficult to adapt to real-time changes in user behavior, limiting the freedom and accuracy of interaction. Therefore, there is an urgent need for a VR interactive control method that can integrate multimodal data, achieve accurate mapping of virtual and physical spaces, and optimize the dynamic response of the device, so as to solve the pain points in the existing technology and improve user experience and interaction efficiency.

[0003] For example, a Chinese patent with authorization announcement number CN109874003B discloses a VR display control method, a VR display control device, and a display device. The VR display control method includes: obtaining the user's action information; rendering the image data according to the action information, wherein the duration from the start time of obtaining the user's action information to the start time of rendering the image according to the action information is a preset duration; displaying according to the rendered image data; wherein, when the duration from the display of the next frame of the image is the preset duration, obtaining the user's next action information. The disclosed embodiment does not need to wait until the Vsync signal arrives before it can start obtaining the user's next action information, thus advancing the moment of obtaining the user's next action information, and also advancing the moment of rendering the image data according to the user's next action information, so that before the end time of displaying the Nth frame of the image, there is more time to render the image data according to the user's next action information.

[0004] The above-mentioned prior art has the problem raised by this background technology: it is difficult to meet the requirements of accuracy, flexibility and immersion of VR interaction in large space scenes. To solve the above problems, this application provides an LBE-enhanced multimodal perception interactive VR display control method and system. Summary of the invention

[0005] The technical problem to be solved by the present invention is to address the deficiencies of the prior art and provide an LBE enhanced multimodal perception interactive VR display control method and system. The multimodal perception data of the user is collected through a VR head display device and a hand controller, and the Kalman filter algorithm is used to fuse the position data and the motion data, calculate the user's continuous motion trajectory, and dynamically adjust the activity range based on the continuous motion trajectory to generate a directional dynamic range. The multimodal data is synchronized, feature extracted, and weighted fused to generate an interactive state vector. According to the interactive state vector, the user's real-time behavior, action intention, and physiological feedback are analyzed to determine the user's operating sensitivity and emotional state, and the response parameters of the VR device are dynamically adjusted.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] The LBE enhanced multimodal perception interactive VR display control method is applied to a VR interactive device in a large space scene, wherein the VR interactive device includes a real-time display VR head display device and a hand controller matched with the VR head display device, wherein the VR head display device is equipped with a space positioning sensor, a voice device and a visual sensor, and the hand controller is equipped with a motion capture sensor and a tactile feedback device, and the method includes:

[0008] Acquire the user's perception data, position data, and action data in a large space through the VR interactive device;

[0009] Positioning is performed according to the position data and motion data, and the range of motion is calculated and adjusted;

[0010] The perception data and action data are integrated and processed synchronously through intelligent algorithms to calculate the interaction state vector;

[0011] The real-time behavior and physiological feedback are analyzed according to the interaction state vector, and the response sensitivity of the VR interaction device is adjusted according to the analysis result.

[0012] The positioning according to the position data and the motion data includes:

[0013] Acquire the user's position data in the physical space through the spatial positioning sensor on the VR head display device, wherein the position data includes current position coordinates and relative displacement information;

[0014] Acquiring user motion data through a motion capture sensor on the hand controller, wherein the motion data includes hand position, direction, acceleration and motion trajectory;

[0015] According to the Kalman filter algorithm, the position data and the motion data are fused and processed to calculate the continuous motion trajectory of the user, wherein the continuous motion trajectory includes the current position, motion direction and motion speed.

[0016] The Kalman filtering algorithm comprises:

[0017] Initialize the user's global state parameters, which include the user's current position, movement speed, movement direction, and the noise covariance matrix of the position data and motion data, wherein the initialization sets the global position reference based on the initial position data obtained by the spatial positioning sensor of the VR head display device, and sets the local motion reference based on the initial motion data obtained by the hand controller;

[0018] According to the global state parameters, an adaptive state prediction model is constructed, and the state prediction model adjusts the state transfer matrix of the Kalman filter through a dynamic scene switching logic, and the dynamic scene switching logic includes:

[0019] The scene in which the user is located is determined according to the global state parameters. When the user is stationary, the position data is used as a reference and the motion data is used for correction. When the user moves or rotates, the motion data is used to provide dynamic updates of direction and acceleration, and the position data is used for drift correction.

[0020] The error between the predicted state and the actual measured value is evaluated according to the state transfer matrix, the noise covariance matrix is ​​compensated according to the error, and the continuous motion trajectory is calculated according to the compensated noise covariance matrix. The calculation formula of the continuous motion trajectory is:

[0021] ,

[0022] in, represents the continuous motion trajectory of the user corresponding to time k, represents the continuous motion trajectory of the user corresponding to the k-1 moment. When k=1, =0, represents the weight factor after compensation based on the noise covariance matrix, Indicates the action data influence weight adjusted according to the dynamic scene switching logic. represents the user position vector corresponding to time k, represents the user position vector corresponding to time k-1, where the position vector includes the current position coordinates, movement speed and direction. Indicates the current position increment, Indicates the acceleration update amount.

[0023] The calculation and adjustment of the activity range includes:

[0024] Generate an initial boundary of the user's activity range according to the actual boundary size of the physical space and the interaction requirements in the virtual environment, wherein the initial boundary includes a dynamic center point and a range radius;

[0025] According to the movement pattern in the user's continuous movement trajectory, adjust the expansion and contraction ratio of the user's activity range;

[0026] Taking the current position of the user in the continuous motion trajectory as the new center point of the activity range, updating the center position of the user's activity range, and based on the motion direction and motion speed at the center position, appropriately extending the boundary of the user's activity range in the first motion direction and appropriately shrinking it in other motion directions to form a directional dynamic range;

[0027] When the boundary of the user's activity range touches the actual boundary of the physical space due to adjustment, conflict handling and range correction are performed.

[0028] The intelligent algorithm comprises:

[0029] Time synchronization of the collected perception data and motion data, and noise filtering of the synchronized multimodal data;

[0030] Extract the angle change of the user's line of sight and the focus area from the visual data, and convert them into the focus area through encoding. Extract the movement speed, direction change and movement amplitude of the action data through trajectory analysis and encode them into behavioral features. Extract voice information from the voice data, and analyze the emotional features in combination with the voice information. Extract the feedback intensity change features from the tactile feedback data.

[0031] A feature fusion network is constructed, and the focus area, behavior feature, emotion feature and feedback intensity change feature are used as input parameters of the feature fusion network. The input parameters are trained through the feature fusion network to output an interaction state vector.

[0032] The extracting of voice information through voice data includes:

[0033] Extract features from speech data, convert the speech data into a two-dimensional matrix, and calculate the feature vector sequence of each frame based on the Mel-frequency cepstral coefficient features of the first N frames and the last N frames of each frame;

[0034] According to the encoder, the feature vector sequence of the speech data is converted into an audio coding feature vector, and the audio coding feature vector is processed by GRU to calculate the speech data content vector;

[0035] The speech data content vector is reconstructed, and the reconstructed speech data content vector is decoded by a decoder to generate a semantic feature vector of the speech data.

[0036] The feature fusion network comprises:

[0037] The input layer is used to format the received multimodal features;

[0038] A feature extraction layer, used to extract high-dimensional features of the output of the input layer;

[0039] A feature weight allocation layer, used for adaptively allocating weights of the high-dimensional features through an attention mechanism, wherein the attention mechanism includes a modal attention mechanism, a channel attention mechanism, and a temporal attention mechanism;

[0040] The feature fusion layer is used to perform weighted fusion processing on the multimodal features after weight assignment, compress and map the fused features through the fully connected layer operation, and generate an interaction state vector.

[0041] The analysis of real-time behavioral and physiological feedback includes:

[0042] Extract the user's action intention and operation trend based on the interaction state vector;

[0043] According to the physiological feedback characteristics in the interaction state vector, the user's emotional state and interaction adaptability are judged;

[0044] Determine whether the user's current operation sensitivity needs to be adjusted based on the user's action intention, operation trend, and emotional state;

[0045] When the user's interactive adaptability shows fatigue or excessive tension, the response strength of the VR interactive device is dynamically adjusted.

[0046] An LBE-enhanced multimodal perception interactive VR display control system, the system comprising a data acquisition module, a range adjustment module, a data processing module and an analysis and response module;

[0047] The data acquisition module is used to collect the user's multimodal perception data through the VR head display device and the hand controller;

[0048] The range adjustment module is configured with a spatial positioning strategy, which is used to perform positioning according to the position data and the action data, and calculate and adjust the user's activity range;

[0049] The data processing module is configured with a multimodal data fusion strategy, which is used to perform synchronous processing, feature extraction and fusion calculation on the collected multimodal data to generate an interactive state vector;

[0050] The analysis and response module is used to analyze the user's real-time behavior and physiological feedback according to the interaction state vector, and dynamically adjust the response parameters of the VR interaction device according to the user state analysis results.

[0051] The range adjustment module comprises:

[0052] The positioning unit is used to fuse the position data with the motion data according to the Kalman filter algorithm to calculate the continuous motion trajectory of the user;

[0053] A dynamic range adjustment unit, used to dynamically adjust the range of motion according to the user's continuous motion trajectory;

[0054] A boundary prompt unit, used to detect the overlap between the activity range boundary and the actual boundary of the physical space, and trigger a prompt when approaching the boundary;

[0055] The spatial positioning strategy includes trajectory calculation logic and range adjustment logic. The trajectory calculation logic is configured in the positioning unit and is used to perform a fusion calculation on the user's current position, movement direction and movement speed. The range adjustment logic is configured in the dynamic range adjustment unit and is used to dynamically update the center point and boundary shape of the activity range based on the user's continuous movement trajectory.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] 1. The present invention uses intelligent algorithms to synchronously process and fuse features of multimodal data, constructs an interactive state vector that can characterize user behavior, emotions, and preferences in real time, realizes comprehensive perception of user operation intentions and physiological states, and improves the accuracy of data processing and the response efficiency of the system.

[0058] 2. The present invention proposes a dynamic activity range adjustment strategy, including updating the center point of the activity range, dynamic boundary expansion and directional adjustment, which can accurately match the virtual interaction area with the user's actual motion state. At the same time, through boundary detection and warning mechanisms, the risk of users exceeding the actual boundaries of the physical space is effectively avoided, ensuring the safety and freedom of interaction in large space scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0060] Figure 1 Schematic diagram of the flow of the LBE enhanced multimodal perception interactive VR display control method in Example 1 of the present invention;

[0061] Figure 2This is a schematic diagram of the motion trajectory extraction process of Example 1 of the present invention;

[0062] Figure 3 This is a flow chart of the state prediction model of Example 1 of the present invention;

[0063] Figure 4 This is a schematic diagram of the process of extracting a region of interest according to Embodiment 1 of the present invention;

[0064] Figure 5 This is a schematic diagram of the behavior feature extraction process in Example 1 of the present invention;

[0065] Figure 6 This is a flow chart of speech feature extraction in Example 1 of the present invention;

[0066] Figure 7 This is a feature fusion network structure diagram of Example 1 of the present invention;

[0067] Figure 8 This is a module diagram of the LBE enhanced multimodal perception interactive VR display control system in Example 2 of the present invention. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0069] Embodiment 1:

[0070] See also Figure 1 The present invention provides an embodiment: an LBE enhanced multi-modal perception interactive VR display control method, the specific steps of the method are as follows:

[0071] S1: Obtain the user's perception data, location data, and action data in a large space;

[0072] In this step, multimodal sensors are used to obtain the user's perception data, location data, and motion data in real time to establish a comprehensive perception of the user's behavior and environment. Multimodal perception can capture the user's all-round information in a large space, providing multi-dimensional data support for subsequent processing. At the same time, it solves the blind spot problem that may be caused by a single data source and improves the integrity of the data and the accuracy of the interactive response.

[0073] S2: Positioning the user's motion trajectory according to the position data and the motion data;

[0074] In this step, by fusing position data and motion data, the user's motion trajectory is located in real time using methods such as Kalman filtering to obtain the user's current position, movement direction, and speed. Accurate motion trajectory positioning can reflect the user's behavioral characteristics in the physical space in real time, reduce virtual interaction offset problems caused by sensor errors or delays, and ensure the synchronization and consistency of virtual space and physical space.

[0075] S3: Adjust the activity range according to the user's movement trajectory;

[0076] In this step, the activity range is dynamically adjusted according to the user's motion trajectory, including updating the center point, boundary shape and size of the range, to ensure that the activity range adapts to the user's motion state and environmental changes. Dynamic adjustment of the activity range can make the user's interaction area match the actual motion situation, which not only improves the freedom of large-space interaction, but also prevents users from exceeding the physical space range through boundary management, thereby improving the safety of interaction.

[0077] S4: Synchronize and fuse the perception data and action data, and calculate the interaction state vector;

[0078] In this step, the collected perception data and action data are time synchronized and noise filtered, and the interaction state vector describing the user's behavior, emotions, and preferences is generated through a fusion algorithm.

[0079] S5: Analyze the user's real-time behavior and physiological feedback according to the interaction state vector;

[0080] In this step, the user's behavior pattern and physiological state are analyzed based on the interaction state vector, including action intention, operation trend, emotional state and fatigue level.

[0081] S6: Adjust the response sensitivity of the VR interactive device according to the analysis results;

[0082] In this step, the response sensitivity of the VR device is dynamically adjusted according to the results of behavioral analysis and physiological feedback, including the action response speed of the hand controller, the tactile feedback strength, and the response time of the virtual interactive object.

[0083] The specific steps of S2 are as follows:

[0084] S2.1: Obtaining the user's position data in the physical space through a spatial positioning sensor on the VR head display device, wherein the position data includes current position coordinates and relative displacement information;

[0085] Spatial positioning sensors can use ultra-wideband (UWB), infrared, lidar or IMU sensors to collect real-time location information of users in physical space. The acquisition frequency is controlled between 30Hz and 100Hz according to actual conditions to ensure the timeliness and accuracy of the data. The current position coordinates are used to define the absolute position of the user in the physical space as a global reference for the motion trajectory. The relative displacement information is recorded through the sensor's time difference measurement algorithm (such as the integral calculation of the IMU) to record the user's displacement change per unit time and generate a local trajectory.

[0086] S2.2: obtaining the user's motion data through a motion capture sensor on the hand controller, wherein the motion data includes hand position, direction, acceleration and motion trajectory;

[0087] Motion capture sensors use IMU sensors, optical tracking or ultrasonic positioning technology to record the user's hand movement information in real time. The motion data can supplement the subtle motion features that spatial positioning sensors cannot capture, such as local motion changes when the user moves their arms quickly or rotates their body. By recording acceleration, direction and trajectory, the user's real movement can be restored more comprehensively.

[0088] S2.3: According to the Kalman filter algorithm, the position data and the motion data are fused and processed to calculate the continuous motion trajectory of the user, wherein the continuous motion trajectory includes the current position, the motion direction and the motion speed.

[0089] The fusion calculation of Kalman filtering can find a balance between the data weights of different sensors and provide a stable and high-frequency motion trajectory output. By dynamically adjusting the weights of the predicted value and the observed value, drift and jitter can be eliminated in real time to ensure the continuity and smoothness of the motion trajectory. The continuous motion trajectory provides the user's precise motion state in the physical space, which is used to dynamically adjust the range of activities, map the user's current position, speed and direction to the virtual space, and ensure the synchronization of virtual and physical movements.

[0090] See also Figure 2 , a schematic diagram of the motion trajectory extraction process of an embodiment of the present invention, the specific steps of S2.3 are as follows:

[0091] S2.3.1: Initialize the user's global state parameters, which include the user's current position, movement speed, movement direction, and the noise covariance matrix of the position data and motion data, wherein the initialization sets the global position reference based on the initial position data obtained by the spatial positioning sensor of the VR head display device, and sets the local motion reference based on the initial motion data obtained by the hand controller;

[0092] The noise covariance matrix initializes the error distribution of the position data according to the noise characteristics of the spatial positioning sensor, and initializes the noise covariance in the motion data according to the characteristics of the motion capture sensor;

[0093] Specifically, initializing the global state parameters is the starting point of the positioning calculation. By clarifying the initial position and motion reference, a global benchmark is established to ensure the accuracy of subsequent motion trajectory calculations. By setting the noise covariance matrix, the sensor error can be preliminarily quantified, providing the necessary error estimation input for subsequent Kalman filter processing.

[0094] S2.3.2: Based on the global state parameters, an adaptive state prediction model is constructed. The state prediction model adjusts the state transfer matrix of the Kalman filter through a dynamic scene switching logic. The dynamic scene switching logic includes:

[0095] The scene where the user is located is determined according to the global state parameters. When the user is stationary, the position data is used as a reference and the motion data is used for correction. When the user moves or rotates quickly, the motion data is used to provide dynamic updates of direction and acceleration, and the position data is used for drift correction.

[0096] The motion characteristics of users in different scenarios (stationary, fast moving, rotating) are different, so it is necessary to dynamically adjust the parameters of the state prediction model to improve the prediction accuracy. By adjusting the state transfer matrix through dynamic scene switching logic, the real-time and accuracy of trajectory calculation can be adaptively optimized. The speed value in the position data is used to determine whether the user is stationary, the acceleration of the motion data is used to determine whether the user is moving fast, and the direction change rate of the hand controller is used to determine whether the user is in a rotating state. When stationary, the reference weight of the position data is strengthened to reduce the drift effect of the motion data. When moving fast, the motion data is used as the main method, and the user's motion direction and speed are updated in real time. When rotating, the direction weight is adjusted to ensure the continuity of the rotation angle change.

[0097] S2.3.3: Evaluate the error between the predicted state and the actual measured value according to the state transfer matrix, compensate the noise covariance matrix according to the error, and calculate the continuous motion trajectory according to the compensated noise covariance matrix. The calculation formula of the continuous motion trajectory is:

[0098] ,

[0099] in, represents the continuous motion trajectory of the user corresponding to time k, represents the continuous motion trajectory of the user corresponding to the k-1 moment. When k=1, =0, represents the weight factor after compensation based on the noise covariance matrix, Indicates the action data influence weight adjusted according to the dynamic scene switching logic. represents the user position vector corresponding to time k, represents the user position vector corresponding to time k-1, where the position vector includes the current position coordinates, movement speed and direction. Indicates the current position increment, Indicates the acceleration update amount;

[0100] In the Kalman filter algorithm, the error between the predicted state and the actual measured value is the key to correcting the state estimation. By evaluating this error, the degree of deviation between the state prediction model and the actual situation can be determined, and the noise covariance matrix reflects the uncertainty in the measurement process. According to the error adjustment matrix, the filter's adaptability to dynamic environments and sensor noise can be adaptively improved. Through the compensated noise covariance matrix, the user's state can be estimated more accurately, thereby calculating an accurate continuous motion trajectory. This is crucial for VR systems to accurately reflect the user's movement in physical space in real time.

[0101] See also Figure 3 , a flow chart of the state prediction model of an embodiment of the present invention, specifically, based on the user's last known state and the state transition rule of the Kalman filter algorithm, the state of the user at the next moment is predicted. The state prediction here combines the user's motion model in space. For example, when the user moves freely or rotates quickly, the model will consider the characteristics of linear motion and angular velocity respectively. Further, the user's historical position information, speed, direction and environmental control parameters are combined to predict the user's possible current position and direction. For example, if the user has moved in a certain direction in the last few steps, it will be inferred that the user continues to move in this direction. The predicted state is compared with the current position, speed and direction measured by the actual sensor. For example, the spatial positioning sensor provides the current position, the motion capture sensor provides the direction and speed, and the difference between the two is calculated. This difference is called "prediction error", which represents the deviation between the Kalman filter algorithm and the real world. If the difference between the user's current position measured by the sensor and the predicted position is small, it means that the prediction model is more accurate, and the last correction result is output; if the difference is large, it may be caused by noise or model deviation, and correction needs to be made according to the error.

[0102] The specific steps of S3 are as follows:

[0103] S3.1: Generate an initial boundary of the user's activity range according to the actual boundary size of the physical space and the interaction requirements in the virtual environment, wherein the initial boundary includes a dynamic center point and a range radius;

[0104] The initial boundary of the activity range is used to define the user's interactive area in the virtual environment, ensuring that the activity range matches the actual boundary of the physical space to prevent the user from crossing the boundary, and providing a benchmark for subsequent dynamic adjustments. The initial boundary sets the dynamic center point and range radius, and the range size can be flexibly adjusted according to the user's current actual location and preset interaction requirements.

[0105] S3.2: adjusting the expansion and contraction ratio of the user's activity range according to the movement pattern in the user's continuous movement trajectory;

[0106] The user's movement pattern changes dynamically, and the range of motion needs to adapt to the user's behavior in real time. For example, when the user moves quickly, a larger range is needed, while when moving slowly or stationary, the range can be reduced to save computing resources. By adjusting the expansion and contraction ratio, the range of motion can be flexibly adapted to the user's movement needs;

[0107] Specifically, the speed and direction change frequency in the user's continuous motion trajectory is used as the basis for judging the motion mode. When maintaining stable motion, the consistency of motion direction and speed stability are calculated, the range expansion logic is triggered, and the radius of the activity range is increased. When accelerating motion or turning frequently, the acceleration fluctuation and direction switching frequency are calculated, the range contraction logic is triggered, and the activity range is reduced to improve the response speed of the system. When pausing or moving slowly, the situation where the speed approaches zero in the motion trajectory is detected, and the range is gradually contracted to the initial boundary. Through dynamic expansion and contraction, the activity range can always match the user's current motion needs, which improves the synchronization between the virtual space and user behavior and effectively saves computing resources and the scheduling cost of interactive objects.

[0108] S3.3: Taking the current position of the user in the continuous motion trajectory as the new center point of the activity range, update the center position of the user's activity range to ensure that the activity range is adjusted in real time around the user's actual motion position. Based on the motion direction and motion speed at the center position, appropriately extend the boundary of the user's activity range in the first motion direction. The first motion direction is the main direction of the user's current motion, which can be calculated by the speed vector in the continuous motion trajectory, and appropriately shrink in other motion directions to form a directional dynamic range.

[0109] The user's current position and movement direction determine the distribution characteristics of the activity range. By taking the current position as the center and the directional dynamic range as the boundary, the user's interaction space in the first movement direction can be optimized while reducing the computational overhead of non-essential areas.

[0110] Specifically, in the first direction of motion, the extension amount is calculated according to the current motion speed and the extension scale factor, and in the secondary direction perpendicular to the first direction of motion, the range is contracted according to the contraction scale factor, wherein the extension scale factor and the contraction scale factor can be calculated by those skilled in the art through a large number of experiments, and the calculation formula of the directional dynamic range is:

[0111] ,

[0112] in, Indicates the direction angle of the current boundary, Indicates the dynamic radius in the direction angle, forming a directional dynamic activity range. Indicates the basic radius, that is, the initial radius of the user's activity range. represents the stretching scale factor, Indicates the user's current speed. represents the first motion direction angle, Represents the cosine value between the current boundary direction angle and the first motion direction angle, which is used to quantify the correlation between the current direction and the first motion direction. Represents the shrinkage factor.

[0113] S3.4: When the user activity range boundary contacts the actual boundary of the physical space due to adjustment, conflict handling and range correction are performed, including:

[0114] Detect whether the future position of the user's continuous motion trajectory exceeds the actual boundary of the physical space;

[0115] If the future location will be beyond the boundary, ensure that the adjusted boundary is consistent with the actual boundary of the physical space by reducing the activity range radius or changing the boundary shape;

[0116] Dynamic boundary lines are displayed in the virtual environment to prompt users that they are approaching the actual boundaries of the physical space, and tactile feedback devices are used to remind users to adjust their movement behavior.

[0117] The user's activity range may be close to the actual boundary of the physical space (such as walls or furniture). If not handled in time, it may cause the risk of users crossing the boundary or colliding. Therefore, it is necessary to ensure that the activity range is consistent with the actual boundary of the physical space through boundary detection and correction logic.

[0118] The specific steps of S4 are as follows:

[0119] S4.1: Time synchronization of the collected visual data, motion data, voice data and tactile feedback data, and noise filtering of the synchronized multimodal data;

[0120] Multimodal data (such as vision, motion, voice, and touch) comes from different sensors. Due to differences in sampling rates and data transmission delays, data alignment issues may occur on the time axis. If these data are not synchronized in time, visual and motion information may not correspond, which in turn affects the fusion effect of multimodal data. Noise filtering is to reduce data noise caused by environmental interference or device jitter during sensor capture and improve data reliability.

[0121] Specifically, a global timestamp mechanism is used to align multimodal data. Each data frame is assigned a unified time stamp, and linear interpolation is used to synchronize data of different frequencies to the same time axis. For example, the visual data (sampling frequency 60Hz) and the motion data (sampling frequency 100Hz) are unified to 60Hz, the motion data values ​​of unsampled points are estimated by interpolation, and the data is preprocessed with a low-pass filter to remove high-frequency noise (such as hand shaking data in motion capture or random noise in visual sensors).

[0122] S4.2: Extract the angle change of the user's line of sight and the focus area from the visual data, and convert them into the focus area through encoding;

[0123] The user's line of sight and gaze focus are important bases for understanding the user's intention, and can reflect the virtual object or area that the user is currently focusing on. This is of great significance for identifying the user's operation target and optimizing the virtual interaction response.

[0124] See also Figure 4 , a schematic diagram of the process of extracting the region of interest according to an embodiment of the present invention, the specific steps of S4.2 are as follows:

[0125] S4.2.1: Use the built-in visual sensor of the VR headset to capture the user's eye features in real time, including the location of the pupil center and corneal reflection point;

[0126] Specifically, eye tracking technology is used to illuminate the eyeball with infrared light, detect the position of the pupil center and the corneal reflection point, and establish an eye movement vector, wherein the eye movement vector is used to describe the movement direction and change trend of the eyeball;

[0127] S4.2.2: Calculate the direction of the sight line vector in the three-dimensional coordinate system of the VR headset device based on the change in the angle of the user's eye movement;

[0128] Specifically, the VR head display device's posture sensor (such as a gyroscope and an accelerometer) is combined to correct the deviation between the eye movement vector and the device's spatial posture to obtain the sight vector;

[0129] S4.2.3: Use the relationship between the eye movement vector and the gaze vector and the position of objects in the virtual scene to determine the specific area of ​​the gaze focus;

[0130] Specifically, the raycasting technology is used to detect the intersection of the sight vector and the virtual object, and the intersection is matched with the boundary of the object in the virtual scene to identify the object or area currently being looked at. If the sight does not match the intersection of any object, it is marked as a blank looking area.

[0131] S4.2.4: Quantify and encode the identified focus points to form the features of the attention area. The encoding content includes:

[0132] The ID (unique identifier) ​​of the gazed object;

[0133] The coordinates of the gaze point in the virtual scene;

[0134] Fixation duration (calculated by accumulating time between frames).

[0135] S4.3: Extract the movement speed, direction change and movement amplitude of the action data through trajectory analysis and encode them into behavioral features;

[0136] The user's motion data directly reflects their behavioral characteristics in the virtual space, such as hand grabbing, sliding or pointing at an object. These characteristics are the core basis for VR devices to generate real-time interactive feedback.

[0137] See also Figure 5 , a flow chart of behavior feature extraction according to an embodiment of the present invention, the specific steps of S4.3 are as follows:

[0138] S4.3.1: Use motion capture sensors to collect the position information of the user's hands in real time and record the continuous path of the movements in space;

[0139] The data stream is saved in the form of a time series. Each frame contains the position, direction vector and timestamp. The trajectory data is interpolated to fill the path gaps caused by different sampling rates and ensure the continuity of the trajectory.

[0140] S4.3.2: Use a sliding time window to record the average speed and speed change trend in each time period, such as detecting a sudden increase or decrease in speed, which indicates a rapid switch or stop of the action, analyzing the spatial displacement of the action trajectory in continuous time intervals, and extracting the user's motion speed characteristics;

[0141] S4.3.3: Continuously analyze the direction of the user's hand trajectory and identify the user's movement direction based on the change angle of the direction vector. If the angle change between two consecutive directions exceeds a set threshold (such as the user changes from a straight-line movement to a sharp turn), it is recorded as a significant direction change. Combined with the movement speed characteristics, further determine whether the direction switch is related to the operation target, such as grabbing a virtual object or pointing to a certain area.

[0142] S4.3.4: Based on trajectory analysis, calculate the motion amplitude characteristics of the user's hand movement, including the path length and the intensity of the acceleration phase. The path length is a direct reflection of the user's operating range and is used to determine whether the action covers a specific target area. The time of the acceleration phase is recorded to extract the characteristics of high-intensity actions such as rapid waving or tapping.

[0143] S4.3.5: Normalize the extracted motion speed features, direction changes, and motion amplitude characteristics to eliminate data differences caused by different user operation habits, and classify the feature data according to the complexity of the operation. For example, define "rapid linear movement" as a simple operation and "rapid direction switching combined with large-amplitude movements" as a complex operation. Finally, encode these characteristics into behavioral features that describe the user's current operation behavior and output feature vectors in a fixed format.

[0144] S4.4: Extract voice information through voice data and analyze emotional characteristics based on voice information;

[0145] The user's voice not only conveys operational instructions, but also contains emotional information, which can help judge the user's current mental state, such as happiness, tension or fatigue. The voice emotional characteristics can supplement behavioral data, help understand the user's current emotional state, and avoid misjudging the user's interaction needs.

[0146] See also Figure 6 , a schematic diagram of the emotion feature extraction process according to an embodiment of the present invention, the specific steps of S4.4 are as follows:

[0147] S4.4.1: Extract features from the speech data. The feature extraction includes speech data emphasis, framing, windowing, Fourier transform, triangular bandpass filtering, and logarithmic calculation. The speech data is converted into a two-dimensional matrix. The feature vector sequence of each frame is calculated based on the Mel cepstral coefficient features of the first N frames and the last N frames of each frame. The feature vector sequence calculation formula of the speech data is:

[0148] ,

[0149] in, represents the feature vector of speech data, ln(•) represents the natural logarithm function, f represents the average frequency of speech data, Represents the nth frame sample of speech data, represents the weighting factor, represents the n-1th frame sample of speech data, F(•) represents fast Fourier transform, Represents the Mel cepstral coefficient features of the first N frames of the nth frame of the speech data. Represents the Mel-frequency cepstral coefficient features of the next N frames of the nth frame of the speech data;

[0150] Pre-emphasize the original speech signal, amplify the high-frequency components, balance the high- and low-frequency components in the speech signal, and enhance the speech recognition.

[0151] The speech data is divided into short-time frames, and the impact of data mutations between frames is reduced by adding windows (such as Hamming windows), making it more suitable for short-time Fourier transform analysis.

[0152] Perform fast Fourier transform (FFT) on the windowed speech signal to convert the time domain signal into a frequency domain signal and extract the frequency component.

[0153] The frequency domain signal is decomposed into multiple sub-bands through a triangular bandpass filter, and its energy spectrum is calculated. The logarithm of the energy value is taken to form a sensitivity feature to the pitch change of the speech signal.

[0154] S4.4.2: According to the encoder, the feature vector sequence of the speech data is converted into an audio coding feature vector, and the audio coding feature vector generates a hidden vector through a gated recurrent unit , another gated recurrent unit calculates in the reverse direction to generate the hidden vector , concatenate the two hidden vectors and obtain the speech data content vector through the fully connected layer;

[0155] The encoder's task is to deeply compress and represent the feature vector sequence of speech data and extract audio coding features related to timing. The gated recurrent unit (GRU) is used to process the time series characteristics of speech. The forward GRU processes the time series relationship from the beginning to the end, and the reverse GRU processes the time dependency from the end to the beginning. The output hidden vector of the bidirectional GRU can capture the dependency of the global time series in speech and overcome the problem of information loss in traditional RNN time series analysis. After the hidden vectors in the forward and reverse directions are concatenated, they are nonlinearly mapped through the fully connected layer to obtain the speech data content vector.

[0156] Bidirectional GRU can capture complex contextual relationships in speech features, avoid the loss of key information, extract more complete speech content features, and provide high-quality basic data for semantic and sentiment analysis.

[0157] S4.4.3: Calculate the attention weight through the attention network, reconstruct the speech data content vector, decode the reconstructed speech data content vector through the decoder, and generate the semantic feature vector of the speech data;

[0158] The task of the attention network is to find the most important part from the audio encoding feature vector and perform weighted reconstruction. The specific process includes:

[0159] Use the attention mechanism to calculate a weight for each time step of the speech feature, and the weight value represents the importance of the time step;

[0160] Perform weighted summation of feature vectors according to weights, reduce the weights of unimportant time steps, and retain only key information;

[0161] The decoder performs denoising and optimization on the reconstructed features to generate a semantic feature vector.

[0162] S4.4.4: Process the semantic feature vector frame by frame through the long short-term memory neural network to extract the time series characteristics related to emotions in the speech signal, perform dimensionality reduction processing on the time series characteristics through the fully connected layer, and map them to the emotion space. During the dimensionality reduction process, calculate the probability distribution of each emotion category (such as anger, happiness, calmness, anxiety) through the Softmax classifier to generate an emotion feature classification vector;

[0163] LSTM can capture dependencies over long time spans and extract dynamic change characteristics related to emotions in speech signals, such as volume change trends and intonation fluctuation patterns. Time series feature extraction combined with emotion classification can comprehensively analyze the emotional information in speech signals, including both categories (such as anger and happiness) and intensity (intense or mild).

[0164] S4.4.5: Analyze the volume, intonation and speech speed characteristics of the semantic feature vector by tensor analysis, calculate the emotion intensity index, and generate an emotion feature output by combining the emotion intensity index and the emotion feature classification vector;

[0165] The physical characteristics of speech (such as volume, intonation, and speaking speed) are quantified using tensor analysis to calculate emotional intensity indicators, including:

[0166] Calculate the average amplitude and fluctuation range of the speech signal to determine the intensity of the tone;

[0167] Analyze frequency change trends and quantify the degree to which the tone rises or falls;

[0168] The speech rate feature is calculated based on the speech frame length and time span to determine the tense or calm state.

[0169] S4.5: Extract feedback intensity variation features from tactile feedback data;

[0170] The intensity of tactile feedback can reflect the user's interaction preferences, such as the grasping force of an object or changes in tactile sensitivity, which is of great significance for adjusting the response parameters of the tactile device.

[0171] By recording the output intensity value of the tactile feedback device and analyzing the intensity change trend over time, the feedback intensity change characteristics can reflect the user's action details and interaction needs, and support personalized tactile feedback optimization.

[0172] S4.6: Construct a feature fusion network, use the focus area, behavior characteristics, emotion characteristics and feedback intensity change characteristics as input parameters of the feature fusion network, train the input parameters through the feature fusion network, and output an interaction state vector.

[0173] See also Figure 7 , a structural diagram of a feature fusion network according to an embodiment of the present invention, wherein the feature fusion network comprises:

[0174] The input layer is used to format the received multimodal features;

[0175] In multimodal perception data, each modality has different feature dimensions, formats, and time axes (e.g., visual data is in the form of images, motion data is in the form of vectors, and speech data is in the form of time series). Direct processing will lead to incompatible features or difficulty in fusion. The input layer standardizes and formats these features to provide a unified data structure for subsequent feature extraction.

[0176] The feature extraction layer is used to extract high-dimensional features of the output of the input layer, including:

[0177] Perform convolution feature extraction on the focus area features of the visual data, extract the spatial features (such as target shape, size and position) within the focus area, and generate a visual feature map;

[0178] The trajectory coding algorithm is used to analyze user behavior characteristics (such as hand movement trajectory) and extract movement patterns, including acceleration changes, movement direction deviation, etc. The time series modeling method (such as LSTM) is combined to extract time correlation and generate movement pattern feature vectors.

[0179] The speech data is processed using a recursive neural network (RNN) to capture the dynamic changes of speech emotion characteristics. Through the stacking structure of a multi-layer recursive network, the frame-by-frame change pattern of speech emotion is extracted to generate a feature vector reflecting the emotion.

[0180] Perform time series analysis on tactile feedback data, extract the variation pattern of tactile response intensity, and generate a feature vector representing the user's tactile preference (such as pressure intensity and feedback time);

[0181] Different modal data contain different forms of information. For example, visual data is mainly spatial features, motion data is time-varying features, and speech and tactile feedback are time series features. Therefore, it is necessary to design a dedicated feature extraction method for each modality to maximize the extraction of its core features.

[0182] A feature weight allocation layer, used to adaptively allocate the weights of the high-dimensional features through an attention mechanism, wherein the attention mechanism includes a modal attention mechanism, a channel attention mechanism, and a temporal attention mechanism. The modal attention mechanism is used to calculate the modal weight according to the global contribution of different modal features. The channel attention mechanism is used to calculate the dynamic weight of the feature channel according to the activation value of each modal feature channel, and to reduce the weight of the low-correlated channels. The temporal attention mechanism is used to calculate the dynamic change weight of the multimodal features in the time series;

[0183] Different modalities have different importance in different interaction scenarios. For example, visual feature maps may be key in object detection, but may be secondary in gesture recognition. Therefore, it is necessary to dynamically adjust the contribution of each modality feature to the final fusion result through a weight distribution mechanism.

[0184] The feature fusion layer is used to perform weighted fusion processing on the multimodal features after weight assignment, compress and map the fused features through the fully connected layer operation, and generate an interaction state vector.

[0185] When fusing multimodal features, it is necessary to retain the core information of each modality while reducing redundancy to ensure that the fused feature vector is compact and efficient, which facilitates subsequent state modeling and response optimization.

[0186] The specific steps of S5 are as follows:

[0187] S5.1: Extract user action intention and operation trend based on interaction state vector;

[0188] S5.2: Determine the user's emotional state and interaction adaptability based on the physiological feedback features in the interaction state vector;

[0189] S5.3: Determine whether the user's current operation sensitivity needs to be adjusted based on the user's action intention, operation trend and emotional state;

[0190] S5.4: Dynamically adjust the response strength of the VR interactive device when the user’s interactive adaptability indicates fatigue or excessive tension.

[0191] Embodiment 2:

[0192] See also Figure 8 , the present invention provides an embodiment: an LBE enhanced multi-modal perception interactive VR display control system, the system comprising a data acquisition module, a range adjustment module, a data processing module and an analysis response module;

[0193] The data acquisition module is used to collect the user's multimodal perception data through the VR head display device and the hand controller;

[0194] The range adjustment module is configured with a spatial positioning strategy, which is used to perform positioning according to the position data and the action data, and calculate and adjust the user's activity range;

[0195] The data processing module is configured with a multimodal data fusion strategy, which is used to perform synchronous processing, feature extraction and fusion calculation on the collected multimodal data to generate an interactive state vector;

[0196] The analysis and response module is used to analyze the user's real-time behavior and physiological feedback according to the interaction state vector, and dynamically adjust the response parameters of the VR interaction device according to the user state analysis results.

[0197] The range adjustment module comprises:

[0198] The positioning unit is used to fuse the position data with the motion data according to the Kalman filter algorithm to calculate the continuous motion trajectory of the user;

[0199] A dynamic range adjustment unit, used to dynamically adjust the range of motion according to the user's continuous motion trajectory;

[0200] A boundary prompt unit, used to detect the overlap between the activity range boundary and the actual boundary of the physical space, and trigger a prompt when approaching the boundary;

[0201] The spatial positioning strategy includes trajectory calculation logic and range adjustment logic. The trajectory calculation logic is configured in the positioning unit and is used to perform a fusion calculation on the user's current position, movement direction and speed. The range adjustment logic is configured in the dynamic range adjustment unit and is used to dynamically update the center point and boundary shape of the activity range based on the user's continuous movement trajectory.

[0202] The data processing module comprises:

[0203] Data synchronization unit, used to align and synchronize the collected multi-modal data in time axis, eliminating data inconsistency caused by sensor acquisition delay and frequency difference;

[0204] A feature extraction unit is used to extract high-dimensional features of each modality from the synchronized multimodal data for subsequent fusion processing;

[0205] The feature fusion unit is used to perform weighted fusion on the extracted multimodal high-dimensional features to generate a unified interaction state vector.

[0206] The analysis response module comprises:

[0207] Behavior analysis unit, which analyzes the user's real-time behavior pattern and operation intention based on the interaction state vector;

[0208] A physiological feedback analysis unit, which analyzes the user's emotional state and interaction adaptability according to the physiological feedback features in the interaction state vector;

[0209] The response optimization unit dynamically adjusts the response parameters of the VR interactive device based on the analysis results of behavioral and physiological feedback.

[0210] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. An LBE enhanced multimodal perception interactive VR display control method is applied to a VR interactive device in a space scene, wherein the VR interactive device comprises a real-time display VR head display device and a hand controller matched with the VR head display device, wherein the VR head display device is equipped with a space positioning sensor, a voice device and a visual sensor, and the hand controller is equipped with a motion capture sensor and a tactile feedback device, and is characterized in that: The method comprises: Acquire the user's perception data, position data and action data in space through the VR interactive device; According to the position data and motion data, positioning is performed through a Kalman filter algorithm, and the range of activity is calculated and adjusted; The perception data and action data are synchronized and integrated through intelligent algorithms to calculate the interaction state vector; Analyze the real-time behavior and physiological feedback according to the interaction state vector, and adjust the response sensitivity of the VR interaction device according to the analysis result; The Kalman filtering algorithm comprises: Initialize the user's global state parameters, which include the user's current position, movement speed, movement direction, and the noise covariance matrix of the position data and motion data, wherein the initialization sets the global position reference based on the initial position data obtained by the spatial positioning sensor of the VR head display device, and sets the local motion reference based on the initial motion data obtained by the hand controller; According to the global state parameters, an adaptive state prediction model is constructed, and the state prediction model adjusts the state transfer matrix of the Kalman filter through a dynamic scene switching logic, and the dynamic scene switching logic includes: The scene in which the user is located is determined according to the global state parameters. When the user is stationary, the position data is used as a reference and the motion data is used for correction. When the user moves or rotates, the motion data is used to provide dynamic updates of direction and acceleration, and the position data is used for drift correction. The error between the predicted state and the actual measured value is evaluated according to the state transfer matrix, the noise covariance matrix is ​​compensated according to the error, and the continuous motion trajectory is calculated according to the compensated noise covariance matrix. The calculation formula of the continuous motion trajectory is: , in, represents the continuous motion trajectory of the user corresponding to time k, represents the continuous motion trajectory of the user corresponding to the k-1 moment. When k=1, =0, represents the weight factor after compensation based on the noise covariance matrix, Indicates the influence weight of action data adjusted according to dynamic scene switching logic. represents the user position vector corresponding to time k, represents the user position vector corresponding to time k-1, where the position vector includes the current position coordinates, movement speed and direction. Indicates the current position increment, Indicates the acceleration update amount.

2. According to claim 1, the LBE enhanced multi-modal perception interactive VR display control method is characterized in that: The positioning according to the position data and the motion data includes: Acquire the user's position data in the physical space through the spatial positioning sensor on the VR head display device, wherein the position data includes current position coordinates and relative displacement information; Acquiring user motion data through a motion capture sensor on the hand controller, wherein the motion data includes hand position, direction, acceleration and motion trajectory; According to the Kalman filter algorithm, the position data and the motion data are fused and processed to calculate the continuous motion trajectory of the user, wherein the continuous motion trajectory includes the current position, motion direction and motion speed.

3. According to claim 1, the LBE enhanced multi-modal perception interactive VR display control method is characterized in that: The calculation and adjustment of the activity range includes: Generate an initial boundary of the user's activity range according to the actual boundary size of the physical space and the interaction requirements in the virtual environment, wherein the initial boundary includes a dynamic center point and a range radius; According to the movement pattern in the user's continuous movement trajectory, adjust the expansion and contraction ratio of the user's activity range; Taking the current position of the user in the continuous motion trajectory as the new center point of the activity range, updating the center position of the user's activity range, based on the motion direction and motion speed, extending the boundary of the user's activity range in the first motion direction and shrinking it in other motion directions at the center position, thus forming a directional dynamic range; When the boundary of the user's activity range touches the actual boundary of the physical space due to adjustment, conflict handling and range correction are performed.

4. According to claim 1, the LBE enhanced multimodal perception interactive VR display control method is characterized in that: The intelligent algorithm comprises: Time synchronization of the collected perception data and motion data, and noise filtering of the synchronized multimodal data; Extract the angle change of the user's line of sight and the focus area from the visual data, and convert them into the focus area through encoding. Extract the movement speed, direction change and movement amplitude of the action data through trajectory analysis and encode them into behavioral features. Extract voice information from the voice data, and analyze the emotional features in combination with the voice information. Extract the feedback intensity change features from the tactile feedback data. A feature fusion network is constructed, and the focus area, behavior feature, emotion feature and feedback intensity change feature are used as input parameters of the feature fusion network. The input parameters are trained through the feature fusion network to output an interaction state vector.

5. According to claim 4, the LBE enhanced multi-modal perception interactive VR display control method is characterized in that: The extracting of voice information through voice data includes: Extract features from speech data, convert the speech data into a two-dimensional matrix, and calculate the feature vector sequence of each frame based on the Mel-frequency cepstral coefficient features of the first N frames and the last N frames of each frame; According to the encoder, the feature vector sequence of the speech data is converted into an audio coding feature vector, and the audio coding feature vector is processed by GRU to calculate the speech data content vector; The speech data content vector is reconstructed, and the reconstructed speech data content vector is decoded by a decoder to generate a semantic feature vector of the speech data.

6. The LBE enhanced multi-modal perception interactive VR display control method according to claim 5 is characterized in that: The feature fusion network comprises: The input layer is used to format the received multimodal features; A feature extraction layer, used to extract high-dimensional features of the output of the input layer; A feature weight allocation layer, used for adaptively allocating weights of the high-dimensional features through an attention mechanism, wherein the attention mechanism includes a modal attention mechanism, a channel attention mechanism, and a temporal attention mechanism; The feature fusion layer is used to perform weighted fusion processing on the multimodal features after weight assignment, compress and map the fused features through the fully connected layer operation, and generate an interaction state vector.

7. The LBE enhanced multi-modal perception interactive VR display control method according to claim 6 is characterized in that: The analysis of real-time behavioral and physiological feedback includes: Extract the user's action intention and operation trend based on the interaction state vector; According to the physiological feedback characteristics in the interaction state vector, the user's emotional state and interaction adaptability are judged; Determine whether the user's current operation sensitivity needs to be adjusted based on the user's action intention, operation trend, and emotional state; When the user's interactive adaptability shows fatigue or excessive tension, the response strength of the VR interactive device is dynamically adjusted.

8. An LBE-enhanced multimodal perception interactive VR display control system, used to implement the LBE-enhanced multimodal perception interactive VR display control method as described in any one of claims 1 to 7, characterized in that: The system includes a data acquisition module, a range adjustment module, a data processing module and an analysis response module; The data acquisition module is used to collect the user's multimodal perception data through the VR head display device and the hand controller; The range adjustment module is configured with a spatial positioning strategy, which is used to perform positioning according to the position data and the action data, and calculate and adjust the user's activity range; The data processing module is configured with a multimodal data fusion strategy, which is used to perform synchronous processing, feature extraction and fusion calculation on the collected multimodal data to generate an interactive state vector; The analysis and response module is used to analyze the user's real-time behavior and physiological feedback according to the interaction state vector, and dynamically adjust the response parameters of the VR interaction device according to the user state analysis results.

9. The LBE enhanced multi-modal perception interactive VR display control system according to claim 8, characterized in that: The range adjustment module comprises: The positioning unit is used to fuse the position data with the motion data according to the Kalman filter algorithm to calculate the continuous motion trajectory of the user; A dynamic range adjustment unit, used to dynamically adjust the range of motion according to the user's continuous motion trajectory; A boundary prompt unit, used to detect the overlap between the activity range boundary and the actual boundary of the physical space, and trigger a prompt when approaching the boundary; The spatial positioning strategy includes trajectory calculation logic and range adjustment logic. The trajectory calculation logic is configured in the positioning unit and is used to perform a fusion calculation on the user's current position, movement direction and movement speed. The range adjustment logic is configured in the dynamic range adjustment unit and is used to dynamically update the center point and boundary shape of the activity range based on the user's continuous movement trajectory.

Citation Information

Patent Citations

  • VR display control method, VR display control device, and display device

    CN109874003B

  • Virtual scene control method and system based on MR large space

    CN118708085A

  • VR interaction method and device based on meta-universe virtual reality technology

    CN118860156A

Cited By

  • Construction method of VR scene analysis based on three-dimensional digital visualization

    CN122799056A