A virtual reality-based interactive gesture recognition method and system

CN122064227BActive Publication Date: 2026-09-22WUHAN HIPAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202512040421.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-09-22
Estimated Expiration
2045-12-31

AI Technical Summary

Technical Problem

[0003]申请号为202111336108.4的发明专利申请中公开了一种基于虚拟现实的手势识别方法和系统,该申请旨在解决“不同设备使用者的手势行为习惯不同,而现有技术未考虑设备使用者自身的手势行为习惯,仅依靠训练集的样本难以将手势识别网络泛化至具体的设备使用者,导致手势识别准确率难以提高

Benefits of technology

[0041]本发明通过多源数据融合获取精准三维坐标与运动时序数据,动态调整权重适配不同环境与运动状态,有效降低干扰影响,经噪声滤除与时序规整消除速度差异带来的偏差,基于球面投影与分段轨迹提取多维度特征,提升手势辨识度,同时结合空间缩放与视角变换校正不同站位、朝向的偏差,确保特征一致性,并采用加权融合的匹配方式适配轨迹与轮廓特征的不同表现,提高识别准确率,通过遮挡状态检测与坐标预测应对关节点遮挡问题,避免识别中断,整体使得虚拟现实交互响应更精准、流畅,适配多样化使用场景,从而进一步提升用户交互体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064227B_ABST
    Figure CN122064227B_ABST
Patent Text Reader

Abstract

The application discloses an interactive gesture recognition method and system based on virtual reality, and relates to the field of human-computer interaction, and comprises the following steps: a collection module is used for collecting three-dimensional space coordinates and motion time sequence data of a user gesture in a preset virtual interaction space, and capturing dynamic information of a gesture joint; a calibration module is used for filtering noise of the collected three-dimensional gesture data, and standardizing time sequence data, so as to correct joint coordinate deviation; the application obtains accurate three-dimensional coordinates and motion time sequence data through multi-source data fusion, dynamically adjusts weights to adapt to different environments and motion states, effectively reduces interference influence, eliminates deviation caused by speed difference through noise filtering and time sequence standardization, extracts multi-dimensional features based on spherical projection and segmented trajectory, and improves gesture recognition degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, specifically to a method and system for recognizing interactive gestures based on virtual reality. Background Technology

[0002] Existing interactive gesture recognition technologies can be mainly divided into sensor-based wearable device recognition schemes and vision-based contactless recognition schemes. Sensor-based wearable device recognition schemes, by using data gloves integrated with components such as accelerometers, gyroscopes, and bending sensors on the user's hands, can accurately collect hand joint motion parameters and transmit them to a processing unit for gesture analysis. These schemes feature high recognition stability and strong resistance to environmental interference. Vision-based contactless recognition schemes, on the other hand, utilize depth cameras or RGB cameras equipped in VR devices to capture hand image information. Gesture recognition is achieved through algorithms such as image segmentation, feature extraction, and pattern matching. This eliminates the need for users to wear additional devices, offering advantages such as natural operation and high degree of interaction freedom, and has become a mainstream research direction in recent years.

[0003] Patent application number 202111336108.4 discloses a gesture recognition method and system based on virtual reality. This application aims to solve the problem that "different device users have different gesture habits, and the existing technology does not take into account the gesture habits of the device users themselves. It is difficult to generalize the gesture recognition network to specific device users by relying solely on training set samples, resulting in difficulty in improving the gesture recognition accuracy. Moreover, during the gesture recognition process, VR devices cannot render the scene in advance where the instruction information obtained from the gesture recognition result is applied, resulting in a worse user experience."

[0004] However, existing contactless gesture recognition technologies for visual images are still susceptible to factors such as changes in ambient lighting, background complexity, and hand occlusion, leading to inaccurate gesture feature extraction, which in turn reduces recognition accuracy and increases the false recognition rate.

[0005] To address this, we propose a virtual reality-based interactive gesture recognition method and system. Summary of the Invention

[0006] In view of the above-mentioned shortcomings of the existing technology, the present invention provides an interactive gesture recognition method and system based on virtual reality, which can effectively solve the problems of the existing technology.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions;

[0008] This invention discloses an interactive gesture recognition system based on virtual reality, comprising:

[0009] The system comprises the following modules: a data acquisition module, a calibration module, and a response module. The acquisition module collects the 3D spatial coordinates and motion timing data of user gestures within a preset virtual interaction space, capturing dynamic information of gesture joints. The calibration module filters noise from the acquired 3D gesture data and standardizes the timing data to correct joint coordinate deviations. The extraction module receives the 3D gesture data output from the calibration module, extracts the 3D contour features and motion trajectory features of the gestures, and constructs multi-dimensional gesture feature vectors. The correction module performs spatial calibration and viewpoint adaptation correction on the extracted multi-dimensional gesture feature vectors based on virtual reality spatial parameters to eliminate feature deviations in the virtual environment. The matching module calls a preset interactive gesture feature template library, performs matching operations between the calibrated feature vectors and the templates, and outputs the matching results. The response module receives the feature matching results, generates corresponding virtual reality interaction commands, and drives virtual scene elements to execute response actions corresponding to the gestures.

[0010] The acquisition module is interactively connected to the calibration module via a local area network. The calibration module is interactively connected to the extraction module via a local area network. The extraction module is interactively connected to the correction module and the matching module via a local area network. The correction module and the matching module are interactively connected to the response module via a local area network.

[0011] Furthermore, the acquisition module acquires data through a fusion acquisition method combining a binocular infrared depth camera and an inertial measurement unit, wherein the three-dimensional spatial coordinates are obtained by fusing binocular visual disparity calculation with inertial measurement unit attitude compensation.

[0012] ;

[0013] In the formula: The final output consists of three-dimensional spatial coordinates; The three-dimensional coordinates are calculated using binocular vision. The three-dimensional coordinates measured and transformed by the inertial measurement unit; For spatial gradient operators; These are dynamic weighting coefficients;

[0014] The motion timing data is processed by timestamp alignment, and the alignment accuracy is controlled within a preset time interval.

[0015] Furthermore, the noise filtering operation for the 3D gesture data in the calibration module follows the following rules:

[0016] ;

[0017] In the formula: This is the filtered 3D gesture data; This is the original collected data; , , These are the window offsets in three-dimensional space; The dynamic window scale at time t; Represents the median operation function;

[0018] The standardization and normalization stage of the time-series data maps gesture motion data of different durations to a unified time dimension interval to eliminate the joint coordinate deviation caused by the difference in gesture execution speed.

[0019] in, ,in Indicates the reference window scale. This represents the speed of the hand gesture at time t. This indicates the reference speed of motion.

[0020] Furthermore, when extracting three-dimensional contour features, the extraction module applies a contour description method based on spherical projection: by constructing a virtual sphere with the center of mass of the hand as the center, the joints on the surface of the hand are projected onto the sphere to obtain the spherical coordinates, and then the angle distribution features and arc length features after spherical projection are extracted.

[0021] Motion trajectory feature extraction is achieved through segmented trajectory feature description: the gesture motion trajectory is divided into multiple trajectory segments according to the points of change in motion direction. Curvature, torsion, and length ratio features of each trajectory segment are extracted. The temporal sequence relationship of each trajectory segment is combined to construct a subset of motion trajectory features. Finally, it is fused with three-dimensional contour features to form a multi-dimensional gesture feature vector. The dimension of the feature vector is dynamically adjusted according to the types of extracted features.

[0022] Furthermore, during the operation of the correction module, the camera coordinate system where the gesture feature vector is located is transformed to the virtual reality world coordinate system. During the transformation process, the scaling factor and offset of the virtual space are introduced for calibration.

[0023] In the perspective adaptation and correction stage, the feature vector is rotated and translated by constructing a perspective transformation matrix to eliminate perspective deviations caused by different positions and gesture orientations of users in the virtual interactive space.

[0024] Furthermore, the calibration and correction process follows:

[0025] ;

[0026] In the formula; These are the feature vectors in the transformed virtual reality world coordinate system; This is the virtual space scaling factor; It is a 3×3 rotation matrix; This is the original feature vector in the camera coordinate system; It is a three-dimensional translation vector; This is the final feature vector after viewpoint correction; Adjust the number of dimensions for the viewpoint; It is the identity matrix; The correction angle for the k-th dimension viewpoint; Let be the unit direction vector of the k-th dimension viewpoint; for The transpose of .

[0027] Furthermore, in the matching module, the matching operation phase follows the following rules:

[0028] ;

[0029] In the formula: This is the final matching result; These are the weighting coefficients; For matching scores; The cosine similarity matching score is used.

[0030] The ;

[0031] ;

[0032] In the formula: To improve the core computational complexity of the dynamic time warping algorithm, namely the optimal temporal alignment cumulative distance between the feature sequence of the trajectory to be matched and the template trajectory feature sequence; The preset maximum threshold for DTW; The minimum value is set to the preset minimum value; m and n are the lengths of the feature sequence of the trajectory to be matched and the feature sequence of the template trajectory, respectively. The weight coefficients for the dynamically regularized path; The i-th feature point in the trajectory feature sequence T The j-th feature point in the template trajectory feature sequence F The Euclidean distance.

[0033] Furthermore, the calibration module incorporates an occlusion detection and compensation unit to detect the occlusion status of hand joints in real time and predict the coordinates of occluded joints based on the movement trends and three-dimensional spatial relationships of adjacent joints.

[0034] ;

[0035] In the formula: The predicted coordinates of the occluded joints; This represents the number of adjacent unobstructed joints. Let be the coordinates of the i-th adjacent unobstructed joint. This refers to the weighting coefficient for the movement trend; The motion trend vector of the occluded joint;

[0036] When the duration of occlusion exceeds a preset threshold, the gesture re-recognition mechanism is activated, which triggers the occlusion detection and compensation unit to run.

[0037] On the other hand, a virtual reality-based interactive gesture recognition method includes:

[0038] Within a pre-defined virtual interactive space, the three-dimensional spatial coordinates and motion time-series data of user gestures are acquired through the fusion of a binocular infrared depth camera and an inertial measurement unit. Simultaneously, the three-dimensional coordinates are obtained through binocular visual parallax calculation and inertial measurement unit attitude compensation. The motion time-series data is aligned by timestamps and the alignment accuracy is controlled. Noise is filtered out from the three-dimensional gesture data based on dynamic window midpoint filtering, and gesture motion time-series data of different durations are mapped to a unified time dimension interval to correct joint coordinate deviations. A virtual sphere is constructed based on the hand centroid, and the angle distribution and arc length features after the joint projection are extracted as three-dimensional contour features. The motion trajectory features are extracted in segments and fused to form a multi-dimensional gesture feature vector. The feature vectors in the camera coordinate system are transformed to the virtual reality world coordinate system and scaling factors and offset calibration are introduced. Rotation and translation are performed through a real-time generated viewpoint transformation matrix to eliminate viewpoint deviations. A pre-defined hierarchical storage interactive gesture feature template library is called, and an improved dynamic time warping and cosine similarity weighted matching algorithm is used to match the calibrated feature vectors with the templates and output the results.

[0039] The system receives feature matching results, generates corresponding virtual reality interaction commands, and drives virtual scene elements to perform response actions corresponding to the gesture.

[0040] Compared with the known prior art, the technical solution provided by this invention has the following beneficial effects:

[0041] This invention acquires accurate 3D coordinates and motion time-series data through multi-source data fusion, dynamically adjusts weights to adapt to different environments and motion states, effectively reducing interference. Noise filtering and time-series regularization eliminate deviations caused by speed differences. Multi-dimensional features are extracted based on spherical projection and segmented trajectory to improve gesture recognition. At the same time, spatial scaling and viewpoint transformation are combined to correct deviations in different positions and orientations, ensuring feature consistency. A weighted fusion matching method is used to adapt to different performances of trajectory and contour features, improving recognition accuracy. Occlusion detection and coordinate prediction are used to address key point occlusion issues and avoid recognition interruptions. Overall, this makes virtual reality interaction response more accurate and smooth, adaptable to diverse usage scenarios, and further enhances the user interaction experience. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0043] Figure 1 This is a schematic diagram of the structure of an interactive gesture recognition method and system based on virtual reality;

[0044] Figure 2 This is a flowchart illustrating an interactive gesture recognition method and system based on virtual reality. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0046] The present invention will be further described below with reference to embodiments.

[0047] Example 1:

[0048] This embodiment provides an interactive gesture recognition system based on virtual reality, such as... Figure 1 As shown, it includes:

[0049] The acquisition module is used to acquire the three-dimensional spatial coordinates and motion timing data of user gestures within a preset virtual interaction space, and to capture dynamic information of gesture joints.

[0050] The acquisition module obtains data through a fusion acquisition method using a binocular infrared depth camera and an inertial measurement unit (IMU). The three-dimensional spatial coordinates are obtained by fusing binocular visual disparity calculation with IMU attitude compensation.

[0051] ;

[0052] In the formula: The final output consists of three-dimensional spatial coordinates; The three-dimensional coordinates are calculated using binocular vision. The three-dimensional coordinates measured and transformed by the inertial measurement unit; For spatial gradient operators; These are dynamic weighting coefficients;

[0053] The above formula combines the data acquired by binocular vision and inertial measurement unit. By dynamically adjusting three weighting coefficients, it fully takes into account the advantages of the two acquisition methods. When the binocular vision data has good signal-to-noise ratio, lighting and background conditions, it receives higher weight. When the inertial measurement unit data has stable motion posture and no violent shaking, it receives higher weight. When the two types of data have large deviations and poor consistency, the deviation correction effect is enhanced by the spatial gradient term. This allows for accurate fusion of the two types of data, effectively addressing the acquisition differences under different environmental lighting and motion states, and improving the accuracy and stability of the three-dimensional spatial coordinate output.

[0054] Motion time series data is processed by timestamp alignment, and the alignment accuracy is controlled within a preset time interval;

[0055] in, The values ​​of are all within the range of (0, 1), and follow the condition that: when the signal-to-noise ratio of the binocular vision acquisition data is higher than the preset threshold, the ambient lighting is stable, and there is no obvious background interference. The larger the value, the lower the accuracy of the binocular vision data when it is affected by sudden changes in illumination or complex background. The smaller the value, the better; when the stability of the motion attitude data of the inertial measurement unit meets the preset requirements and there is no severe hand shaking. The larger the value, the greater the data fluctuation caused by rapid hand movements or external vibrations in the inertial measurement unit. The smaller the value, the better; when the deviation between the binocular vision data and the inertial measurement unit data exceeds the preset deviation range, and the consistency between the two types of data is poor. The larger the value, the better when the deviation between the two types of data is within a preset reasonable range and the data consistency is good. The smaller the value;

[0056] The calibration module is used to filter out noise from the acquired 3D gesture data and standardize and normalize the temporal data to correct the joint coordinate deviation.

[0057] The noise filtering operation for 3D gesture data in the calibration module follows the following rules:

[0058] ;

[0059] In the formula: This is the filtered 3D gesture data; This is the original collected data; , , These are the window offsets in three-dimensional space; The dynamic window scale at time t; Represents the median operation function;

[0060] The above formula filters out noise from the original 3D gesture data through median operation. By introducing the time-varying 3D spatial window offset and dynamic window scale, it avoids the problem of incomplete filtering or loss of details that occurs when fixed window filtering is used at different motion speeds. At the same time, it is combined with time series data standardization and regularization to map gesture data of different durations to a unified time dimension and eliminate the deviation caused by the difference in gesture execution speed.

[0061] In the standardization and normalization stage of time-series data, gesture motion data of different durations are mapped to a unified time dimension interval to eliminate the joint coordinate deviation caused by the difference in gesture execution speed.

[0062] in, ,in Indicates the reference window scale. This represents the speed of the hand gesture at time t. Indicates the reference speed of motion;

[0063] The calibration module has a built-in occlusion detection and compensation unit, which is used to detect the occlusion status of hand joints in real time and predict the coordinates of occluded joints by analyzing the motion trends and three-dimensional spatial relationships of adjacent joints.

[0064] ;

[0065] In the formula: The predicted coordinates of the occluded joints; This represents the number of adjacent unobstructed joints. Let be the coordinates of the i-th adjacent unobstructed joint. This refers to the weighting coefficient for the movement trend; The motion trend vector of the occluded joint;

[0066] When a hand joint is detected to be occluded, the above formula calculates the average coordinates of adjacent unoccluded joints and combines them with the motion trend vector obtained by fitting the historical motion trajectories of adjacent joints to predict the coordinates of the occluded joint. The motion trend weight coefficient is dynamically adjusted according to the consistency of motion of adjacent unoccluded joints and the fitting degree of historical trajectory. The more consistent the motion of adjacent joints and the better the trajectory fitting degree, the larger the coefficient. If the duration of occlusion exceeds a preset threshold, the gesture re-recognition mechanism is activated. This fully utilizes the spatial position relationship and motion rules of adjacent joints, effectively compensates for the lack of joint coordinates caused by occlusion, and ensures the integrity of gesture data and the accuracy of subsequent processing.

[0067] When the duration of occlusion exceeds a preset threshold, the gesture re-recognition mechanism is activated, that is, the occlusion detection and compensation unit is triggered to run.

[0068] in, ∈(0,1), the higher the consistency of motion between adjacent unoccluded joints and the better the fit of historical motion trajectories, the better. The larger the value, the greater the motion dispersion of adjacent joints, and the more likely the trajectory fitting error exceeds the preset range. The smaller the value; It is obtained by fitting the historical motion trajectories of adjacent joints;

[0069] The extraction module is used to receive the 3D gesture data output by the calibration module, extract the 3D contour features and motion trajectory features of the gesture, and construct a multi-dimensional gesture feature vector.

[0070] When extracting 3D contour features, the extraction module applies a contour description method based on spherical projection: by constructing a virtual sphere with the hand's centroid as the center, the hand's surface joints are projected onto the sphere to obtain spherical coordinates, and then the angle distribution features and arc length features after spherical projection are extracted.

[0071] Motion trajectory feature extraction is achieved through segmented trajectory feature description: the gesture motion trajectory is divided into multiple trajectory segments according to the points of change in motion direction. Curvature, torsion and length ratio features of each trajectory segment are extracted. The temporal relationship of each trajectory segment is combined to construct a subset of motion trajectory features. Finally, it is fused with the three-dimensional contour features to form a multi-dimensional gesture feature vector. The dimension of the feature vector is dynamically adjusted according to the types of extracted features.

[0072] The calibration module is used to perform spatial calibration and viewpoint adaptation correction on the extracted multi-dimensional gesture feature vectors based on virtual reality spatial parameters, so as to eliminate feature deviations in the virtual environment.

[0073] During the calibration module's operation phase, the camera coordinate system containing the gesture feature vector is transformed to the virtual reality world coordinate system. During the transformation process, a scaling factor and offset of the virtual space are introduced for calibration.

[0074] In the perspective adaptation and correction stage, the feature vector is rotated and translated by constructing a perspective transformation matrix to eliminate perspective deviations caused by different positions and gesture orientations of users in the virtual interactive space.

[0075] The transformation matrix is ​​generated in real time based on preset parameters of the virtual reality space to ensure consistent expression of feature vectors under different viewpoints.

[0076] The calibration and correction process follows:

[0077] ;

[0078] In the formula; These are the feature vectors in the transformed virtual reality world coordinate system; This is the virtual space scaling factor; It is a 3×3 rotation matrix; This is the original feature vector in the camera coordinate system; It is a three-dimensional translation vector; This is the final feature vector after viewpoint correction; Adjust the number of dimensions for the viewpoint; It is the identity matrix; The correction angle for the k-th dimension viewpoint; Let be the unit direction vector of the k-th dimension viewpoint; for The transpose of the matrix;

[0079] The above formula first converts the original feature vector in the camera coordinate system into the feature vector in the virtual reality world coordinate system through scaling factor, rotation matrix and three-dimensional translation vector. The scaling factor is dynamically adjusted according to the scale ratio between the virtual scene and the real acquisition space and the distance between the user's hand and the acquisition device. Then, the feature vector is rotated and translated through a multi-dimensional perspective transformation matrix. The perspective correction angle is determined by the user's current position and gesture orientation. The transformation matrix is ​​generated in real time according to the preset parameters of the virtual reality space. It can simultaneously adapt to the scale difference of the virtual space and the user's different positions and gesture orientations, ensuring the consistent expression of the feature vector in different spaces and perspectives, thereby effectively eliminating spatial deviation and perspective interference in the virtual environment.

[0080] in, The value range is a preset interval (0, 2]. The larger the scale ratio between the virtual scene and the real-world acquisition space, and the greater the distance between the user's hand and the acquisition device, the better. The larger the value, the smaller the scale ratio between the virtual scene and the real-world acquisition space, and the closer the user's hand is to the acquisition device. The smaller the value;

[0081] , , All parameters are calculated and updated in real time based on preset parameters of the virtual reality space;

[0082] Calculated from the user's current position and the direction of their gesture;

[0083] The matching module is used to call the preset interactive gesture feature template library, perform matching operations between the calibrated feature vectors and the templates, and output the matching results;

[0084] In the matching operation phase of the matching module, the following rules apply:

[0085] ;

[0086] In the formula: This is the final matching result; These are the weighting coefficients; For matching scores; The cosine similarity matching score is used.

[0087] ;

[0088] ;

[0089] In the formula: To improve the core computational complexity of the dynamic time warping algorithm, namely the optimal temporal alignment cumulative distance between the feature sequence of the trajectory to be matched and the template trajectory feature sequence; The preset maximum threshold for DTW; The minimum value is set to the preset minimum value; m and n are the lengths of the feature sequence of the trajectory to be matched and the feature sequence of the template trajectory, respectively. The weight coefficients for the dynamically regularized path; The i-th feature point in the trajectory feature sequence T The j-th feature point in the template trajectory feature sequence F The Euclidean distance;

[0090] The above formula comprehensively improves the matching score of the dynamic time warping algorithm and the cosine similarity matching score. The influence of the two types of scores is balanced by dynamic weight coefficients. When the recognition of the gesture motion trajectory is high and the temporal stability is strong, the weight of the dynamic time warping algorithm is higher. When the contour feature discrimination is more significant and the trajectory has slight jitter or incompleteness, the weight of cosine similarity is higher. The dynamic time warping algorithm optimizes the alignment effect by dynamically adjusting the path weight coefficients and combining the spatial distance and temporal motion trend consistency between the feature point to be matched and the template feature point. With the help of the hierarchical storage of the interactive gesture feature template library, it can not only accurately handle the alignment problem of temporal features, but also make full use of the discrimination of contour features, ultimately improving the matching accuracy in different gesture scenarios.

[0091] The preset interactive gesture feature template library adopts a hierarchical storage structure and is classified and stored according to the motion type and contour feature category of the gesture.

[0092] ∈ (0,1), the value is larger when the recognizability of the gesture motion trajectory is higher than that of the contour feature and the trajectory temporal stability is stronger, and the value is smaller when the discriminativeness of the contour feature is more significant and the trajectory has slight jitter or incompleteness. The preset value range is (0,2], when the feature points of the trajectory to be matched With template trajectory feature points The closer the spatial distance and the higher the consistency of the movement trend of the corresponding temporal positions, the larger the value; the farther the spatial distance between the two points and the more obvious the temporal phase deviation, the smaller the value.

[0093] The response module is used to receive feature matching results, generate corresponding virtual reality interaction commands, and drive virtual scene elements to perform response actions corresponding to gestures.

[0094] The acquisition module is interconnected with the calibration module via a local area network. The calibration module is interconnected with the extraction module via a local area network. The extraction module is interconnected with the correction module and the matching module via a local area network. The correction module and the matching module are interconnected with the response module via a local area network.

[0095] In this embodiment, the acquisition module operates within a preset virtual interactive space, acquiring the three-dimensional spatial coordinates and motion timing data of the user's gestures, and capturing dynamic information of the gesture joints. The calibration module runs in sequence to filter out noise from the acquired three-dimensional gesture data and standardize the timing data to correct joint coordinate deviations. The extraction module further receives the three-dimensional gesture data output by the calibration module, extracts the three-dimensional contour features and motion trajectory features of the gestures, and constructs a multi-dimensional gesture feature vector. Then, the calibration module performs spatial calibration and viewpoint adaptation correction on the extracted multi-dimensional gesture feature vectors based on virtual reality spatial parameters to eliminate feature deviations in the virtual environment. The matching module calls a preset interactive gesture feature template library, performs matching operations between the calibrated feature vectors and the templates, and outputs the matching results. Finally, the response module receives the feature matching results, generates corresponding virtual reality interactive commands, and drives virtual scene elements to perform response actions corresponding to the gestures.

[0096] In the above embodiments, the system can accurately capture the three-dimensional information and motion trajectory of gestures, effectively filter out noise and correct deviations, adapt to different perspectives and scene changes, improve the accuracy and stability of gesture recognition, and even in the presence of occlusion, shaking and other situations, it can reliably match gesture features and quickly generate corresponding interaction commands, making the virtual scene more responsive and natural, and greatly optimizing the user's interactive experience in virtual reality.

[0097] Referring to the system in the above embodiments, an application example of the system is shown:

[0098] In a VR adventure game, players need to pick up treasure chests in the virtual scene by making a "clenched fist and then raised" gesture. The game uses this system, and the specific application process is as follows:

[0099] After players enter the preset VR interactive space, the system's data acquisition module works in conjunction with a binocular infrared depth camera and an inertial measurement unit. Due to the stable lighting in the game scene, the absence of significant background interference, and the absence of severe hand tremors when players perform gestures, the data collected by the two types of devices are highly consistent. After fusion calculation, the three-dimensional spatial coordinates of each joint of the player's hand are accurately obtained. At the same time, the motion timing data is processed through timestamp alignment, and the error is controlled within a preset small interval.

[0100] The collected raw data is processed by the calibration module. First, noise from the environment is eliminated through dynamic window mid-range filtering. The filter window size is adjusted according to the player's gesture movement speed, ultimately yielding clear 3D gesture data. Then, the system uniformly maps the duration of each player's gesture to the same time dimension, effectively eliminating joint coordinate deviations caused by differences in execution speed. During this process, the player's hand is briefly and slightly obscured by virtual vegetation. The system accurately predicts the coordinates of the obscured joint by analyzing the movement trends and spatial relationships of adjacent unobscured joints, ensuring data integrity.

[0101] The extraction module constructs a virtual sphere with the hand's center of mass as the center of the sphere based on the calibrated data. After projecting the hand's surface joints onto the sphere, it extracts the angle distribution features and arc length features that can clearly represent the fist-clenching shape. At the same time, the motion trajectory of "clenching fist - lifting" is divided into two segments according to the points of directional change, and the curvature, torsion, and length ratio features of each segment are extracted. Combined with temporal relationship and contour features, a multi-dimensional gesture feature vector is formed.

[0102] The correction module first transforms the feature vector from the camera coordinate system to the VR world coordinate system. Based on the scale ratio between the game scene and the real-world acquisition space, it calculates the appropriate scaling factor and translation amount. Then, through the real-time generated view transformation matrix, it rotates and translates the feature vector, successfully eliminating the view deviation caused by the player's slightly off-center standing position and ensuring that the feature vector expression is consistent.

[0103] The matching module calls the preset interactive gesture feature template library. Because the gesture has high recognition of motion trajectory and stable timing, the system adopts improved dynamic time warping and cosine similarity weighted calculation to finally obtain a matching result that meets the preset threshold and confirms that the player's gesture is a "pick up" command.

[0104] After receiving the matching result, the response module immediately generates the corresponding VR interaction command, drives the virtual treasure chest to pop up and open, and automatically stores the items in the treasure chest into the player's VR backpack, completing the entire interaction process.

[0105] Example 2:

[0106] At the implementation level, based on Example 1, this example refers to... Figure 2 A more detailed description of the virtual reality-based interactive gesture recognition system in Example 1 is provided below:

[0107] A virtual reality-based interactive gesture recognition method, comprising:

[0108] Within a pre-defined virtual interactive space, the three-dimensional spatial coordinates and motion timing data of the user's gestures are collected by fusing binocular infrared depth cameras and inertial measurement units. Simultaneously, the three-dimensional coordinates are obtained by fusing binocular visual parallax calculation with inertial measurement unit attitude compensation. The motion timing data is aligned according to timestamps and the alignment accuracy is controlled.

[0109] The noise is filtered out of the 3D gesture data by using dynamic window mid-range filtering, and the timing data of gesture motion of different durations are mapped to a unified time dimension interval to correct the deviation of the joint coordinates.

[0110] A virtual sphere is constructed based on the hand's centroid. The angle distribution and arc length features after the joint points are projected are extracted as three-dimensional contour features. The motion trajectory features are extracted in segments and then fused to form a multi-dimensional gesture feature vector.

[0111] The feature vectors in the camera coordinate system are transformed to the virtual reality world coordinate system and scaling factors and offset calibration are introduced. Rotation and translation are performed through the real-time generated view transformation matrix to eliminate view deviation.

[0112] The system calls a pre-defined hierarchical storage interactive gesture feature template library, uses an improved dynamic time warping and cosine similarity weighted matching algorithm to perform matching operations on the calibrated feature vectors and templates, and outputs the results.

[0113] The system receives feature matching results, generates corresponding virtual reality interaction commands, and drives virtual scene elements to perform response actions corresponding to the gesture.

[0114] In summary, the system and method in the above embodiments obtain accurate 3D coordinates and motion time-series data through multi-source data fusion, dynamically adjust weights to adapt to different environments and motion states, effectively reduce interference, eliminate deviations caused by speed differences through noise filtering and time-series regularization, extract multi-dimensional features based on spherical projection and segmented trajectory to improve gesture recognition, and combine spatial scaling and viewpoint transformation to correct deviations in different positions and orientations to ensure feature consistency. Furthermore, a weighted fusion matching method is used to adapt to different performances of trajectory and contour features, improving recognition accuracy. Occlusion detection and coordinate prediction are used to address key point occlusion issues and avoid recognition interruptions. Overall, the virtual reality interaction response is more accurate and smooth, adapting to diverse usage scenarios, thereby further enhancing the user interaction experience.

[0115] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A virtual reality-based interactive gesture recognition system, characterized in that, include: The acquisition module is used to acquire the three-dimensional spatial coordinates and motion timing data of user gestures within a preset virtual interaction space, and to capture dynamic information of gesture joints. The calibration module is used to filter out noise from the acquired 3D gesture data and standardize and normalize the temporal data to correct the joint coordinate deviation. The extraction module is used to receive the 3D gesture data output by the calibration module, extract the 3D contour features and motion trajectory features of the gesture, and construct a multi-dimensional gesture feature vector. The calibration module is used to perform spatial calibration and viewpoint adaptation correction on the extracted multi-dimensional gesture feature vectors based on virtual reality spatial parameters, so as to eliminate feature deviations in the virtual environment. During the operation of the correction module, the camera coordinate system where the gesture feature vector is located is transformed to the virtual reality world coordinate system. During the transformation process, the scaling factor and offset of the virtual space are introduced for calibration. In the viewpoint adaptation and correction stage, the feature vector is rotated and translated by constructing a viewpoint transformation matrix to eliminate the viewpoint deviation caused by the user's different positions and gesture orientations in the virtual interactive space. The calibration and correction process in the calibration module follows the following: ; In the formula; These are the feature vectors in the transformed virtual reality world coordinate system; This is the scaling factor for the virtual space. It is a 3×3 rotation matrix; This is the original feature vector in the camera coordinate system; It is a three-dimensional translation vector; This is the final feature vector after viewpoint correction; Adjust the number of dimensions for the viewpoint; It is the identity matrix; The correction angle for the k-th dimension viewpoint; Let be the unit direction vector of the k-th dimension viewpoint; for The transpose of the matrix; The matching module is used to call the preset interactive gesture feature template library, perform matching operations between the calibrated feature vectors and the templates, and output the matching results; The response module is used to receive feature matching results, generate corresponding virtual reality interaction commands, and drive virtual scene elements to perform response actions corresponding to gestures.

2. The interactive gesture recognition system based on virtual reality according to claim 1, characterized in that, The acquisition module obtains data through a fusion acquisition method using a binocular infrared depth camera and an inertial measurement unit (IMU). The three-dimensional spatial coordinates are obtained by fusing binocular visual disparity calculation with IMU attitude compensation. ; In the formula: The final output consists of three-dimensional spatial coordinates; The three-dimensional coordinates are calculated using binocular vision. The three-dimensional coordinates measured and transformed by the inertial measurement unit; For spatial gradient operators; These are dynamic weighting coefficients; The motion timing data is processed by timestamp alignment, and the alignment accuracy is controlled within a preset time interval.

3. The interactive gesture recognition system based on virtual reality according to claim 1, characterized in that, The noise filtering operation for 3D gesture data in the calibration module follows the following rules: ; In the formula: This is the filtered 3D gesture data; This is the original collected data; , , These are the window offsets in three-dimensional space; The dynamic window scale at time t; Represents the median operation function; The standardization and normalization stage of the time-series data maps gesture motion data of different durations to a unified time dimension interval to eliminate the joint coordinate deviation caused by the difference in gesture execution speed. in, ,in Indicates the reference window scale. This represents the speed of the hand gesture at time t. This indicates the reference speed of motion.

4. The interactive gesture recognition system based on virtual reality according to claim 1, characterized in that, When extracting three-dimensional contour features, the extraction module applies a contour description method based on spherical projection: by constructing a virtual sphere with the center of mass of the hand as the center, the joints on the surface of the hand are projected onto the sphere to obtain the spherical coordinates, and then the angle distribution features and arc length features after spherical projection are extracted. Motion trajectory feature extraction is achieved through segmented trajectory feature description: the gesture motion trajectory is divided into multiple trajectory segments according to the points of change in motion direction. Curvature, torsion, and length ratio features of each trajectory segment are extracted. The temporal sequence relationship of each trajectory segment is combined to construct a subset of motion trajectory features. Finally, it is fused with three-dimensional contour features to form a multi-dimensional gesture feature vector. The dimension of the feature vector is dynamically adjusted according to the types of extracted features.

5. The interactive gesture recognition system based on virtual reality according to claim 1, characterized in that, The matching operation phase in the matching module follows the following principle: ; In the formula: This is the final matching result; These are the weighting coefficients; For matching scores; The cosine similarity matching score is used. The ; ; In the formula: To improve the core computational complexity of the dynamic time warping algorithm, namely the optimal temporal alignment cumulative distance between the feature sequence of the trajectory to be matched and the template trajectory feature sequence; The preset maximum threshold for DTW; This is the preset minimum value; m and n are the lengths of the feature sequence of the trajectory to be matched and the feature sequence of the template trajectory, respectively. The weight coefficients for the dynamically regularized path; The i-th feature point in the trajectory feature sequence T The j-th feature point in the template trajectory feature sequence F The Euclidean distance.

6. The interactive gesture recognition system based on virtual reality according to claim 1, characterized in that, The calibration module has a built-in occlusion detection and compensation unit, which is used to detect the occlusion status of hand joints in real time and predict the coordinates of occluded joints by analyzing the movement trends and three-dimensional spatial relationships of adjacent joints. ; In the formula: The predicted coordinates of the occluded joints; This represents the number of adjacent unobstructed joints. Let be the coordinates of the i-th adjacent unobstructed joint. This refers to the weighting coefficient for the movement trend; The motion trend vector of the occluded joint; When the duration of occlusion exceeds a preset threshold, the gesture re-recognition mechanism is activated, which triggers the occlusion detection and compensation unit to run.

7. The interactive gesture recognition system based on virtual reality according to claim 1, characterized in that, The acquisition module is interactively connected to the calibration module via a local area network. The calibration module is interactively connected to the extraction module via a local area network. The extraction module is interactively connected to the correction module and the matching module via a local area network. The correction module and the matching module are interactively connected to the response module via a local area network.

8. A method for recognizing interactive gestures based on virtual reality, wherein the method is an implementation method of the interactive gesture recognition system based on virtual reality as described in any one of claims 1-7, characterized in that, include: Within a pre-defined virtual interactive space, the three-dimensional spatial coordinates and motion timing data of the user's gestures are collected by fusing binocular infrared depth cameras and inertial measurement units. Simultaneously, the three-dimensional coordinates are obtained by fusing binocular visual parallax calculation with inertial measurement unit attitude compensation. The motion timing data is aligned according to timestamps and the alignment accuracy is controlled. The noise is filtered out of the 3D gesture data by using dynamic window mid-range filtering, and the timing data of gesture motion of different durations are mapped to a unified time dimension interval to correct the deviation of the joint coordinates. A virtual sphere is constructed based on the hand's centroid. The angle distribution and arc length features after the joint points are projected are extracted as three-dimensional contour features. The motion trajectory features are extracted in segments and then fused to form a multi-dimensional gesture feature vector. The feature vectors in the camera coordinate system are transformed to the virtual reality world coordinate system and scaling factors and offset calibration are introduced. Rotation and translation are performed through the real-time generated view transformation matrix to eliminate view deviation. The system calls a pre-defined hierarchical storage interactive gesture feature template library, uses an improved dynamic time warping and cosine similarity weighted matching algorithm to perform matching operations on the calibrated feature vectors and templates, and outputs the results. The system receives feature matching results, generates corresponding virtual reality interaction commands, and drives virtual scene elements to perform response actions corresponding to the gesture.

Citation Information

Patent Citations

  • Gesture recognition method based on virtual reality and system thereof

    CN114035687A

  • Method and system for achieving virtual touch calibration

    CN103941851A

  • Virtual reality parallax correction

    CN109660783A