A high-precision gesture tracking method and system based on hand-handle fusion
By employing online extrinsic parameter calibration and multi-sensor tight coupling optimization, combined with self-occlusion neural network compensation, the problems of hand and controller extrinsic parameter estimation and self-occlusion in VR/AR systems were solved, achieving high-precision, low-latency hand pose estimation and improving user experience and system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-10
AI Technical Summary
In existing VR/AR systems, the extrinsic parameters of the hand and controller are unknown, making it difficult to tightly couple low-frequency visual data with high-frequency IMU data. The self-occlusion problem is severe in the handheld state, resulting in insufficient accuracy and continuity of hand posture estimation.
The grip state is determined by the head-mounted display camera and the IMU module of the hand controller. Online extrinsic parameter calibration and kinematic constraint calculation are performed. The hand posture data is tightly coupled and fused with factor graphs, and a complete posture is generated through neural network compensation when self-occlusion occurs.
It achieves high-precision, low-latency hand pose estimation, adapts to different users and grip postures, requires no offline calibration, improves user experience and system stability, and ensures the continuity and immersion of interaction.
Smart Images

Figure CN122363519A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of virtual reality and augmented reality technology, and in particular to a high-precision gesture tracking method and system based on the fusion of hand and controller. Background Technology
[0002] In VR / AR systems, natural and accurate hand interaction is key to improving the user experience. Currently, the mainstream technologies fall into two categories: one is camera-based hand pose estimation (gesture recognition), and the other is controller tracking based on handheld devices (with IMU and infrared dot sensors).
[0003] Hand pose estimation (such as vision-based algorithms) can provide natural and intuitive interaction, but it has the following shortcomings: 1) The output frame rate is low (usually 30Hz), making it difficult to capture fast motion, resulting in visual jitter and latency; 2) When the hand is holding an object (such as a handle or tool), the hand and the object are severely self-occluded, and a large number of joints are not visible, causing the algorithm to fail or its accuracy to drop sharply.
[0004] Handle tracking (such as inside-out solutions) uses an infrared light spot observed by the headset and combines it with a high-frequency IMU (such as 500Hz) to provide robust, low-latency 6-DoF pose. However, it cannot sense the specific posture of the hand (such as finger bending), and the interaction method is limited to button presses.
[0005] Existing technologies attempt to combine the two, such as simply using the hand's pose as the overall hand pose when holding a controller. However, they neglect the subtle movements and fixed offsets of the hand itself (especially the wrist) relative to the controller, resulting in a visual "disconnect" between the hand model and the real hand, thus disrupting immersion. Therefore, existing technologies mainly suffer from the following shortcomings: 1) Unknown / Time-varying external parameters: The relative positional relationship (external parameter) between the hand (wrist joint) and the handle is unknown and varies from person to person and from grip to grip, and cannot be obtained through offline calibration.
[0006] 2) Data asynchrony and frequency mismatch: Low frame rate visual hand gestures and high frame rate IMU data are difficult to integrate effectively, resulting in inconsistent system output.
[0007] 3) Self-occlusion failure: When the hand is holding the handle, the accuracy of visual hand tracking drops significantly or even fails due to self-occlusion, and it cannot provide reliable posture constraints. Summary of the Invention
[0008] This application provides a high-precision gesture tracking method and system based on hand and controller fusion to solve the following technical problems: in existing VR / AR systems, there are unknown situations in the estimation of extrinsic parameters of hands and controllers, making it difficult to tightly couple low-frequency vision with high-frequency IMU, as well as the self-occlusion problem in handheld state.
[0009] The embodiments of this application adopt the following technical solutions: On one hand, this application provides a high-precision gesture tracking method based on hand and controller fusion, including: judging the grip state between the controller and hand based on the head-mounted display camera and the controller's built-in IMU module, and determining the grip state result; if the grip state result between the controller and hand is a grip state with no or low occlusion, then performing optimal solution calculations under relevant kinematic constraints on the controller and hand based on the visual wrist node pose and the controller IMU data to obtain grip state extrinsic parameters; based on the grip state extrinsic parameters, the visual wrist node pose and the controller IMU data are combined to obtain the optimal solution calculations under relevant kinematic constraints on the controller and hand. The U-data and infrared spot data are subjected to tight-coupled fusion processing based on factor graphs to obtain a wrist 6-DoF pose with high frequency and low latency. If the grip state result is a high-occlusion grip state, a self-occlusion pose compensation network is used to perform regression mapping processing on the infrared spot data and the controller IMU data for relevant spatiotemporal feature sequences to generate hand pose data based on complete hand joints. Based on the grip state result and the corresponding wrist 6-DoF pose and hand pose data, the gesture tracking between the controller and the hand is dynamically switched to output the final hand pose data.
[0010] This application achieves high-precision, low-latency hand pose estimation in handheld controller scenarios through online extrinsic parameter calibration, multi-sensor tight coupling optimization, and self-occlusion neural network compensation. Moreover, it eliminates the need for cumbersome offline calibration by the user; the system automatically completes calibration the moment the hand is grasped and adapts to different users and grip postures, greatly enhancing the user experience. Through tight coupling fusion, the 30Hz visual wrist pose is boosted to 500Hz, significantly reducing jitter and latency, making virtual hand movements smoother and more responsive. Even when the hand grips an object causing visual impairment, the hand pose compensation network driven by controller data still outputs complete and reasonable hand poses, ensuring the continuity of interaction. Simultaneously, during rapid hand movements or partial occlusion, IMU data provides strong motion priors, ensuring the overall stability of the system. Finally, it allows users to enjoy the convenience of natural gesture interaction while simultaneously gripping the controller for precise control; the system seamlessly integrates and accurately represents hand details, greatly enhancing the immersion and functionality of VR / AR applications.
[0011] In one feasible implementation, before determining the grip state between the handle and the hand based on the multi-source data collected by the head-mounted display camera and the built-in IMU module of the handle, the method further includes: acquiring images containing infrared light spots of the hand and the handle through the head-mounted display camera to obtain hand posture images and handle infrared light spot images; acquiring handle IMU data from the built-in IMU module of the handle; wherein the handle IMU data is high-frequency IMU data and includes at least: acceleration data and gyroscope data; determining the head-mounted display's own pose in the current head-mounted display state based on the head-mounted display VIO; performing recognition processing on the hand posture image related to the 3D joint position of the hand through a hand posture estimation network to obtain the visual wrist node pose; parsing the handle infrared light spot image for 3D point cloud information to obtain infrared light spot data; and performing spatiotemporal alignment processing on the head-mounted display's own pose, the visual wrist node pose, the handle IMU data, and the infrared light spot data according to the spatiotemporal feature sequence to generate the multi-source data.
[0012] In one feasible implementation, based on the head-mounted display camera and the built-in IMU module of the controller, the grip state between the controller and the hand is determined from the collected multi-source data, specifically including: performing a spatiotemporal consistency judgment on the hand geometric features in the visual wrist node posture and the controller light spot distribution features in the infrared light spot data to obtain a judgment result; if the judgment result is consistent, the grip state between the controller and the hand is in a grip state; if the judgment result is inconsistent, the grip state between the controller and the hand is in a non-grip state; if the grip state result is in a grip state, the hand joint features in the hand geometric features are identified based on an occlusion threshold using the head-mounted display camera to obtain a hand occlusion result; wherein, the hand occlusion result includes: a high occlusion result and a no or low occlusion result; based on the hand occlusion result, the grip state results in the grip state are classified to obtain a grip state with no or low occlusion and a grip state with high occlusion.
[0013] In one feasible implementation, before calculating the optimal solution for the grip state extrinsic parameters between the handle and hand under relevant kinematic constraints based on the visual wrist node posture and the handle's IMU data, the method further includes: initiating online extrinsic parameter calibration control when the grip state result is identified as the initial grip state; and real-time acquisition and processing of the motion state of the handle's built-in IMU module and the wrist node based on the transient rigid body formed between the wrist node and the handle; according to... The wrist angular velocity based on the wrist node is obtained. ;in, The angular velocity of the handle is based on the built-in IMU module; Rotational extrinsic parameters in the extrinsic parameter state; wrist angular velocity. High-frequency interpolation obtained from the visual wrist pose node via B-spline; based on The wrist acceleration based on the wrist node is obtained. ;in, The translational extrinsic parameter in the extrinsic parameter state; It is the gravity vector; The accelerometer of the built-in IMU module has zero bias; The kinematic constraints between the hand and the handle are obtained based on the hand controller acceleration, the hand controller angular velocity, the wrist acceleration, and the rotational constraints, translational and gravitational constraints between the hand controller acceleration, the hand controller angular velocity, the wrist acceleration, and the wrist angular velocity.
[0014] In one feasible implementation, if the grip state between the handle and the hand is a grip state with no or low occlusion, then based on the visual wrist node posture and the handle IMU data, the optimal solution calculation under relevant kinematic constraints is performed on the grip state extrinsic parameters between the handle and the hand to obtain the grip state extrinsic parameters. Specifically, this includes: calculating the residuals of the joint extrinsic parameter calibration between the hand and the handle through the kinematic constraints to obtain the minimum sum of residuals; and performing the optimal solution calculation on the online extrinsic parameter state fused with the hand and the handle through the minimum sum of residuals to obtain the optimal grip state extrinsic parameters and the corresponding IMU zero bias in a grip state with no or low occlusion.
[0015] In one feasible implementation, before performing factor-map-based tight-coupling fusion processing on the visual wrist node pose, the handle IMU data, and the infrared spot data according to the grip state extrinsic parameters to obtain a wrist 6-DoF pose with high frequency and low latency, the method further includes: determining the visual wrist node pose of the visual front end as the observed value of the visual factor and constructing a reprojection error; using the high-frequency handle IMU data, performing pre-integration calculation between adjacent visual keyframes to construct a high-frequency motion constraint based on the IMU pre-integration factor; determining the infrared spot distribution data in the infrared spot data as sparse visual features, and projecting the sparse visual features onto the wrist node coordinate system using the grip state extrinsic parameters to obtain the constrained infrared spot factor; and constructing an optimized factor map under tight-coupling fusion of the hand and handle based on the visual factor, the IMU pre-integration factor, and the infrared spot factor. In one feasible implementation, based on the grip state extrinsic parameters, the visual wrist node pose, the handle IMU data, and the infrared spot data are subjected to tight coupling fusion processing based on a factor graph to obtain a wrist 6-DoF pose with high frequency and low latency. Specifically, this includes: when the grip state result is a grip state with no or low occlusion, determining the current wrist node pose, velocity, and IMU zero bias as wrist node state variables based on the wrist node pose; and optimizing the wrist node state variables based on the output of a relevant sliding window through the optimization factor graph to obtain the wrist 6-DoF pose with high frequency and low latency.
[0016] In one feasible implementation, if the gripping state result is a highly occluded gripping state, then a self-occlusion posture compensation network is used to perform regression mapping processing on the infrared spot data and the controller IMU data for relevant spatiotemporal feature sequences to generate hand posture data based on complete hand joints. Specifically, this includes: constructing the infrared spot data and the controller IMU data into a spatiotemporal feature sequence based on the input of the self-occlusion posture compensation network; wherein, the self-occlusion posture compensation network is an end-to-end neural network; extracting the spatiotemporal features in the spatiotemporal feature sequence through a temporal convolutional encoder; and inputting the spatiotemporal features into a decoder of a fully connected layer to regress the 3D position of the hand joints based on the wrist coordinate system; and training the motion posture mapping of the 3D position of the hand joints for relevant complete hand joints based on data samples of normal hand movement data and corresponding controller movement data, and outputting the hand posture data based on self-occlusion posture compensation.
[0017] In one feasible implementation, the gesture tracking between the controller and the hand is dynamically switched based on the grip state result and the corresponding wrist 6-DoF pose and hand pose data, and the final hand pose data is output. Specifically, this includes: if the grip state result is a high-occlusion grip state, the final hand pose data is determined as the hand pose data based on self-occlusion pose compensation; if the grip state result is a no-occlusion or low-occlusion grip state, the final hand pose data is determined as the wrist 6-DoF pose output based on the optimized factor map; if the grip state result switches from a high-occlusion grip state to a no-occlusion or low-occlusion grip state, the final hand pose data is switched from the hand pose data to the wrist 6-DoF pose accordingly; based on the final hand pose data, continuous gesture tracking control between the controller and the hand is completed, and high-precision, low-latency hand pose estimation is achieved in handheld controller scenarios.
[0018] On the other hand, this application also provides a high-precision gesture tracking system based on hand and controller fusion. The high-precision gesture tracking system can be executed by at least one processor, so that at least one processor can execute the high-precision gesture tracking method based on hand and controller fusion described in any of the above embodiments.
[0019] This application discloses a high-precision gesture tracking method and system based on hand and controller fusion. Compared with the prior art, the embodiments of this application have the following beneficial technical effects: 1. High-precision online calibration of hand-handle extrinsic parameters: The system automatically completes the calibration the moment the user holds the hand, eliminating the need for tedious offline calibration. It can also adapt to different users and grip postures, greatly improving the user experience.
[0020] 2. Improve wrist tracking accuracy and frequency: Through tight coupling fusion, the visual wrist pose is increased from 30Hz to 500Hz, significantly reducing jitter and latency, making the virtual hand movement smoother and more responsive.
[0021] 3. Solving the self-occlusion problem: When the hand is holding an object and visual loss occurs, the hand posture compensation network driven by the handle data can still output complete and reasonable hand posture, ensuring the continuity of interaction.
[0022] 4. Strong system robustness: When the hand moves rapidly or is partially occluded, the IMU data provides strong motion priors, ensuring the overall stability of the system.
[0023] 5. Enhanced naturalness of interaction: Users can enjoy the convenience of natural gesture interaction, and when precise control is needed, they can hold the controller. The system can seamlessly connect and accurately represent hand details, greatly enhancing the immersion and functionality of VR / AR applications. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart illustrating a high-precision gesture tracking method based on hand and controller fusion is provided in this application embodiment; Figure 2 This application provides an overall architecture diagram of a high-precision gesture tracking system as an embodiment. Figure 3 A schematic diagram of an online calibration principle for hand-handle extrinsic parameters provided in this application embodiment; Figure 4 A self-occlusion posture compensation network structure diagram provided in this application embodiment; Figure 5 A schematic diagram of a tightly coupled fusion factor graph provided in an embodiment of this application; Figure 6 This is a schematic diagram of a high-precision gesture tracking system based on the fusion of hand and controller, provided as an embodiment of this application. Detailed Implementation
[0025] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0026] It should be noted that the embodiments of this application provide a high-precision gesture tracking system based on hand and handle fusion, including: 1) a data input module: receiving multi-source data, including: a) an image captured by the head-mounted display camera containing infrared light spots of the hand and handle; b) data from the high-frequency IMU (accelerometer, gyroscope) built into the handle; c) the head-mounted display's own pose (provided by the head-mounted display VIO).
[0027] 2) Visual front-end module: Performs parallel processing on images. On one hand, it runs a hand pose estimation network (such as MediaPipe, FrankMocap, etc.) to output the preliminary 3D hand joint (including wrist node) positions at 30Hz. On the other hand, it detects and tracks infrared light spots on the handle to provide sparse 3D point cloud information.
[0028] 3) Online calibration module for hand-handle extrinsic parameters for "grip state" detection: Based on the spatiotemporal consistency of hand geometry and handle light spot distribution, it determines whether the hand is in a grip state. When entering the "grip state" for the first time, or when the extrinsic parameters change, the hand-handle extrinsic parameters (RHS, tHS) are optimized online using visual wrist node pose (low-frequency data) and handle IMU data (high-frequency data), where H is the handle coordinate system and S is the wrist node coordinate system.
[0029] 4) Tightly Coupled Fusion and Pose Optimization Module: After confirming the "grip state" and knowing the external parameters, a unified factor graph or filtering framework is constructed to tightly couple and fuse the low-frequency visual wrist node pose, the high-frequency handle IMU data, and the observation of the handle infrared light spot, and output a high frame rate (e.g., 500Hz) and low latency wrist 6-DoF pose.
[0030] 5) Self-Occlusion Pose Compensation Network: This network is specifically designed to address the issue of visual front-end failure when the hand is severely occluded (e.g., when the hand is holding an object). This module utilizes the IMU and infrared spot data from the handle to directly regress the complete hand pose (including finger joints) through a pre-trained neural network.
[0031] 6) State Management Module: Based on the "grip state" and visual tracking quality, the above modules are dynamically switched to smoothly output the final hand posture.
[0032] This application provides a high-precision gesture tracking method based on hand and controller fusion, such as... Figure 1 As shown, the high-precision gesture tracking method based on hand and controller fusion specifically includes steps S101-S105: S101, based on the built-in IMU module of the head-mounted display camera and the controller, judges the grip state between the controller and the hand based on the collected multi-source data, and determines the grip state result.
[0033] Specifically, the system first uses the head-mounted display's camera to capture images including the hand and the infrared light spots on the controller, obtaining hand posture images and controller infrared light spot images. Then, it acquires controller IMU data from the built-in IMU module. This controller IMU data is high-frequency IMU data and includes at least acceleration data and gyroscope data. Finally, based on the head-mounted display's VIO (Voice over IoV), the head-mounted display's own pose in the current state is determined.
[0034] Furthermore, a hand pose estimation network is used to identify the positions of 3D hand joints in the hand pose image to obtain the visual wrist node pose. Then, the infrared spot image of the handle is processed to analyze the 3D point cloud information to obtain infrared spot data.
[0035] Furthermore, based on the spatiotemporal feature sequence, the head-mounted display's own pose, the visual wrist node pose, the controller IMU data, and the infrared spot data are spatiotemporally aligned to generate multi-source data.
[0036] Furthermore, the spatiotemporal consistency of the hand geometric features in the visual wrist node pose and the handle spot distribution features in the infrared spot data is judged to obtain the judgment result.
[0037] Furthermore, Figure 2 This application provides an overall architecture diagram of a high-precision gesture tracking system, as shown in the embodiments below. Figure 2 As shown, if the judgment result is consistent, the grip state between the handle and the hand is in a gripping state. If the judgment result is inconsistent, the grip state between the handle and the hand is in a non-gripping state.
[0038] Furthermore, if the gripping state result indicates that the hand is in a gripping state, the head-mounted display camera performs feature occlusion recognition based on an occlusion threshold on the hand joint features within the hand's geometric features to obtain the hand occlusion result. The hand occlusion result includes: a high occlusion result and a result with no or low occlusion.
[0039] Furthermore, based on the hand occlusion results, the holding state results are classified to obtain holding states with no or low occlusion and holding states with high occlusion.
[0040] S102. If the grip state between the handle and the hand is in a grip state with no or low obstruction, then based on the visual wrist node posture and the handle IMU data, perform optimal solution calculations for the grip state under relevant kinematic constraints to obtain the grip state extrinsic parameters.
[0041] Specifically, when the grip state is identified as the initial grip state, online extrinsic parameter calibration control is initiated. Then, based on the transient rigid body formed between the wrist node and the handle, the motion state of the handle's built-in IMU module and the wrist node is acquired and processed in real time.
[0042] In one embodiment, online extrinsic parameter calibration is initiated when a hand is detected gripping the handle for the first time. This calibration does not rely on the handle's own output pose, but rather on the gripping state. Figure 3 A schematic diagram of online calibration principle for hand-handle extrinsic parameters is provided for embodiments of this application, such as... Figure 3 As shown, the wrist node S and the handle H form a transient rigid body. The angular velocity measured by the handle IMU... and acceleration angular velocity of the wrist node and acceleration There is a fixed rotational and translational relationship.
[0043] Furthermore, according to The wrist angular velocity based on the wrist node is obtained. .in, The angular velocity of the handle is based on the built-in IMU module; The rotational extrinsic parameter is in the extrinsic parameter state. Wrist angular velocity. The high-frequency interpolation is obtained from the visual wrist node's pose node through numerical differentiation or B-spline. In other words, under rotational constraints, the rotational speed of the wrist node should be equal to the rotational speed of the handle IMU after being rotated using extrinsic parameters.
[0044] Furthermore, according to The wrist acceleration based on the wrist node is obtained. ;in, The translational extrinsic parameter in the extrinsic parameter state; This is the gravity vector. Zero bias for the accelerometer with built-in IMU module; For controller acceleration based on the built-in IMU module, It can be obtained from the second-order differential of the pose at the visual wrist node. That is, under translation and gravity constraints, the relationship between the acceleration of the wrist node and the acceleration measured by the handle IMU is calculated.
[0045] Furthermore, based on the rotational constraints and translational and gravitational constraints between the handle acceleration, handle angular velocity, wrist acceleration, and wrist angular velocity, the kinematic constraints between the hand and the handle are obtained.
[0046] Furthermore, such as Figure 3 As shown, residuals are calculated by calibrating the joint extrinsic parameters between the hand and the controller using kinematic constraints, and the sum of the minimized residuals is obtained. Then, by minimizing the sum of the minimized residuals, the optimal solution is calculated for the online extrinsic parameter state under the fusion of the hand and the controller, obtaining the optimal grip state extrinsic parameters and the corresponding IMU zero bias under unobstructed or low-occlusion grip states.
[0047] As a feasible implementation method, a nonlinear optimization problem can be constructed first within the sliding window. The online extrinsic state variables under hand-handle fusion are: The residual terms are constructed using the aforementioned kinematic constraints, and then the optimal grip state extrinsic parameters and IMU zero bias are solved by minimizing the sum of the residuals. This is analogous to multi-sensor calibration methods in VI-SLAM. This application innovatively uses the visual output joint pose as a "virtual sensor" for joint calibration with the handle IMU.
[0048] S103. Based on the extrinsic parameters of the grip state, the visual wrist node pose, handle IMU data, and infrared spot data are subjected to tight coupling fusion processing based on factor graph to obtain a wrist 6-DoF pose with high frequency and low latency.
[0049] Specifically, firstly, the pose of the visual wrist node at the visual front end is determined as the observed value of the visual factor, and a reprojection error is constructed. Then, using high-frequency handle IMU data, pre-integration is performed between adjacent visual keyframes to construct a high-frequency motion constraint based on the IMU pre-integration factor. Next, the infrared spot distribution data in the infrared spot data is determined as sparse visual features, and the sparse visual features are projected onto the wrist node coordinate system using the grip state extrinsic parameters to obtain the constrained infrared spot factor.
[0050] Furthermore, based on visual factors, IMU pre-integration factors, and infrared spot factors, an optimized factor map under tight coupling fusion of hand and handle is constructed.
[0051] In one embodiment, Figure 5 This is a schematic diagram of a tightly coupled fusion factor graph provided in an embodiment of this application, such as... Figure 5 As shown, in a non-severely occluded grip state, to obtain the optimal wrist node pose, a tightly coupled optimization framework needs to be constructed: First, the system state is defined as the current wrist node pose, velocity, IMU zero bias, etc., that is, the wrist node state variables based on the wrist node pose are defined first. Then, factor graph construction is performed, including: 1) Visual factors: using the 30Hz wrist node pose output by the visual front end as the observation value, constructing reprojection error or direct 3D position error. 2) IMU pre-integration factors: using the 500Hz IMU data of the handle, pre-integration is performed between two adjacent visual keyframes, and then high-frequency motion constraints are constructed. 3) Infrared spot factors: using the infrared spot observed on the handle as sparse visual features, using the known hand-handle grip extrinsic parameters (grip state extrinsic parameters), it is projected onto the wrist node coordinate system for constraint, to further improve positioning accuracy and robustness.
[0052] Furthermore, when the grip state is unobstructed or minimally occluded, the current wrist node pose, velocity, and IMU bias are determined as wrist node state variables based on the wrist node pose. Finally, through optimization factor graphs, the wrist node state variables are optimized based on the output of a sliding window to obtain a high-frequency, low-latency wrist 6-DoF pose. In other words, the final smooth, high-frequency, low-latency wrist 6-DoF pose is output through sliding window optimization.
[0053] S104. If the grip state result is a highly occluded grip state, then through the self-occlusion posture compensation network, the infrared spot data and the handle IMU data are processed by regression mapping of relevant spatiotemporal feature sequences to generate hand posture data based on complete hand joints.
[0054] Specifically, the infrared spot data and the handheld IMU data need to be constructed into a spatiotemporal feature sequence based on the input of a self-occlusion posture compensation network. The self-occlusion posture compensation network is an end-to-end neural network. Then, a temporal convolutional encoder is used to extract the spatiotemporal features from the spatiotemporal feature sequence; and these features are input into a fully connected layer decoder to regress the 3D positions of the hand joints based on the wrist coordinate system.
[0055] Furthermore, based on data samples of normal hand movement data and corresponding handle movement data, motion posture mapping training is performed on the 3D positions of hand joints to output hand posture data based on self-occlusion posture compensation.
[0056] In one embodiment, Figure 4A self-occlusion pose compensation network structure diagram is provided for embodiments of this application, such as... Figure 4 As shown, when the hand grips an object (such as a handle), i.e., the gripping state is a highly occluded gripping state, a large number of hand joints are obscured, making the visual output of the head-mounted display unreliable. Therefore, this application designs an end-to-end neural network to solve this problem. The network architecture includes: 1) Input: The network input does not depend on images, but on infrared light spots at fixed positions on the controller (which can provide sparse 6-DoF controller pose) and the high-frequency IMU data sequence built into the controller (angular velocity and acceleration of the past N moments), which then form a spatiotemporal feature sequence.
[0057] 2) Network Structure: An encoder based on Temporal Convolutional Network (TCN) or Long Short-Term Memory Network (LSTM) is used to extract the spatiotemporal features of the input sequence. Subsequently, the features are input into a fully connected layer decoder to directly regress the 3D positions of 21 or more hand joints (relative to the wrist coordinate system or world coordinate system).
[0058] 3) Training Strategy: The network is trained on a large amount of "clean" normal hand motion data. During training, occlusion situations when the hand is holding an object are simulated (such as randomly discarding some joint information), and corresponding handle motion data are input. The network learns the mapping relationship of inferring the complete hand posture from the handle motion, thereby obtaining hand posture data based on self-occlusion posture compensation.
[0059] As a feasible implementation, when the visual front end detects that the confidence level of the hand key points is lower than the occlusion threshold and the "grip state" is true, the system triggers a "self-occlusion mode". At this time, the hand pose is no longer provided by the visual front end, but is generated in real-time by the self-occlusion pose compensation network, i.e., the subsequently generated hand pose data. The input to the self-occlusion pose compensation network is the handle IMU data from the previous N frames and the handle pose calculated from infrared light points, and the output is complete hand joint data. When the hand leaves the object or the occlusion disappears, the system smoothly switches back to the vision-based tightly coupled fusion mode, that is, switches to a high-frequency, low-latency wrist 6-DoF pose.
[0060] S105. Based on the grip state results and the corresponding wrist 6-DoF pose and hand posture data, dynamically switch the gesture tracking between the controller and the hand, and output the final hand posture data.
[0061] Specifically, if the grip state result is a high-occlusion grip state, the final hand pose data will be determined as the hand pose data based on self-occlusion pose compensation. If the grip state result is a no-occlusion or low-occlusion grip state, the final hand pose data will be determined as the wrist 6-DoF pose based on the optimized factor map output.
[0062] Furthermore, if the grip state result changes from a high-occlusion grip state to a no-occlusion or low-occlusion grip state, the final hand pose data will be switched from the hand pose data to the wrist 6-DoF pose.
[0063] Furthermore, based on the final hand pose data, continuous tracking control of gestures between the controller and the hand is achieved, and high-precision, low-latency hand pose estimation is realized in handheld controller scenarios.
[0064] As a feasible implementation method, this application can dynamically switch the above-mentioned hand posture calculation method according to the "holding state" and the quality of visual tracking, so as to smoothly output the final hand posture.
[0065] As a feasible implementation method, this application utilizes the kinematic relationship between the visual wrist node pose and the high-frequency IMU data of the handle to solve the extrinsic parameters between the two online during dynamic gripping, without relying on the handle's own visual pose tracking or offline calibration. Furthermore, to address the self-occlusion problem when the hand is holding an object, a novel neural network architecture is proposed. This network takes handle IMU and infrared spot data as input and directly regresses the complete pose of the occluded hand, achieving robust hand tracking under visual failure. Simultaneously, the low-frequency visual wrist node, high-frequency handle IMU, and infrared spot data are tightly coupled and fused within a unified optimization framework, fully leveraging the advantages of each sensor to output high-precision, high-frequency, and low-latency 6-DoF wrist pose.
[0066] In addition, embodiments of this application also provide a high-precision gesture tracking system based on hand and controller fusion, such as Figure 6 As shown, the high-precision gesture tracking system 600 specifically includes: The grip state detection module 610 is used to determine the grip state between the handle and the hand based on the head-mounted display camera and the built-in IMU module of the handle, and to determine the grip state result. The hand-handle extrinsic parameter online calibration module 620 is used to calculate the optimal solution of the grip state extrinsic parameters between the handle and the hand under relevant kinematic constraints based on the visual wrist node posture and the handle IMU data if the grip state result between the handle and the hand is a grip state with no or low obstruction. The tightly coupled fusion and posture optimization module 630 is used to perform factor graph-based tightly coupled fusion processing on the visual wrist node posture, handle IMU data and infrared spot data according to the grip state extrinsic parameters to obtain a wrist 6-DoF pose with high frequency and low latency. The self-occlusion posture compensation network module 640 is used to perform regression mapping processing on the infrared spot data and the handle IMU data on the spatiotemporal feature sequence if the grip state result is a high occlusion grip state, and generate hand posture data based on complete hand joints. The state management module 650 is used to dynamically switch the gesture tracking between the handle and the hand based on the grip state result and the corresponding wrist 6-DoF pose and hand posture data, and output the final hand posture data.
[0067] This application achieves high-precision, low-latency hand pose estimation in handheld controller scenarios through online extrinsic parameter calibration, multi-sensor tight coupling optimization, and self-occlusion neural network compensation. Moreover, it eliminates the need for cumbersome offline calibration by the user; the system automatically completes calibration the moment the hand is grasped and adapts to different users and grip postures, greatly enhancing the user experience. Through tight coupling fusion, the 30Hz visual wrist pose is boosted to 500Hz, significantly reducing jitter and latency, making virtual hand movements smoother and more responsive. Even when the hand grips an object causing visual impairment, the hand pose compensation network driven by controller data still outputs complete and reasonable hand poses, ensuring the continuity of interaction. Simultaneously, during rapid hand movements or partial occlusion, IMU data provides strong motion priors, ensuring the overall stability of the system. Finally, it allows users to enjoy the convenience of natural gesture interaction while simultaneously gripping the controller for precise control; the system seamlessly integrates and accurately represents hand details, greatly enhancing the immersion and functionality of VR / AR applications.
[0068] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0069] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0070] The above description is merely an embodiment of this application and is not intended to limit this application. For those skilled in the art, various modifications and variations can be made to the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of this application should be included within the scope of the claims of this application.
Claims
1. A high-precision gesture tracking method based on hand and controller fusion, characterized in that, The method includes: Based on the built-in IMU module of the head-mounted display camera and the controller, the grip state between the controller and the hand is determined by the collected multi-source data, and the grip state result is determined. If the grip state between the handle and the hand is a grip state with no or low obstruction, then based on the visual wrist node posture and the handle IMU data, the optimal solution calculation under relevant kinematic constraints is performed between the handle and the hand to obtain the grip state extrinsic parameters. Based on the grip state extrinsic parameters, the visual wrist node pose, the handle IMU data, and the infrared spot data are subjected to tight coupling fusion processing based on factor graph to obtain a wrist 6-DoF pose with high frequency and low latency. If the gripping state result is a high-occlusion gripping state, then through a self-occlusion posture compensation network, regression mapping processing of relevant spatiotemporal feature sequences is performed on the infrared spot data and the handle IMU data to generate hand posture data based on complete hand joints. Based on the grip state results and the corresponding wrist 6-DoF pose and hand posture data, the gesture tracking between the controller and the hand is dynamically switched, and the final hand posture data is output.
2. The high-precision gesture tracking method based on hand and controller fusion according to claim 1, characterized in that, Before determining the grip state between the controller and hand based on the built-in IMU module of the head-mounted display camera and the controller, the method further includes: The head-mounted display camera captures images including infrared dots on the hand and the handle, resulting in hand posture images and handle infrared dot images. Collect IMU data from the IMU module built into the controller; wherein the controller IMU data is high-frequency IMU data and includes at least: acceleration data and gyroscope data; Based on the head-mounted display's VIO, determine the head-mounted display's own pose in the current head-mounted display state; The hand pose estimation network is used to identify the positions of 3D hand joints in the hand pose image to obtain the visual wrist node pose. The infrared spot image of the handle is analyzed and processed to obtain infrared spot data; Based on the spatiotemporal feature sequence, the head-mounted display's own pose, the visual wrist node pose, the controller IMU data, and the infrared spot data are spatiotemporally aligned to generate the multi-source data.
3. The high-precision gesture tracking method based on hand and controller fusion according to claim 2, characterized in that, Based on the head-mounted display's camera and the built-in IMU module of the controller, the system analyzes the collected multi-source data to determine the grip state between the controller and the hand, thus identifying the grip state result, which specifically includes: The spatiotemporal consistency of the hand geometric features in the visual wrist node posture and the handle light spot distribution features in the infrared light spot data is judged to obtain the judgment result; If the judgment result is consistent, the grip state between the handle and the hand is in a gripping state; if the judgment result is inconsistent, the grip state between the handle and the hand is in a non-gripping state. If the gripping state result is that the hand is in a gripping state, then the head-mounted display camera performs feature occlusion recognition on the hand joint features in the hand geometry features based on an occlusion threshold to obtain the hand occlusion result; wherein, the hand occlusion result includes: high occlusion result and no or low occlusion result; Based on the hand occlusion results, the holding state results in the holding state are classified to obtain the holding state with no or low occlusion and the holding state with high occlusion.
4. The high-precision gesture tracking method based on hand and controller fusion according to claim 1, characterized in that, Before obtaining the grip state extrinsic parameters by performing optimal solution calculations for the kinematic constraints between the handle and the hand based on the visual wrist node pose and the handle IMU data, the method further includes: When the gripping state is identified as the first gripping state, online external parameter calibration control is initiated; Based on the transient rigid body formed between the wrist node and the handle, the motion state of the handle's built-in IMU module and the wrist node is collected and processed in real time. according to The wrist angular velocity based on the wrist node is obtained. ;in, The angular velocity of the handle is based on the built-in IMU module; Rotational extrinsic parameters in the extrinsic parameter state; wrist angular velocity. High-frequency interpolation obtained from the visual wrist node pose node via B-spline; according to The wrist acceleration based on the wrist node is obtained. ;in, The translational extrinsic parameter in the extrinsic parameter state; It is the gravity vector; The accelerometer of the built-in IMU module has zero bias; The acceleration of the handle is based on the built-in IMU module; Based on the rotational constraints and translational and gravitational constraints between the handle acceleration, the handle angular velocity, the wrist acceleration, and the wrist angular velocity, the kinematic constraints between the hand and the handle are obtained.
5. A high-precision gesture tracking method based on hand and controller fusion according to claim 4, characterized in that, If the grip state between the handle and the hand is an unobstructed or low-obstruction grip state, then based on the visual wrist node pose and the handle IMU data, the optimal solution calculation under relevant kinematic constraints is performed to obtain the grip state extrinsic parameters, specifically including: Using the aforementioned kinematic constraints, residuals are calculated based on the joint extrinsic parameter calibration between the hand and the handle, resulting in the minimum sum of residuals. By minimizing the sum of residuals, the optimal solution is calculated for the online extrinsic state under hand and controller fusion, resulting in the optimal grip state extrinsic parameters and the corresponding IMU zero bias under no or low occlusion grip conditions.
6. The high-precision gesture tracking method based on hand and controller fusion according to claim 1, characterized in that, Before performing factor-map-based tight-coupling fusion processing on the visual wrist node pose, the handle IMU data, and the infrared spot data according to the grip state extrinsic parameters to obtain a wrist 6-DoF pose with high frequency and low latency, the method further includes: The pose of the visual wrist node at the visual front end is determined as the observed value of the visual factor, and the reprojection error is constructed. By using the high-frequency IMU data from the handle, pre-integration calculations are performed between adjacent visual keyframes to construct a high-frequency motion constraint based on the IMU pre-integration factor. The infrared spot distribution data in the infrared spot data is determined as sparse visual features, and the sparse visual features are projected onto the wrist node coordinate system through the grip state extrinsic parameters to obtain the constrained infrared spot factor. Based on the visual factor, the IMU pre-integration factor, and the infrared spot factor, an optimized factor map under tight coupling fusion of hand and handle is constructed.
7. A high-precision gesture tracking method based on hand and controller fusion according to claim 6, characterized in that, Based on the grip state extrinsic parameters, the visual wrist node pose, the handle IMU data, and the infrared spot data are subjected to factor graph-based tight coupling fusion processing to obtain a wrist 6-DoF pose with high frequency and low latency, specifically including: When the holding state result is a holding state with no or low occlusion, the current wrist node pose, velocity, and IMU zero bias are determined as wrist node state variables based on wrist node pose. The wrist node state variables are optimized based on the output of the relevant sliding window using the optimization factor graph to obtain the wrist 6-DoF pose with high frequency and low latency.
8. A high-precision gesture tracking method based on hand and controller fusion according to claim 1, characterized in that, If the gripping state result indicates a highly occluded gripping state, then a self-occlusion posture compensation network is used to perform regression mapping processing on the infrared spot data and the handle IMU data for relevant spatiotemporal feature sequences to generate hand posture data based on complete hand joints, specifically including: The infrared spot data and the handle IMU data are used to construct a spatiotemporal feature sequence based on the input of the self-occlusion attitude compensation network; wherein, the self-occlusion attitude compensation network is an end-to-end neural network. The temporal features are extracted from the spatiotemporal feature sequence by a temporal convolutional encoder; and the spatiotemporal features are input into the decoder of the fully connected layer to regress the 3D position of the hand joints based on the wrist coordinate system. Based on data samples of normal hand movement data and corresponding handle movement data, the 3D position of the hand joints is trained to perform motion posture mapping of the complete hand joints, and the hand posture data based on self-occlusion posture compensation is output.
9. A high-precision gesture tracking method based on hand and controller fusion according to claim 1, characterized in that, Based on the grip state results and the corresponding wrist 6-DoF pose and hand posture data, the gesture tracking between the controller and the hand is dynamically switched, and the final hand posture data is output, specifically including: If the gripping state result is a highly occluded gripping state, then the final hand posture data is determined as the hand posture data based on self-occlusion posture compensation; If the gripping state result is a gripping state with no or low occlusion, then the final hand posture data is determined as the wrist 6-DoF pose based on the output of the optimized factor map; If the grip state result changes from a high-occlusion grip state to a no-occlusion or low-occlusion grip state, then the final hand posture data is switched from the hand posture data to the wrist 6-DoF pose. Based on the final hand posture data, continuous tracking and control of gestures between the controller and the hand can be achieved, and high-precision, low-latency hand posture estimation can be realized in handheld controller scenarios.
10. A high-precision gesture tracking system based on hand and controller fusion, characterized in that, The system is executable by at least one processor, which enables the at least one processor to perform a high-precision gesture tracking method based on hand and handle fusion as described in any one of claims 1-9.