A three-dimensional first-person multi-modal hand trajectory prediction method and system

CN122799488APending Publication Date: 2026-09-22SHANGHAI MAJIKE IND INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610780022.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]本发明的目的是为了解决现有三维第一人称多模态手部轨迹预测方法在多模态推理过程中过度依赖历史轨迹、忽视视觉信息以及现有基准无法评估视觉模态依赖性等技术问题,本发明提供了一种基于视觉因果干预的三维第一人称多模态手部轨迹预测方法

Benefits of technology

(1)通过结构因果模型分析,首次系统性地识别了现有3DEMHTF模型过度依赖历史轨迹的技术缺陷,提出视觉因果干预方法阻断后门路径,建立从视觉信息到未来轨迹的正确因果路径,有效缓解了模型多模态推理能力退化的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799488A_ABST
    Figure CN122799488A_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional first-person multi-modal hand trajectory prediction method and system, relates to the first-person hand movement prediction field in computer vision, and utilizes a structural causal model to model a three-dimensional first-person multi-modal hand trajectory prediction task, and identifies a backdoor path introduced by a historical trajectory in an existing model; according to the do-operator principle, a visual causal intervention module is constructed by forcibly constructing part of the missing in the historical trajectory, the model is forced to utilize visual information to infer a future trajectory, so that the backdoor path is blocked, a correct causal path from the visual information to the future trajectory is established, and the problem of degradation of multi-modal reasoning capability of the existing model is effectively solved; and the application can be applied to application occasions, such as extended reality, human-computer interaction, remote education, robot operation and the like, which need to predict hand movement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and multimodal learning technology, specifically to a three-dimensional first-person multimodal hand trajectory prediction method based on visual causal intervention. Background Technology

[0002] In recent years, with the rapid development of emerging applications such as extended reality, human-computer interaction, distance education, and robot operation, first-person visual understanding has gradually become one of the important research directions in the field of computer vision. Within first-person visual understanding, 3D first-person multimodal hand trajectory prediction (3DEMHTF) is a key subtask. It aims to combine first-person visual observation with historical 3D hand trajectories to accurately predict future hand movements, and is a core technology for realizing natural human-computer interaction, decision support, and human behavior prediction.

[0003] Existing 3DEMHTF methods can be broadly categorized into two types: those based on state-space models and those based on diffusion models. The former, such as USST, introduces an uncertainty-aware state-space Transformer to model long-term hand motion sequences; the latter, such as MMTwin's dual-diffusion architecture, decouples camera motion from hand motion modeling, further improving prediction accuracy. Although these methods have achieved significant quantitative improvements on standard benchmarks, their evaluations are all based on the assumption of complete historical trajectories, failing to reveal whether the models truly utilize visual context for multimodal reasoning.

[0004] Through in-depth causal analysis, this invention reveals a severe visual bypass phenomenon in existing 3DEMHTF models: when the visual encoder is removed from the model and only historical trajectory input is retained, the model's predictive performance does not significantly degrade; on the contrary, it still maintains strong predictive accuracy. This counterintuitive phenomenon indicates that historical trajectories actually act as confounding factors in existing models, introducing a backdoor path H→F directly from historical trajectory H to future trajectory F. This backdoor path allows the model to almost completely bypass the visual inference process, degenerating the 3D multimodal hand trajectory prediction task into a pure time series extrapolation problem, severely weakening the model's multimodal inference ability and limiting its generalization in complex scenarios with rich visual changes and missing historical trajectories. Furthermore, existing evaluation benchmarks H2O-PT and EgoPAT3D-DT are both built on the assumption of complete historical trajectories, failing to quantify the model's true dependence on trajectory modalities, and even more so failing to reveal whether the model truly understands the visual context. Therefore, it is urgent to propose new evaluation benchmarks and metrics to systematically evaluate the degree to which 3DEMHTF models utilize visual modalities. Summary of the Invention

[0005] The purpose of this invention is to address the technical problems of existing 3D first-person multimodal hand trajectory prediction methods, such as over-reliance on historical trajectories, neglect of visual information, and the inability of existing benchmarks to assess visual modality dependence during multimodal inference. This invention provides a 3D first-person multimodal hand trajectory prediction method based on visual causal intervention. This invention utilizes a structural causal model to perform causal analysis on the task, identifying backdoor paths introduced by historical trajectories. Based on the do-operator principle, a visual causal intervention module is constructed by forcibly creating partial missing data in historical trajectories, forcing the model to use visual information to infer future trajectories. This blocks backdoor paths and establishes the correct causal path from visual information to future trajectories, effectively mitigating the degradation problem of the model's multimodal inference ability. Simultaneously, this invention proposes biased evaluation benchmark datasets H2O-CR and EgoPAT3D-CR based on coverage reduction, as well as average coverage degradation and average final degradation evaluation metrics. These can systematically evaluate the model's dependence on visual modality, providing a more comprehensive evaluation method for 3DEMHTF research.

[0006] A three-dimensional first-person multimodal hand trajectory prediction method includes: Acquire multimodal input, which includes visual input V and historical three-dimensional hand trajectory H; Visual causal intervention based on the do-operator is performed on the historical 3D hand trajectory H, and the post-intervention trajectory is obtained by introducing partial missing data. This is to block the backdoor path H→F from the historical trajectory to the future trajectory, forcing the prediction model to rely on visual information to infer the future trajectory. Visual anchoring information is extracted from the visual input V to generate a visual reconstruction trajectory. ; The post-intervention trajectory With the visual reconstruction trajectory The fusion is performed to obtain the causal alignment trajectory Hcomp; Multimodal reasoning is performed based on the visual input V and the causal alignment trajectory Hcomp to predict the future hand trajectory F.

[0007] The visual causal intervention is based on a structural causal model, which identifies the backdoor path H→F directly from the historical 3D hand trajectory H to the future hand trajectory F. In the structural causal model, visual information V and historical trajectory H guide future trajectory prediction through human intention I as an intermediary variable. The correct causal structure is (V,H)→I→F. The backdoor path H→F bypasses the intermediary variable I, allowing the prediction model to rely solely on the historical trajectory H for prediction, thus reducing the multimodal hand trajectory prediction task to a time series extrapolation problem.

[0008] The visual causal intervention is performed based on the do-operator principle, and the intervention operation is formally represented as follows: The missing information includes: defining a coverage ratio α∈[0,1] as a control parameter, specifying the percentage of historical trajectory path points provided to the prediction model out of the total observation window path points; constructing a binary mask. The post-intervention trajectory , where Tobs is the number of observation window frames, and ⊙ represents element-wise multiplication.

[0009] The missing portion of the introduction employs an end-truncation strategy: (1-α) proportion of trajectory points are continuously removed from the end of the observation window, retaining the first α×Tobs of the initial trajectory points. The binary mask Mα is at the index... The value is 1 when the time is right and 0 otherwise. The end-truncation strategy eliminates the most recent motion information that is closest to the prediction window in time, so that the prediction model cannot predict the future trajectory by temporal extrapolation, thereby forcing the prediction model to make up for the missing trajectory segment by using the complete visual input V.

[0010] The step of extracting visual anchoring information from visual input V includes: using a pre-trained hand keypoint detector to detect hand keypoints in each frame of RGB image within the observation window, and outputting the two-dimensional coordinates of multiple hand keypoints in the pixel coordinate system; selecting a wrist keypoint from the multiple hand keypoints as a representative of the hand position, wherein the wrist keypoint is located at the base of the hand, and provides stable hand position tracking under different hand poses and occlusion conditions.

[0011] The generation of the visual reconstruction trajectory includes: acquiring a depth map aligned with the RGB image from a first-person camera, querying the depth value at the pixel coordinates of the wrist key points; back-projecting the two-dimensional wrist coordinates to the three-dimensional camera coordinate system using the camera intrinsic parameter matrix to obtain the visually anchored three-dimensional trajectory points; the back-projection process is entirely based on first-person visual observation and depth information, independent of the historical three-dimensional hand trajectory H, providing independent visual evidence for the visual causal intervention.

[0012] The generation of the visually reconstructed trajectory further includes: applying a Savitzky-Golay filter to the sequence of visually anchored 3D trajectory points obtained by backprojection along the temporal dimension for smoothing. The Savitzky-Golay filter uses the least squares method to perform polynomial fitting on the data points within the sliding window, preserving motion dynamic features while suppressing keypoint detection noise, thus obtaining the visually reconstructed trajectory. The preferred window size for the Savitzky-Golay filter is 7.

[0013] Among them, the post-intervention trajectory With visual reconstruction trajectory The fusion process includes: selective fusion using the binary mask Mα, and using the intervened trajectory at positions where the mask value is 1. The original trajectory points in the image are used to reconstruct the trajectory using the vision at positions with a mask value of 0. The causal alignment trajectory is obtained by visually anchoring trajectory points in the data. .

[0014] This includes a debiasing evaluation step based on a coverage reduction protocol: on an existing first-person hand trajectory prediction dataset, historical trajectories are truncated from the end of the observation window according to different coverage α values ​​in the evaluation coverage ratio set A, while the visual input V and future trajectory annotations are fully preserved, thus constructing a debiasing evaluation benchmark dataset; a unified degradation operation MD(ε)=(1 / (|A|-1))·Σ{αi∈A{αmax}}(ε(αi)-ε(αmax)) / (αmax-αi) is defined, and when the degradation operation MD is combined with the average displacement error ADE, the average coverage degradation MCD is obtained, and when the degradation operation MD is combined with the final displacement error FDE, the average final degradation MFD is obtained, which is used to quantify the degree of dependence of the prediction model on the visual modality under different degrees of historical trajectory loss.

[0015] This invention also relates to a three-dimensional first-person multimodal hand trajectory prediction system, comprising: A multimodal input module is used to acquire multimodal inputs, including visual input V and historical 3D hand trajectory H; The visual causal intervention module is used to perform visual causal intervention on the historical three-dimensional hand trajectory H, and obtain the post-intervention trajectory by introducing partial missing data. This is to block the backdoor path H→F from the historical trajectory to the future trajectory, forcing the prediction model to rely on visual information to infer the future trajectory. The visual reconstruction module is used to extract visual anchoring information from the visual input V and generate a visual reconstruction trajectory. ; The causal alignment and fusion module is used to integrate the post-intervention trajectory With the visual reconstruction trajectory The fusion is performed to obtain the causal alignment trajectory Hcomp; The trajectory prediction module is used to perform multimodal reasoning based on the visual input V and the causal alignment trajectory Hcomp to predict the future hand trajectory F.

[0016] The technical solution adopted in this invention is as follows: The beneficial effects of this invention are as follows: (1) Through structural causal model analysis, the technical defects of existing 3DEMHTF models that rely too much on historical trajectories were systematically identified for the first time. A visual causal intervention method was proposed to block the backdoor path and establish the correct causal path from visual information to future trajectory, which effectively alleviated the problem of degradation of the model's multimodal reasoning ability.

[0017] (2) The proposed visual causal intervention module is a lightweight preprocessing plugin that can be seamlessly integrated into various existing three-dimensional first-person multimodal hand trajectory prediction architectures, including USST and MMTwin, without modifying the internal structure of the prediction model. It has good versatility, lightweight and portability.

[0018] (3) New debiased evaluation benchmark datasets H2O-CR and EgoPAT3D-CR were constructed, and MCD and MFD evaluation indicators were proposed. These can systematically evaluate the dependence of the model on the visual modality and provide a more comprehensive evaluation method for research in the 3DEMHTF field.

[0019] (4) Experimental results show that the method of the present invention achieves state-of-the-art performance on standard benchmarks H2O-PT and EgoPAT3D-DT as well as coverage reduction benchmarks H2O-CR and EgoPAT3D-CR, especially under the condition of missing trajectory, it has a significant improvement in robustness compared with the baseline method. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall framework of the method proposed in this invention, which sequentially shows four stages: multimodal input, visual causal intervention module, causal alignment trajectory fusion, and future trajectory prediction.

[0021] Figure 2 This is a visualization of the prediction results of the proposed method under different historical trajectory coverage, demonstrating the robustness advantage of VCI over the baseline method in the case of trajectory missing. Detailed Implementation

[0022] like Figure 1 A three-dimensional first-person multimodal hand trajectory prediction method based on visual causal intervention includes the following steps: Step 1: Constructing the multimodal input. Given an observation window containing Tobs frames, the model receives visual input V={v1,v2,...,vTobs} consisting of RGB images and the corresponding historical 3D hand trajectories H={x1,x2,...,xTobs}, where... Let Xpred be the coordinate position of the hand in three-dimensional space at time t; the goal is to predict the hand trajectory Xpred={xTobs+1,...,xTobs+Tpred} in the future Tpred frame.

[0023] Step 2: Conduct causal analysis and problem identification based on the structural causal model. Identify backdoor paths H→F in the existing model that directly lead from historical trajectories to future trajectories, which result in spurious correlations.

[0024] Step 3: Visual causal intervention based on the do-operator. By forcibly introducing partial missing data into historical trajectories, spurious temporal dependencies are disrupted, forcing the model to rely on visual information to infer future motion.

[0025] Step 4: Hand keypoint detection. 21 hand keypoints were extracted from the RGB image using MediaPipe, and wrist keypoints were selected for stable hand position tracking.

[0026] Step 5: Depth-guided 3D reconstruction. Using the depth map and camera intrinsics, the 2D wrist detection is back-projected onto the 3D camera coordinate system to obtain the visually anchored 3D path points.

[0027] Step 6: Temporal Filtering. Apply a Savitzky-Golay filter to temporally smooth the reconstructed trajectory, generating a complete visual reconstructed trajectory.

[0028] Step 7: Causal alignment trajectory fusion. The visually reconstructed trajectory is fused with a portion of the historical trajectory using a binary mask to generate a causal alignment trajectory.

[0029] Step 8: Future Trajectory Prediction. Input the causal aligned trajectory along with the complete visual input into the prediction model, perform multimodal inference, and output the predicted future hand trajectory.

[0030] As a preferred technical solution, step 1 further includes, in more detail, a formal definition of the multimodal input: Given an observation window containing Tobs frames, the visual input acquired from a first-person camera is defined as an RGB image sequence V = {v1, v2, ..., vTobs}, where each frame vt is the RGB image at time t. Simultaneously, aligned historical 3D hand trajectories H = {x1, x2, ..., xTobs} are acquired, where... Let represent the coordinates of the hand and wrist position in three-dimensional space at time t. Based on the above multimodal input, the prediction objective of the method is to output the three-dimensional hand trajectory Xpred={xTobs+1,...,xTobs+Tpred} for the future Tpred frame. This invention uses a structural causal model (SCM) to formally model the above task, where V represents visual information, H represents historical trajectory, I represents human intention, and F represents the predicted future trajectory. The correct causal structure can be represented as: , I serves as a mediating variable, integrating visual information V with historical trajectory H to guide future trajectory prediction, thereby achieving true multimodal reasoning.

[0031] As a preferred technical solution, step 2 further includes causal analysis and backdoor path identification: Through controlled modal perturbation experiments with the visual encoder removed while keeping historical trajectory input unchanged, the behavior of existing 3D first-person multimodal hand trajectory prediction models was analyzed: If the model truly uses visual information V and historical trajectory H to jointly infer intent I, then the absence of V should lead to a significant decrease in prediction performance; however, empirical results show that the existing model still maintains strong prediction accuracy when visual input is missing, indicating that the future trajectory F can be inferred almost independently of V. This phenomenon reveals that the existing model utilizes a shortcut dependency introduced by historical trajectory H, i.e., a backdoor path. , This backdoor bypasses the visual-trajectory joint mediation inference performed by Intent I, allowing the model to predict future trajectories F solely based on historical trajectories H, thus reducing the 3D first-person multimodal hand trajectory prediction task to a time series extrapolation problem. Furthermore, due to this backdoor, the model's prediction performance significantly degrades when historical trajectories H are missing or incomplete, further limiting the model's generalization ability in complex scenarios such as missing historical trajectories.

[0032] As a preferred technical solution, step 3 further includes, in more detail, the principle of visual causal intervention based on the do-operator: Based on the do-operator principle in causal inference, the post-intervention trajectory is obtained by forcibly introducing partial missing data into the historical trajectory H. This disrupts the shortcut dependency from H to F; the intervention operation is formally represented as: , This formalization shows that after performing do-operator intervention on the historical trajectory H, the predicted distribution of the future trajectory F is... Equivalent to the intervention trajectory Conditional distribution under visual input V This effectively blocks backdoor paths, forcing the model to rely on visual information V to infer future motion. Through this intervention, the correct causal path V→I→F from visual information V to future trajectory F is re-established, achieving correct multimodal reasoning.

[0033] As a preferred technical solution, steps 3 and 7 further include, in more detail, coverage reduction intervention: Step 3.1: Define the coverage ratio α∈[0,1] as a control parameter to specify the percentage of historical trajectory path points provided to the model out of the total observation window path points; construct a binary mask Mα, which belongs to the domain: , Each component of the mask is defined according to the following formula: when index i satisfies The value is 1 when the time is right and 0 otherwise; partial historical trajectory after intervention. for: ; Step 3.2: The intervention employs an end-truncation strategy, which involves continuously removing a (1-α) proportion of path points from the end of the observation window, retaining the first α×Tobs of the initial path points. Compared to the random masking strategy, the end-truncation strategy eliminates the most recent motion information that is temporally closest to the prediction window, making it more challenging for the model to predict future trajectories through simple temporal extrapolation. This forces the model to utilize the complete visual input V to compensate for the missing trajectory segments, thereby truly leveraging visual information for multimodal inference.

[0034] As a preferred technical solution, step 4 includes, in more detail, the specific implementation of hand key point detection: Step 4.1: Use the pre-trained MediaPipe hand detector to detect hand key points in each frame of RGB image vt within the observation window, and output the two-dimensional coordinates of 21 hand key points in the pixel coordinate system, including wrist key points, the base of the five fingers, the first joint, the second joint, and the fingertip key points.

[0035] Step 4.2: Select wrist key points from 21 key points As a representative of hand position, ut and vt represent the horizontal and vertical coordinates of the wrist keypoint in the pixel coordinate system, respectively. The reason for selecting the wrist keypoint is that, located at the base of the hand, it provides more stable hand position tracking than the knuckle keypoints under different hand poses, different hand orientations, different hand opening and closing states, and under slight occlusion, thus ensuring the stability of subsequent 3D reconstruction.

[0036] As a preferred technical solution, step 5 further includes the specific steps of depth-guided 3D reconstruction: Step 5.1: Obtain a depth map Dt from the first-person camera that is aligned with the RGB image frame vt. Query the depth value dt=Dt(ut,vt) at the wrist pixel coordinates (ut,vt), where Dt represents the query function of the depth map at time t, and dt represents the actual physical depth value corresponding to the wrist position.

[0037] Step 5.2: Using the focal length (fx, fy) and principal point (cx, cy) in the camera intrinsic parameter matrix, back-project the 2D wrist detection onto the 3D camera coordinate system according to the following formula to obtain the 3D path points for visual anchoring: , Where (fx, fy) are the focal lengths of the camera in the horizontal and vertical directions, respectively, and (cx, cy) are the pixel coordinates of the camera's principal point. This represents the three-dimensional path point obtained through visual reconstruction at time t.

[0038] Step 5.3: The three-dimensional reconstruction process is entirely based on first-person visual observation and depth information, independent of historical trajectory patterns, providing independent visual evidence for subsequent visual causal intervention and avoiding dependence on historical trajectory patterns.

[0039] As a preferred technical solution, step 6 further includes the specific steps of time-series filtering: Step 6.1: Reconstruct the 3D hand trajectory sequence obtained in Step 5. Along the time-series dimension, a Savitzky-Golay filter with a window size of w is used for smoothing. This filter employs least squares to perform polynomial fitting on the data points within the window, thereby suppressing noise while preserving dynamic motion characteristics. The filtering formula is: , Where SG() represents the Savitzky-Golay filter function, and ⌊⌋ is the floor function. The filtering is performed independently on the three spatial dimensions of x, y, and z.

[0040] Step 6.2: After filtering, the complete visual reconstruction trajectory is obtained: ; Step 6.3: Perform sensitivity analysis on the Savitzky-Golay filter window size w, with a candidate value set of {3,5,7,9,11}. Experimental results show that w is preferably 7: when w is too small (e.g., w=3, w=5), filtering is insufficient, keypoint detection noise is not effectively suppressed, leading to a decrease in prediction accuracy; when w is too large (e.g., w=9, w=11), oversmoothing occurs, blurring key fast motion details, which also leads to a degradation in prediction performance; when w=7, detection noise is effectively suppressed while retaining key motion dynamic features, achieving optimal performance on the standard benchmark.

[0041] As a preferred technical solution, step 7 further includes, in more detail, causal alignment trajectory fusion and integration with the existing architecture: Step 7.1: Given a portion of the historical trajectory after truncating the terminal (1-α) proportion path points according to the coverage ratio α. and the visual reconstruction trajectory obtained in step 6 By selectively fusing the original path points preserved by binary mask Mα with the visually reconstructed path points, the causal aligned trajectory Hcomp is obtained: , Where ⊙ represents element-wise multiplication, and the mask Mα controls the fusion method: the original historical trajectory is used at positions with a mask value of 1 (i.e., the initial α×Tobs path points). The actual path points in the image are used to visually reconstruct the trajectory at the positions with a mask value of 0 (i.e., the (1-α)×Tobs path points at the end). Visual anchor path points in the image.

[0042] Step 7.2: Embed the Visual Causal Intervention Module (VCI) as a preprocessing plugin into an existing 3D first-person multimodal hand trajectory prediction architecture, including but not limited to USST and MMTwin, without modifying the internal structure of the prediction model. This modular design has the following advantages: (i) versatility, making the visual causal intervention method widely applicable to different 3D first-person multimodal hand trajectory prediction methods; (ii) lightweight, as the VCI module is only a preprocessing step and does not introduce additional trainable parameters; (iii) portability, allowing seamless integration with future more advanced prediction models and continuous benefit from advancements in prediction models.

[0043] As a preferred technical solution, step 8 includes, in more detail, the specific process of predicting the future trajectory: The causal aligned trajectory Hcomp obtained in step 7, together with the complete visual input V, is used as input to a 3D first-person multimodal hand trajectory prediction model. This prediction model infers human intention I based on V and Hcomp, and then outputs a 3D hand trajectory prediction result F for future Tpred frames. Specifically, the prediction model uses a visual encoder to extract visual features from the visual input V and a trajectory encoder to extract trajectory features from the causal aligned trajectory Hcomp. These two features are then fused and decoded to output the future trajectory prediction result. Because the historical trajectory is interfered with, the prediction model cannot rely solely on trajectory extrapolation to generate the prediction result; it must combine visual features for multimodal inference to recover the correct causal path (V,H)→I→F.

[0044] As a preferred technical solution, the method also includes constructing a biased evaluation benchmark dataset and evaluation metrics: Step 9.1: Coverage Reduction Protocol. Based on the existing first-person 3D hand trajectory prediction datasets H2O-PT and EgoPAT3D-DT, the biased evaluation benchmark datasets H2O-CR and EgoPAT3D-CR are constructed according to the coverage reduction protocol. Specifically, the evaluation coverage ratio set A={0.5,0.6,0.7,0.8,0.9,1.0} is defined. For each sample, the historical trajectory is truncated from the end of the observation window at different α values, while the original visual input V and the future trajectory annotation are fully preserved. In this way, the robustness of the model under different degrees of missing historical trajectories is systematically evaluated.

[0045] Step 9.2: Degradation Evaluation Metrics. Based on the unified formal definition of degradation under coverage reduction, two complementary metrics are constructed to evaluate trajectory accuracy degradation and endpoint stability degradation. Specifically, the unified degradation operation MD() is defined as follows: , The MD() operation quantifies the normalized degradation rate of the error metric ε as coverage decreases by comparing the rate of change of the error relative to the full coverage baseline for each reduced coverage level. When MD is combined with the average displacement error (ADE), the mean coverage degradation (MCD) is obtained; when MD is combined with the final displacement error (FDE), the mean final degradation (MFD) is obtained. ; Where MCD reflects the robustness of the model's predictions at all time points, MFD assesses the stability of the endpoint prediction, A is the set of assessment coverage proportions, and αmax=1.0 corresponds to complete trajectory coverage. The lower the MCD and MFD values, the more gradual the performance degradation of the model in the case of missing trajectories, and the more fully it utilizes the visual modality; conversely, the higher the MCD and MFD values, the stronger the model's dependence on the continuity of historical trajectories, and the weaker its multimodal inference ability.

[0046] This invention also relates to a three-dimensional first-person multimodal hand trajectory prediction system, comprising: A multimodal input module is used to acquire multimodal inputs, including visual input V and historical 3D hand trajectory H; The visual causal intervention module is used to perform visual causal intervention on the historical three-dimensional hand trajectory H, and obtain the post-intervention trajectory by introducing partial missing data. This is to block the backdoor path H→F from the historical trajectory to the future trajectory, forcing the prediction model to rely on visual information to infer the future trajectory. The visual reconstruction module is used to extract visual anchoring information from the visual input V and generate a visual reconstruction trajectory. ; The causal alignment and fusion module is used to integrate the post-intervention trajectory With the visual reconstruction trajectory The fusion is performed to obtain the causal alignment trajectory Hcomp; The trajectory prediction module is used to perform multimodal reasoning based on the visual input V and the causal alignment trajectory Hcomp to predict the future hand trajectory F.

[0047] The present invention will be further described below with reference to specific embodiments, but these should not be construed as limiting the scope of protection of the present invention.

[0048] First scenario case: Performance was evaluated on the H2O-PT dataset, which contains 571 first-person RGB-D video sequences with high-quality 3D hand trajectory annotations. This invention integrates the proposed visual causal intervention module as a preprocessing plugin into the state-of-the-art models USST and MMTwin, trained using a single NVIDIA-A6000 GPU, with all training hyperparameters consistent with the original implementation. Performance evaluation uses average displacement error and final displacement error, in meters. The comparison results between this invention and other existing methods are shown in Table 1.

[0049] Table 1. Performance comparison on the H2O-PT and H2O-CR datasets. Second scenario example: Performance evaluation was conducted on the EgoPAT3D-DT dataset, which provides diverse operational scenarios across visible and unseen object categories. Using MMTwin as a baseline, this invention integrates Visual Anchoring (VCI) into MMTwin and evaluates its performance on EgoPAT3D-DT and EgoPAT3D-CR under standard and coverage reduction settings. Results show that VCI consistently improves performance for both visible and unseen scenarios, and visual anchoring helps the model adapt more effectively to new object scenarios. Comparison results between this invention and other existing methods are shown in Table 2.

[0050] Table 2 shows the performance comparison on the EgoPAT3D-DT and EgoPAT3D-CR datasets. Third scenario case: Ablation experiments and parameter sensitivity analysis. Ablation analysis was performed on each component of the VCI module on the EgoPAT3D-DT dataset. The results show that hand keypoint detection and 3D reconstruction together constitute the geometric basis of visual anchoring, causing a slight decrease in ADE / FDE from the baseline. The temporal filtering module contributed the most significant performance improvement, proving its necessity for eliminating keypoint detection noise and preserving motion dynamic features. Further sensitivity analysis was performed on the window size w of the Savitzky-Golay filter (w∈{3,5,7,9,11}). The results show that the performance is optimal when w=7. When w is too small, the filtering is insufficient and the detection noise is not effectively suppressed. When w is too large, oversmoothing occurs, blurring the fast motion details of key points. The specific ablation experiments and parameter sensitivity analysis results are shown in Table 3.

[0051] Table 3. Sensitivity analysis of component ablation experiments and time-series filtering window size w. Fourth scenario case: Robustness Validation. The performance of VCI and baseline models was compared on the EgoPAT3D-CR dataset under different levels of missing historical trajectories. As the coverage α gradually decreased from 1.0 to 0.5, the baseline methods (USST, MMTwin) showed significant performance degradation, with ADE increasing rapidly with decreasing coverage, indicating that these models heavily rely on complete historical trajectories and struggle to generalize in situations with sparse historical observations. In contrast, the model equipped with VCI exhibited significantly improved robustness: the performance curve remained essentially flat across all coverage levels, showing only slight degradation even under severe trajectory loss (α=0.5). This stability demonstrates that VCI enables the model to effectively utilize visual cues to compensate for missing motion history, avoiding over-reliance on historical trajectories and maintaining reliable predictive performance even under challenging conditions with incomplete input.

[0052] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A three-dimensional first-person multimodal hand trajectory prediction method, characterized in that, include: Acquire multimodal input, which includes visual input V and historical three-dimensional hand trajectory H; Visual causal intervention based on the do-operator is performed on the historical 3D hand trajectory H, and the post-intervention trajectory is obtained by introducing partial missing data. This is to block the backdoor path H→F from the historical trajectory to the future trajectory, forcing the prediction model to rely on visual information to infer the future trajectory. Visual anchoring information is extracted from the visual input V to generate a visual reconstruction trajectory. ; The post-intervention trajectory With the visual reconstruction trajectory The fusion is performed to obtain the causal alignment trajectory Hcomp; Multimodal reasoning is performed based on the visual input V and the causal alignment trajectory Hcomp to predict the future hand trajectory F.

2. The method according to claim 1, wherein, The visual causal intervention is based on a structural causal model, which identifies the backdoor path H→F directly from the historical 3D hand trajectory H to the future hand trajectory F. In the structural causal model, visual information V and historical trajectory H guide future trajectory prediction through human intention I as an intermediary variable. The correct causal structure is (V,H)→I→F. The backdoor path H→F bypasses the intermediary variable I, allowing the prediction model to rely solely on the historical trajectory H for prediction, thus reducing the multimodal hand trajectory prediction task to a time series extrapolation problem.

3. The method according to claim 1, wherein, The visual causal intervention is performed based on the do-operator principle, and the intervention operation is formally represented as follows: The missing information includes: defining a coverage ratio α∈[0,1] as a control parameter, specifying the percentage of historical trajectory path points provided to the prediction model out of the total observation window path points; constructing a binary mask. The post-intervention trajectory , where Tobs is the number of observation window frames, and ⊙ represents element-wise multiplication.

4. The method according to claim 3, wherein, The missing portion of the introduction employs an end-truncation strategy: (1-α) proportion of trajectory points are continuously removed from the end of the observation window, retaining the first α×Tobs of the initial trajectory points. The binary mask Mα is at the index... The value is 1 when the time is right and 0 otherwise. The end-truncation strategy eliminates the most recent motion information that is closest to the prediction window in time, so that the prediction model cannot predict the future trajectory by temporal extrapolation, thereby forcing the prediction model to make up for the missing trajectory segment by using the complete visual input V.

5. The method according to claim 1, wherein, The extraction of visual anchoring information from visual input V includes: using a pre-trained hand keypoint detector to detect hand keypoints in each frame of RGB image within the observation window, and outputting the two-dimensional coordinates of multiple hand keypoints in the pixel coordinate system; selecting wrist keypoints from the multiple hand keypoints as representatives of the hand position, wherein the wrist keypoints are located at the base of the hand, providing stable hand position tracking under different hand poses and occlusion conditions.

6. The method according to claim 5, wherein, The generation of the visual reconstruction trajectory includes: acquiring a depth map aligned with the RGB image from a first-person camera, querying the depth value at the pixel coordinates of the wrist key points; using the camera intrinsic parameter matrix to back-project the two-dimensional wrist coordinates to the three-dimensional camera coordinate system to obtain the visually anchored three-dimensional trajectory points; the back-projection process is entirely based on first-person visual observation and depth information, independent of the historical three-dimensional hand trajectory H, providing independent visual evidence for the visual causal intervention.

7. The method according to claim 6, wherein, The generation of the visually reconstructed trajectory further includes: applying a Savitzky-Golay filter to the sequence of visually anchored 3D trajectory points obtained by backprojection along the temporal dimension for smoothing. The Savitzky-Golay filter uses the least squares method to perform polynomial fitting on the data points within the sliding window, preserving motion dynamic features while suppressing keypoint detection noise, thus obtaining the visually reconstructed trajectory. The preferred window size for the Savitzky-Golay filter is 7.

8. The method according to claim 3, wherein, The post-intervention trajectory With visual reconstruction trajectory The fusion process includes: selective fusion using the binary mask Mα, and using the intervened trajectory at positions where the mask value is 1. The original trajectory points in the image are used to reconstruct the trajectory using the vision at positions with a mask value of 0. The causal alignment trajectory is obtained by visually anchoring trajectory points in the data. .

9. The method according to claim 1, wherein, It also includes a debiasing evaluation step based on the coverage reduction protocol: on the existing first-person hand trajectory prediction dataset, historical trajectories are truncated from the end of the observation window according to different coverage α values ​​in the evaluation coverage ratio set A, while the visual input V and future trajectory annotations are completely preserved, and a debiasing evaluation benchmark dataset is constructed; a unified degradation operation MD(ε)=(1 / (|A|-1))·Σ{αi∈A{αmax}}(ε(αi)-ε(αmax)) / (αmax-αi) is defined, when the degradation operation MD is combined with the average displacement error ADE, the average coverage degradation MCD is obtained, and when the degradation operation MD is combined with the final displacement error FDE, the average final degradation MFD is obtained, which is used to quantify the degree of dependence of the prediction model on the visual modality under different degrees of historical trajectory missingness.

10. A three-dimensional first-person multimodal hand trajectory prediction system, characterized in that, include: A multimodal input module is used to acquire multimodal inputs, including visual input V and historical 3D hand trajectory H; The visual causal intervention module is used to perform visual causal intervention on the historical three-dimensional hand trajectory H, and obtain the post-intervention trajectory by introducing partial missing data. This is to block the backdoor path H→F from the historical trajectory to the future trajectory, forcing the prediction model to rely on visual information to infer the future trajectory. The visual reconstruction module is used to extract visual anchoring information from the visual input V and generate a visual reconstruction trajectory. ; The causal alignment and fusion module is used to integrate the post-intervention trajectory With the visual reconstruction trajectory The fusion is performed to obtain the causal alignment trajectory Hcomp; The trajectory prediction module is used to perform multimodal reasoning based on the visual input V and the causal alignment trajectory Hcomp to predict the future hand trajectory F.