Motion prediction method and system based on multi-modal data fusion and dynamic intention

By using multimodal data fusion and dynamic intent modeling, the limitations of single-modal data in motion prediction are overcome, enabling a deeper understanding of the behavior of moving subjects and more accurate predictions. This improves the interpretability and credibility of predictions and makes the method applicable to multiple application scenarios.

CN121767718APending Publication Date: 2026-03-31周知源
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing motion prediction technologies have limited capabilities in single-modal data dimensions, making it difficult to fully capture complex motion characteristics. Furthermore, they lack a deep understanding of the subjective intentions of the moving subject, resulting in insufficient prediction accuracy and interpretability, and a lack of universality.

Method used

We employ a multimodal data fusion and dynamic intent approach. By collecting and extracting features from visual, sensor, and intent data, we combine a trajectory prediction network and an intent analysis network, utilize an inverse reinforcement learning framework to model intent, and perform secondary fusion to output the final prediction result.

Benefits of technology

It achieves more accurate and reliable motion prediction in complex and uncertain scenarios, improves the interpretability and credibility of prediction results, and can be applied in multiple fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767718A_ABST
    Figure CN121767718A_ABST
Patent Text Reader

Abstract

The invention discloses a motion prediction method and system based on multi-modal data fusion and dynamic intention. The method comprises the following steps: carrying out acquisition and time synchronization processing on multi-source heterogeneous data of a motion subject; performing feature extraction on the original data of different modes to obtain a visual feature vector, a sensor feature vector and an intention feature vector, and performing fusion to obtain a comprehensive feature vector; outputting a trajectory prediction result based on the comprehensive feature vector; outputting a quantized intention probability vector based on the intention data; performing secondary fusion on the trajectory prediction result and the intention probability vector to obtain a corrected final prediction result; and outputting a final prediction result and a corresponding intention analysis report. According to the motion prediction method and system based on multi-modal data fusion and dynamic intention provided by the invention, dynamic intention modeling is introduced, so that the motivation behind the motion behavior can be understood in a deeper level, and more accurate and more reliable prediction can be made in a complex and uncertain scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent motion prediction technology, specifically relating to a motion prediction method and system based on multimodal data fusion and dynamic intent. Background Technology

[0002] Motion prediction technology, as a core component of the cross-application of artificial intelligence and big data in sports, healthcare, transportation, and other fields, is experiencing a phase of rapid development. This technology aims to make high-precision predictions of the future behavior, trajectory, or state of a moving subject by analyzing historical and real-time data. Its application value has expanded beyond traditional sports analysis to multiple cutting-edge fields such as athlete injury risk management, medical rehabilitation guidance, pedestrian avoidance in autonomous driving, and even gaming and smart home interaction. This technology provides not only insights into the present but also forward-looking guidance for the future, offering unprecedented decision support for coaches, doctors, system designers, and others.

[0003] The limited dimensions of single-modal data make it difficult to comprehensively capture complex motion characteristics, thus limiting prediction accuracy. To compensate for the shortcomings of single data sources, some technical solutions have begun to integrate multimodal data. However, most of these solutions remain at the level of simply splicing or merging different data sources, failing to delve into the deeper meaning behind the data, particularly the "subjective intent" of the moving subject. Furthermore, current solutions are generally applied to specific scenarios, such as motion analysis for specific sports (e.g., pickleball) or early warning for specific states (e.g., sports injury risk), lacking universality and difficult to transfer to other fields. Existing solutions generally lack interpretability; the absence of intent information significantly reduces the reliability and credibility of the prediction results. Summary of the Invention

[0004] This invention provides a motion prediction method and system based on multimodal data fusion and dynamic intent to solve the technical problem of inaccuracy of traditional early warning methods mentioned above. Specifically, the technical solution is as follows:

[0005] A motion prediction method based on multimodal data fusion and dynamic intent includes:

[0006] The multi-source heterogeneous data of the moving subject is collected and time-synchronized, and the multi-source heterogeneous data includes at least visual data, sensor data and intent data.

[0007] Feature extraction is performed on the raw data of different modalities to obtain visual feature vectors, sensor feature vectors and intent feature vectors, and the visual feature vectors, sensor feature vectors and intent feature vectors are fused to obtain a comprehensive feature vector;

[0008] The trajectory prediction network outputs trajectory prediction results based on the comprehensive feature vector.

[0009] The intent analysis network outputs a quantized intent probability vector based on intent data.

[0010] The trajectory prediction result and the intention probability vector are fused twice to obtain the corrected final prediction result;

[0011] Output the final prediction result and the corresponding intent analysis report.

[0012] Furthermore, a comprehensive feature vector is obtained by fusing the visual feature vector, the sensor feature vector, and the intent feature vector through a multi-head cross-attention mechanism.

[0013] Furthermore, the intent analysis subnetwork infers the potential reward function from historical expert demonstration data based on an inverse reinforcement learning framework and outputs a quantized intent probability vector.

[0014] Furthermore, the inverse reinforcement learning framework employs maximum entropy inverse reinforcement learning, with the optimization objective being:

[0015]

[0016] Where w is the weight of the reward function, and φ(ξ) is the feature vector of the expert demonstration trajectory ξ.

[0017] Furthermore, the specific method for performing a secondary fusion of the trajectory prediction result and the intention probability vector to obtain the corrected final prediction result is as follows:

[0018] Using the intent probability vector as weights, the confidence levels of multiple possible trajectories in the trajectory prediction result are proportionally weighted and corrected to obtain the final prediction result.

[0019] Furthermore, the trajectory prediction network is based on a Transformer model or a conditional generative adversarial network.

[0020] A motion prediction system based on multimodal data fusion and dynamic intent includes:

[0021] The data acquisition and synchronization module is used to acquire and time-synchronize multi-source heterogeneous data of the moving subject. The multi-source heterogeneous data includes at least visual data, sensor data and intent data.

[0022] The feature extraction and fusion module is used to extract features from the raw data of different modalities to obtain visual feature vectors, sensor feature vectors and intent feature vectors, and to fuse the visual feature vectors, sensor feature vectors and intent feature vectors to obtain a comprehensive feature vector.

[0023] The hierarchical prediction module includes an independent trajectory prediction network and an intent analysis network. The trajectory prediction network is used to output multiple possible future trajectories and their confidence levels based on the comprehensive feature vector. The intent analysis network is used to output a quantized intent probability vector based on intent data.

[0024] The correction fusion module is used to perform a secondary fusion of the trajectory prediction result and the intention probability vector to obtain a corrected final prediction result;

[0025] The output module is used to output the final prediction result and the corresponding intent analysis report.

[0026] Furthermore, the feature extraction and fusion module fuses the visual feature vector, the sensor feature vector, and the intent feature vector through a multi-head cross-attention mechanism to obtain a comprehensive feature vector.

[0027] Furthermore, the intent analysis network infers the potential reward function from historical expert demonstration data and outputs a quantized intent probability vector based on an inverse reinforcement learning framework.

[0028] Furthermore, the correction fusion module uses the intent probability vector as weight to perform probability-weighted correction on the confidence of multiple possible trajectories in the trajectory prediction result, thereby obtaining the final prediction result.

[0029] The motion prediction method and system based on multimodal data fusion and dynamic intent provided by this invention can gain a deeper understanding of the motivations behind motion behavior by introducing dynamic intent modeling, thereby making more accurate and reliable predictions in complex and uncertain scenarios.

[0030] The core concept of the motion prediction method and system based on multimodal data fusion and dynamic intent provided by this invention lies in elevating the dimension of motion prediction from the traditional "prediction result" (such as trajectory or action) to "prediction result and cause," that is, from "data-based" prediction to "data-and-intent-based" prediction. By integrating more comprehensive heterogeneous data sources and introducing an independent intent analysis sub-network, and using methods such as inverse reinforcement learning to model the intrinsic motivation of the moving subject, the intent information is finally fused with the motion trajectory prediction in a secondary manner, thereby fundamentally solving the pain points of existing technologies in terms of interpretability, accuracy, and subjective intent grasp. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram of a motion prediction method based on multimodal data fusion and dynamic intent according to this application. Detailed Implementation

[0033] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0034] like Figure 1 The image shows a motion prediction method based on multimodal data fusion and dynamic intent according to this application, comprising:

[0035] S1: Collect and time-synchronize multi-source heterogeneous data of the moving subject. This multi-source heterogeneous data includes at least visual data, sensor data, and intent data. S2: Extract features from the raw data of different modalities to obtain visual feature vectors, sensor feature vectors, and intent feature vectors. Fuse these vectors to obtain a comprehensive feature vector. S3: Output trajectory prediction results based on the comprehensive feature vector using a trajectory prediction network. S4: Output a quantized intent probability vector based on the intent data using an intent analysis network. S5: Perform a secondary fusion of the trajectory prediction result and the intent probability vector to obtain a corrected final prediction result. S6: Output the final prediction result and the corresponding intent analysis report. This application's motion prediction method based on multimodal data fusion and dynamic intent, by introducing dynamic intent modeling, enables a deeper understanding of the motivations behind motion behavior, thereby making more accurate and reliable predictions in complex and uncertain scenarios. The steps described above are detailed below.

[0036] For step S1: Collect and time-synchronize multi-source heterogeneous data of the moving subject. The multi-source heterogeneous data includes at least visual data, sensor data and intent data.

[0037] The accuracy of the predictions primarily depends on high-quality and multi-dimensional input data. This application's data acquisition includes not only traditional visual information (attributed to posture, movement, and expression via multi-angle cameras), wearable sensor information (such as accelerometers, gyroscopes, and heart rate sensors), and environmental sensor information (such as force sensors on smart climbing walls), but more importantly, it innovatively introduces intent data. This intent data is acquired through biometric technologies (such as eye tracking or facial expression capture), providing quantifiable input for subsequent intent modeling.

[0038] For step S2: extract features from the raw data of different modalities to obtain visual feature vectors, sensor feature vectors and intent feature vectors, and fuse the visual feature vectors, sensor feature vectors and intent feature vectors to obtain a comprehensive feature vector.

[0039] Preprocessing and feature extraction are performed on data from different modalities and sampling rates. For a video sequence V = {V1, ..., V...} L}, using a convolutional neural network (CNN) encoder E vid V for each frame t Convert to visual feature vector v t =E vid (V t For sensor time series data S={S1,…,S…} L Using Transformer encoder E sen Convert it into a time-series sensor feature vector s t =E sen (S t For the intent data I = {I1, ..., I...} L}, using a dedicated intent encoder E int Convert it into an intention feature vector i t =E int (I t ).

[0040] In the embodiments of this application, a multi-head cross-attention mechanism is used to fuse visual feature vectors, sensor feature vectors, and intent feature vectors to obtain a comprehensive feature vector. This mechanism enables deep interaction and fusion of cross-modal information. Its core idea is to use features from one modality as the query (Q) and features from another modality as the key (K) and value (V), thereby highlighting the features most critical to the prediction result. The fused comprehensive feature vector F t The calculation formula is as follows:

[0041]

[0042] For example, the intent feature i t As a query Q i Visual features v t and sensor features t The concatenated vector serves as the key K. vs Sum V vs Then the attention in cross-modal fusion can be expressed as: F t =Multi-HeadAttention(Q i ,K vs V vs )

[0043] For step S3: The trajectory prediction network outputs the trajectory prediction result based on the comprehensive feature vector.

[0044] For step S4: The intent analysis network outputs a quantized intent probability vector based on the intent data.

[0045] This application employs a hierarchical prediction model:

[0046] Basic trajectory prediction: Use a trajectory prediction network, such as a sequence prediction network P. traj (Based on the Transformer model or its variants, or a conditional generative adversarial network) the fused comprehensive feature vector F t Processing is performed to predict the future T of the moving subject. pred Frame motion trajectory

[0047] Dynamic intent modeling: Introducing a separate "intent analysis network" P int This subnetwork utilizes the Inverse Reinforcement Learning (IRL) framework. The intention analysis subnetwork infers the latent reward function R(s,a) from historical expert demonstration data based on the IRL framework and outputs a quantified intention probability vector. This reward function can be regarded as a mathematical representation of the motion subject's intention.

[0048] IRL assumes that the expert (the agent) follows an optimal policy π, but the reward function is unknown. * The process aims to find a weight vector w to approximate the reward function R(x,u) = w using methods such as Maximum Entropy (MaxEnt). T φ(x,u) is used to make the expected features of the corresponding strategy match the expected features of the expert demonstration as closely as possible.

[0049] Optimization objective: This process solves the following optimization problem to infer the reward function that best explains the observed trajectory:

[0050]

[0051] Where w is the weight of the reward function, and φ(ξ) is the feature vector of the expert (human) demonstration trajectory ξ.

[0052] IRL (Intent Level Query) yields a quantified intent model, whose output is a probability distribution of different intents, i.e., an intent probability vector.

[0053] For step S5: The trajectory prediction result and the intention probability vector are fused twice to obtain the corrected final prediction result.

[0054] In the embodiments of this application, the specific method for obtaining the corrected final prediction result by performing a secondary fusion of the trajectory prediction result and the intention probability vector is as follows: using the intention probability vector as the weight, the confidence of multiple possible trajectories in the trajectory prediction result is probability-weighted and corrected to obtain the final prediction result.

[0055] Specifically, the basic trajectory prediction results Intent probability output by the intent analysis network A second fusion is performed to obtain the final corrected prediction result. A specific fusion method can be probabilistic weighted fusion. For example, if the basic trajectory prediction model outputs N trajectories T1, T2, ..., T... N and their corresponding confidence levels C1, C2, ..., C N The intent analysis subnetwork outputs M types of intents I1, I2, ..., I... M and its corresponding probability P(I) m The final corrected trajectory confidence C′ n It can be represented as C′ n =C n ·P(I m |Trajectory T n This integration incorporates the intrinsic motivations of the participants into the prediction, making the final prediction results more reasonable, credible, and accurate.

[0056] For step S6: Output the final prediction result and the corresponding intent analysis report.

[0057] The output of this application goes beyond just the final prediction result; it also includes an interpretable report that provides valuable decision-making support for users. The report comprises multimodal future trajectory prediction, confidence analysis, and intent analysis. For example, in pedestrian trajectory prediction applications within autonomous driving systems, the system collects pedestrian trajectory data using sensors such as cameras and radar, while simultaneously employing biometric technology to capture signals indicating whether a pedestrian intends to cross the road. After this data is fused and modeled, the system can not only predict the pedestrian's future trajectory but also output an interpretable report such as, "It is predicted that the pedestrian will cross the road in 2 seconds, with a strong intent (90% probability); it is recommended to immediately slow down and stop to let them pass," thereby enabling safer and more proactive avoidance maneuvers.

[0058] The core of this application's method is to treat "intent" as a quantifiable modal information and use it as a crucial input to drive the prediction model. In existing technologies, intent is often considered a "black box" that is difficult to quantify and predict. However, this application utilizes advanced methods such as inverse reinforcement learning to model intent and deeply integrate it with motion trajectory prediction. This approach, which combines physical behavior with intrinsic motivation, breaks through the limitations of traditional "black box" prediction. By providing intent analysis reports, the prediction results become understandable and verifiable, thereby increasing the system's transparency and user trust. This allows the technology to transform from a simple "result provider" into an "intelligent decision supporter."

[0059] This application also discloses a motion prediction system based on multimodal data fusion and dynamic intent, used to implement the aforementioned method. The motion prediction system based on multimodal data fusion and dynamic intent includes: a data acquisition and synchronization module, a feature extraction and fusion module, a hierarchical prediction module, a correction fusion module, and an output module.

[0060] The data acquisition and synchronization module is used to acquire and time-synchronize multi-source heterogeneous data of the moving subject. This multi-source heterogeneous data includes at least visual data, sensor data, and intent data. The feature extraction and fusion module extracts features from the raw data of different modalities, obtaining visual feature vectors, sensor feature vectors, and intent feature vectors, and then fuses these vectors to obtain a comprehensive feature vector. The hierarchical prediction module includes independent trajectory prediction and intent analysis networks. The trajectory prediction network outputs multiple possible future trajectories and their confidence levels based on the comprehensive feature vector, while the intent analysis network outputs a quantified intent probability vector based on intent data. The correction and fusion module performs a secondary fusion of the trajectory prediction result and the intent probability vector to obtain a corrected final prediction result. The output module outputs the final prediction result and the corresponding intent analysis report.

[0061] In the embodiments of this application, the feature extraction and fusion module fuses the visual feature vector, sensor feature vector and intent feature vector through a multi-head cross-attention mechanism to obtain a comprehensive feature vector.

[0062] In the embodiments of this application, the intent analysis network infers the potential reward function from historical expert demonstration data based on the inverse reinforcement learning framework and outputs a quantified intent probability vector.

[0063] In the embodiments of this application, the correction fusion module uses the intent probability vector as weight to perform probability-weighted correction on the confidence of multiple possible trajectories in the trajectory prediction result, so as to obtain the final prediction result.

[0064] For specific details of the motion prediction system based on multimodal data fusion and dynamic intent in this application, please refer to the aforementioned motion prediction method based on multimodal data fusion and dynamic intent, which will not be repeated here.

[0065] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any way, and all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A motion prediction method based on multimodal data fusion and dynamic intent, characterized in that, include: The multi-source heterogeneous data of the moving subject is collected and time-synchronized, and the multi-source heterogeneous data includes at least visual data, sensor data and intent data. Feature extraction is performed on the raw data of different modalities to obtain visual feature vectors, sensor feature vectors and intent feature vectors, and the visual feature vectors, sensor feature vectors and intent feature vectors are fused to obtain a comprehensive feature vector; The trajectory prediction network outputs trajectory prediction results based on the comprehensive feature vector. The intent analysis network outputs a quantized intent probability vector based on intent data. The trajectory prediction result and the intention probability vector are fused twice to obtain the corrected final prediction result; Output the final prediction result and the corresponding intent analysis report.

2. The motion prediction method based on multimodal data fusion and dynamic intent according to claim 1, characterized in that, A comprehensive feature vector is obtained by fusing the visual feature vector, the sensor feature vector, and the intent feature vector through a multi-head cross-attention mechanism.

3. The motion prediction method based on multimodal data fusion and dynamic intent according to claim 1, characterized in that, The intent analysis subnetwork infers the potential reward function from historical expert demonstration data and outputs a quantized intent probability vector based on an inverse reinforcement learning framework.

4. The motion prediction method based on multimodal data fusion and dynamic intent according to claim 3, characterized in that, The aforementioned inverse reinforcement learning framework employs maximum entropy inverse reinforcement learning, with the optimization objective being: Where w is the weight of the reward function, and φ(ξ) is the feature vector of the expert demonstration trajectory ξ.

5. The motion prediction method based on multimodal data fusion and dynamic intent according to claim 1, characterized in that, The specific method for performing a secondary fusion of the trajectory prediction result and the intent probability vector to obtain the corrected final prediction result is as follows: Using the intent probability vector as weights, the confidence levels of multiple possible trajectories in the trajectory prediction result are proportionally weighted and corrected to obtain the final prediction result.

6. The motion prediction method based on multimodal data fusion and dynamic intent according to claim 1, characterized in that, The trajectory prediction network is based on the Transformer model or a conditional generative adversarial network.

7. A motion prediction system based on multimodal data fusion and dynamic intent, characterized in that, include: The data acquisition and synchronization module is used to acquire and time-synchronize multi-source heterogeneous data of the moving subject. The multi-source heterogeneous data includes at least visual data, sensor data and intent data. The feature extraction and fusion module is used to extract features from the raw data of different modalities to obtain visual feature vectors, sensor feature vectors and intent feature vectors, and to fuse the visual feature vectors, sensor feature vectors and intent feature vectors to obtain a comprehensive feature vector. The hierarchical prediction module includes an independent trajectory prediction network and an intent analysis network. The trajectory prediction network is used to output multiple possible future trajectories and their confidence levels based on the comprehensive feature vector. The intent analysis network is used to output a quantized intent probability vector based on intent data. The correction fusion module is used to perform a secondary fusion of the trajectory prediction result and the intention probability vector to obtain a corrected final prediction result; The output module is used to output the final prediction result and the corresponding intent analysis report.

8. The motion prediction system based on multimodal data fusion and dynamic intent according to claim 7, characterized in that, The feature extraction and fusion module fuses the visual feature vector, the sensor feature vector, and the intent feature vector through a multi-head cross-attention mechanism to obtain a comprehensive feature vector.

9. The motion prediction system based on multimodal data fusion and dynamic intent according to claim 7, characterized in that, The intent analysis network infers the potential reward function from historical expert demonstration data and outputs a quantified intent probability vector based on an inverse reinforcement learning framework.

10. The motion prediction system based on multimodal data fusion and dynamic intent according to claim 7, characterized in that, The correction fusion module uses the intent probability vector as weights to perform probability-weighted correction on the confidence levels of multiple possible trajectories in the trajectory prediction result, thereby obtaining the final prediction result.