A trajectory prediction method based on high-order kinematic prior and global-local feature fusion

CN122813892APending Publication Date: 2026-09-25HUNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610979297.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0009]本发明的目的在于克服现有技术的上述不足,提供一种基于高阶运动学先验与全局-局部特征融合的轨迹预测方法,通过显式建模高阶运动学先验、全局 -局部联合结构化校正以及渐进式稳定训练策略,解决现有技术中历史表征物理一致性不足、轨迹细化结构意识缺失、多阶段训练优化不稳定的技术问题,提升复杂交通场景下轨迹预测的精度、物理合理性与训练鲁棒性

Benefits of technology

[0020]1、本发明通过动力学增强的历史运动表征显式整合了加速度和航向角变化等高阶运动学先验,克服了传统方法仅依赖几何坐标而导致的运动状态建模不充分缺陷,增强了模型对智能体潜在运动倾向的感知能力,使生成的预测轨迹更符合实际物理规律。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813892A_ABST
    Figure CN122813892A_ABST
Patent Text Reader

Abstract

The application discloses a trajectory prediction method based on high-order kinematics prior and global-local feature fusion, which explicitly integrates high-order kinematics prior such as acceleration and change of heading angle through history motion representation enhanced by dynamics, overcomes the defect of insufficient modeling of motion state caused by the fact that traditional methods only rely on geometric coordinates, enhances the perception ability of the model to the potential motion tendency of the intelligent agent, and makes the generated prediction trajectory more in line with the actual physical law; through joint modeling of the local motion mode and the global structured feature of the anchor point trajectory by the future motion perception refinement module, the prediction fidelity and trajectory smoothness in the fine stage are greatly improved; and through the introduced gradual zero initialization future coding strategy, the common optimization instability problem in the trajectory prediction task is effectively solved, noise interference is relieved, and the model is guided to realize faster and more accurate deterministic convergence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving technology, and in particular to a trajectory prediction method based on high-order kinematic priors and global-local feature fusion. Background Technology

[0002] In autonomous driving and intelligent transportation systems, accurate and stable trajectory prediction is fundamental for vehicles to make forward-looking decisions and plan their subsequent motion, especially in complex, high-frequency urban road scenarios. The goal of trajectory prediction algorithms is to generate safe and reasonable driving trajectories for a future period based on historical observations of the vehicle and its surrounding environment. With the development of deep learning technology, trajectory prediction has evolved from traditional physical models such as Kalman filtering to complex models based on graph neural networks and Transformer architectures, effectively capturing complex interactions between agents and maps, and between agents themselves.

[0003] Although existing trajectory prediction methods (such as HPNet) have made some progress in handling spatiotemporal interactions, they still exhibit significant limitations in real-world, complex traffic scenarios.

[0004] First, at the feature representation level, the historical motion representation of mainstream algorithms is often oversimplified, relying too much on geometric coordinate sequences and lacking explicit modeling of higher-order kinematic properties with clear physical meanings, such as acceleration and heading angle changes. This results in insufficient coherence and physical consistency in the representation when the model performs highly dynamic maneuvers.

[0005] Secondly, in the trajectory refinement stage, existing solutions usually only regard the coarse predicted trajectory as a set of discrete coordinate points and are limited to local coordinate offset regression. They fail to effectively capture the local motion patterns (such as midpoint motion intensity and turning trend) and global structured attributes (such as straightness and global displacement) contained in the trajectory. This directly limits its prediction accuracy and physical rationality in highly interactive scenarios.

[0006] Furthermore, since the prediction results in the early stages of training generally contain great uncertainty and noise, directly using these unreliable prediction trajectories as prior inputs for subsequent refinement stages will seriously interfere with the overall optimization process of the model, leading to slow training convergence and reduced model robustness.

[0007] In summary, there is an urgent need for a novel trajectory prediction method that can explicitly utilize kinematic priors to enhance historical perception, achieve refined correction of trajectory structure through global-local joint modeling, and maintain high optimization stability during training, so as to meet the real-time and high reliability requirements of autonomous driving systems in complex dynamic environments. Summary of the Invention

[0008] (a) Technical problems to be solved

[0009] The purpose of this invention is to overcome the above-mentioned shortcomings of the prior art and provide a trajectory prediction method based on higher-order kinematic priors and global-local feature fusion. By explicitly modeling higher-order kinematic priors, global-local joint structured correction, and progressive stable training strategies, this invention solves the technical problems of insufficient physical consistency of historical representation, lack of awareness of trajectory refinement structure, and instability of multi-stage training optimization in the prior art, thereby improving the accuracy, physical rationality, and training robustness of trajectory prediction in complex traffic scenarios.

[0010] (II) Technical Solution

[0011] This invention discloses a trajectory prediction method based on higher-order kinematic priors and global-local feature fusion, characterized by the following steps:

[0012] S1: Dynamic Enhancement of Historical Motion Representation: Obtain the historical trajectory coordinate sequence of the agent, calculate and extract high-order motion features including displacement, acceleration, and heading angle changes, and construct a three-dimensional motion feature vector. The dynamically enhanced historical motion embedding is obtained through feature mapping. ;

[0013] S2: Spatiotemporal context encoding and feature initialization: Encoding the vectorized road map to obtain the initial map embedding. Combined with the historical movement embedding Initial prediction labels carrying spatiotemporal context information are generated through time-pattern graph attention. ;

[0014] S3: Triple Factorization Spatiotemporal Interaction Mechanism: The social dependencies between agents are captured through the agent interaction attention module, and historical prediction attention module embeds historical motion. Associating with the predicted state, the pattern interaction attention module models the competition and correlation of multiple prediction patterns, and generates multiple coarse prediction trajectory anchor points through the interaction embedding output by the triple factorization attention module.

[0015] S4: Future Motion Perception Structured Feature Extraction: For each coarse predicted trajectory anchor point, local motion pattern features and global trajectory structure features are extracted and encoded to generate a future motion perception descriptor. ;

[0016] S5: Predictive Label Reconstruction and Feature Fusion: Extracting Geometric Embedding of Coarse Predicted Trajectory Anchor Points With the future motion perception descriptor Weighted fusion is performed to obtain the reconstructed prediction labels. ; where the future motion perception descriptor The corresponding projection matrix is ​​initialized to a zero matrix, and during training, a preset period threshold is applied. Maintain freeze suppression until a preset periodic threshold is reached. Post-activation weight update;

[0017] S6: Secondary Interactive Modeling and Trajectory Residual Decoding: Reconstructing Predictive Markers Spatiotemporal alignment is performed again, and the final multimodal predicted trajectory and corresponding confidence score are output after residual decoding.

[0018] Preferably, further detailed descriptions of steps 1-6 of the present invention can be found in the following detailed embodiments.

[0019] (III) Beneficial Effects

[0020] 1. This invention explicitly integrates higher-order kinematic priors such as acceleration and heading angle changes through dynamically enhanced historical motion representation, overcoming the shortcomings of traditional methods that rely solely on geometric coordinates for insufficient motion state modeling. It enhances the model's ability to perceive the agent's potential motion tendencies, making the generated predicted trajectory more consistent with actual physical laws.

[0021] 2. The future motion perception refinement module designed in this invention achieves a qualitative leap from simple coordinate point offset regression to multi-scale structured correction by jointly modeling the local motion patterns and global structured features of anchor point trajectories, which greatly improves the prediction fidelity and trajectory smoothness in the refinement stage.

[0022] 3. The progressive zero-initialization future encoding strategy introduced in this invention effectively solves the common optimization instability problem in trajectory prediction tasks. By suppressing unreliable prediction features in the early stage of training and dynamically activating refined branches as the process progresses, it successfully alleviates noise interference and guides the model to achieve faster and more accurate deterministic convergence.

[0023] 4. Excellent prediction performance and significant engineering value. In the Argoverse dataset test, this invention reduced minADE and minFDE to 0.7583 and 1.0901 respectively, and achieved a MissRate of 0.103. In the joint prediction task of the INTERACTION dataset, this invention reduced minJointADE from 0.1739 to 0.1674 and minJointFDE from 0.5577 to 0.5512, fully demonstrating the robustness of this invention in large-scale, high-frequency interactive urban traffic environments and its broad engineering application prospects.

[0024] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0025] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0026] Figure 1 This is a schematic diagram of the overall architecture of the HFMNet motion-aware dynamic trajectory prediction framework according to a preferred embodiment of the present invention. Detailed Implementation

[0027] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.

[0028] Furthermore, unless otherwise defined, the technical or scientific terms used in this application description shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "upper," "lower," "left," "right," "center," "vertical," "horizontal," "inner," and "outer," etc., used in this application description to indicate relative direction or positional relationship are used only to indicate relative orientation or positional relationship, and do not imply that the device or component must have a specific orientation, or be constructed and operated in a specific orientation. When the absolute position of the described object changes, its relative positional relationship may also change accordingly, and therefore should not be construed as a limitation on this application. The terms "first," "second," "third," and similar terms used in this application description are used only for descriptive purposes to distinguish different components, and should not be construed as indicating or implying relative importance. The terms "a," "one," or "the," etc., used in this application description should not be construed as an absolute limitation on quantity, but should be construed as indicating the existence of at least one. The terms "including," "comprising," etc., used in this application description mean that the element or object preceding the word covers the element or object listed after the word and its equivalents, without excluding other elements or objects.

[0029] It should also be noted that, unless otherwise explicitly specified and limited, the terms such as “installation,” “connection,” and “linkage” used in the description of this application should be interpreted broadly. For example, a connection can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can also refer to the internal connection of two components. Those skilled in the art can understand its specific meaning in this application according to the specific circumstances.

[0030] like Figure 1 As shown, the HFMNet (Hierarchical Fusion Motion Network) proposed in this invention is a motion-aware dynamic trajectory prediction framework. It aims to improve prediction accuracy and physical consistency in complex interactive scenarios by explicitly modeling the physical kinematic priors and structured features of the predicted trajectory. The overall architecture consists of a dynamics enhancement encoding module, a spatiotemporal context encoding module, and a structured future refinement module. Through a closed-loop prediction pipeline of "dynamics enhancement representation → multi-scale feature interaction → progressive structured refinement," it can not only explicitly extract high-order motion dynamics information contained in historical trajectories but also perform joint correction using the global structure and local motion patterns of the generated candidate predicted trajectory anchor points.

[0031] Spatio-temporal context interaction refers to the process by which a model jointly models using motion change information in the time dimension and environmental interaction information in the spatial dimension; Kinematic Prior refers to prior constraint information derived from the laws of real physical motion, such as velocity continuity, turning inertia, and acceleration variation.

[0032] The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion in this embodiment is characterized by the following steps:

[0033] S1: Kinetic Enhancement of Historical Motion Representation (KHE):

[0034] Obtain the historical trajectory coordinate sequence of the agent, calculate and extract high-order motion features including displacement, acceleration, and heading angle changes, and construct a three-dimensional motion feature vector. The dynamically enhanced historical motion embedding is obtained through feature mapping. .

[0035] This step is used to construct a historical trajectory representation that includes higher-order dynamic features, aiming to solve the problem that traditional trajectory prediction methods rely solely on the original coordinate sequence, resulting in insufficient kinematic property modeling.

[0036] Specifically, let's say an intelligent agent The observation history trajectory sequence is ,in Let represent the two-dimensional position coordinates at time t, and let i be the index of the agent. This refers to the historical observation duration.

[0037] In order to fully capture the motion trends of the agent and enhance the physical consistency of the prediction, this embodiment not only focuses on the geometric evolution of the position, but also explicitly calculates features with clear physical meaning, such as displacement vector, acceleration, and heading angle changes.

[0038] set up The time step between adjacent observations is based on historical trajectory sequences. The displacement is calculated using the following formula. Heading angle and change in heading angle :

[0039]

[0040] in, This represents the displacement between adjacent time steps. This represents the heading angle estimated from the direction of motion. The angle normalization function normalizes the angle difference to... Interval.

[0041] It should be noted that, in this embodiment, the heading angle refers to the direction angle of the current velocity direction of the agent relative to the global coordinate system during the movement, which is used to characterize the direction of movement of the vehicle or target; the heading angle change refers to the difference in heading angle between adjacent time steps, which is used to reflect the turning trend and the intensity of direction change during the target's movement.

[0042] Furthermore, this embodiment estimates the acceleration metric value by measuring the change in speed. The displacement, acceleration, and heading change characteristics mentioned above are then concatenated to construct a three-dimensional motion feature vector. :

[0043]

[0044] Finally, the motion feature vector is mapped to a high-dimensional latent embedding space using a multilayer perceptron (MLP) to obtain the dynamically enhanced historical motion embedding. :

[0045]

[0046] It should be noted that the three-dimensional motion feature vector is generated through a multilayer perceptron. Mapping to a high-dimensional embedding space, the Multilayer Perceptron (MLP) can be replaced by any of the following: a Convolutional Neural Network, a Recurrent Neural Network, or a Transformer encoder, all of which can achieve equivalent feature mapping functionality.

[0047] The resulting KHE representation not only contains basic spatial location information, but also enhances the model's ability to perceive complex maneuvers such as turning and acceleration in highly dynamic scenarios by explicitly modeling physical laws, thus providing a more physically grounded historical trajectory prior for subsequent spatiotemporal interaction mechanisms.

[0048] S2: Spatiotemporal context encoding and feature initialization:

[0049] Encoding the vectorized road map yields the initial map embedding. Combined with the historical movement embedding Initial prediction labels carrying spatiotemporal context information are generated through time-pattern graph attention. ;

[0050] This step is used to integrate structured map constraints and historical motion information to provide an initial state representation for subsequent multimodal prediction models.

[0051] Specifically, to capture the constraints of road topology on agent behavior, a vectorized map representation method is adopted, treating each segment of the lane centerline as a graph node. For each segment, a graph containing the segment length is constructed. and semantic attributes Feature vectors (such as intersection signs, traffic control signals, etc.) :

[0052] This feature is projected using a multilayer perceptron (MLP) to obtain the initial map embedding. Furthermore, it utilizes a self-attention mechanism to model the contextual dependencies between road segments.

[0053] At the same time, the system maintains a set of learnable mode embeddings. Used to represent Different potential predictable motion patterns.

[0054] In order to incorporate historical motion features into the prediction model, each agent... At a moment of historical observation Constructing initial prediction labels This initialization process is implemented through a temporal-to-mode graph attention layer, enabling pattern embeddings to adaptively aggregate specific time windows. Embedded historical motion perception within And explicitly consider the spatiotemporal edge features that include relative pose and time interval. :

[0055]

[0056] This step effectively integrates map topological constraints with dynamic historical representations, generating predictive labels. It carries rich spatiotemporal context information, serving as the initial input state for subsequent complex interaction modeling and multimodal trajectory decoding.

[0057] S3: Triple factorization spatiotemporal interaction mechanism:

[0058] The social dependencies between agents are captured by the agent interaction attention module, and historical prediction attention module embeds historical motion. Associating with the predicted state, the pattern interaction attention module models the competition and correlation of multiple prediction patterns, and generates multiple coarse prediction trajectory anchor points through the interaction embedding output by the triple factorization attention module.

[0059] This step is used to model the structured interaction relationships between historical motion, surrounding agents, and road topology after acquiring the dynamically enhanced historical motion representation (KHE) and map embedding, in order to capture complex spatiotemporal dependencies.

[0060] Specifically, the Triple Factorized Attention Mechanism is an attention structure that models agent interaction, historical prediction, and mode interaction separately. It is used to simultaneously capture the relationships between agents, between the past and the future, and between different prediction modes. This mechanism is implemented through three sub-modules: Agent Interaction Attention, Historical Prediction Attention, and Mode Interaction Attention.

[0061] First, the agent interaction attention module utilizes a graph-based attention mechanism to process the historical motion embeddings of all agents, capturing the social dependencies between agents and their interaction characteristics. The calculation formula is as follows:

[0062]

[0063] in To enhance the dynamics of agent j by embedding historical motion, The projection matrix is ​​the value. The weights are calculated using scaled dot product attention.

[0064] Secondly, the historical prediction attention module enables each prediction tag to actively query the historical motion embeddings of all agents within the observation period. The aim is to inject explicit motion features into the predicted state, and its attention weight calculation formula is as follows:

[0065]

[0066] in The query vector generated for predicting the tags, The key vectors generated for embedding historical motion. As the embedding dimension, the aggregated update prediction is labeled as .

[0067] Finally, the pattern interaction attention module allows different pattern labels of the same agent to attend to each other across different prediction hypotheses, in order to model competition and correlation between multimodal paths, and update the prediction labels. It is expressed as follows:

[0068] in For pattern For pattern Attention weights This is the corresponding value vector.

[0069] It should be noted that the functions implemented by the agent interaction attention module, the historical prediction attention module, and the pattern interaction attention module in this embodiment can be implemented by any one of the standard self-attention mechanism, cross-attention mechanism, or graph convolutional neural network (GCN).

[0070] Through triple factorization interaction, deep motion features with social consciousness and temporal coherence can be extracted from high-dimensional representations. Based on the interaction embedding output by the triple factorization attention module, multiple coarse predicted trajectory anchors are generated. ,in To predict duration, To predict the location coordinates.

[0071] S4: Future motion modeling and feature extraction based on the FMAR module:

[0072] For each coarsely predicted trajectory anchor point, local motion pattern features and global trajectory structure features are extracted and encoded to generate a future motion perception descriptor. .

[0073] This step is used to extract multidimensional motion patterns from the initially generated candidate trajectories, aiming to overcome the problem that traditional prediction methods only use discrete coordinates and ignore the inherent structured features of the trajectory.

[0074] First, the system generates a series of coarse predicted trajectories as motion anchors based on the interaction embedding output by the triple factorized attention module. .

[0075] Next, the Future Motion Refinement (FMAR) module captures high-order motion patterns for each anchor point by jointly modeling local motion patterns and the global trajectory structure. At the local level, this involves predicting the midpoint moment. Extracting the amplitude of motion for the core With change of heading At the global level, extract the displacement vector of the entire trajectory. And the straightness coefficient, which reflects the smoothness of the path. The relevant core kinematic characteristics are calculated using the following formulas:

[0076]

[0077] in, This represents the total length of the trajectory.

[0078] It should be noted that the straightness coefficient is an indicator used to measure the overall smoothness and curvature change of the trajectory, and is usually calculated from the relationship between the actual path length of the trajectory and the straight-line distance between the start and end points.

[0079] Finally, this step uses a multilayer perceptron (MLP) to encode these local motion pattern features and global trajectory structure features into a future motion perception descriptor. :

[0080]

[0081] generated descriptor The dynamic characteristics and spatiotemporal consistency of the candidate trajectory are encapsulated, serving as a key guiding signal for subsequent refined trajectory characterization and reconstruction.

[0082] S5: Predictive Label Reconstruction and Feature Fusion

[0083] Extracting the geometric embedding of coarse predicted trajectory anchor points With the future motion perception descriptor Weighted fusion is performed to obtain the reconstructed prediction labels. ; where the future motion perception descriptor The corresponding projection matrix is ​​initialized to a zero matrix, and during training, a preset period threshold is applied. Maintain freeze suppression until a preset periodic threshold is reached. Then activate weight update.

[0084] This step is used to reconstruct the query embedding during the trajectory refinement stage, enhancing the representational power of the predicted labels by fusing geometric path information with structured motion priors.

[0085] Specifically, the system extracts two complementary feature representations from the coarsely predicted trajectory anchor points: a geometric structure embedding based on spatial coordinate encoding and another based on spatial coordinate encoding. and multi-scale future motion perception descriptors extracted through the FMAR module To minimize uncertainty interference during the refinement process, the fusion mechanism of this invention strictly excludes the predicted labels in the initial stage, retaining only the path coordinates and descriptor terms with kinematic priors for joint modeling. The calculation formula for feature fusion is as follows:

[0086] in, and The projection matrix is ​​learnable. The recoded prediction label.

[0087] To ensure optimization stability during the initial training phase, this invention further implements a Progressive Zero-Initialized Future Encoding (PZFE) strategy. This strategy initializes the projection matrix involved in the future motion descriptor branch as a zero matrix:

[0088] During training, this future-aware branch reaches a preset periodic threshold. The system is currently in a suppressed state. It should be noted that the preset periodic threshold... The weights are dynamically adjusted based on the dataset size and model convergence speed. Activation methods include catastrophe activation and linear growth activation. Zero matrix initialization ensures the initial contribution of this term is zero, thus avoiding early inaccurate anchor point trajectories that introduce prediction noise and allowing the model backbone to prioritize learning stable historical interaction features. Once the model training has converged to a stable stage, the system automatically activates the weight update for this branch, enabling the model to adaptively integrate the local motion patterns and global structural priors contained in future motion descriptors, thereby achieving dynamic updates and high-quality reconstruction of predicted labels.

[0089] S6: Secondary Interactive Modeling and Trajectory Residual Decoding

[0090] Reconstruct the prediction label Spatiotemporal alignment is performed again, and the final multimodal predicted trajectory and corresponding confidence score are output after residual decoding.

[0091] This step is used to perform deep spatiotemporal alignment on the reconstructed prediction markers and decode the final trajectory, aiming to further improve the physical consistency and environmental adaptability of the prediction results through secondary refined modeling.

[0092] Specifically, the updated prediction label generated in step S5 Instead of being directly input into the decoding network, it is re-input into another round of triple factorized attention interaction mechanism to achieve the final fine-tuning of the trajectory.

[0093] In this process, the predicted labels can be deeply aligned with historical motion representations (KHE) and the dynamic states of surrounding agents based on the integrated motion-aware priors and global structural features. This fully leverages the coupling information between historical motion inertia and future prediction trends to generate refined predicted labels with high contextual consistency. .

[0094] Finally, the trajectory decoding network is used. Refined prediction labeling Perform residual analysis to obtain the final predicted trajectory. From the initial trajectory anchor point It is formed by adding the predicted future motion residual displacement, and its calculation formula is as follows:

[0095] At the same time, the system utilizes a probabilistic decoding network Simultaneously estimate the confidence scores for each motion pattern. This is used to evaluate the rationality of multimodal trajectory prediction.

[0096]

[0097] By introducing a secondary interaction decoding strategy, the present invention ensures that the structured features captured by the FMAR module are fully aligned with the complex social game environment, thereby effectively correcting the kinematic bias in the initial prediction and finally outputting a multimodal trajectory set with high accuracy and in compliance with vehicle dynamics constraints.

[0098] In terms of experimental data, this invention demonstrates superior performance. In the Argoverse dataset leaderboard test, this invention reduces minADE and minFDE to 0.7583 and 1.0901 respectively, achieving a MissRate (MR) of 0.103, ranking 9th on the test set of this dataset. In the joint prediction task of the INTERACTION dataset, this invention also performs excellently, reducing minJointADE from 0.1739 to 0.1674 and minJointFDE from 0.5577 to 0.5512, fully demonstrating the excellent robustness and broad engineering application prospects of this invention in dealing with large-scale, high-frequency interactive urban traffic environments.

[0099] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A trajectory prediction method based on high-order kinematic priors and global-local feature fusion, characterized in that, Specifically, the following steps are included: S1: Dynamic Enhancement of Historical Motion Representation: Obtain the historical trajectory coordinate sequence of the agent, calculate and extract high-order motion features including displacement, acceleration, and heading angle changes, and construct a three-dimensional motion feature vector. The dynamically enhanced historical motion embedding is obtained through feature mapping. ; S2: Spatiotemporal context encoding and feature initialization: Encoding the vectorized road map to obtain the initial map embedding. Combined with the historical movement embedding Initial prediction labels carrying spatiotemporal context information are generated through time-pattern graph attention. ; S3: Triple Factorization Spatiotemporal Interaction Mechanism: The social dependencies between agents are captured through the agent interaction attention module, and historical prediction attention module embeds historical motion. Associating with the predicted state, the pattern interaction attention module models the competition and correlation of multiple prediction patterns, and generates multiple coarse prediction trajectory anchor points through the interaction embedding output by the triple factorization attention module. S4: Future Motion Perception Structured Feature Extraction: For each coarse predicted trajectory anchor point, local motion pattern features and global trajectory structure features are extracted and encoded to generate a future motion perception descriptor. ; S5: Predictive Label Reconstruction and Feature Fusion: Extracting Geometric Embedding of Coarse Predicted Trajectory Anchor Points With the future motion perception descriptor Weighted fusion is performed to obtain the reconstructed prediction labels. ; where the future motion perception descriptor The corresponding projection matrix is ​​initialized to a zero matrix, and during training, a preset period threshold is applied. Maintain freeze suppression until a preset periodic threshold is reached. Post-activation weight update; S6: Secondary Interactive Modeling and Trajectory Residual Decoding: Reconstructing Predictive Markers Spatiotemporal alignment is performed again, and the final multimodal predicted trajectory and corresponding confidence score are output after residual decoding.

2. The trajectory prediction method based on high-order kinematic priors and global-local feature fusion as described in claim 1, characterized in that, The aforementioned step S1 specifically includes: Acquiring intelligent agents Observational historical trajectory sequence ,in Let represent the two-dimensional position coordinates at time t, and let i be the index of the agent. This refers to the historical observation duration; Let Δt be the time step between adjacent observations, based on historical trajectory sequences. The displacement is calculated using the following formula. Heading angle and change in heading angle : in, This represents the displacement between adjacent time steps. This represents the heading angle estimated from the direction of motion. The angle normalization function normalizes the angle difference to... interval; Acceleration measurement based on velocity change The displacement, acceleration, and heading angle changes are concatenated to construct a three-dimensional motion feature vector. : Three-dimensional motion feature vectors are processed using a multilayer perceptron. Mapping to a high-dimensional embedding space yields a dynamically enhanced historical motion embedding. .

3. The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion as described in claim 1, characterized in that, Step S2 specifically includes: A vectorized map representation is used, with the center line segment of each lane as a graph node, constructing a map containing the road segment length. and semantic attributes eigenvectors : Where m represents the map dimension and j represents the index of the road segment; Initial map embedding obtained by multilayer perceptron projection Furthermore, it models the contextual dependencies between road segments through self-attention; Set up K groups of learnable pattern embeddings , corresponding to K potential movement patterns; For each intelligent agent By using a time-pattern graph attention layer, pattern embedding can aggregate time windows of historical observation time h. Embedded historical movements Combined with spatiotemporal edge features including relative pose and time interval Generate initial prediction labels : 。 4. The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion as described in claim 3, characterized in that, The specific process of the triple factorization attention mechanism in step S3 is as follows: The agent interaction attention module utilizes a graph-based attention mechanism to process the historical motion embeddings of all agents, capturing the social dependencies between agents and obtaining interaction features. : in, To enhance the dynamics of agent j by embedding historical motion, The projection matrix is ​​the value. These are the weights calculated using scaled dot product attention; The historical prediction attention module enables each prediction tag to query the historical motion embeddings of all agents within the observation period. Motion features are injected into the predicted state, and its attention weights satisfy: in The query vector generated for predicting the tags, The key vectors generated for embedding historical motion. As the embedding dimension, the aggregated update prediction is labeled as ; The pattern interaction attention module allows different pattern labels of the same agent to attend to each other across different prediction hypotheses, in order to model the competition and correlation of multiple prediction patterns and update the prediction labels. for: in, For pattern For pattern Attention weights For the corresponding value vector; The interactive embedding output by the triple factorization attention module generates multiple coarse prediction trajectory anchor points. ,in To predict duration, To predict the location coordinates.

5. The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion as described in claim 4, characterized in that, Step S4 specifically includes: The future motion perception refinement module, at the local level, predicts the midpoint moment for each anchor point. Extracting the amplitude of motion for the core With change of heading At the global level, extract the displacement vector of the entire trajectory. And the straightness coefficient, which reflects the smoothness of the path. The formulas for calculating local motion pattern features and global trajectory structure features are as follows: in, This represents the total length of the trajectory. The aforementioned local motion pattern features and global trajectory structure features are encoded using a multilayer perceptron to obtain a future motion perception descriptor. : 。 6. The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion as described in claim 1, characterized in that, Step S5 specifically includes: Extracting geometric embeddings from coarse-predicted trajectory anchors With future motion perception descriptor The reconstructed prediction labels are generated by fusing them according to the following formula. : in For the geometric feature projection matrix, and Projection matrix for future motion features; In the initial training phase, the future motion feature projection matrix is ​​used. Initialize as a zero matrix: During training, when the number of training rounds is less than or equal to the preset period threshold... At that time, freeze The weights are updated to suppress future motion perception branches; when the number of training rounds exceeds a preset period threshold... After that, activate The weight updates enable the model to adaptively integrate future motion descriptors. The implied local motion patterns and global structural priors.

7. The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion as described in claim 1, characterized in that, Step S6 specifically includes: Reconstruct the prediction label The data is then re-input into the triple factorization attention module for secondary spatiotemporal interaction alignment, generating refined prediction labels with high contextual consistency. ; Fine-grained prediction labels are generated through a trajectory decoding network. Perform residual analysis to ultimately predict the trajectory. For coarse prediction trajectory anchor points Sum of residual displacement: Simultaneously, confidence scores for each motion pattern are estimated using a probabilistic decoding network. : 。 8. The trajectory prediction method based on higher-order kinematic priors and global-local feature fusion as described in claim 2, characterized in that, The three-dimensional motion feature vector is generated through a multilayer perceptron. Mapped into a high-dimensional embedding space, the multilayer perceptron can be replaced by a convolutional neural network, a recurrent neural network, or other equivalent linear or nonlinear mapping networks.

9. The trajectory prediction method based on high-order kinematic priors and global-local feature fusion as described in claim 1, characterized in that, The preset period threshold The weights are dynamically adjusted based on the size of the dataset and the model's convergence speed; the activation methods for the weights include catastrophe activation and linear growth activation.

10. The trajectory prediction method based on high-order kinematic priors and global-local feature fusion according to claim 1, characterized in that, The functions of the agent interaction attention module, the historical prediction attention module, and the pattern interaction attention module in step S3 can be implemented using any one of the standard self-attention mechanism, cross-attention mechanism, or graph convolutional neural network.