Multi-modal vehicle trajectory prediction method considering driving intention and interaction

By combining road structure and vehicle interaction models and using a multimodal decoder to predict vehicle trajectories, the problem of existing methods failing to effectively model road geometric constraints and interaction relationships is solved, achieving accurate and diverse predictions in complex traffic scenarios.

CN121246833AActive Publication Date: 2026-01-02BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511686637.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-01-02
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

Existing vehicle trajectory prediction methods fail to effectively model road geometric constraints and dynamic interactions between vehicles, resulting in insufficient prediction accuracy in complex traffic scenarios and an inability to accurately describe the uncertainties under diverse driving behaviors.

Method used

By introducing a trajectory encoder, intent recognition module, interaction force module, intent interaction fusion module, and spatial relationship module, and combining prior knowledge of road structure and physical interaction force model, the future trajectory of the target vehicle is predicted, and an expert hybrid multimodal decoder is used for decoding.

Benefits of technology

It improves the accuracy and adaptability of vehicle trajectory prediction, enables multimodal prediction in complex traffic environments, and enhances the reliability and practical value of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121246833A_ABST
    Figure CN121246833A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal vehicle trajectory prediction method considering a driving intention and an interaction effect. The method comprises the steps of obtaining a historical driving trajectory of a target vehicle and a historical driving trajectory of a surrounding environment vehicle; inputting the two historical driving tracks into a vehicle track prediction model, and predicting, generating and obtaining a future driving track of the target vehicle; the vehicle trajectory prediction model comprises a trajectory encoder, an intention recognition module, an interaction force module, an intention interaction fusion module, a spatial relationship module and a multi-mode decoder. According to the invention, vehicle intention and vehicle interaction trajectory prediction can be effectively identified in combination with road structure priori knowledge; according to the method, the intention and interaction of the vehicle are depicted based on the interaction force and potential field modeling of the class physics in the prediction process, so that the adaptability and prediction precision of a complex traffic scene are improved; in addition, by introducing an expert hybrid multi-mode decoding mechanism, the reliability and the practical value of trajectory prediction can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle trajectory prediction technology, and in particular to a multimodal vehicle trajectory prediction method that takes into account driving intentions and interactions. Background Technology

[0002] Vehicle trajectory prediction is a key technology in autonomous driving and intelligent transportation systems, and its accuracy directly affects driving safety and decision-making efficiency. In recent years, with the rapid development of deep learning technology, data-driven trajectory prediction methods have become the mainstream research approach.

[0003] However, current vehicle trajectory prediction methods do not model road geometric constraints to predict the target vehicle's lane-changing intentions, nor do they model the dynamic interaction relationships between surrounding vehicles and the target vehicle to obtain the target vehicle's overall interaction characteristics. Furthermore, they do not consider fusing the target vehicle's lane-changing intentions with its overall interaction characteristics, nor do they model the global spatial dependencies in complex traffic scenarios. They also cannot achieve decoupled prediction of future trajectories corresponding to different driving operations. Therefore, existing vehicle trajectory prediction methods cannot achieve accurate prediction of vehicle trajectories.

[0004] For example, some potential field-based trajectory prediction methods typically calculate interaction intensity only based on the relative positions and geometric relationships between vehicles, without modeling road geometry, lane boundaries, or lane attributes. Therefore, they cannot infer the vehicle's lane-changing intentions from a road semantic level. In multi-lane or complex driving environments, these methods often fail to adequately respond to changes in vehicle intentions, significantly limiting prediction accuracy.

[0005] Meanwhile, while some modeling methods based on spatiotemporal attention or neighborhood masks can extract the correlation between vehicles from historical trajectories, the interaction relationships are usually learned implicitly by neural networks, without incorporating dynamic factors such as relative speed and relative acceleration of vehicles. This results in the lack of interpretability of the obtained interaction features, making it difficult to accurately describe the dynamic interaction changes under different traffic conditions.

[0006] Existing graph-based modeling methods often focus on local neighbor relationships, resulting in vehicle relationship graphs that only contain local spatial information and lack the ability to represent the global dependencies between multiple vehicles in the overall traffic scene. Furthermore, these methods often fail to adequately characterize the patterns of vehicle interaction over time, making it difficult to form a stable and effective scene understanding in complex traffic environments.

[0007] Furthermore, some methods rely on a fixed number of intent branches or a single predictor head to output future trajectories, failing to form independent trajectory patterns for different driving operations and thus struggling to reflect the uncertainties brought about by diverse driving behaviors. In scenarios with multiple possible actions such as lane changes, acceleration / deceleration, and following other vehicles, these methods often fail to provide reasonable multimodal prediction results. Summary of the Invention

[0008] This invention aims to address, at least to some extent, the technical problems in related technologies. Therefore, the objective of this invention is to provide a multimodal vehicle trajectory prediction method that considers driving intentions and interactions. This method can effectively identify vehicle intentions and interactions through prior knowledge of road structure. During the prediction process, this method uses physics-like interaction forces and potential field modeling to characterize vehicle intentions and interactions, thereby improving adaptability to complex traffic scenarios and prediction accuracy. Furthermore, by introducing an expert hybrid multimodal decoding mechanism, this method can significantly improve the reliability and practical value of trajectory prediction.

[0009] To achieve the above objectives, the present invention is implemented through the following technical solution:

[0010] A multimodal vehicle trajectory prediction method considering driving intention and interaction includes:

[0011] Obtain the historical driving trajectory of the target vehicle and the historical driving trajectory of vehicles in the surrounding environment;

[0012] Two historical driving trajectories are input into a vehicle trajectory prediction model to predict and generate the future driving trajectory of the target vehicle. The vehicle trajectory prediction model includes a trajectory encoder, an intent recognition module, an interaction force module, an intent interaction fusion module, a spatial relationship module, and a multimodal decoder.

[0013] The trajectory encoder is used to encode two historical driving trajectories to obtain high-level trajectories of the two historical driving trajectories. The high-level trajectories are vehicle driving behavior data and motion trend data extracted from the historical driving trajectories.

[0014] The intent recognition module is used to determine the tunneling code of the target vehicle based on the high-level trajectory of the target vehicle;

[0015] The interaction module is used to determine the overall interaction characteristics of the target vehicle based on the high-level trajectory of two historical driving trajectories. The overall interaction characteristics characterize the dynamic interaction relationship between the target vehicle and the surrounding vehicles.

[0016] The intent interaction fusion module is used to fuse the tunneling code with the overall interaction features of the target vehicle to obtain the interaction code;

[0017] The spatial relationship module is used to determine the graph structure encoding in the traffic scenario;

[0018] The multimodal decoder is used to fuse interactive coding and graph structure coding to obtain a unified trajectory of the target vehicle. It also uses two hybrid expert networks to predict the target vehicle's lateral and longitudinal movement categories in the lane, respectively. The predicted movement categories and the vectors corresponding to the unified trajectory are concatenated to generate a probability distribution of the target vehicle's future position based on the concatenated vector, thereby achieving the prediction of the target vehicle's future driving trajectory.

[0019] Preferably, the step of the intent recognition module determining the tunneling code of the target vehicle based on the high-level trajectory of the target vehicle includes:

[0020] Determine the lane potential field, boundary potential field, number of vehicles, average speed and vehicle type distribution in each lane for the target vehicle, so as to determine the high-dimensional potential field of each lane.

[0021] Based on the high-dimensional potential field of each lane, the penetration rate of each lane is determined, and the penetration rate characterizes the probability of a target vehicle changing lanes to the corresponding lane.

[0022] The penetration rates of each lane are aggregated to obtain a penetration vector, which is then sequentially input into a multilayer perceptron and a long short-term memory network for processing to obtain a tunneling vector. The tunneling vector is then fused with the high-level trajectory of the target vehicle to obtain the tunneling code.

[0023] Preferably, the step of the interaction module determining the overall interaction characteristics of the target vehicle based on the high-level trajectory of two historical driving trajectories includes:

[0024] Determine the interaction forces between the target vehicle and vehicles in the surrounding environment;

[0025] The unnormalized attention coefficients between the target vehicle and surrounding vehicles are determined based on the high-level trajectories of the two historical driving trajectories, in order to determine the normalized attention weights.

[0026] The overall interaction characteristics of the target vehicle are determined based on the interaction force and the normalized attention weight.

[0027] Preferably, the step of determining the interaction forces between the target vehicle and vehicles in the surrounding environment includes:

[0028] Determine the Euclidean distance between the target vehicle's historical driving trajectory and the historical driving trajectories of vehicles in the surrounding environment;

[0029] The desired safe distance between the target vehicle and surrounding vehicles is obtained by processing the relative motion characteristics between the target vehicle and surrounding vehicles using a multilayer perceptron; the relative motion characteristics include relative position, relative speed, relative acceleration, and the type of surrounding vehicles.

[0030] The interaction force between the target vehicle and surrounding vehicles is determined based on the Euclidean distance and the desired safe distance.

[0031] Preferably, the step of fusing the tunneling code with the overall interaction features of the target vehicle to obtain the interaction code by the intent interaction fusion module includes:

[0032] The learnable weight matrix of multiple attention heads is used to perform weight operations on the overall interaction features of tunneling coding and target vehicle to obtain query parameters, key parameters and value parameters corresponding to multiple attention heads;

[0033] The output of each attention head is obtained by performing scaled dot product attention calculation based on the query parameters, key parameters, and value parameters of each attention head. The outputs of each attention head are then concatenated and linearly transformed to obtain the interactive code.

[0034] Preferably, the step of the spatial relationship module determining the graph structure encoding in the traffic scene includes:

[0035] Determine multiple original relationship matrices in a traffic scenario. These original relationship matrices include a speed difference matrix, an acceleration difference matrix, and an action type matrix between any two vehicles in the traffic scenario. The action types include following, parallel, and cutting in.

[0036] Preliminary interaction features are obtained by processing multiple original relation matrices through a two-layer convolutional network, batch normalization, and regularization.

[0037] The initial interaction features are input into the gated loop unit to time-encode the interaction sequence of each vehicle, where the interaction sequence is a sequence formed by the change of the interaction relationship between vehicles over time.

[0038] The time-encoded result output by the gated recurrent unit is concatenated with multiple original relation matrices to construct node features. The node features are then processed through a two-layer graph attention network to obtain the graph structure encoding, where the nodes are vehicles in the traffic scene and the graph is used to represent the traffic scene.

[0039] Preferably, each hybrid expert network includes three sub-expert networks; wherein, the first hybrid expert network predicts the target vehicle's lateral movement categories in the lane, including lane keeping, left lane change, and right lane change; and the second hybrid expert network predicts the target vehicle's longitudinal movement categories in the lane, including acceleration, speed maintenance, and deceleration.

[0040] Preferably, the step of generating a probability distribution of the future position of the target vehicle based on the concatenated vector includes: inputting the concatenated vector into a long short-term memory network sequence decoder, and processing the decoding result through a multilayer perceptron to generate a probability distribution of the future position of the target vehicle.

[0041] Preferably, the historical driving trajectory includes the vehicle's two-dimensional coordinates, speed, acceleration, and lane index information.

[0042] Preferably, before the trajectory encoder performs encoding, the method further includes screening the vehicles in the surrounding environment, wherein the screening conditions at least meet the following conditions:

[0043] The distance from the target vehicle in the longitudinal direction shall not exceed ±90 feet;

[0044] Within the two adjacent lanes to the left and right of the lane where the target vehicle is located.

[0045] This invention has at least the following technical effects:

[0046] This invention provides a multimodal vehicle trajectory prediction method considering driving intentions and interactions. The method first processes the historical trajectory of the target vehicle, the historical trajectories of surrounding vehicles, and road structure information as input. An intention recognition module is used to establish road geometric constraints to predict the target vehicle's lane-changing intention and obtain a tunneling code. Subsequently, an interaction force module models the dynamic interactions between the target vehicle and surrounding vehicles, fusing this model with the tunneling code representing the vehicle's lane-changing intention to characterize the mutual influence relationships between vehicles. Building upon this, a spatial relationship module is further employed to model the global spatial dependencies in the traffic scene, improving the scene understanding capability in complex environments such as multi-vehicle interactions. Finally, a multimodal decoder based on an expert hybrid mechanism is introduced to decouple and predict the future trajectories corresponding to different driving operations. Therefore, this invention can achieve accurate prediction of vehicle trajectories.

[0047] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0048] Figure 1 This is a flowchart of a multimodal vehicle trajectory prediction method that considers driving intentions and interactions, according to an embodiment of the present invention.

[0049] Figure 2 This is a flowchart illustrating the multimodal vehicle trajectory prediction process according to an embodiment of the present invention. Detailed Implementation

[0050] The following describes this embodiment in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention.

[0051] The following description, with reference to the accompanying drawings, describes a multimodal vehicle trajectory prediction method that considers driving intentions and interactions.

[0052] Figure 1 This is a flowchart illustrating a multimodal vehicle trajectory prediction method that considers driving intentions and interactions, as described in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0053] Step S101: Obtain the historical driving trajectory of the target vehicle and the historical driving trajectory of vehicles in the surrounding environment.

[0054] Step S102: Input the two historical driving trajectories into the vehicle trajectory prediction model to predict and generate the future driving trajectory of the target vehicle; the vehicle trajectory prediction model includes a trajectory encoder, an intent recognition module, an interaction force module, an intent interaction fusion module, a spatial relationship module, and a multimodal decoder.

[0055] In this embodiment, let T h and T f These represent the historical time length and the predicted time length, respectively. Given the historical driving trajectory of the target vehicle. And the historical driving trajectories X of N surrounding vehicles. sur ={X1,X2,…,X N},in, Indicates the target vehicle is in tT h Historical driving trajectory at any moment, X N This represents the historical driving trajectory of the Nth vehicle in the surrounding environment, where N represents the total number of vehicles in the surrounding environment. The observation of each vehicle k at time t is... The observation includes two-dimensional coordinates. speed acceleration Lane Index in, Let x and y represent the x and y coordinates of vehicle k at time t, indicating its position. The goal of the vehicle trajectory prediction model is to predict the target vehicle's position in the future. f The trajectory at each time step: Among them, Y tar Indicates the future trajectory of the target vehicle. This indicates that the target vehicle will be at a future time t+T. f The trajectory of time.

[0056] As mentioned above, historical driving trajectories include the vehicle's two-dimensional coordinates, speed, acceleration, and lane index information. By inputting the historical driving trajectories of the target vehicle and surrounding vehicles into the vehicle trajectory prediction model, the future trajectory of the target vehicle can be predicted. The specific implementation process of each part of the vehicle trajectory prediction model is described in detail below.

[0057] In one embodiment of the present invention, a trajectory encoder is used to encode two historical driving trajectories to obtain high-level trajectories for the two historical driving trajectories. The high-level trajectories are vehicle driving behavior data and motion trend data extracted from the historical driving trajectories. An intent recognition module is used to determine the tunneling code of the target vehicle based on the high-level trajectory of the target vehicle. An interaction force module is used to determine the overall interaction features of the target vehicle based on the high-level trajectories of the two historical driving trajectories. The overall interaction features characterize the dynamic interaction relationship between the target vehicle and surrounding vehicles. An intent interaction fusion module is used to fuse the tunneling code with the overall interaction features of the target vehicle to obtain an interaction code. A spatial relationship module is used to determine the graph structure code in the traffic scene. A multimodal decoder is used to fuse the interaction code and the graph structure code to obtain a unified trajectory of the target vehicle. Two hybrid expert networks are used to predict the action categories of the target vehicle in the lateral and longitudinal directions of the lane, respectively. The predicted action categories and the vectors corresponding to the unified trajectory are concatenated to generate a probability distribution of the future position of the target vehicle based on the concatenated vector, thereby realizing the prediction of the future driving trajectory of the target vehicle.

[0058] This method also includes screening for vehicles in the surrounding environment before the trajectory encoder performs encoding.

[0059] Specifically, the surrounding environment of the target vehicle can be set to vehicles within ±90 feet of the target vehicle in the longitudinal direction, and within the two adjacent lanes to the left and right of the target vehicle's lane. Then, given the target vehicle's historical driving trajectory X... i and the historical driving trajectory of the vehicle in the surrounding environment X sur The trajectory encoder uses an MLP (Multilayer Perceptron) and an LSTM (Long Short-Term Memory Network) to encode historical driving trajectories, obtaining the high-level trajectory h of the target vehicle. i High-level trajectory H of vehicles in the surrounding environment N ={h1,h2,…,h j ,…,h N}, where h j and h N These represent the high-level trajectories of the j-th and N-th surrounding vehicles, respectively. The high-level trajectories are vehicle driving behavior data, motion trend data, or lane-change intention data extracted from historical driving trajectories.

[0060] In one embodiment, the step of the intent recognition module determining the tunneling code of the target vehicle based on the high-level trajectory of the target vehicle includes: determining the lane potential field, boundary potential field, number of vehicles, average speed, and vehicle type distribution of each lane of the target vehicle, so as to determine the high-dimensional potential field of each lane; determining the penetration rate of each lane based on the high-dimensional potential field of each lane, wherein the penetration rate characterizes the probability of the target vehicle changing lanes to the corresponding lane; aggregating the penetration rates of each lane to obtain a penetration vector, and sequentially inputting it into a multilayer perceptron and a long short-term memory network for processing to obtain a tunneling vector; and fusing the tunneling vector with the high-level trajectory of the target vehicle to obtain the tunneling code.

[0061] To capture structured prior information about the environment, which is crucial for intent reasoning, an intent recognition module is introduced, which is as follows: Figure 2 The diagram consists of two parts: a potential field modeling unit (used for potential field calculation) and a penetration rate calculator (i.e., penetration rate calculation) used for lane change intention inference.

[0062] In this embodiment, the overall potential field consists of two parts: the lane potential field and the boundary potential field, which respectively characterize the geometric properties of the current lane and the adjacent lanes. The lane potential field is defined as follows:

[0063]

[0064] Among them, E R Let A represent the lane potential field of the current lane, m represent the m-th lane, and A represent the lane potential field of the current lane. m These are learnable weighting factors that represent the importance of each lane line, y l,m y represents the lateral position of the boundary of the m-th lane. v σ represents the lateral position of the target vehicle, and is an adjustable parameter for the distribution range of the control field. This potential field reflects the degree of alignment and proximity of the vehicle to the current lane centerline.

[0065] The boundary potential field is defined as:

[0066]

[0067] Among them, E B Let |y| represent the boundary potential field of the current lane, where η is a learnable scaling parameter. l,m -y v || represents the distance between the target vehicle and the nearest lane edge. This function decays smoothly with distance, thus encouraging the target vehicle to stay away from the lane edge.

[0068] To integrate dynamic road information at a higher level, this embodiment also extracts lane-level traffic statistics, including the number of vehicles C in the m-th lane. m Average speed and vehicle type distribution T m(Statistics are collected over historical time slices, such as cars, trucks, and buses). These features, along with the calculated potential field, are input into the MLP and embedded into a high-dimensional potential field:

[0069]

[0070] Among them, E l The potential field is high-dimensional. Vehicles typically have a natural tendency to stay in the center of the lane, minimizing deviations unless necessary. This phenomenon can be analogized to a particle trapped near the minimum of a potential well, where the particle is most stable and has the lowest energy. In this analogy, the lane center corresponds to the lowest point of the potential field, the potential well, which exerts an attractive force, pulling the vehicle back to the centerline. This is achieved by weighting the decoder in the direction of minimum potential energy, thus supporting the vehicle's trajectory along the centerline.

[0071] Conversely, deviating from the center of the lane or entering an adjacent lane increases the risk, such as a collision or leaving the passable area. These danger zones can be viewed as potential energy barriers that repel vehicles and simulate the cost of unsafe maneuvering. This reflects the principle of physics: particles generally cannot overcome high-energy barriers unless they possess sufficient kinetic energy. Therefore, lane boundaries and adjacent lanes are embedded as high-potential-energy regions, forming a repulsive gradient, and the learned potential field penalty suppresses unsafe or unnecessary lane-changing behavior.

[0072] To further refine this representation, a tunneling analogy, inspired by quantum mechanics, is introduced to more realistically model lane-changing intentions. In quantum mechanics, a particle lacking sufficient energy may still penetrate a potential energy barrier with a certain probability, determined by the penetration rate. Similarly, a driver may intentionally cross lane boundaries, even with risks, the probability depending on driving behavior and surrounding traffic constraints.

[0073] The lane-changing intention (penetration rate) of target vehicle i towards lane m is defined as:

[0074]

[0075] Among them, T crs,m Let α represent the lane-changing intention of target vehicle i towards lane m, i.e., the penetration rate of the m-th lane. α is a learnable scaling factor, and M is an effective quality term discretely assigned based on vehicle type. It is an adjustable parameter, inspired by Planck's constant, used to control tunneling sensitivity, h i E represents the high-level trajectory of the target vehicle. l,m It is the potential barrier of lane m, that is, the high-dimensional potential field of lane m.

[0076] The penetration rates of all adjacent lanes c-1, c+1, and the current lane c are aggregated into a penetration vector, reflecting the probability of a vehicle entering each lane. Then, the penetration vector is processed using LSTM and MLP to obtain the context tunneling vector representation:

[0077] t i =LSTM(MLP([T crs,c-1 ,T crs,c ,T crs,c+1 ])) (5)

[0078] Among them, T crs,c-1 T crs,c T crs,c+1 These represent the penetration rates of adjacent lanes c-1, c+1, and the current lane c, respectively.

[0079] Finally, the high-level trajectory h of the target vehicle is obtained through an attention mechanism. i The tunneling vector is fused with the tunneling vector, and the contributions of each vector are adaptively weighted based on the context to obtain the final tunneling code.

[0080] For example, if the target vehicle is traveling at a constant speed in the center of the lane, then the high-level trajectory h of the target vehicle is... i Very stable, t characterizing tunneling intent i It's also very weak. At this point, the system might place more emphasis on historical driving trajectories, predicting that the vehicle will stay in its lane. That is, it assigns a high-level trajectory h. i Higher weighting. For example, if the target vehicle has started to drift to the right, the target vehicle's high-level trajectory h... i The system exhibits significant lateral movement, while the right lane is relatively empty (leading to a strong tunneling intent). In this situation, both historical driving trajectories and lane-changing intentions are crucial, and the system assigns high weights to both. Therefore, the final tunneling code can be obtained by adaptively weighting the contributions based on the context.

[0081] In one embodiment, the step of the interaction force module determining the overall interaction characteristics of the target vehicle based on the high-level trajectories of two historical driving trajectories includes: determining the interaction force between the target vehicle and surrounding vehicles; determining the unnormalized attention coefficient between the target vehicle and surrounding vehicles based on the high-level trajectories of the two historical driving trajectories in order to determine the normalized attention weight; and determining the overall interaction characteristics of the target vehicle based on the interaction force and the normalized attention weight.

[0082] The step of determining the interaction force between the target vehicle and surrounding vehicles includes: determining the Euclidean distance between the historical driving trajectory of the target vehicle and the historical driving trajectory of the surrounding vehicles; processing the relative motion characteristics between the target vehicle and the surrounding vehicles using a multilayer perceptron to obtain the expected safe distance between the target vehicle and the surrounding vehicles; the relative motion characteristics include relative position, relative speed, relative acceleration, and the type of surrounding vehicles; and determining the interaction force between the target vehicle and the surrounding vehicles based on the Euclidean distance and the expected safe distance.

[0083] Specifically, after capturing the intent representation through the potential field, the dynamic interaction relationship between vehicles, i.e., the interaction intensity between vehicles, can be further modeled through the interaction force module. Accurately capturing the spatiotemporal interaction between vehicles helps to more accurately predict the future trajectory of the target vehicle. Here, physical interaction modeling is combined with attention mechanisms that consider human perception to quantify the interaction intensity between vehicles. This process consists of two stages: the first is to calculate the interaction force between vehicles based on the encoded state; the second is to introduce an attention weighting mechanism that is affected by the human visual range and temporal context.

[0084] First, interaction force is defined as:

[0085]

[0086] Among them, f ij This represents the interaction force between the target vehicle i and the surrounding environment vehicle j. A and B are two learnable parameters, and d... ij It is the Euclidean distance between the target vehicle i and the surrounding environment vehicle j, which is calculated based on the two-dimensional coordinates of the target vehicle and the surrounding environment vehicles. ij n is calculated from the relative motion characteristics (relative position, vehicle type, relative speed, relative acceleration) of the interacting vehicles using MLP. ij It is a unit vector pointing from surrounding vehicle j to target vehicle i.

[0087] Furthermore, it was noted that drivers typically pay more attention to nearby vehicles within their forward fan-shaped visual range. Therefore, an attention weighting mechanism was introduced to reflect spatial and perceptual preferences in interaction modeling.

[0088] Since human drivers don't pay equal attention to all vehicles, they generally focus more on vehicles ahead and those closer to them. Therefore, by normalizing the attention weight α... ij To adjust the interaction force calculated above.

[0089] Specifically, the unnormalized attention coefficient e between the target intelligent agent (target vehicle i) and the surrounding intelligent agents (surrounding environment vehicles j) ij The calculation is as follows:

[0090] eij =LeakyReLU(W attn [h i ,h j (7)

[0091] Where LeakyReLU represents an activation function, h i and h j These are high-level trajectories of the target vehicle and the historical driving trajectories of vehicles in the surrounding environment, respectively. attn It is a trainable weight matrix.

[0092] Furthermore, the normalized attention weight α ij Obtained through the softmax function:

[0093]

[0094] Among them, W i It is a sector-shaped attention matrix determined by the current vehicle speed and lane. ik This represents the unnormalized attention coefficient between the target vehicle i and the surrounding environment vehicles k.

[0095] Ultimately, the overall interaction characteristics F of target vehicle i i It is obtained by calculating the interaction force through aggregated attention weighting, and modulated by considering the perceptual distance decay function:

[0096]

[0097] Where β and γ are adjustable parameters that control the decay rate, and the denominator is a decay function that takes into account the sensing distance.

[0098] In this embodiment, the combined modeling method can simultaneously capture physical interaction and cognitively consistent human attention dynamics, thereby generating a more accurate interaction representation that conforms to human driving behavior.

[0099] In one embodiment, the step of fusing tunneling code with the overall interaction features of the target vehicle to obtain interaction code by the intent interaction fusion module includes: using learnable weight matrices of multiple attention heads to perform weight calculations on the tunneling code and the overall interaction features of the target vehicle to obtain query parameters, key parameters, and value parameters corresponding to multiple attention heads; performing scaled dot product attention calculations based on the query parameters, key parameters, and value parameters of each attention head to obtain the output result of each attention head; and concatenating and linearly transforming the output results of each attention head to obtain the interaction code.

[0100] In complex driving environments, the behavior of a target vehicle is influenced not only by the interactions of vehicles in its current surroundings, but also by time dependence and lane-changing intentions. Therefore, this embodiment designs an intention-based interaction fusion module to integrate the overall interaction characteristics F of the target vehicle. i With tunnel coding The contributions of both are integrated and adaptively adjusted based on the temporal context and surrounding traffic dynamics.

[0101] Specifically, a learnable weight matrix of an attention head is used to perform weight operations on the overall interaction features of tunneling coding and target vehicles to obtain the query parameter Q, key parameter K, and value parameter V corresponding to the attention head.

[0102]

[0103] Among them, W Q W K W V ∈R D×D There are three learnable weight matrices, where D is the feature dimension of the input vector, and the features are defined over a historical time length T. h In the above, R is a real number.

[0104] Each attention head independently computes scaled dot product attention:

[0105]

[0106] Where h = 1, ..., H, h represents the h-th attention head, H represents the total number of attention heads, and Q h K h V h These represent the query parameters, key parameters, and value parameters corresponding to the h-th attention head, respectively. h This represents the output of the h-th attention head. Each attention head has a dimension of d. k =D / H. The outputs of all attention heads are concatenated and linearly transformed to obtain the fused interactive code O. i :

[0107] O i =W O ·Concat(head1,…,head H (12)

[0108] Among them, W O ∈R D×D To output the projection matrix, Concat represents the concatenation function, and head... H This represents the output of the Hth attention head.

[0109] In one embodiment, the step of the spatial relationship module in determining the graph structure encoding in a traffic scene includes: determining multiple original relationship matrices in the traffic scene, the multiple original relationship matrices including a speed difference matrix, an acceleration difference matrix, and an action type matrix between any two vehicles in the traffic scene, wherein the action types include following, parallel, and cutting in; processing the multiple original relationship matrices through a two-layer convolutional network, batch normalization, and regularization to obtain preliminary interaction features; inputting the preliminary interaction features into a gated recurrent unit to perform temporal encoding on the interaction sequence of each vehicle, wherein the interaction sequence is a sequence formed by the change of the interaction relationship between vehicles over time; concatenating the temporal encoding result output by the gated recurrent unit with the multiple original relationship matrices to construct node features, and processing the node features through a two-layer graph attention network to obtain graph structure encoding, wherein the nodes are vehicles in the traffic scene, and the graph is used to represent the traffic scene.

[0110] By combining lane-level intent representation with interaction perception, a spatial relationship module was designed to capture higher-order spatial interactions that cannot be fully represented by pairwise relationships. This module consists of three stages:

[0111] The first stage involves concatenating multiple original relationship matrices (such as speed difference, acceleration difference, and action type) and processing them through a two-layer convolutional network, batch normalization, and dropout (regularization) to obtain preliminary interaction features. These preliminary interaction features comprehensively encode the motion differences between vehicles (speed difference, acceleration difference, relative distance, direction angle, etc.), action semantics (action type embedding), and spatial context information (relative lane and orientation relationships), providing a multimodal interaction representation for subsequent temporal and graph structure modeling.

[0112] The second stage involves inputting the preliminary interaction features into a GRU (Gated Recurrent Unit) to temporally encode the interaction sequence for each vehicle. For each vehicle, its preliminary interaction features across consecutive time frames constitute an interaction sequence, reflecting the temporal changes in the relationship between the vehicle and its neighbors. By inputting this interaction sequence into the GRU, the model can encode temporal dependencies, capture the dynamic patterns of how the intensity of interactions between vehicles changes over time, and provide temporal context for subsequent modeling of higher-order spatial relationships.

[0113] The third stage involves concatenating the GRU output with the original relation matrix to construct node features. These features include the temporal dynamics between the target vehicle and its neighboring vehicles, the original relation features, and spatial semantic context information, simultaneously reflecting the temporal dependency and spatial topology of the interaction relationships. This is then processed through a two-layer Graph Attention Network (GAT) to obtain the final graph structure encoding O. gThe graph structure encoding captures the dynamic influence intensity of different neighboring vehicles on the target vehicle at the node level, and models the multi-vehicle group collaborative features at the global level, that is, the higher-order structure in the traffic scene that goes beyond pairwise interactions, thus forming a spatiotemporal joint representation of structure awareness.

[0114] Furthermore, the interaction code O output by the intent interaction fusion module i The graph structure encoding O output by the spatial relations module g A unified trajectory I is generated through Transformer (a deep learning model based on attention mechanisms) fusion, simultaneously encoding temporal interaction dependencies and higher-order spatial relationships. Encoding temporal interaction dependencies refers to encoding the dynamic changes in the target vehicle's behavior over time and the instantaneous interactions between vehicles. Encoding higher-order spatial relationships refers to encoding complex group spatial constraints in traffic scenarios that go beyond pairwise relationships, such as group spatial constraints between three or more vehicles. For example, if target vehicle A wants to change lanes to the right, it is influenced not only by vehicle B in the right lane but also by vehicle C in front of vehicle B. Vehicle A's decision is actually influenced by the global situation of a small group consisting of vehicles A, B, and C; this complex relationship constitutes a higher-order spatial relationship.

[0115] The decoder first extracts the final hidden state I0 from the unified trajectory I, which is used to predict the target vehicle's lateral and longitudinal movement probabilities M1 = {m} through two independent MoE modules (hybrid expert networks). p |p=1,2,…,6},m p Let be the probability of the p-th action. Each MoE module contains three sub-expert networks, each specializing in predicting different action categories (e.g., lane keeping, left lane change, right lane change, and acceleration, speed maintenance, deceleration). The predicted action category is encoded using one-hot encoding, then mapped to an embedding vector and concatenated with a unified trajectory I.

[0116] The concatenated fused vector is input into an LSTM sequence decoder, and then an MLP is used to generate the probability distribution of future positions. The vehicle trajectory prediction model outputs a two-dimensional Gaussian distribution parameter at each future time step. This indicates the uncertainty and correlation of the predicted trajectory. Predicted location. for:

[0117]

[0118] in, The two-dimensional Gaussian distribution representing the predicted target vehicle position. and Let be the mean and variance of the future positions, respectively. Let Ω be the correlation coefficient. The parameter of the two-dimensional Gaussian distribution is denoted as Ω.

[0119] Based on the historical driving trajectory X of the target vehicle in the traffic scene i Future trajectory Y i The posterior probability P(Y) i |X i )for:

[0120] P(Y i |X i )=∑ p P(m p |X i )P Ω (Y i |m p ,X i (14)

[0121] Wherein, P(m) p |X i ) represents the historical driving trajectory X of a given target vehicle. i Under the given conditions, the target vehicle actually takes driving actions m p The probability, P Ω (Y i |m p ,X i ) represents the historical driving trajectory X of a given target vehicle. i And assume that the target vehicle determines to take driving action m p Under the given conditions, the probability of the future trajectory Yi occurring.

[0122] The training process of this model is briefly described below.

[0123] (1) Data preparation and preprocessing

[0124] Dataset preparation: Real-world vehicle trajectory datasets such as the NGSIM dataset are used, which contain trajectory records of various vehicles.

[0125] Preprocessing: 1) Group by vehicle number and frame, and store all trajectory data for each vehicle. 2) Obtain the complete trajectory of the current vehicle and find the initial trajectory time of the current trajectory point. 3) Calculate the lateral / longitudinal lane change action type based on the vehicle's historical trajectory (±40 frames, i.e., 4s). 4) Obtain all vehicle trajectory data within the current time frame and divide them into the current lane, left lane, and right lane. For each vehicle in the lane, calculate its longitudinal distance from the current vehicle. If the absolute value of the distance is less than 90 feet, calculate its index in the grid and store the vehicle's number in the corresponding column of the current trajectory point data. 5) Store the complete trajectory of each vehicle by dataset and vehicle ID (identity identifier), filter the first 20 frames (2s) and the last 50 frames (5s), retain the trajectory data of the past 3s and the next 5s, and take its dataset number, vehicle number, frame number, X and Y coordinates, vehicle speed, vehicle acceleration, lane ID, and vehicle type. Then save the training, testing, and validation sets.

[0126] (2) Track encoder encoding

[0127] For the historical driving trajectory of the target vehicle and the historical driving trajectory of vehicles in the surrounding environment, their coordinates, vehicle speed, vehicle acceleration, lane ID, and vehicle type are mapped into a latent vector sequence through LSTM and MLP.

[0128] (3) Training process

[0129] 1) For a batch of target vehicles' historical driving trajectories and the historical driving trajectories of surrounding vehicles, trajectory codes are generated using the intent recognition module, interaction force module, intent interaction fusion module, and spatial relationship module; 2) The trajectory codes are input into a multimodal decoder based on driving action MoE to predict the target vehicle's trajectory in the next 3 seconds; 3) These modules are pre-trained using the RMSE (root mean square error) loss function; 4) After training for 4 rounds, the model is trained using the NLL (negative log-likelihood) loss function until it converges.

[0130] In summary, the vehicle trajectory prediction method proposed in this invention fully integrates prior knowledge of road structure with physical interaction force and potential field modeling, thereby making the prediction results more consistent with actual traffic constraints. This invention effectively characterizes the interaction between vehicle intent and vehicle dynamics by introducing a force field mechanism, resulting in high prediction accuracy and stability in highway scenarios. Furthermore, this invention employs an expert hybrid decoding approach for multimodal prediction of future trajectories, ensuring good separability and diversity among the prediction modes. This not only covers reasonable possible trajectories under different driving intentions but also avoids the problems of prediction meanization and distribution ambiguity in traditional methods. Because this invention incorporates environmental constraints and interaction modeling in its method design, its prediction results possess better interpretability and robustness, enhancing the reliability and practical value of autonomous driving systems in safety decision-making and path planning.

[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal vehicle trajectory prediction method considering driving intention and interaction, characterized in that, include: Obtain the historical driving trajectory of the target vehicle and the historical driving trajectory of vehicles in the surrounding environment; Two historical driving trajectories are input into a vehicle trajectory prediction model to predict and generate the future driving trajectory of the target vehicle. The vehicle trajectory prediction model includes a trajectory encoder, an intent recognition module, an interaction force module, an intent interaction fusion module, a spatial relationship module, and a multimodal decoder. The trajectory encoder is used to encode two historical driving trajectories to obtain high-level trajectories of the two historical driving trajectories. The high-level trajectories are vehicle driving behavior data and motion trend data extracted from the historical driving trajectories. The intent recognition module is used to determine the tunneling code of the target vehicle based on the high-level trajectory of the target vehicle; The interaction module is used to determine the overall interaction characteristics of the target vehicle based on the high-level trajectory of two historical driving trajectories. The overall interaction characteristics characterize the dynamic interaction relationship between the target vehicle and the surrounding vehicles. The intent interaction fusion module is used to fuse the tunneling code with the overall interaction features of the target vehicle to obtain the interaction code; The spatial relationship module is used to determine the graph structure encoding in the traffic scenario; The multimodal decoder is used to fuse interactive coding and graph structure coding to obtain a unified trajectory of the target vehicle. Two hybrid expert networks are used to predict the target vehicle's lateral and longitudinal movement categories in the lane, respectively. The vectors corresponding to the predicted movement categories and the unified trajectory are concatenated to generate a probability distribution of the target vehicle's future position based on the concatenated vector, thereby realizing the prediction of the target vehicle's future driving trajectory.

2. The method as described in claim 1, characterized in that, The step of the intent recognition module determining the tunneling code of the target vehicle based on the high-level trajectory of the target vehicle includes: Determine the lane potential field, boundary potential field, number of vehicles, average speed and vehicle type distribution in each lane for the target vehicle, so as to determine the high-dimensional potential field of each lane. Based on the high-dimensional potential field of each lane, the penetration rate of each lane is determined, and the penetration rate characterizes the probability of a target vehicle changing lanes to the corresponding lane. The penetration rates of each lane are aggregated to obtain a penetration vector, which is then sequentially input into a multilayer perceptron and a long short-term memory network for processing to obtain a tunneling vector. The tunneling vector is then fused with the high-level trajectory of the target vehicle to obtain the tunneling code.

3. The method as described in claim 1, characterized in that, The step of the interaction module determining the overall interaction characteristics of the target vehicle based on the high-level trajectory of two historical driving trajectories includes: Determine the interaction forces between the target vehicle and vehicles in the surrounding environment; The unnormalized attention coefficients between the target vehicle and surrounding vehicles are determined based on the high-level trajectories of the two historical driving trajectories, in order to determine the normalized attention weights. The overall interaction characteristics of the target vehicle are determined based on the interaction force and the normalized attention weight.

4. The method as described in claim 3, characterized in that, The steps for determining the interaction forces between the target vehicle and vehicles in the surrounding environment include: Determine the Euclidean distance between the target vehicle's historical driving trajectory and the historical driving trajectories of vehicles in the surrounding environment; The desired safe distance between the target vehicle and surrounding vehicles is obtained by processing the relative motion characteristics between the target vehicle and surrounding vehicles using a multilayer perceptron; the relative motion characteristics include relative position, relative speed, relative acceleration, and the type of surrounding vehicles. The interaction force between the target vehicle and surrounding vehicles is determined based on the Euclidean distance and the desired safe distance.

5. The method as described in claim 1, characterized in that, The step of fusing tunneling coding with the overall interaction features of the target vehicle to obtain interaction coding in the intent interaction fusion module includes: The learnable weight matrix of multiple attention heads is used to perform weight calculations on the overall interaction features of tunneling coding and target vehicles to obtain query parameters, key parameters and value parameters corresponding to multiple attention heads; The output of each attention head is obtained by performing scaled dot product attention calculation based on the query parameters, key parameters, and value parameters of each attention head. The outputs of each attention head are then concatenated and linearly transformed to obtain the interactive code.

6. The method as described in claim 1, characterized in that, The steps for the spatial relationship module to determine the graph structure encoding in a traffic scenario include: Determine multiple original relationship matrices in the traffic scene. The multiple original relationship matrices include the speed difference matrix, acceleration difference matrix, and action type matrix between any two vehicles in the traffic scene. The action types include following, parallel, and cutting in. Preliminary interaction features are obtained by processing multiple original relation matrices through a two-layer convolutional network, batch normalization, and regularization. The initial interaction features are input into the gated loop unit to time-encode the interaction sequence of each vehicle, where the interaction sequence is a sequence formed by the change of the interaction relationship between vehicles over time. The time-encoded result output by the gated recurrent unit is concatenated with multiple original relation matrices to construct node features. The node features are then processed through a two-layer graph attention network to obtain the graph structure encoding, where the nodes are vehicles in the traffic scene and the graph is used to represent the traffic scene.

7. The method as described in claim 1, characterized in that, Each hybrid expert network comprises three sub-expert networks; wherein, the first hybrid expert network predicts the target vehicle's lateral movement categories in the lane, including lane keeping, left lane change, and right lane change; and the second hybrid expert network predicts the target vehicle's longitudinal movement categories in the lane, including acceleration, speed maintenance, and deceleration.

8. The method as described in claim 7, characterized in that, The steps for generating the probability distribution of the future position of the target vehicle based on the concatenated vector include: inputting the concatenated vector into a long short-term memory network sequence decoder, and processing the decoding result through a multilayer perceptron to generate the probability distribution of the future position of the target vehicle.

9. The method according to any one of claims 1 to 8, characterized in that, The historical driving trajectory includes the vehicle's two-dimensional coordinates, speed, acceleration, and lane index information.

10. The method according to any one of claims 1 to 8, characterized in that, Before the trajectory encoder performs encoding, the method further includes screening the surrounding environment vehicles, wherein the screening criteria at least meet the following conditions: The distance from the target vehicle in the longitudinal direction shall not exceed ±90 feet; Within the two adjacent lanes to the left and right of the lane where the target vehicle is located.

Citation Information

Patent Citations

  • Time-space interaction vehicle trajectory prediction method based on speed perception

    CN120611190A

  • Vehicle multi-modal trajectory prediction method based on improved attention network

    CN120672802A