Vehicle track prediction method based on multi-image attention fusion
By building a multi-graph attention fusion network, combined with a graph attention network and global pooling strategy, the problem of insufficient accuracy and robustness of vehicle trajectory prediction in complex traffic scenarios is solved, and more efficient multi-dimensional spatial information fusion and correlation improvement are achieved.
Patent Information
- Application Number
- CN202510071528.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing vehicle trajectory prediction methods are insufficient in the accuracy and robustness of the prediction results in complex traffic scenarios, and traditional graph neural networks lack efficient multi-dimensional spatial information fusion mechanisms, ignoring the correlation and interaction between different graph sequences.
The vehicle trajectory prediction method based on multi-graph attention fusion is adopted, and the vehicle information interaction topology diagram is constructed by constructing a timing chart of topology diagram and a vehicle information interaction topology diagram, combined with the graph attention network and global pooling strategy, feature extraction and information fusion are carried out to fully capture spatial and spatial and temporal information.
It significantly improves the correlation and interaction intensity between the topological structure diagram and the vehicle interaction diagram, improves the accuracy and robustness of vehicle trajectory prediction, and enhances the model's adaptability in complex traffic scenarios.
Smart Images

Figure BDA0005245838820000021 
Figure BDA0005245838820000031 
Figure BDA0005245838820000041
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation vehicle trajectory prediction, and in particular to a vehicle trajectory prediction method based on multi-image attention fusion. Background Art
[0002] At present, vehicle trajectory prediction mainly relies on historical trajectory data, using models such as long short-term memory network (LSTM), recurrent neural network (RNN), and graph neural network to predict the future position of vehicles. Graph neural network can use graph structure to model the spatial information of vehicles and roads, and extract the interactive relationship between vehicles and the environment.
[0003] However, in actual driving, the trajectory of a vehicle is not only affected by its own historical information, but also by the dynamic changes of other vehicles and the road environment around it. For example, complex spatial interactions such as the spatial topological structure of the road, the relative positions between vehicles, speed changes, and acceleration and deceleration behaviors lead to insufficient accuracy and robustness of prediction results in complex traffic scenarios. The current graph structure design method cannot effectively take into account both the topological relationship of the road network and the flow information of the behavior between vehicles, resulting in limitations in the model's understanding of real information. In addition, when dealing with complex traffic scenarios, traditional graph neural networks lack an efficient fusion mechanism for multidimensional spatial information, ignoring the correlation and interaction between different graph sequences, resulting in reduced utilization efficiency of spatiotemporal information, limiting the model's adaptability and prediction accuracy in complex traffic scenarios. Therefore, it is urgent to introduce more efficient graph structure design and spatiotemporal information capture mechanisms to improve the performance of trajectory prediction. Summary of the invention
[0004] A vehicle trajectory prediction method based on multi-image attention fusion includes the following steps:
[0005] S1. Constructing a timing diagram for predicting a topological structure diagram of a traffic scene in which a vehicle is located and a timing diagram for predicting a topological diagram of information interaction between the vehicle and surrounding vehicles, including the following steps:
[0006] S1.1. Establish a time sequence diagram of the topological structure diagram of the traffic scene where the predicted vehicle is located. The expression is as follows:
[0007]
[0008] In the formula, G ro Represents the total set of topological structure graphs and time series graphs of the traffic scene where the predicted vehicle is located; T represents the time length of the data set, T∈(0,+∞); G ro,t It represents the topological structure diagram of the traffic scene within a 50m diameter circle with the predicted vehicle as the center at the time t of the scene, and the circle is taken as the sensing range; V ro,tRepresents G ro,t The node set at time t, V ro,t The set range is (0, nt-ro), nt-ro represents the maximum number of nodes in the traffic scene topology graph at time t, where nt-ro∈(0,+∞); E ro,t Represents G ro,t The set of node-edge relationships at time t, which means the summary of the connection relationships of the roads in the traffic scene at this time, E ro,t The set range is [0,2×nt-ro]; x i ro,t represents V at time t ro,t The characteristics of the node set are due to the different time G ro,t It is represented by dynamic road information, which reflects the dynamic change of information of traffic scene roads, i∈(1~nt-ro); and V is the node set ro,t The first and second child nodes of represent two roads in the set of perceived roads in the traffic scene. If the two roads are connected or can be switched, they are included in the edge relationship set of the road connection. represents the The feature set of a node set, Indicates V ro,t One of the nodes below, It means one of the road information in the traffic scene at this time. The following contains four characteristic information of the road, which are the nodes at time t The rear drive road number of the road Indicates the number of the connecting road ahead in the direction of vehicle travel; the node at time t The road's predecessor road number Indicates the number of the connecting road behind the vehicle in the direction of travel; the node at time t Road centerline information for roads Indicates the node road at this time The coordinate information of the center line; the node at time t Road ID code Indicates the current node road Serial number;
[0009] S1.2. Establish a time sequence diagram of the topological diagram of the information interaction between the predicted vehicle and the surrounding vehicles, and its expression is as follows:
[0010]
[0011] In the formula, G vorepresents the total set of topological time-series diagrams of information interaction between the constructed predicted vehicle and surrounding vehicles; T represents the time length of the data set, T∈(0,+∞); G vo,t V represents the vehicle interaction information topology diagram of the traffic scene predicted by the vehicle within the perception range set by S1.1 at time t; vo,t Represents G vo,t The node set at time t means the vehicle information set within the sensing range at this time, V vo,t The set range is (0, nt-vo), nt-vo represents the maximum number of nodes in the vehicle information interaction topology graph at time t, where nt-vo∈(0,+∞); E vo,t Represents G vo,t The set of node-edge relationships at time t, E vo,t The collection range is [0,2×nt-vo]; G ro,t V at time t ro,t The feature set of a node set, The collection range is (0,nt-vo); represents V at time t vo,t The characteristics of the node set are due to the different time G vo,t It is expressed as dynamic vehicle interaction information, which reflects the change of interaction information between the predicted vehicle and surrounding vehicles, where,, where,k∈(1,nt-vo); and Represents the node set V vo,t The first and second child nodes of represent the two vehicles in the traffic scene within the sensing range at this time. If there is an interaction relationship between the vehicles, they are included in the edge set of vehicle interaction. Indicates V vo,t One of the nodes below, represents the The feature set of the node set; The following contains 5 characteristic information of the vehicle, which are the nodes at time t Vehicle coding characteristic information of the vehicle Represents this node Vehicle number; Node Vehicle coordinate feature information Represents the coordinate information of this vehicle; node Vehicle speed characteristic information Represents this node Vehicle speed; node Vehicle acceleration characteristic information Represents this node The acceleration of the vehicle; the node Vehicle yaw angle characteristic information Represents this node The vehicle's yaw angle.
[0012] S2. Constructing a multi-image attention fusion network in a vehicle trajectory prediction method including the following steps:
[0013] S2.1, unify the feature information in S1.1 and S1.2;
[0014] The characteristic information includes: number information, coordinate information and vehicle physical information;
[0015] The steps to unify the size are as follows:
[0016] S2.1.1. Encode the numbering information in S1.1 and S1.2. The expression is as follows:
[0017]
[0018] In the formula, Represents node features The 1×32 vector obtained by encoding is Represents node features The 1×32 vector obtained by encoding is Roembeding and Voembeding represent the road number embedding layer and vehicle number embedding layer respectively;
[0019] S2.1.2, encode the coordinate information and vehicle physical information in S1.1 and S1.2, and the expression is as follows:
[0020]
[0021] In the formula, Indicates FC1 1-32 Fully connected layer pair The 1×32 vector obtained by encoding the node features, F veh-coordiantes Indicates FC2 2-32 The fully connected layer is used to calculate the vehicle coordinate information The obtained encoding vector; The centerline information dimension of each road in the lower sensing range of S1.1 will change dynamically at different times. For subsequent feature processing, it is necessary to use TP function and FC3 240-32 Fully connected layer pair The dimension is normalized, where the TP function will Feature information with a unified dimension of 120 In the TP(·) function, An represents the range of the input set A, which is Dynamic range of dimension;;F center-lines Indicates FC3 120-32 Fully connected layer pair Encoding; An in the TP(·) function represents the range of the input set A;
[0022] S2.2, the graph set G obtained by combining S1.1 and S1.2 vo With G ro In the random time interval θ, ts is the sequence length, ts∈(0,+∞), and the time series graph G is taken ro / vo,θ ,G ro / vo,θ+1 ,...,G ro / vo,θ+ts As input, it is defined as a graph sequence The importance of the interaction between the predicted vehicle and the surrounding vehicles and the mutual influence of the environmental spatial information and road information within the input step are adaptively captured by the graph attention network. The steps are as follows:
[0023] S2.2.1. The expression of the graph attention convolution process of the time series graph of the traffic scene topology graph where the vehicle is located is as follows:
[0024]
[0025] In the formula, LakyReLu is the activation function; Represents node v i Update features under the k'th attention head; Node v representing the k'th head i and neighbor node v j The attention coefficient between (k') The dimension of the k'th head is F'×F weight matrix, F' is the new feature dimension of the mapping, and F is the original feature dimension; N(i) represents the node v i The set of neighbor nodes; W (k') h i Represents node v i Feature transformation under the k'th head; W (k') h i ||W (k') h j Represents node v i With neighbor node v j Feature splicing of a (k')T represents the learnable attention weight vector of the k'th attention head; || represents concatenation; h′ i represents the node features finally updated by the graph attention network; H i Represents the feature matrix of all nodes in the final updated graph; where node vi Represents all nodes in the time series diagram of the traffic scene topology diagram where the predicted vehicle is located;
[0026] S2.2.2. The expression of the graph attention convolution process of the time series graph of the topological graph of the information interaction between the vehicle and the surrounding vehicles is as follows:
[0027]
[0028] In the formula, Represents node v k Update features under the k'th attention head; Node v representing the k'th head k and neighbor node v jk The attention coefficient between (k') Indicates that the dimension of the k'th head is F'×F weight matrix, F' is the new feature dimension of the mapping, and F is the original feature dimension; N(k) represents the node v k The set of neighbor nodes; W (k') h k Represents node v k Feature transformation under the k'th head; W (k') h k ||W (k') h jk Represents node v k With neighbor node v jk Feature splicing of a (k')T represents the learnable attention weight vector of the kth attention head, with dimension 2F'; || represents concatenation, h′ i represents the node features finally updated by the graph attention network; H k Represents the feature matrix of all nodes in the final updated graph; where node v k Represents all nodes in the time series diagram of the topological diagram of the information interaction between the predicted vehicle and surrounding vehicles;
[0029] S2.3. For input In each picture Perform a global pooling operation, where The details are as follows:
[0030] S2.3.1. The expression of the global pooling process of the time sequence diagram of the traffic scene topology diagram where the predicted vehicle is located is as follows:
[0031]
[0032] In the formula, p is in the graph sequence The index in V pi yes The set number of nodes in the graph, hpi Representation diagram The global feature representation of h pi The feature dimension is determined by the output channel number dimension of the GAT convolution setting. The convolution output channel number dimension of the two images is set to Gc-d.
[0033] S2.3.2. The expression of the global pooling process of the time sequence diagram of the predicted vehicle and surrounding vehicle information interaction topology diagram is as follows:
[0034]
[0035] In the formula, p is in the graph sequence The index in V pk yes The set number of nodes in the graph, h pk Representation diagram The global feature representation of h pk The feature dimension is determined by the output channel number dimension of the GAT convolution setting. The convolution output channel number dimension of the two images is set to Gc-d.
[0036] S2.3.3. Graph Sequence The global feature information of the road topology graph sequence is obtained through graph attention convolution and global pooling at each step Global feature information of topological graph sequence interacting with vehicle information
[0037] S2.4, the global feature information of S2.3 is fused through the co-attention network to obtain a fused feature sequence;
[0038] Here are the steps:
[0039] S2.4.1. Linearly transform the query, key, and value matrices under the number of attention heads; the expression is as follows:
[0040] The query, key, and value matrices are linearly transformed under the number of attention heads; the expression is as follows:
[0041]
[0042] Where ch is the number of shared attention heads, ch∈(0,+∞); Q ch , K ch , V ch They are the linear changes of Q, K and V matrices under each attention head; is the weight matrix of each attention head that can be learned; both are matrix;
[0043] S2.4.2. Under each attention head, through Qch and The attention score matrix is obtained by matrix multiplication. The scaling factor is applied to reduce the elements in the attention score matrix to avoid the element values in the attention score matrix being too large. The elements of the attention score matrix are further normalized by the softmax function. The content of the attention score matrix elements represents V ch The weight of each element in the matrix is obtained by comparing the attention score matrix element with V ch Perform matrix multiplication to obtain the fusion information of each attention head;
[0044] Its expression is as follows:
[0045]
[0046] In the formula, Attention ch-weight Represents the attention score matrix; Attention ch represents the fusion information under the attention head of ch; d ck is a scaling factor to prevent the matrix dot product value from being too large; softmax represents a normalization function;
[0047] S2.4.3. Collect the fusion information of all attention heads to obtain the final fusion feature sequence * ;
[0048] Its expression is as follows:
[0049] CoattentInfo=Concat(Attention 1 ,...,Attention ch )W O
[0050] In the formula, CoattentInfo represents the fusion feature sequence under multi-head co-attention * ; Concat represents the matrix concatenation process; W O Represents the learnable output matrix; CoattentInfo,
[0051] S2.4.4. Accelerate model training speed for fusion feature sequence * The feature information is used to reduce the dimension.
[0052] Its expression is as follows:
[0053] CoattentInfo'=DeFc Gc-d-32 (CoattentInfo)
[0054] In the formula, CoattentInfo' represents the DeFc Gc-d-32The fully connected layer is a fusion feature sequence after the dimension reduction of each step feature in the CoattentInfo sequence; DeFc Gc-d-32 The input dimension is the output channel dimension Gc-d of the GAT convolution setting set in S2.2, and the output dimension is 32.
[0055] S3. Based on the fused feature sequence information obtained in S2, a multi-image attention fusion network is constructed as the prediction network in the vehicle trajectory prediction method. The specific steps are as follows:
[0056] S3.1, convolve the fusion feature sequence information obtained in S2.4 through the SSM recursive hidden state equation sequence convolution kernel under Mamba to obtain a time convolution sequence. The specific process is as follows:
[0057] S3.1.1. Map the input sequence to the output sequence by establishing the SSM recursive hidden state equation. The expression of the SSM recursive hidden state continuity equation is as follows:
[0058]
[0059] In the formula, x(t) represents continuous input; y(t) represents continuous output; h(t) is the state representation matrix; A and B are the first state matrix and the second state matrix; C is the projection matrix; A, B, and C matrices are all learnable matrices. seq represents the continuity length of x(t), and f_num represents the feature dimension of x(t);
[0060] However, the time series is discrete, so the above formula needs to be discretized using the zero-order hold technique. The expression is as follows:
[0061]
[0062] in, is the discretization matrix; I is the identity matrix; Δ is used to convert continuous parameters A and B into discretized matrices and The learnable parameters of
[0063] Discretize the above matrix and Substituting the SSM recursive hidden state continuity equation into the SSM recursive hidden state discretization equation, the expression is as follows:
[0064]
[0065] In the formula, h t Represents the hidden state of the last step of the input sequence; h t-1 Represented as ht The hidden state of the previous step; x t is the corresponding discrete input; t For the corresponding discrete output; ls represents the step size of the input sequence, and fd represents the feature dimension of each step input;
[0066] S3.1.2, Substitute the fusion feature sequence CoattentInfo' obtained in S2.4 as input into the SSM recursive hidden state discretization equation. Due to the recursive law of the equation output, the output result can be used only C is the matrix set of three matrices expressed in the form of product of the input, The matrix set of three matrices C is represented as a sequential convolution kernel, and the expression is as follows:
[0067]
[0068] In the formula, Represents the sequence convolution kernel, ls represents the sequence step dimension; CoattentInfo' passes The output can be obtained as a time convolution sequence * y op ,
[0069] S3.2, FC 32-2 The fully connected layer converts the temporal convolution sequence obtained in S3.1 into * y op The sequence feature dimension is reduced to a time convolution sequence y with a feature dimension of 2 op ', 2 is the trajectory dimension, FC 32-2 Represents a fully connected layer with an input dimension of 32 and an output dimension of 2;
[0070] The above process expression is as follows:
[0071] y op '=FC 32-2 (y op );
[0072] S3.3, in order to dynamically adjust the step size of model prediction, the time convolution sequence y in S3.2 op 'Transpose to get (y op ') T , y op 'The sequence feature dimension is obtained by transposing (y op ') T With fully connected layer FC ts-preAdjust the current sequence step dimension ts to the prediction step dimension pre of this model, pre∈(0,+∞), and then transpose the sequence step dimension and sequence feature dimension to obtain the predicted sequence y pre ,
[0073] The above process expression is as follows:
[0074]
[0075] S4. Construct the weight reduction parameter of the predicted trajectory output window loss and train the constructed multi-image attention fusion network model to obtain the model parameters; the model training loss expression is as follows:
[0076]
[0077] Among them, prewin represents the prediction window length, and its value is consistent with the prediction step dimension pre; a i” represents the loss ratio under the i-th prediction window; τ is the smoothing degree, τ∈(0,+∞); loss i” Represents the predicted value in the "i" window and the true value y i” The error value of ; Loss is the total loss between the predicted value and the true value.
[0078] Beneficial effects of the present invention:
[0079] 1. The present invention introduces a graph attention network and a global pooling strategy to extract features from road topology graph sequences and vehicle information interaction graph sequences, fully capturing the importance of each node in the spatial dimension and the spatiotemporal dynamic changes between sequences.
[0080] 2. The present invention organically integrates the road network topological relationship and the spatiotemporal flow characteristics of vehicle behavior by using the multi-graph attention information fusion method, significantly improving the correlation and interaction intensity between the topological structure graph sequence and the vehicle interaction graph sequence.
[0081] 3. The present invention uses the Mamba neural network to effectively make up for the deficiencies in the use of spatiotemporal information in static graphs, and through a reasonable loss-reducing strategy, it optimizes the later prediction errors while enhancing the early prediction accuracy, thereby efficiently predicting future vehicle trajectories. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a flow chart of the steps of the present invention;
[0083] Figure 2 A schematic diagram of the perception range of the predicted vehicle setting of the present invention;
[0084] Figure 3 A schematic diagram of road information within the sensing range of the present invention;
[0085] Figure 4 It is a diagram of the road topology structure of the present invention;
[0086] Figure 5 A schematic diagram of vehicle information interaction within the sensing range of the present invention;
[0087] Figure 6 It is a topological diagram of information interaction between the predicted vehicle and surrounding vehicles of the present invention;
[0088] Figure 7 It is a GAT convolution pooling feature flow chart of the topological structure sequence of the traffic scene of the present invention;
[0089] Figure 8 It is a GAT convolution pooling feature flow chart of the topological graph sequence of the vehicle and surrounding vehicle information interaction of the present invention;
[0090] Fig. 9 This is a flow chart of information fusion of the multi-image feature co-attention mechanism of the present invention;
[0091] Fig.10 This is a flow chart of the Mamba prediction trajectory network output of the present invention. DETAILED DESCRIPTION
[0092] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0093] like Figure 1 As shown, a vehicle trajectory prediction method based on multi-image attention fusion includes the following steps:
[0094] S1. Constructing a time sequence diagram for predicting a topological structure diagram of a traffic scene in which a vehicle is located and a time sequence diagram for predicting a topological diagram of information interaction between the vehicle and surrounding vehicles;
[0095] The following steps are involved:
[0096] S1.1. Establish a time sequence diagram of the topological structure diagram of the traffic scene where the predicted vehicle is located. The expression is as follows:
[0097]
[0098] In the formula, G roThe total set of topological structure diagrams of the traffic scene where the predicted vehicle is located is represented by the time series diagram. This set is a summary of the topological structure diagrams of the traffic scene where the predicted vehicle is located, which are constructed based on the traffic data extracted from the open source dataset Argoverse through Python code. T represents the time length of the dataset, T∈(0,+∞); combined with the attached Figure 2 and 3 , G ro,t It represents the topological structure of the traffic scene within a 50m diameter circle with the predicted vehicle as the center at time t, and the circle is used as the sensing range; combined with the attached Figure 4 , V ro,t Represents G ro,t The node set at time t means the road information appearing in the traffic scene at this time, V ro,t The set range is (0, nt-ro), nt-ro represents the maximum number of nodes in the traffic scene topology graph at time t, specifically the maximum number of perceived roads in the traffic scene at this time, where nt-ro∈(0,+∞); E ro,t Represents G ro,t The set of node-edge relationships at time t, which means the summary of the connection relationships of the roads in the traffic scene at this time, E ro,t The set range is [0,2×nt-ro]; Represents G ro,t V at time t ro,t The feature set of the node set means the summary of the feature information of the roads in the traffic scene at this time. The collection range is (0, nt-ro); x i ro,t represents V at time t ro,t The characteristics of the node set are due to the different time G ro,t It is represented by dynamic road information, which reflects the dynamic change of information of traffic scene roads, i∈(0~nt-ro); Represents the node set V ro,t The child nodes of and V is the node set ro,t The first and second child nodes of represent two roads in the set of perceived roads in the traffic scene. If the two roads are connected or can be switched, they are included in the edge relationship set of the road connection. represents the The feature set of a node set, Indicates V ro,t One of the nodes below, It means one of the road information in the traffic scene at this time. The following contains four characteristic information of the road, which are the nodes at time t The rear drive road number of the road Indicates the number of the connecting road ahead in the direction of vehicle travel; the node at time t The road's predecessor road number Indicates the number of the connecting road behind the vehicle in the direction of travel; the node at time t Road centerline information for roads Indicates the node road at this time The coordinate information of the center line; the node at time t Road ID code Indicates the current node road Serial number;
[0099] S1.2. Establish a time sequence diagram of the topological diagram of the information interaction between the predicted vehicle and the surrounding vehicles, and its expression is as follows:
[0100]
[0101] In the formula, G vo It represents the total set of the constructed topological time-series diagrams of the information interaction between the predicted vehicle and the surrounding vehicles. This set is a summary of the topological structure diagrams of the information interaction between the predicted vehicle and the vehicles within the sensing range constructed based on the traffic data extracted from the open source dataset through Python code, where T represents the time length of the dataset, T∈(0,+∞); combined with the attached Figure 5 and 6 , G vo,t V represents the vehicle interaction information topology diagram of the traffic scene predicted by the vehicle within the perception range set by S1.1 at time t; vo,t Represents G vo,t The node set at time t means the vehicle information set within the sensing range at this time, V vo,t The set range is (0, nt-vo), nt-vo represents the maximum number of nodes in the vehicle information interaction topology graph at time t, specifically the maximum number of vehicles in the traffic scene within the perception range at this time, where nt-vo∈(0,+∞); E vo,t Represents G vo,t The set of node-edge relationships at time t, which means whether there is an interactive relationship between vehicles within the sensing range, E vo,t The collection range is [0,2×nt-vo]; G ro,t V at time t ro,t The feature set of the node set is the summary of the feature information of each vehicle within the sensing range. The collection range is (0,nt-vo); represents V at time t vo,t The characteristics of the node set are due to the different time G vo,t It is represented as dynamic vehicle interaction information, which reflects the change of interaction information between the predicted vehicle and surrounding vehicles, where k∈(1,nt-vo); and Represents the node set V vo,t The first and second child nodes of represent two vehicles in the traffic scene within the sensing range at this time. If there is an interaction relationship between the vehicles, they are included in the edge set of vehicle interaction. In this implementation case, the interaction behavior is simplified to that all vehicles within the sensing range have interaction behavior. Indicates V vo,t One of the nodes below, It means the information of one of the vehicles in the current traffic scene; represents the The feature set of the node set; The following contains 5 characteristic information of the vehicle, which are the nodes at time t Vehicle coding characteristic information of the vehicle Represents this node Vehicle number; Node Vehicle coordinate feature information Represents the coordinate information of this vehicle; node Vehicle speed characteristic information Represents this node Vehicle speed; node Vehicle acceleration characteristic information Represents this node The acceleration of the vehicle; the node Vehicle yaw angle characteristic information Represents this node The vehicle's yaw angle.
[0102] S2. Constructing a multi-image attention fusion network in a vehicle trajectory prediction method including the following steps:
[0103] S2.1, unify the feature information in S1.1 and S1.2 to enable the subsequent GAT graph convolution operation;
[0104] The characteristic information includes: number information, coordinate information and vehicle physical information;
[0105] The steps to unify the size are as follows:
[0106] S2.1.1、Combined with attachment Figure 3 With attached Figure 5The number information in S1.1 and S1.2 is encoded and its expression is as follows:
[0107]
[0108] In the formula, Represents node features The 1×32 vector obtained by encoding is Represents node features The 1×32 vector obtained by encoding is Roembeding and Voembeding represent the road number embedding layer and vehicle number embedding layer respectively;
[0109] S2.1.2, encode the coordinate information and vehicle physical information in S1.1 and S1.2, and the expression is as follows:
[0110]
[0111] In the formula, Indicates FC1 1-32 Fully connected layer pair The 1×32 vector obtained by encoding the node features, FC1 1-32 The input dimension is 1 and the output dimension is 32. F veh-coordiantes Indicates FC2 2-32 The fully connected layer is used to calculate the vehicle coordinate information The resulting encoding vector, FC2 2-32 The input dimension is 2 and the output dimension is 32, because The centerline information dimension of each road in the lower sensing range of S1.1 will change dynamically at different times. For subsequent feature processing, it is necessary to use TP function and FC3 240-32 Fully connected layer pair The dimension is normalized, where the TP function will Feature information with a unified dimension of 120 In the TP(·) function, An represents the range of the input set A, which is The dynamic range of dimension, the information in set A is arranged in order from small to large through Python code, that is, The information is arranged regularly. When the input dimension is greater than 120, the information is truncated from small to large to 120 dimensions. When the input dimension is less than 120, the last bit of the input information is padded with 0 to 120 dimensions. F center-lines Indicates FC3 120-32 Fully connected layer pair The 1×32 vector obtained by encoding, FC3 120-32 The input dimension is 120 and the output dimension is 32. It can be obtained that S1.1 and S1.2 The node feature scales are 1×128 and 1×192;
[0112] S2.2, combined with attachment Figure 7 and 8 , the graph set G obtained by combining S1.1 and S1.2 vo With G ro In the random time interval θ, ts is the sequence length, ts∈(0,+∞), and the time series graph G is taken ro / vo,θ ,G ro / vo,θ+1 ,...,G ro / vo,θ+ts As input, it is defined as a graph sequence The importance of the interaction between the predicted vehicle and the surrounding vehicles and the mutual influence of the environmental spatial information and road information within the input step are adaptively captured by the graph attention network. The steps are as follows:
[0113] The timing diagrams include: G ro The total set of topological structure diagrams and time sequence diagrams representing the traffic scene where the measured vehicle is located and G vo Represents the total set of topological timing diagrams of information interaction between the constructed test vehicle and surrounding vehicles;
[0114] S2.2.1. The expression of the graph attention convolution process of the time series graph of the traffic scene topology graph where the vehicle is located is as follows:
[0115]
[0116] In the formula, LakyReLu is the activation function; Represents node v i Update features under the k'th attention head, k'∈(0,+∞); Node v representing the k'th head i and neighbor node v j The attention coefficient between (k') Indicates that the dimension of the k'th head is F'×F weight matrix, F' is the new feature dimension of the mapping, F'∈(192,+∞), F is the original feature dimension, F∈(128,192); N(i) represents the node v i The set of neighbor nodes; W (k') h i Represents node v i Feature transformation under the k'th head; W (k') h i ||W (k') h jRepresents node v i With neighbor node v j Feature splicing of a (k ' )T represents the learnable attention weight vector of the kth attention head, with dimension 2F'; || represents concatenation, h′ i represents the node features finally updated by the graph attention network; H i Represents the feature matrix of all nodes in the final updated graph; where node v i Represents all nodes in the time series diagram of the traffic scene topology diagram where the predicted vehicle is located;
[0117] S2.2.2. The expression of the graph attention convolution process of the time series graph of the topological graph of the information interaction between the vehicle and the surrounding vehicles is as follows:
[0118]
[0119] In the formula, Represents node v k Update features under the k'th attention head, k'∈(0,+∞); Node v representing the k'th head k and neighbor node v jk The attention coefficient between (k') Indicates that the dimension of the k'th head is F'×F weight matrix, F' is the new feature dimension of the mapping, F'∈(192,+∞), F is the original feature dimension, F∈(128,192); N(k) represents the node v k The set of neighbor nodes; W (k ' ) h k Represents node v k Feature transformation under the k'th head; W (k') h k ||W (k') h jk Represents node v k With neighbor node v jk Feature splicing of a (k')T represents the learnable attention weight vector of the kth attention head, with dimension 2F'; || represents concatenation, h′ i represents the node features finally updated by the graph attention network; H k Represents the feature matrix of all nodes in the final updated graph; where node v k Represents all nodes in the time series diagram of the topological diagram of the information interaction between the predicted vehicle and surrounding vehicles;
[0120] S2.3. For input In each picture Perform a global pooling operation, where The details are as follows:
[0121] S2.3.1. The expression of the global pooling process of the time sequence diagram of the traffic scene topology diagram where the predicted vehicle is located is as follows:
[0122]
[0123] In the formula, p is in the graph sequence The index in V pi yes The set number of nodes in the graph, h pi Representation diagram The global feature representation of h pi The feature dimension is determined by the output channel number dimension set by the GAT convolution. The convolution output channel number dimension of the two images is set to Gc-d, Gc-d∈(k'×192,+∞), then
[0124] S2.3.2. The expression of the global pooling process of the time sequence diagram of the predicted vehicle and surrounding vehicle information interaction topology diagram is as follows:
[0125]
[0126] In the formula, p is in the graph sequence The index in V pk yes The set number of nodes in the graph, h pk Representation diagram The global feature representation of h pk The feature dimension is determined by the output channel number dimension set by the GAT convolution. The convolution output channel number dimension of the two images is set to Gc-d, Gc-d∈(k'×192,+∞), then
[0127] S2.3.3. Graph Sequence The global feature information of the road topology graph sequence is obtained through graph attention convolution and global pooling at each step Global feature information of topological graph sequence interacting with vehicle information
[0128] S2.4, combined with Fig. 9 As shown, the global feature information of S2.3 is fused through the co-attention network to obtain a fused feature sequence;
[0129] The global feature information of the road topology sequence is used as the query matrix The global feature information of the vehicle information interaction topology sequence is used as the key matrix With value matrix It fully captures the interaction information between the predicted vehicle space and the scene road as well as the interaction information between the surrounding vehicles, and at the same time further extracts the feature information of the two input graph sequences in the spatiotemporal dimensions.
[0130] Here are the steps:
[0131] S2.4.1. Linearly transform the query, key, and value matrices under the number of attention heads; the expression is as follows:
[0132]
[0133] Where ch is the number of shared attention heads, ch∈(0,+∞); Q ch , K ch , V ch They are the linear changes of Q, K and V matrices under each attention head; is the weight matrix of each attention head that can be learned; both are matrix;
[0134] S2.4.2. Under each attention head, through Q ch and The attention score matrix is obtained by matrix multiplication. The scaling factor is applied to reduce the elements in the attention score matrix to avoid the element values in the attention score matrix being too large. The elements of the attention score matrix are further normalized by the softmax function. The content of the attention score matrix elements represents V ch The weight of each element in the matrix is obtained by comparing the attention score matrix element with V ch Perform matrix multiplication to obtain the fusion information of each attention head.
[0135] Its expression is as follows:
[0136]
[0137] In the formula, Attention ch-weight Represents the attention score matrix; Attention ch represents the fusion information under the attention head of ch; d ck is a scaling factor to prevent the matrix dot product value from being too large; softmax represents a normalization function;
[0138] S2.4.3. Collect the fusion information of all attention heads to obtain the final fusion feature sequence * .
[0139] Its expression is as follows:
[0140] CoattentInfo=Concat(Attention 1 ,...,Attention ch )W O
[0141] In the formula, CoattentInfo represents the fusion feature sequence under multi-head co-attention * ; Concat represents the matrix concatenation process; W O Represents the learnable output matrix; CoattentInfo,
[0142] S2.4.4. Accelerate model training speed for fusion feature sequence * The feature information is used to reduce the dimension.
[0143] Its expression is as follows:
[0144] CoattentInfo'=DeFc Gc-d-32 (CoattentInfo)
[0145] In the formula, CoattentInfo' represents the DeFc Gc-d-32 The fully connected layer is a fusion feature sequence after the dimension reduction of each step feature in the CoattentInfo sequence; DeFc Gc-d-32 The input dimension is the output channel dimension Gc-d of the GAT convolution setting set in S2.2, and the output dimension is 32.
[0146] S3, based on the fusion feature sequence information obtained in S2, combined with the attached Fig.10 As shown in the figure, a multi-image attention fusion network is constructed as the prediction network in the vehicle trajectory prediction method. The specific steps are as follows:
[0147] S3.1, convolve the fusion feature sequence information obtained in S2.4 through the SSM recursive hidden state equation sequence convolution kernel under Mamba to obtain a time convolution sequence. The specific process is as follows:
[0148] S3.1.1. Map the input sequence to the output sequence by establishing the SSM recursive hidden state equation. The expression of the SSM recursive hidden state continuity equation is as follows:
[0149]
[0150] In the formula, x(t) represents continuous input; y(t) represents continuous output; h(t) is the state representation matrix; A and B are the first state matrix and the second state matrix; C is the projection matrix; A, B, and C matrices are all learnable matrices. seq represents the continuity length of x(t), and f_num represents the feature dimension of x(t);
[0151] However, the time series is discrete, so the above formula needs to be discretized using the zero-order hold technique. The expression is as follows:
[0152]
[0153] in, is the discretization matrix; I is the identity matrix; Δ is used to convert continuous parameters A and B into discretized matrices and The learnable parameters of
[0154] Discretize the above matrix and Substituting the SSM recursive hidden state continuity equation into the SSM recursive hidden state discretization equation, the expression is as follows:
[0155]
[0156] In the formula, h t Represents the hidden state of the last step of the input sequence; h t-1 Represented as h t The hidden state of the previous step, where the initial hidden state is 0; x t is the corresponding discrete input; t For the corresponding discrete output; ls represents the step size of the input sequence, and fd represents the feature dimension of each step input;
[0157] S3.1.2, Substitute the fusion feature sequence CoattentInfo' obtained in S2.4 as input into the SSM recursive hidden state discretization equation. Due to the recursive law of the equation output, the output result can be used only C is the matrix set of three matrices expressed in the form of product of the input, The matrix set of three matrices C is represented as a sequential convolution kernel, and the expression is as follows:
[0158]
[0159] In the formula, Represents the sequence convolution kernel, ls represents the sequence step dimension; CoattentInfo' passes The output can be obtained as a time convolution sequence * y op ,
[0160] S3.2, FC32-2 The fully connected layer converts the temporal convolution sequence obtained in S3.1 into * y op The sequence feature dimension is reduced to a time convolution sequence y with a feature dimension of 2 op ', 2 is the trajectory dimension, FC 32-2 Represents a fully connected layer with an input dimension of 32 and an output dimension of 2;
[0161] The above process expression is as follows:
[0162] y op '=FC 32-2 (y op );
[0163] S3.3, in order to dynamically adjust the step size of model prediction, the time convolution sequence y in S3.2 op 'Transpose to get (y op ') T , y op 'The sequence feature dimension is obtained by transposing (y op ') T With fully connected layer FC ts-pre Adjust the current sequence step dimension ts to the prediction step dimension pre of this model, pre∈(0,+∞), and then transpose the sequence step dimension and sequence feature dimension to obtain the predicted sequence y pre ,
[0164] The above process expression is as follows:
[0165]
[0166] S4. Construct the weight reduction parameter of the predicted trajectory output window loss and train the constructed multi-image attention fusion network model to obtain the model parameters; the model training loss expression is as follows:
[0167]
[0168] Among them, prewin represents the prediction window length, and its value is consistent with the prediction step dimension pre; a i” represents the loss ratio in the i-th prediction window; τ is the smoothness, τ∈(0,+∞), in this implementation case, τ is 10, and the larger the τ value, the smoother the decrease between weights; loss i” Represents the predicted value in the "i" window and the true value y i”The error value of the prediction value; Loss is the total loss between the predicted value and the true value. The model parameters can be updated using the back propagation algorithm through the Adam optimizer. After multiple rounds of optimization, the trained model can be obtained.
[0169] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A vehicle trajectory prediction method based on multi-image attention fusion, characterized by: The following steps are involved: S1. Constructing a time sequence diagram for predicting a topological structure diagram of a traffic scene in which a vehicle is located and a time sequence diagram for predicting a topological diagram of information interaction between the vehicle and surrounding vehicles; S2. Constructing a multi-image attention fusion network in the vehicle trajectory prediction method; S3, based on the fused feature sequence information obtained in S2, a multi-image attention fusion network is constructed as the prediction network in the vehicle trajectory prediction method; S4. Construct the weight reduction parameter of the predicted trajectory output window loss and train the constructed multi-image attention fusion network model to obtain the model parameters to complete the model construction.
2. The vehicle trajectory prediction method based on multi-image attention fusion according to claim 1 is characterized in that: The steps of constructing a timing diagram for predicting a topological structure diagram of a traffic scene in which a vehicle is located and a timing diagram for predicting a topological diagram of information interaction between the vehicle and surrounding vehicles are as follows: S1.
1. Establish a time sequence diagram of the topological structure diagram of the traffic scene where the predicted vehicle is located. The expression is as follows: In the formula, G ro Represents the total set of topological structure graphs and time series graphs of the traffic scene where the predicted vehicle is located; T represents the time length of the data set, T∈(0,+∞); G ro,t It represents the topological structure diagram of the traffic scene within a 50m diameter circle with the predicted vehicle as the center at the time t of the scene, and the circle is taken as the sensing range; V ro,t Represents G ro,t The node set at time t, V ro,t The set range is (0, nt-ro), nt-ro represents the maximum number of nodes in the traffic scene topology graph at time t, where nt-ro∈(0,+∞); E ro,t Represents G ro,t The set of node-edge relationships at time t, E ro,t The set range is [0,2×nt-ro]; Represents G ro,t V at time t ro,t The feature set of a node set, The collection range is (0, nt-ro); x i ro,t represents V at time t ro,t The characteristics of the node set, due to different time G ro,t It is represented by dynamic road information, which reflects the dynamic change of information of traffic scene roads, i∈(0~nt-ro); Represents the node set V ro,t The child nodes of and V is the node set ro,t The first and second child nodes of represent two roads in the set of perceived roads in the traffic scene. If the two roads are connected or can be switched, they are included in the edge relationship set of the road connection. represents the The feature set of a node set, Indicates V ro,t One of the nodes below, It means one of the road information in the traffic scene at this time. The following contains four characteristic information of the road, which are the nodes at time t The rear drive road number of the road Indicates the number of the connecting road ahead in the direction of vehicle travel; the node at time t The road's predecessor road number Indicates the number of the connecting road behind the vehicle in the direction of travel; the node at time t Road centerline information for roads Indicates the node road at this time The coordinate information of the center line; the node at time t Road ID code Indicates the current node road Serial number; S1.
2. Establish a time sequence diagram of the topological diagram of the information interaction between the predicted vehicle and the surrounding vehicles, and its expression is as follows: In the formula, G vo represents the total set of topological time-series diagrams of information interaction between the constructed predicted vehicle and surrounding vehicles; T represents the time length of the data set, T∈(0,+∞); G vo,t V represents the vehicle interaction information topology diagram of the traffic scene predicted by the vehicle within the perception range set by S1.1 at time t; vo,t Represents G vo,t The node set at time t, V vo,t The set range is (0, nt-vo), nt-vo represents the maximum number of nodes in the vehicle information interaction topology graph at time t, where nt-vo∈(0,+∞); E vo,t Represents G vo,t The set of node-edge relationships at time t, E vo,t The set range is [0,2×nt-vo]; G ro,t V at time t ro,t The feature set of a node set, The collection range is (0,nt-vo); represents V at time t vo,t The characteristics of the node set, due to different time G vo,t It is represented as dynamic vehicle interaction information, which reflects the change of interaction information between the predicted vehicle and surrounding vehicles, where k∈(1,nt-vo); and Represents the node set V vo,t The first and second child nodes of , if there is an interaction relationship between vehicles, are included in the edge set of vehicle interaction; Indicates V vo,t One of the nodes below, It means the information of one of the vehicles in the current traffic scene; represents the The feature set of the node set; The following contains 5 characteristic information of the vehicle, which are the nodes at time t Vehicle coding characteristic information of the vehicle Represents this node Vehicle number; Node Vehicle coordinate feature information Represents the coordinate information of this vehicle; node Vehicle speed characteristic information Represents this node Vehicle speed; node Vehicle acceleration characteristic information Represents this node The acceleration of the vehicle; the node Vehicle yaw angle characteristic information Represents this node The vehicle's yaw angle.
3. The vehicle trajectory prediction method based on multi-image attention fusion according to claim 1 is characterized in that: The steps of constructing a multi-graph attention fusion network in the vehicle trajectory prediction method of the multi-graph attention fusion network are as follows: S2.1, unify the feature information in S1.1 and S1.2; The characteristic information includes: number information, coordinate information and vehicle physical information; The steps to unify the size are as follows: S2.1.
1. Encode the numbering information in S1.1 and S1.
2. The expression is as follows: In the formula, Represents node features The 1×32 vector obtained by encoding is Represents node features The 1×32 vector obtained by encoding is Roembeding and Voembeding represent the road number embedding layer and vehicle number embedding layer respectively; S2.1.2, encode the coordinate information and vehicle physical information in S1.1 and S1.2, and the expression is as follows: In the formula, F fi Indicates FC1 1-32 Fully connected layer pair The 1×32 vector obtained by encoding the node features, FC1 1-32 The input dimension is 1 and the output dimension is 32. F veh-coordiantes Indicates FC2 2-32 The fully connected layer is used to calculate the vehicle coordinate information The resulting encoding vector, FC2 2-32 The input dimension is 2 and the output dimension is 32, because The centerline information dimension of each road in the lower sensing range of S1.1 will change dynamically at different times. For subsequent feature processing, it is necessary to use TP function and FC3 240-32 Fully connected layer pair The dimension is normalized, where the TP function will Feature information with a unified dimension of 120 In the TP(·) function, An represents the range of the input set A, which is The dynamic range of dimension, the information in set A is arranged in order from small to large through Python code, that is, The information is arranged regularly. When the input dimension is greater than 120, the information is truncated from small to large to 120 dimensions. When the input dimension is less than 120, the last bit of the input information is padded with 0 to 120 dimensions. F center-lines Indicates FC3 120-32 Fully connected layer pair The 1×32 vector obtained by encoding, FC3 120-32 The input dimension is 120 and the output dimension is 32. S2.2, the graph set G obtained by combining S1.1 and S1.2 vo With G ro In the random time interval θ, ts is the sequence length, ts∈(0,+∞), and the time series graph G is taken ro / vo,θ ,G ro / vo,θ+1 ,...,G ro / vo,θ+ts As input, it is defined as a graph sequence The importance of the interaction between the predicted vehicle and the surrounding vehicles and the mutual influence of the environmental spatial information and road information within the input step are adaptively captured by the graph attention network. The steps are as follows: The timing diagrams include: G ro The total set of topological structure diagrams and time sequence diagrams representing the traffic scene where the measured vehicle is located and G vo Represents the total set of topological timing diagrams of information interaction between the constructed test vehicle and surrounding vehicles; S2.2.
1. The expression of the graph attention convolution process of the time series graph of the traffic scene topology graph where the vehicle is located is as follows: In the formula, LakyReLu is the activation function; Represents node v i Update features under the k'th attention head, k'∈(0,+∞); Node v representing the k'th head i and neighbor node v j The attention coefficient between (k') Indicates that the dimension of the k'th head is F'×F weight matrix, F' is the new feature dimension of the mapping, F'∈(192,+∞), F is the original feature dimension, F∈(128,192); N(i) represents the node v i The set of neighbor nodes of (k') h i Represents node v i Feature transformation under the k'th head; W (k') h i ||W (k') h j Represents node v i With neighbor node v j Feature splicing of a (k')T represents the learnable attention weight vector of the kth attention head, with dimension 2F'; ‖ represents concatenation, h i ' represents the node features finally updated by the graph attention network; H i Represents the feature matrix of all nodes in the final updated graph; where node v i Represents all nodes in the time series diagram of the traffic scene topology diagram where the predicted vehicle is located; S2.2.
2. The expression of the graph attention convolution process of the time series graph of the topological graph of the information interaction between the vehicle and the surrounding vehicles is as follows: In the formula, Represents node v k Update features under the k'th attention head, k'∈(0,+∞); Node v representing the k'th head k and neighbor node v jk The attention coefficient between (k') Indicates that the dimension of the k'th head is F'×F weight matrix, F' is the new feature dimension of the mapping, F'∈(192,+∞), F is the original feature dimension, F∈(128,192); N(k) represents the node v k The set of neighbor nodes of (k ' ) h k Represents node v k Feature transformation under the k'th head; W (k') h k ||W (k') h jk Represents node v k With neighbor node v jk Feature splicing of a (k')T represents the learnable attention weight vector of the kth attention head, with dimension 2F'; ‖ represents concatenation, h i ' represents the node features finally updated by the graph attention network; H k Represents the feature matrix of all nodes in the final updated graph; where node v k Represents all nodes in the time series diagram of the topological diagram of the information interaction between the predicted vehicle and surrounding vehicles; S2.
3. For input In each picture Perform a global pooling operation, where The details are as follows: S2.3.
1. The expression of the global pooling process of the time sequence diagram of the traffic scene topology diagram where the predicted vehicle is located is as follows: In the formula, p is in the graph sequence The index in V pi yes The set number of nodes in the graph, h pi Representation diagram The global feature representation of h pi The feature dimension is determined by the output channel number dimension set by the GAT convolution. The convolution output channel number dimension of the two images is set to Gc-d, Gc-d∈(k'×192,+∞), then S2.3.
2. The expression of the global pooling process of the time sequence diagram of the predicted vehicle and surrounding vehicle information interaction topology diagram is as follows: In the formula, p is in the graph sequence The index in V pk yes The set number of nodes in the graph, h pk Representation diagram The global feature representation of h pk The feature dimension is determined by the output channel number dimension of the GAT convolution setting. The convolution output channel number dimension of the two images is set to Gc-d, Gc-d∈(k'×192,+∞), then h pk S2.3.
3. Graph Sequence The global feature information of the road topology graph sequence is obtained through graph attention convolution and global pooling at each step Global feature information of topological graph sequence interacting with vehicle information S2.4, the global feature information of S2.3 is fused through the co-attention network to obtain a fused feature sequence; The global feature information of the road topology sequence is used as the query matrix The global feature information of the vehicle information interaction topology sequence is used as the key matrix With value matrix Fully capture the interaction information between the predicted vehicle space and the scene road and the surrounding vehicles, and further extract the feature information of the two input image sequences in the spatiotemporal dimension; Here are the steps: S2.4.
1. Linearly transform the query, key, and value matrices under the number of attention heads; the expression is as follows: Where ch is the number of shared attention heads, ch∈(0,+∞); Q ch , K ch , V ch They are the linear changes of Q, K and V matrices under each attention head; is the weight matrix of each attention head that can be learned; both are matrix; S2.4.
2. Under each attention head, through Q ch and The attention score matrix is obtained by matrix multiplication, and the scaling factor is applied to reduce the elements in the attention score matrix; the elements of the attention score matrix are further normalized by the softmax function, and the content of the attention score matrix elements represents V ch The weight of each element in the matrix is obtained by comparing the attention score matrix element with V ch Perform matrix multiplication to obtain the fusion information of each attention head; Its expression is as follows: In the formula, Attention ch-weight Represents the attention score matrix; Attention ch represents the fusion information under the attention head of ch; d ck is a scaling factor to prevent the matrix dot product value from being too large; softmax represents a normalization function; S2.4.
3. Collect the fusion information of all attention heads to obtain the final fusion feature sequence * ; Its expression is as follows: CoattentInfo=Concat(Attention1,...,Attention ch )W O In the formula, CoattentInfo represents the fusion feature sequence under multi-head co-attention * ; Concat represents the matrix concatenation process; W O Represents the learnable output matrix; CoattentInfo, S2.4.
4. Accelerate model training speed for fusion feature sequence * Reduce the dimension of the feature information; Its expression is as follows: CoattentInfo'=DeFc Gc-d-32 (CoattentInfo) In the formula, CoattentInfo' represents the DeFc Gc-d-32 The fully connected layer is a fusion feature sequence after the dimension reduction of each step feature in the CoattentInfo sequence; DeFc Gc-d-32 The input dimension is the output channel dimension Gc-d of the GAT convolution setting set in S2.2, and the output dimension is 32.
4. The vehicle trajectory prediction method based on multi-image attention fusion according to claim 1 is characterized in that: The fused feature sequence information obtained by S2 is used to construct a multi-image attention fusion network as the prediction network in the vehicle trajectory prediction method. The specific steps are as follows: S3.1, convolve the fusion feature sequence information obtained in S2.4 through the SSM recursive hidden state equation sequence convolution kernel under Mamba to obtain a time convolution sequence. The specific process is as follows: S3.1.1, map the input sequence to the output sequence by establishing the SSM recursive hidden state equation; The expression of the SSM recursive hidden state continuity equation is as follows: In the formula, x(t) represents continuous input; y(t) represents continuous output; h(t) is the state representation matrix; A and B are the first state matrix and the second state matrix; C is the projection matrix; A, B, and C matrices are all learnable matrices. seq represents the continuity length of x(t), and f_num represents the feature dimension of x(t); However, the time series is discrete, so the above formula needs to be discretized using the zero-order hold technique. The expression is as follows: in, is the discretization matrix; I is the identity matrix; Δ is used to convert continuous parameters A and B into discretized matrices and The learnable parameters of Discretize the above matrix and Substituting the SSM recursive hidden state continuity equation into the SSM recursive hidden state discretization equation, the expression is as follows: In the formula, h t Represents the hidden state of the last step of the input sequence; h t-1 Represented as h t The hidden state of the previous step, where the initial hidden state is 0; x t is the corresponding discrete input; t For the corresponding discrete output; ls represents the step size of the input sequence, and fd represents the feature dimension of each step input; S3.1.2, Substitute the fusion feature sequence CoattentInfo' obtained in S2.4 as input into the SSM recursive hidden state discretization equation. Due to the recursive law of the equation output, the output result can be used only C is the matrix set of three matrices expressed in the form of product of the input, The matrix set of three matrices C is represented as a sequential convolution kernel, and the expression is as follows: In the formula, Represents the sequence convolution kernel, ls represents the sequence step dimension; CoattentInfo' passes The output can be obtained as a time convolution sequence * y op , S3.2, FC 32-2 The fully connected layer converts the temporal convolution sequence obtained in S3.1 into * y op The sequence feature dimension is reduced to a time convolution sequence y with a feature dimension of 2 op ', 2 is the trajectory dimension, FC 32-2 Represents a fully connected layer with an input dimension of 32 and an output dimension of 2; The above process expression is as follows: y op '=FC 32-2 (y op ); S3.3, in order to dynamically adjust the step size of model prediction, the time convolution sequence y in S3.2 op 'Transpose to get (y op ') T , y op 'The sequence feature dimension is obtained by transposing (y op ') T With fully connected layer FC ts-pre Adjust the current sequence step dimension ts to the prediction step dimension pre of this model, pre∈(0,+∞), and then transpose the sequence step dimension and sequence feature dimension to obtain the predicted sequence y pre , The above process expression is as follows:
5. A vehicle trajectory prediction method based on multi-image attention fusion according to claim 1, characterized in that: The weight reduction parameter of the predicted trajectory output window loss is constructed and the constructed multi-image attention fusion network model is trained to obtain the model parameters, wherein the model training loss expression is as follows: Among them, prewin represents the prediction window length, and its value is consistent with the prediction step dimension pre; a i” represents the loss ratio under the i-th prediction window; τ is the smoothing degree, τ∈(0,+∞); loss i” Represents the predicted value in the "i" window and the true value y i” The error value of ; Loss is the total loss between the predicted value and the true value.
Citation Information
Cited By
PyraMama-based multi-modal trajectory prediction method
CN121278361A
Automatic driving decision-making method and system based on multi-source perception
CN121316910A
Vehicle track reconstruction method and system based on millimeter wave radar
CN121682446A