Vehicle-around intention prediction method based on multi-modal space-time fusion and related equipment
Through the multimodal space-time fusion of weekly vehicles, the timing and spatial characteristics of the vehicle are extracted using the LSTM and GraphSAGE networks, and combined with the graph attention network and gated fusion mechanism, the problem that the autonomous driving system is difficult to accurately predict the intentions of surrounding vehicles in the highway ramp fusion scenario is solved, and the accuracy and safety of decision planning are improved.
Patent Information
- Application Number
- CN202510424629.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In the highway ramp junction scenario, it is difficult for the autonomous driving system to accurately predict the driving intentions of surrounding vehicles, resulting in the safety and fluency during the incoming process.
The weekly vehicle intention prediction method based on multimodal space-time fusion is adopted, and the timing and spatial characteristics of the vehicle are extracted using long-term and short-term memory networks and GraphSAGE networks, and combined with the graph attention network and gated fusion mechanism, the driving intention of the conflicting vehicle is predicted.
It improves the decision-making and planning capabilities of the autonomous driving system in complex scenarios, ensuring the safety and fluency of the inclusion process.
Smart Images

Figure CN120472656A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving technology, and in particular to a surrounding vehicle intention prediction method based on multimodal spatiotemporal fusion and related equipment. Background Art
[0002] On highways, ramp merging is a highly dynamic and uncertain traffic scenario. Ramps typically bring vehicles from branch lines or urban roads into the main lane of the highway. Vehicles must match speed and merge with the main lane traffic within a shorter acceleration lane. Figure 1 A typical highway ramp merge diagram is shown. The scene can be deconstructed into four functional segments:
[0003] 1) Ramp Connection (80 meters): This section is where the vehicle enters the ramp from a feeder road or main road. The vehicle's speed is relatively low during this section, and its primary task is to assess the main road traffic flow and surrounding vehicle behavior to prepare for the subsequent merge.
[0004] 2) Acceleration and merging section (200 meters): This section, located midway along the ramp, is crucial for the driver's vehicle to accelerate to match the speed of the main road traffic. During this phase, the driver's vehicle gradually increases speed to safely merge into the main road. During the acceleration phase, the driver must constantly monitor the speed and position of surrounding vehicles to ensure they find the right moment to merge.
[0005] 3) Merging Section (60 meters): The merging section is where the vehicle actually meets the main road traffic flow. The vehicle begins to merge with other vehicles on the main road. At this point, the vehicle must not only pay attention to its relative position to the surrounding vehicles on the main road, but also maintain a safe distance from them during the merging process.
[0006] 4) Main Road Access: After merging, the vehicle enters the subsequent lane of the main road. The goal during this phase is to maintain a safe distance from other vehicles on the main road, avoiding collisions and congestion, while merging smoothly into the main road flow.
[0007] In the high-speed ramp merging scenario, the interaction between different types of vehicles directly affects the safety and smoothness of the merging process. Based on the position and function of the vehicle during the merging process, this application mainly divides the merging vehicles into the following three categories: Figure 2 As shown:
[0008] 1) Autonomous driving vehicle on the ramp (main vehicle)
[0009] As the active driver of the lane-merging maneuver, the autonomous vehicle on the ramp has multiple tasks. First, it needs to perceive the dynamic information of the ramp segment in real time, including vehicle position, speed, acceleration, and road geometry. It also uses sensor data and vehicle-to-vehicle information to understand the dynamic behavior of vehicles in the main lane.
[0010] The main vehicle determines the driving intentions of the surrounding conflicting vehicles. Based on the predicted intentions of the surrounding vehicles, the main vehicle needs to make a game decision with the conflicting vehicle in the target lane of the main road. The so-called game decision refers to the selection of an optimal lane-merging strategy that can avoid collisions and ensure smooth traffic flow through interaction and collaboration when there is a potential risk of conflict. This process usually involves multi-objective optimization, including balancing factors such as driving safety, traffic efficiency, and ride comfort. Through the decision-making and planning module, the autonomous driving system comprehensively considers various sensor information, prediction results, and vehicle dynamics constraints to calculate the appropriate acceleration and steering wheel angle to achieve future trajectory planning and control, thereby smoothly completing the lane-merging operation from the ramp to the main lane.
[0011] 2) Conflicting vehicles in the target lane of the main road
[0012] Conflicting vehicles in the main road's target lane are those that present a potential risk of interaction and conflict with the main vehicle on the ramp during the merging process. These vehicles typically maintain a predetermined speed in the main lane, but when faced with an oncoming ramp vehicle, their behavior may change, such as actively slowing down, yielding prematurely, or failing to adjust their speed in time, resulting in a crossover trajectory. The presence of conflicting vehicles increases the uncertainty of the merging process, and fluctuations in their motion state (such as speed, acceleration, and intended trajectory) are key considerations for autonomous driving systems during their decision-making.
[0013] 3) Vehicles traveling in the inner lane of the main road
[0014] Vehicles in the inner lanes of the main road are typically located farther from the merging area, and their driving status and dynamics have less direct impact on merging vehicles. For model simplicity and focus, this application does not consider the dynamic information of vehicles in the inner lanes of the main road when designing the surrounding vehicle driving intention prediction model. Although these vehicles may affect the overall balance of traffic flow on a larger scale, their direct interference in merging decisions is relatively limited. Therefore, in this application, they are treated as background information and not as a primary decision variable.
[0015] In the highway ramp merging scenario, the main vehicle not only needs to accurately perceive its own status and road environment, but also must make timely and accurate judgments on the potential behavior of other vehicles on the main road lane to ensure the safety and efficiency of the lane change operation. Figure 2The schematic diagram shown shows a typical ramp-merging situation, in which the host vehicle is attempting to merge from the ramp into the target lane of the main road and faces the potential interaction impact of at least two conflicting vehicles (conflicting vehicle A and conflicting vehicle B) on the main road. Figure 2 Several possible vehicle decision-making methods are marked in the figure. For example, the main vehicle can choose to "merge directly" or "wait to avoid", while the main road vehicle may "keep straight" or "change lanes to avoid".
[0016] from Figure 2 As can be seen, conflicting vehicle A may change lanes to leave sufficient space for the lead vehicle to merge, or it may choose to continue straight and occupy the target lane; conflicting vehicle B may maintain speed in its own lane or actively slow down to yield. Meanwhile, the lead vehicle on the ramp can choose to merge directly or slow down and wait based on traffic flow conditions and risk assessment results. Clearly, the combination of these vehicle decisions is extremely diverse. If the autonomous driving system cannot promptly predict the true intentions of the surrounding vehicles, it can easily lead to collision risks or cause traffic flow instability. Only by accurately identifying these intentions can the autonomous driving system make targeted game decisions and trajectory planning. Currently, there is a lack of an accurate and effective identification solution. Summary of the Invention
[0017] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a vehicle intention prediction method based on multimodal spatiotemporal fusion and related equipment.
[0018] The first technical solution adopted by the present invention is:
[0019] A method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion includes the following steps:
[0020] Get historical trajectory data;
[0021] Input the acquired historical trajectory data into the trained vehicle intention prediction model and output the prediction results;
[0022] The vehicle intention prediction model includes two parallel feature extraction branches: one is the temporal feature extraction branch, which uses a long short-term memory network to encode the historical trajectory of each vehicle and dynamically captures the temporal characteristics of the trajectory through an attention mechanism; the other is the spatial feature extraction branch, which uses a GraphSAGE network to model the spatial interaction relationship between multiple vehicles. Then, combined with the attention mechanism of the graph attention network, adaptive weights are assigned to different vehicles, enabling the model to highlight the importance of key interfering vehicles.
[0023] The spatiotemporal features extracted by the two branches are input into the fusion module of the model. The fusion module of the model uses a gated fusion mechanism to adaptively weight the features from the temporal feature extraction branch and the spatial feature extraction branch, and further improves the stability of network training through batch normalization. The fused feature information is finally input into the decoder module, and the Softmax function is used to output the probability distribution of various driving intentions of the conflicting vehicles in the future.
[0024] Furthermore, the temporal feature extraction branch includes an LSTM feature extraction module, a multi-head attention module and a Dropout module;
[0025] In the LSTM feature extraction module, an independent LSTM encoder is constructed for each vehicle trajectory. The input of each LSTM encoder is one-dimensional time series trajectory data. After LSTM encoding, the motion pattern feature h of the vehicle trajectory is obtained. i ;
[0026] In the multi-head attention module, the motion pattern feature h of each vehicle is i As the attention query vector (Query), the LSTM output features of all vehicles are used as the key (Key) and value (Value) to calculate the contextual features c of the temporal interaction between vehicles i ;
[0027] The features processed by the multi-head attention module are passed through the Dropout module to prevent network overfitting and improve the robustness of the temporal features.
[0028] Furthermore, in the spatial feature extraction branch, a graph structure is used to express the interaction relationship between vehicles. Vehicles are regarded as nodes, and the interaction relationship between vehicles is regarded as the edges connecting these nodes, that is, G = (V, E), where V = {n1, n2, ..., n i} is a collection of vehicle nodes, is the set of edges that represent the relationships between vehicles;
[0029] Each node n in the node set i Corresponding to a feature vector X i , the feature vector includes the relative position, velocity, and acceleration of the vehicle at the current moment; the initial feature matrix H0 of the node is calculated:
[0030]
[0031] Furthermore, the GraphSAGE network works as follows:
[0032] For node v i , its neighbor node set N(v i ) means all the iConnected nodes, build a K-layer GraphSAGE network on each node, and calculate the features of k-layer node v When , firstly, the neighbor node features of k-1 layer sampling are aggregated by the aggregation function Perform aggregation and obtain aggregation results Then perform vector concatenation operation on it and the feature of node v at the k-1 layer, and finally generate the feature matrix of node v at the k layer through nonlinear mapping
[0033] Furthermore, the graph attention network works as follows:
[0034] The adjacency matrix is constructed based on the relative relationship between the nodes of the spatiotemporal trajectory, and its elements are recorded as Represents node v j (t) with node v j (τ) Whether there is an interactive relationship; when any of the following conditions is met, let Otherwise set it to 0:
[0035] 1) Spatial proximity at the same time: If the Euclidean distance between two different nodes at the same time is less than the preset threshold d0, then the two objects are considered to be close to each other at the same time point and there is a possibility of direct interaction;
[0036] 2) Continuous trajectories of the same object at adjacent moments: If the node index i is the same and the timestamps differ by 1, then the two nodes are considered to be the extension of the trajectory of the same physical object at consecutive time steps, and the node interaction relationship is obtained. In this way, the connection relationship between the nodes will change with time and the relationship between the nodes, and the edge adjacency matrix A is obtained accordingly;
[0037] The node feature matrix H and the edge adjacency matrix A are used as the input of the graph attention network. The attention coefficient between nodes is calculated through the feature matrix H and the adjacency matrix A. The attention coefficient of each pair of neighbor nodes is normalized using the softmax function after being processed by the LeakyReLU activation function. The normalized attention coefficient matrix is used to perform weighted summation on the node feature matrix H to obtain the updated feature representation of the node, and finally the node feature matrix processed by the attention mechanism is obtained.
[0038] Furthermore, the input of the fusion module includes: the feature vector H extracted by the temporal feature extraction branch Temporal , the feature vector H extracted by the spatial feature extraction branch Spatial , and the feature vector H generated by the residual network Residual ;
[0039] Concatenate the three feature vectors to obtain the original feature representation before fusion
[0040]
[0041] The gating factor is obtained through the fully connected layer and the Softmax activation function:
[0042]
[0043] Where W g 、b g are the weights and bias parameters to be learned in the network, and the Softmax operation ensures The obtained gating factor is used to perform weighted fusion on the three-way features to obtain the fused features:
[0044]
[0045] The fused features are then normalized through the batch normalization layer:
[0046]
[0047] Where E[·] represents the mean of the feature, Var[·] represents the variance of the feature, ε is a decimal to prevent division by zero, and γ and β are learnable scaling and translation parameters.
[0048] Furthermore, the decoder module is composed of fully connected layers, which gradually abstract the high-dimensional feature vector into low-dimensional feature expressions corresponding to the number of driving intention categories through layer-by-layer linear mapping and nonlinear activation function processing; the calculation method of each fully connected layer is:
[0049] h l+1 =ReLU(W l h l +b l )
[0050] After being processed by multiple fully connected layers, the feature dimension is gradually reduced and further abstracted, and finally a vector with the same dimension as the number of intent categories is output:
[0051] z=[z LC ,z Dec ,z Keep ,z Acc ]
[0052] Each dimension corresponds to the probability of the possible driving intention of the vehicle: z LC Indicates lane change to avoid, z Dec Indicates slowing down and going straight, z Keep Indicates uniform straight line movement, z Acc Indicates acceleration and going straight;
[0053] The Softmax function is used to normalize the above output:
[0054]
[0055] Where, P i represents the probability that the conflicting vehicle may adopt the i-th driving intention in the future, z i is the original output value of the network for the i-th intent category.
[0056] Furthermore, the historical trajectory data includes the vehicle positions, speeds, and accelerations of the primary vehicle and the conflicting vehicle;
[0057] The obtaining of historical trajectory data includes:
[0058] Select the vehicle history trajectory of the preset time length in the past as the original input data X of the model i :
[0059] X i ={(x i,t ,y i,t ,v i,t ,a i,t )}
[0060] Where x i,t ,y i,t They represent the horizontal and vertical coordinate positions of vehicle i at time t, v i,t and a i,t are the speed and acceleration of the vehicle at time t, respectively.
[0061] The second technical solution adopted by the present invention is:
[0062] An electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement a method for predicting vehicle intention based on multimodal spatiotemporal fusion as described above.
[0063] The third technical solution adopted by the present invention is:
[0064] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a method for predicting surrounding vehicle intentions based on multimodal spatiotemporal fusion as described above.
[0065] The fourth technical solution adopted by the present invention is:
[0066] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method for predicting surrounding vehicle intentions based on multimodal spatiotemporal fusion.
[0067] The beneficial effects of the present invention are as follows: the present invention fully integrates spatiotemporal information through the close coordination of the temporal feature extraction branch and the spatial feature extraction branch, can accurately predict the driving intentions of conflicting vehicles, and provides an efficient and robust solution for the research of autonomous driving technology in multi-vehicle interaction scenarios, and provides solid data and theoretical support for the decision-making planning of autonomous driving systems in complex scenarios such as merging into highway ramps. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present invention or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0069] Figure 1 This is a schematic diagram of the highway ramp merging;
[0070] Figure 2 It is a schematic diagram of the position and function of each vehicle when the vehicle merges into the highway ramp;
[0071] Figure 3 This is an overall framework diagram of the vehicle intention prediction model in an embodiment of the present invention;
[0072] Figure 4 Schematic diagram of a temporal feature extraction branch in an embodiment of the present invention;
[0073] Figure 5 Schematic diagram of a spatial feature extraction branch in an embodiment of the present invention;
[0074] Figure 6 is a schematic diagram of a fusion module in an embodiment of the present invention;
[0075] Figure 7 is a schematic diagram of a decoder module in an embodiment of the present invention;
[0076] Figure 8 This is a radar chart comparison of the performance of each model in the embodiment of the present invention;
[0077] Figure 92 is a schematic diagram of comparison of confusion matrices in an embodiment of the present invention;
[0078] Figure 10 1 is a comparison diagram of the F1 score box plot in an embodiment of the present invention;
[0079] Figure 11 It is a schematic diagram of the overall pseudo code of the weekly vehicle intention prediction model in an embodiment of the present invention. DETAILED DESCRIPTION
[0080] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application. For the step numbers in the following embodiments, they are provided only for the convenience of explanation and are not intended to limit the order of the steps. The order of execution of the steps in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0081] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise clearly defined, words such as setting, installing, and connecting should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0082] In the description of this application, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on this application.
[0083] In the description of this application, "several" means one or more, "many" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The terms "first" and "second" are used solely to distinguish technical features and are not to be construed as indicating or implying relative importance, or as implicitly specifying the number or order of the technical features indicated.
[0084] In the description of this application, "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0085] Explanation of terms:
[0086] GraphSAGE: Graph Sample and Aggregated, a graph neural network model for graph node embedding learning;
[0087] Example 1
[0088] like Figure 3 As shown, this embodiment provides a method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion, comprising the following steps:
[0089] S1. Obtain historical trajectory data;
[0090] S2. Input the obtained historical trajectory data into the trained vehicle intention prediction model and output the prediction results;
[0091] Among them, the weekly vehicle intention prediction model includes two parallel feature extraction branches: one is the temporal feature extraction branch, which uses a long short-term memory network to encode the historical trajectory of each vehicle and dynamically captures the temporal characteristics of the trajectory through an attention mechanism; the other is the spatial feature extraction branch, which uses a GraphSAGE network to model the spatial interaction relationship between multiple vehicles, and then combines the attention mechanism of the graph attention network to assign adaptive weights to different vehicles, so that the model can highlight the importance of key interfering vehicles; the fusion module of the model uses a gated fusion mechanism to adaptively weight the features from the temporal feature extraction branch and the spatial feature extraction branch; the fused feature information is finally input into the decoder module, which outputs the probability distribution of various driving intentions of the conflicting vehicles in the future.
[0092] This embodiment proposes a driving intention prediction network (i.e., a surrounding vehicle intention prediction model) that combines temporal features and spatial interaction information to effectively identify the potential driving intentions of surrounding vehicles in the highway ramp merging scenario. Figure 3The overall framework of the proposed network is presented. It consists of two core branches: a temporal extraction branch and a spatial extraction branch. In the temporal extraction branch, a long short-term memory (LSTM) network combined with an attention mechanism not only captures the dynamic changes in vehicle trajectories but also uses self-attention to weight the importance of each time step, highlighting features at key moments. This design effectively addresses the issues of information redundancy and gradients in long-sequence data. Simultaneously, the spatial extraction branch utilizes GraphSAGE and a graph attention network (GAT) to construct a vehicle-to-vehicle interaction graph. Through neighborhood sampling and adaptive weight adjustment, it accurately depicts the relative positions, speed differences, and lane relationships between vehicles, reflecting the potential interference that conflicting vehicles may cause to the primary vehicle. Subsequently, the model employs a gated fusion mechanism to adaptively integrate temporal and spatial features. This is supplemented by a residual network and batch normalization to ensure stable propagation of the fused features within the deep network. Finally, after a fully connected layer and softmax activation, the model outputs a probability distribution for each driving intention.
[0093] The following is a detailed explanation of the vehicle intention prediction model of this embodiment with reference to the accompanying drawings and specific implementation methods.
[0094] (1) Model input
[0095] The input module is the front end of the network, responsible for data collection and processing. This study targets a highway ramp merge scenario, with historical trajectory data of the vehicle and surrounding vehicles serving as the network input. Considering that vehicle behavior over a longer historical window may reveal more interactive information, this network selects the past four seconds of vehicle trajectories as raw input data.
[0096] Each vehicle trajectory data includes its position (horizontal and vertical coordinates), speed, acceleration and heading angle in the road coordinate system. Assuming the sampling frequency is 10Hz, each vehicle generates 40 moments of state data within 4 seconds. These trajectory data form a time series input. Assuming that there are N vehicles in the scene (including the main vehicle and surrounding conflicting vehicles), the trajectory sequence of each vehicle in the historical time window is defined as formula (3-1), where x i,t ,y i,t They represent the horizontal and vertical coordinate positions of vehicle i at time t, v i,t and a i,t are the speed and acceleration of the vehicle at time t, respectively.
[0097] X i ={(x i,t ,y i,t ,v i,t ,a i,t )},t∈[-4s,0],i=1,2,...,N (1)
[0098] (2) Temporal feature extraction branch
[0099] In autonomous driving scenarios, accurately identifying the intentions of surrounding vehicles often requires reasoning based on their historical trajectories over a period of time. Therefore, effectively extracting temporal features from trajectory data is crucial for achieving high-precision predictions. While traditional long short-term memory (LSTM) networks offer advantages in modeling temporal dependencies, they still have several shortcomings when dealing with long-term dependencies:
[0100] First, LSTMs are limited in their ability to emphasize key information. Although they employ gating mechanisms to control the flow of information, when time series are long, the model often struggles to effectively capture key features at specific moments, relying more heavily on input from more recent moments. Second, LSTMs still lack the ability to memorize extremely long trajectory data. While memory cells extend information storage time, implicit long-term dependencies can still be easily diluted or lost in extremely long time series, hindering the model's grasp of global information. Finally, LSTMs struggle to maintain global attention across the entire time series. They recursively propagate hidden states, relying primarily on information from the previous time step for updates and lacking the ability to comprehensively process the global historical trajectory. Consequently, when the vehicle's trajectory spans a long period, important clues may be overlooked in repeated local updates.
[0101] To address the above issues, this embodiment introduces a multi-head attention mechanism (MHA) based on the LSTM structure, allowing the model to dynamically adjust its focus on the input time series, enabling it to capture global information over a long time span, highlight key information, and improve the accuracy of driving intention recognition. Figure 4 As shown in the figure, this branch mainly consists of three modules: LSTM feature extraction module, multi-head attention module and Dropout module.
[0102] 2.1) LSTM feature extraction module
[0103] This module constructs an independent LSTM encoder for each vehicle i trajectory. The input of each LSTM encoder at time t is After LSTM encoding, the network obtains the hidden state vector of the vehicle trajectory. These hidden states represent the motion pattern characteristics of the vehicle trajectory within the historical window. The LSTM hidden layer dimension is set to 128. The key calculation process of LSTM encoding includes the recursive calculation of the forget gate, input gate, candidate state, unit state update, and output gate. The specific formula is shown in (2):
[0104]
[0105] Where ht-1 ,C t-1 are the hidden state and unit state at the previous moment respectively; f t ,i t ,o t Represent the gate values of the forget gate, input gate, and output gate respectively; σ is the sigmoid function; ⊙ represents the element-wise product; W f ,W i ,W C ,W o and b f ,b i ,b C ,b o is the parameter matrix and bias vector that can be learned during LSTM network training. For each vehicle trajectory sequence, through the above recursive operation, the hidden state sequence is finally obtained That is the output of the LSTM temporal encoding module.
[0106] 2.2) Multi-head Attention Module
[0107] Although the LSTM module can effectively capture the temporal features of a single vehicle, it cannot explicitly model the interaction between vehicles. To further enhance the contribution differences of different historical moments in the vehicle trajectory, a multi-head attention layer based on the self-attention mechanism is added after the LSTM encoding. For each target vehicle’s final hidden feature h i As the attention query vector (Query); the LSTM output features h of all other vehicles j Together they serve as the key vector (Key) and the value vector (Value). By calculating the similarity between the target vehicle query vector Q and the key vectors K of all other vehicles, the model can determine the attention heads of each vehicle and the target vehicle. The attention heads are then weighted and summed with the corresponding value vectors V to ultimately obtain the weighted features of the target vehicle.
[0108] The calculation of each attention head is shown in formula (3), and the final fusion interaction feature H of each vehicle is obtained i Attention , see formula (4). In the formula, W O is a learnable weight matrix, h is the number of attention heads, and Concat(·) represents feature concatenation. This study sets the number of multi-head attention mechanisms to 3, with each head corresponding to a feature dimension of 32. The final output dimension is the same as the LSTM output to facilitate subsequent feature fusion.
[0109]
[0110]
[0111] 2.3) Dropout Module
[0112] Finally, the interaction features processed by the attention mechanism are passed through the Dropout module (Dropout probability is set to 0.3) to prevent the network from overfitting, and the output is (It is also the output of the entire temporal feature extraction branch ).
[0113] The above network structure design effectively captures the temporal information of Zhouche's own historical trajectory through the LSTM network, explicitly models the interaction relationship between vehicles through the multi-head attention mechanism, dynamically assigns different weights to the data at specific moments in the historical trajectory, and uses the Dropout module to further improve the accuracy and generalization ability of feature expression.
[0114] (3) Spatial feature extraction branch
[0115] When a vehicle merges into a highway ramp, its driving behavior is not only affected by its own historical motion trajectory, but also deeply affected by the behavior status of surrounding vehicles. There is a complex spatial interaction relationship between the main vehicle and the conflicting vehicle, which is reflected in the interaction between vehicles, relative position relationship and state change. Traditional methods usually analyze the motion trajectory of a single vehicle in isolation, while ignoring the interactivity between vehicles, making it difficult to fully and accurately characterize the complexity of the traffic scene. To solve this problem, this embodiment proposes a spatial feature extraction module with a graph neural network as the core, specifically using GraphSAGE and GAT to effectively model and extract spatial interaction features between vehicles.
[0116] like Figure 5 As shown in the figure, the main function of this module is to use a graph structure to express the interaction relationship between vehicles. In this embodiment, vehicles are regarded as "nodes" and the interaction relationship between vehicles is regarded as the "edges" connecting these nodes, thereby constructing a spatiotemporal correlation graph of the interaction relationship between multiple vehicles. The dynamic interaction characteristics of multiple vehicles are represented in this paper using an undirected graph network structure, that is, G = (V, E), where V = {n1, n2, ..., n i}, i=1,2,...,N is the set of vehicle "nodes", Is the set of “edges” that are associated between vehicles. Each node n in the node set i Corresponding to a feature vector X i , which is the model input. The initial feature matrix H0 of the node is calculated:
[0117]
[0118] 3.1) Graph SAGE obtains the node feature matrix H G-SAGE
[0119] For node n i, its neighbor node set N(n i ) means all i Connected nodes, a k-layer Graph SAGE network is constructed on each node, where Represents node n i At the kth layer, the Graph SAGE algorithm aggregates the features of each node with the mean of the features of all its neighboring nodes, ensuring that each node initially contains partial information about its neighboring nodes. This approach captures more local information when sampling the same node, enhancing the model's expressiveness.
[0120] In calculating the node n of the kth layer i Features When first passing the pooling aggregation function AGGREGATE pool Neighbor node features sampled at k-1 layers Perform aggregation and obtain aggregation results Then add it to node n i Features at the k-1th layer Perform vector splicing operations and finally generate node n through nonlinear mapping i Features at layer k As shown in formula (3-6).
[0121] Where W is the parameter matrix and b is the bias matrix.
[0122]
[0123] 3.2) Get the adjacency matrix A
[0124] In this embodiment, an adjacency matrix is constructed based on the relative relationship between the nodes of the spatiotemporal trajectory, and its elements are recorded as Represents node n i (t) with node n j (τ) Whether there is an interactive relationship. In order to uniformly characterize the connection between two nodes, the following situations are proposed. When any one of them is met, let Otherwise set it to 0.
[0125] 1) Spatial proximity at the same time
[0126] If the Euclidean distance between two different nodes at the same time is less than the given threshold d0 (50m), it is considered that the two objects are close to each other at the same time point and there is a possibility of direct interaction.
[0127] 2) Continuous trajectories of the same object at adjacent moments
[0128] If the node index i is the same and the timestamps differ by 1, then the two nodes are considered to be extensions of the trajectory of the same physical object at consecutive time steps. In this case, it is also considered that there is a possibility of direct interaction between the two objects. The node interaction relationship is obtained as shown in Equation (7). In this way, the connection relationship between nodes will change over time and the relationship between nodes. Based on this, we obtain the adjacency matrix A of the "edge", as shown in Equation (8):
[0129]
[0130] Based on the above node feature processing and edge improvement, this study transforms the node feature matrix H G-SAGE The adjacency matrix A of the edges is used as the input of GAT to capture the mutual influence between vehicles.
[0131] 3.3) Attention coefficient normalization
[0132] First, through the node feature matrix H G-SAGE And the adjacency matrix A calculates the attention coefficient e between nodes ij , and after processing with the LeakyReLU activation function, the softmax function is used to normalize the attention coefficient of each pair of neighbor nodes to obtain the normalized attention coefficient α ij , see formula (9), where the indicator function (usually expressed as 1 {Aij=1} ) is used to ensure that when calculating the softmax normalization, only the nodes that actually have an adjacency relationship are calculated:
[0133]
[0134] Next, use α ij For the node feature matrix H G-SAGE Perform weighted summation to obtain the updated feature representation of the node, see formula (10). The final H GAT This is the node feature matrix after GAT processing, and is also the output H of the spatial extraction branch. Spatial .
[0135]
[0136] Overall, the spatial feature extraction module, through the multi-layered aggregation structure of the Graph SAGE network, deeply captures the spatial interaction characteristics between vehicles. Using GAT, these features are further refined and enhanced, ultimately yielding high-quality feature vectors representing the spatial interactions between vehicles. Compared to traditional methods, this branch's structural design and feature aggregation mechanism more accurately capture the details of vehicle interactions in complex traffic environments.
[0137] (4) Feature fusion optimization
[0138] In the task of predicting driving intention, a vehicle's driving behavior is often influenced by both its historical motion trajectory (temporal features) and the interactions of surrounding vehicles (spatial features). To fully utilize the multi-source feature information extracted by these two branches, this study designed a feature fusion module that efficiently integrates the features of the temporal branch, the spatial branch, and the residual network to improve the accuracy of the final driving intention prediction.
[0139] like Figure 6 , this module receives feature inputs from three different branches, namely the feature vector H extracted by the temporal feature branch Temporal , the feature vector H extracted by the spatial feature branch Spatial , and the feature vector H generated by the residual network Residual Among them, the residual network is used to directly transfer the features of the lower layer or the previous moment to the fusion layer to enhance the stability of feature transfer and avoid information attenuation caused by deep stacking. The residual connection structure can be expressed as:
[0140] H Residual =X i +MLP(X i ) (11)
[0141] Among them, X i The three sets of features represent the dynamic trends of the vehicle's historical trajectory, the spatial interaction between the vehicle and its surroundings, and low-level, fine-grained spatial feature information.
[0142] To adaptively integrate these heterogeneous features, this module adopts a gated fusion mechanism. The core idea is to assign learnable weights to the temporal and spatial branches, thereby dynamically adjusting the importance of the two types of features. First, the three feature vectors are concatenated to obtain the original feature representation before fusion, as shown in Equation (12):
[0143]
[0144] Then, the gating factor is obtained through the fully connected layer and the Softmax activation function. The calculation formula is shown in (13), where W g 、b g are the weights and bias parameters to be learned in the network, and the Softmax operation ensures Finally, the obtained gating factor is used to perform weighted fusion on the three-way features to obtain the fused features, as shown in formula (14):
[0145]
[0146] The fused features are then normalized by the batch normalization layer, as shown in the following formula (15): Fusion This is the final output of the feature fusion module. Where E[·] represents the mean of the feature, Var[·] represents the variance of the feature, ∈ is a decimal to prevent division by zero, and γ and β are learnable scaling and translation parameters.
[0147]
[0148] Through the combination of the above-mentioned gated fusion mechanism and batch normalization, the fusion module can effectively aggregate feature information from different sources, suppress redundancy and noise, enhance feature expression capabilities, and provide high-quality feature support for the subsequent decoder module to accurately predict the driving intention of the surrounding vehicle.
[0149] (5) Decoder module and intent classification output
[0150] After feature fusion, the model needs to map the high-dimensional fused vector to specific intent categories. The decoder is the key module for this task. Its core concept is to use fully connected layers to perform dimensionality reduction or linear transformation on the fused vector and, through activation functions, map the result to a set of interpretable outputs, such as the different possible driving intentions of the vehicle. The decoder module consists of several layers of fully connected networks, culminating in a softmax function to generate a probability distribution for each intent, thus achieving the final classification output.
[0151] like Figure 7 As shown in Figure 1, the core of this module is a multi-layer fully connected network. Each layer of the fully connected network further abstracts high-dimensional features through linear mapping and nonlinear activation functions to extract high-order semantic information. Specifically, this decoder module is composed of a multi-layer fully connected neural network. Through layer-by-layer linear mapping and nonlinear activation function processing, the high-dimensional feature vector is gradually abstracted into a low-dimensional feature expression corresponding to the number of driving intent categories. The calculation method of each fully connected layer can be expressed by formula (16):
[0152] h l+1 =ReLU(W l h l +b l ) (16)
[0153] After processing through multiple fully connected layers, the network gradually reduces the feature dimension and further abstracts it, and finally outputs a vector z=[z LC ,z Dec ,z Keep ,z Acc ], where each dimension corresponds to the probability of the possible driving intention of the vehicle: z LC Indicates lane change to avoid, zDec Indicates slowing down and going straight, z Keep Indicates uniform straight line movement, z Acc In order to convert the output feature vector into the probability of each driving intention category, this module further uses the Softmax function to normalize the above output, as shown in formula (17):
[0154]
[0155] In order to convert the output feature vector into the probability of each driving intention category, this module further uses the Softmax function to normalize the above output, see formula (17). i represents the probability that the conflicting vehicle may adopt the i-th driving intention in the future, z i is the original output value of the network for the i-th intention category. The normalization effect of the Softmax function ensures that the sum of the probabilities of the four driving intentions is 1. The pseudo code is as follows Figure 11 shown.
[0156] (6) Model training and result analysis
[0157] 6.1) Experimental platform construction and training details
[0158] The experiments in this example were conducted in a high-performance computing environment to ensure efficient training and inference of the dual-branch vehicle intention recognition model based on multimodal spatiotemporal fusion. The experiments used an Intel Xeon Gold 6226R @ 2.90GHz CPU, an NVIDIA RTX 3090 (24GB of video memory) GPU, and 128GB of DDR4 memory. The model was trained using Python 3.8, PyTorch 1.10.1, and CUDA 11.3.
[0159] After completing data preparation, this study input the divided training set, validation set, and test set into the dual-branch spatiotemporal model for training, and set key training parameters to ensure efficient convergence and generalization performance of the model. Specifically, the training process uses a batch size of 128, which can fully utilize the parallel computing power of the GPU and help stabilize gradient updates. The optimizer uses Adam, and its initial learning rate is 1×10 -31 , with β1 = 0.9 and β2 = 0.999 to achieve fast convergence in non-convex deep networks.
[0160] In addition, to avoid overfitting and falling into local optimality, when the performance of the validation set does not improve significantly within 5 consecutive training rounds, the learning rate will be automatically decayed to 0.1 times the original value; the entire training process will be carried out for a maximum of 50 rounds, and an early stopping strategy will be adopted. When the model performance on the validation set tends to converge, the training will be terminated while retaining the optimal model parameters.
[0161] 6.2) Baseline model selection and evaluation indicators
[0162] 6.2.1) Baseline model selection
[0163] The Bi-Branch Spatio-Temporal Fusion Model (Bi-STFM) proposed in this embodiment is collectively referred to as Bi-STFM. To verify the effectiveness of the proposed model, this study conducted comparative experiments with the following three baseline models:
[0164] 1. LSTM: Only LSTM is used to process time series information without considering the spatial interaction between vehicles, which is used to measure the effect of pure time series modeling.
[0165] 2. GCN model: It can capture spatial relationships in graph-structured data, but its ability to capture temporal features is relatively limited. It is suitable as a benchmark for modeling spatial features alone.
[0166] 3. ST-GCN model: Spatio-Temporal Graph Convolutional Networks (ST-GCN) combines graph convolution in the spatial domain with convolution operations in the temporal domain, and is widely used in prediction tasks in dynamic interactive environments.
[0167] The above models cover different feature modeling mechanisms, including pure temporal feature modeling, pure spatial feature modeling, and simple fusion of spatiotemporal features. The structural differences between these models and the Bi-STFM model proposed in this paper fully demonstrate the contribution of model structure innovation to performance improvement.
[0168] 6.2.2) Model Evaluation Metrics
[0169] In the multi-classification scenario of intent recognition, in order to comprehensively evaluate the performance of each model, this study uses the following evaluation indicators: Accuracy, Recall, and F1-score. For the convenience of explanation, assume that there are k types of intent labels, denoted by TP k is the number of true positive samples for the kth class (i.e. the true class is k and the prediction is also k), FP k is the number of false positive samples (the actual category is not k but is predicted to be k), FN kis the number of false negative samples (the true category is k but predicted to be other categories), and TN is ignored k (Weaker association with the definition of each category in multi-classification scenarios).
[0170] 1) Accuracy: In multi-classification, accuracy can be defined as the ratio of correctly predicted samples to the total number of samples. If N is the total number of samples, the more intents the model correctly identifies, the higher the accuracy. See formula (18):
[0171]
[0172] 2) Recall rate: For a specific category k, the recall rate represents the proportion of samples with the true label k that are correctly predicted as k. Its physical meaning is: for a certain type of driving intention such as accelerating and cutting in, whether the model can capture as many positive examples as possible among all samples that actually have this intention. See formula (19):
[0173]
[0174] 3) Precision: For category k, precision indicates how many of the samples predicted as k are actually positive examples belonging to category k. A high precision means that the model rarely makes misjudgments when determining a certain type of intent. See formula (20):
[0175]
[0176] 4) F1 coefficient: F1 is the harmonic mean between precision and recall, which is used to balance the two to obtain the overall F1 performance and measure the balance of the model in identifying different driving intentions. See formula (21):
[0177]
[0178] 5) Security: Confusion matrix M∈R 4×4 The element M in ij represents the probability that the true intention is category i but the prediction is category j. Under this definition, we consider predictions on and below the diagonal (i.e., i ≥ j) as safe predictions, while predictions above the diagonal (i.e., i < j) are considered unsafe. This indicator design can intuitively reflect that when the model makes conservative predictions, the overall safety value is higher. Therefore, it can be defined as the safety probability Safety, see formula (22):
[0179]
[0180] 6.3) Comparative Analysis of Intent Recognition Results
[0181] This example uses a 15-second window for vehicle intent recognition, with a recognition frequency of 0.1 Hz. Each recognition cycle can predict the intent of approximately 10 vehicles. After collecting and organizing the data, Table 1 provides an in-depth analysis of the performance differences between the various models and explains the reasons for the performance improvements based on differences in model structure.
[0182] 1) Overall performance comparison analysis
[0183] First, from the overall performance indicators in Table 1, we can see that the Bi-STFM model outperforms the LSTM, ST-GCN, and GCN baseline models in terms of recall, precision, F1 score, accuracy, and safety. For example, in the "acceleration and lane-cutting" intention recognition, the Bi-STFM model's F1 score reached 92.98%, which is higher than LSTM's 86.82%, ST-GCN's 91.34%, and GCN's 88.63%. This difference illustrates the overall performance advantage of the Bi-STFM model in intent prediction. In order to more intuitively display the performance indicators of each model, the data in the table was processed and converted into a radar chart to more clearly compare the performance of each model on different evaluation indicators. See Figure 8 shown.
[0184] Table 1 Algorithm model performance comparison table
[0185]
[0186]
[0187] from Figure 8 As can be seen from the radar chart comparison, compared with models that use only temporal or spatial features, spatiotemporal fusion methods (Bi-STFM and ST-GCN) generally show higher prediction accuracy. In particular, the Bi-STFM model, by fusing temporal and spatial features, demonstrates higher prediction accuracy and greater safety. Compared with traditional LSTM and GCN models, Bi-STFM performs better in complex environments. This is consistent with the multimodal spatiotemporal feature fusion concept proposed in Section 3.3 of this article: combining LSTM with the attention mechanism to extract temporal features, supplemented by GraphSAGE and GAT to construct detailed spatial interaction features, and adaptively integrating these two types of features through a gated fusion strategy, allowing the model to more comprehensively and meticulously capture dynamic driving intentions in complex environments.
[0188] 2) Driving Intention Recognition Confusion Matrix Analysis
[0189] The confusion matrix intuitively shows the predictions between different categories. By observing the degree of confusion between different categories, we can deeply analyze the model's ability to distinguish between different driving intentions. Figure 9From the above, we can see that the classification errors of the models are obviously different: the diagonal value of the Bi-STFM model is significantly higher than that of other models (such as Figure 9 (a) in the figure), while the non-diagonal element values are extremely low, indicating that the model effectively reduces the misjudgment rate between categories, especially in important behaviors such as acceleration, lane-cutting, and lane-changing. Figure 9 (b)) and GCN (as shown in Figure 9 The model shown in (d) shows significant confusion when predicting the intentions of "lane change to avoid" and "accelerate to cut in." This reflects that single-feature modeling methods (only temporal or only spatial) cannot fully capture the complexity of vehicle intentions.
[0190] This obvious difference once again verifies the effectiveness of the dynamic adaptive attention mechanism for spatial features and the multi-head self-attention mechanism for temporal features adopted by Bi-STFM: this structure can capture richer vehicle interaction and motion features, and improve the ability to distinguish different driving intentions.
[0191] 3) F1 score box plot analysis
[0192] Figure 10 Box plots of the F1 scores of various models across multiple experiments are shown. As can be seen, due to the lack of multidimensional information complementarity, the single model's F1 score distribution exhibits significant fluctuations, with a wide range between the upper and lower quartiles, indicating poor performance stability under different operating conditions. The Bi-STFM model, on the other hand, exhibits higher F1 scores across multiple experiments, with a narrower range between the upper and lower quartiles, demonstrating its greater stability and robustness.
[0193] In summary, the selection of baseline models, the construction of evaluation metrics, and the comprehensive comparative analysis of experimental results fully verify the superior performance of the Bi-STFM model proposed in this study in the driving intention recognition task, reflecting the innovation and effectiveness of the model structure design proposed in Section 3.3 of this paper:
[0194] (1) Single temporal or spatial models have certain limitations due to their single information source and are difficult to cope with the uncertainty brought about by multi-vehicle interactions in complex traffic scenarios.
[0195] (2) The dual-branch prediction network based on multimodal spatiotemporal fusion proposed in this paper greatly improves the accuracy and robustness of driving intention prediction by effectively integrating temporal dynamics and spatial interaction features.
[0196] (3) The synergy between the gated fusion mechanism and the overall network architecture is the key to achieving performance breakthroughs, which is not only intuitively reflected in the confusion matrix and F1 score box plot, but also provides improvement directions for subsequent model optimization.
[0197] Example 2
[0198] An embodiment of the present invention further provides an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the following Figure 1 A vehicle intention prediction method based on multimodal spatiotemporal fusion is shown.
[0199] It is understood that the memory may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the server, etc.
[0200] The processor may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various components within the server. It executes various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory, as well as accessing data stored in memory. Optionally, the processor may be implemented using at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may integrate one or a combination of a central processing unit (CPU) and a modem. The CPU primarily processes the operating system and application programs, while the modem handles wireless communications. It is understood that the modem may not be integrated into the processor and may be implemented separately via a single chip.
[0201] Since the electronic device is an electronic device corresponding to a method for predicting surrounding vehicle intentions based on multimodal spatiotemporal fusion in an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0202] Example 3
[0203] An embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A vehicle intention prediction method based on multimodal spatiotemporal fusion is shown in FIG.
[0204] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0205] Since the storage medium is a storage medium corresponding to a method for predicting surrounding vehicle intentions based on multimodal spatiotemporal fusion in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0206] Example 4
[0207] In some possible implementations, various aspects of the method of the embodiments of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to cause the computer device to execute the steps of a method for predicting surrounding vehicle intentions based on multimodal spatiotemporal fusion according to various exemplary embodiments of the present application described above in this specification. The executable computer program code or "code" for executing each embodiment may be written in a high-level programming language such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0208] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0209] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0210] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. A vehicle intention prediction method based on multimodal spatiotemporal fusion, characterized in that: The following steps are involved: Get historical trajectory data; Input the acquired historical trajectory data into the trained vehicle intention prediction model and output the prediction results; The vehicle intention prediction model includes two parallel feature extraction branches: one is the temporal feature extraction branch, which uses a long short-term memory network to encode the historical trajectory of each vehicle and dynamically captures the temporal characteristics of the trajectory through an attention mechanism; the other is the spatial feature extraction branch, which uses a GraphSAGE network to model the spatial interaction relationship between multiple vehicles. Then, combined with the attention mechanism of the graph attention network, adaptive weights are assigned to different vehicles, enabling the model to highlight the importance of key interfering vehicles. The model's fusion module uses a gated fusion mechanism to perform adaptive weighted fusion of features from the temporal feature extraction branch and the spatial feature extraction branch; the fused feature information is finally input into the decoder module, which outputs the probability distribution of various driving intentions of the conflicting vehicles in the future.
2. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that: The temporal feature extraction branch includes an LSTM feature extraction module, a multi-head attention module and a Dropout module; In the LSTM feature extraction module, an independent LSTM encoder is constructed for each vehicle trajectory. The input of each LSTM encoder is one-dimensional time series trajectory data. After LSTM encoding, the motion pattern feature h of the vehicle trajectory is obtained i ; In the multi-head attention module, the motion pattern feature h of each vehicle is i As the attention query vector, the LSTM output features of all vehicles are used as keys and values to calculate the contextual features c of the temporal interaction between vehicles i ; The features processed by the multi-head attention module are passed through the Dropout module to prevent the network from overfitting.
3. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that: In the spatial feature extraction branch, a graph structure is used to express the interaction relationship between vehicles. Vehicles are regarded as nodes, and the interaction relationship between vehicles is regarded as the edges connecting these nodes, that is, G = (V, E), where V = {n1, n2, ..., n i } is a collection of vehicle nodes, is the set of edges that represent the relationships between vehicles; Each node n in the node set i Corresponding to a feature vector X i , the feature vector includes the relative position, velocity, and acceleration of the vehicle at the current moment; the initial feature matrix H0 of the node is calculated:
4. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 3 is characterized in that: The GraphSAGE network works as follows: For node v i , its neighbor node set N(v i ) means all the i Connected nodes, build a K-layer GraphSAGE network on each node, and calculate the features of k-layer node v When , firstly, the neighbor node features of k-1 layer sampling are aggregated by the aggregation function Perform aggregation and obtain aggregation results Then perform vector concatenation operation on it and the feature of node v at the k-1 layer, and finally generate the feature matrix of node v at the k layer through nonlinear mapping 5. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 3 is characterized in that: The graph attention network works as follows: The adjacency matrix is constructed based on the relative relationship between the nodes of the spatiotemporal trajectory, and its elements are recorded as Represents node v j (t) with node v j (τ) Whether there is an interactive relationship; when any of the following conditions is met, let Otherwise set it to 0: 1) Spatial proximity at the same time: If the Euclidean distance between two different nodes at the same time is less than the preset threshold d0, then the two objects are considered to be close to each other at the same time point and there is a possibility of direct interaction; 2) Continuous trajectories of the same object at adjacent moments: If the node index i is the same and the timestamps differ by 1, then the two nodes are considered to be the extension of the trajectory of the same physical object at consecutive time steps, and the node interaction relationship is obtained. In this way, the connection relationship between the nodes will change with time and the relationship between the nodes, and the edge adjacency matrix A is obtained accordingly; The node feature matrix H and the edge adjacency matrix A are used as the input of the graph attention network. The attention coefficient between nodes is calculated through the feature matrix H and the adjacency matrix A. The attention coefficient of each pair of neighbor nodes is normalized using the softmax function after being processed by the LeakyReLU activation function. The normalized attention coefficient matrix is used to perform weighted summation on the node feature matrix H to obtain the updated feature representation of the node, and finally the node feature matrix processed by the attention mechanism is obtained.
6. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that: The input of the fusion module includes: the feature vector H extracted by the temporal feature extraction branch Temporal , the feature vector H extracted by the spatial feature extraction branch Spatial , and the feature vector H generated by the residual network Residual ; Concatenate the three feature vectors to obtain the original feature representation before fusion The gating factor is obtained through the fully connected layer and the Softmax activation function: Where W g 、b g are the weights and bias parameters to be learned in the network, and the Softmax operation ensures The obtained gating factor is used to perform weighted fusion on the three-way features to obtain the fused features: The fused features are then normalized through the batch normalization layer: Where E[·] represents the mean of the feature, Var[·] represents the variance of the feature, ε is a decimal to prevent division by zero, and γ and β are learnable scaling and translation parameters.
7. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that: The decoder module consists of fully connected layers. Through layer-by-layer linear mapping and nonlinear activation function processing, the high-dimensional feature vector is gradually abstracted into low-dimensional feature expressions corresponding to the number of driving intention categories. The calculation method of each fully connected layer is: h l+1 =ReLU(W l h l +b l ) After being processed by multiple fully connected layers, the feature dimension is gradually reduced and further abstracted, and finally a vector with the same dimension as the number of intent categories is output: z=[z LC ,With Dec ,With Keep ,With Acc ] Each dimension corresponds to the probability of the possible driving intention of the vehicle: z LC Indicates lane change to avoid, z Dec Indicates slowing down and going straight, z Keep Indicates uniform straight line movement, z Acc Indicates acceleration and going straight; The Softmax function is used to normalize the above output: Where, P i represents the probability that the conflicting vehicle may adopt the i-th driving intention in the future, z i is the original output value of the network for the i-th intent category.
8. The method for predicting surrounding vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that: The historical trajectory data includes the vehicle position, speed, and acceleration of the main vehicle and the conflicting vehicle; The obtaining of historical trajectory data includes: Select the vehicle history trajectory of the preset time length in the past as the original input data X of the model i : X i ={(x i,t ,y i,t ,v i,t ,a i,t )} Where x i,t ,y i,t They represent the horizontal and vertical coordinate positions of vehicle i at time t, v i,t and a i,t are the speed and acceleration of the vehicle at time t, respectively.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Vehicle track prediction method based on intention perception space-time attention network
CN117141518A
Trajectory prediction method for multi-dimensional spatio-temporal feature fusion for automatic driving
CN118296090A
Method for driving behavior modeling based on spatio-temporal information fusion
US12240470B1
Cited By
Internet of vehicles cooperative attack anomaly detection method and device
CN120785662A
A method and device for detecting anomalies in vehicle-to-everything (V2X) collaborative attacks
CN120785662B
Vehicle driving intention recognition and trajectory prediction method for automatic driving
CN120840639A
Traffic participant intention prediction method and device based on multi-modal time sequence attention
CN120932467A
Traffic participant intention prediction method and device based on multi-modal time attention
CN120932467B