A multi-modal spatio-temporal fusion based surrounding vehicle intention prediction method and related device

By employing a multimodal spatiotemporal fusion method for predicting vehicle intentions, and utilizing long short-term memory networks and graph neural networks to extract vehicle temporal and spatial features, this method solves the problem of autonomous driving systems struggling to identify vehicle intentions in high-speed ramp merging scenarios. It achieves accurate prediction of conflicting vehicle intentions, thereby improving the safety of the merging process and the stability of traffic flow.

CN120472656BActive Publication Date: 2026-07-21SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2025-04-07
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In high-speed ramp merging scenarios, autonomous driving systems struggle to accurately identify the driving intentions of surrounding vehicles, impacting the safety and smoothness of the merging process.

Method used

A multimodal spatiotemporal fusion-based vehicle intention prediction method is adopted. By combining long short-term memory network and graph neural network, the temporal and spatial features of the vehicle are extracted. The features are weighted and fused using attention mechanism and gating fusion mechanism. Finally, the probability distribution of driving intention is output through decoder module.

Benefits of technology

It achieves efficient and robust prediction of the driving intentions of conflicting vehicles, provides decision support for autonomous driving systems in complex scenarios, and improves the safety of the merging process and the smoothness of traffic flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472656B_ABST
    Figure CN120472656B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal space-time fusion's week car intention prediction method and related equipment, wherein method includes: obtaining historical trajectory data;The obtained historical trajectory data is input into the trained week car intention prediction model, and output prediction result;Wherein, model includes two branches: one is time sequence feature extraction branch, long short-term memory network is used to encode the historical trajectory of each vehicle, and the time sequence feature of track is dynamically captured by attention mechanism;Two is space feature extraction branch, uses GraphSAGE network to model the space interaction between multiple vehicles, then combines the attention mechanism of graph attention network, and different vehicles are assigned adaptive weight, so that model can highlight the importance of key interference vehicle;From the feature of time sequence feature extraction branch and space feature extraction branch, adaptively weighted combination is carried out, and the probability distribution of each type of driving intention at future time is obtained according to the fused feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, and in particular to a method and related equipment for predicting vehicle intentions based on multimodal spatiotemporal fusion. Background Technology

[0002] In a highway environment, ramp merging presents a highly dynamic and unpredictable traffic scenario. Ramps typically bring vehicles from side roads or urban roads into the main highway lanes, requiring them to match speeds and merge with the main lane traffic within a short acceleration lane. Figure 1 This diagram illustrates a typical highway ramp merging scenario, which can be broken down into four functional segments:

[0003] 1) Ramp connection section (80 meters): This refers to the section where main vehicles enter the ramp from branch roads or main roads. On this section, the speed of main vehicles is relatively low, and the main task is to assess the traffic flow and surrounding vehicle behavior on the main road to prepare for subsequent merging.

[0004] 2) Acceleration and Connection Section (200 meters): This section is located in the middle of the ramp and is a crucial area for the lead vehicle to accelerate to match the speed of traffic on the main road. During this stage, the lead vehicle gradually increases its speed to safely merge into the main road. During the acceleration phase, the lead vehicle needs to constantly pay attention to the speed and position of surrounding vehicles to ensure it finds a suitable opportunity to merge.

[0005] 3) Merging Section (60 meters): The merging section is where the main traffic flow actually intersects with the traffic flow on the main road, and vehicles begin to merge with other vehicles on the main road. At this time, the main traffic flow not only needs to pay attention to its relative position to the surrounding vehicles on the main road, but also needs to maintain a safe distance from the surrounding vehicles during the merging process.

[0006] 4) Main Road Entry Section: After merging, the main vehicle enters the subsequent lane of the main road. The task at this stage is to maintain a safe distance from other vehicles on the main road to avoid collisions and congestion, while smoothly integrating into the main road traffic flow.

[0007] In highway ramp merging scenarios, the interaction between different types of vehicles directly affects the safety and smoothness of the merging process. Based on the vehicle's position and function during the merging process, this application mainly classifies merging vehicles into the following three categories: Figure 2 As shown:

[0008] 1) Autonomous vehicles (main vehicle) on the ramp

[0009] As the main entity actively completing lane-changing operations, autonomous vehicles on ramps undertake multiple tasks. First, they need to perceive the dynamic information within their own ramp segment in real time, including vehicle position, speed, acceleration, and road geometry. At the same time, they need to obtain the dynamic behavior of vehicles in the main lane through sensor data and vehicle-to-everything (V2X) information.

[0010] The driver determines the driving intentions of surrounding vehicles involved in the collision. Based on the predicted intentions of these vehicles, the driver needs to engage in a game-theoretic decision-making process with the vehicles in the target lane on the main road. Game-theoretic decision-making refers to selecting an optimal lane-changing strategy that avoids collisions while ensuring smooth traffic flow, through interaction and coordination, in the presence of potential conflict risks. This process typically involves multi-objective optimization, including balancing driving safety, traffic efficiency, and passenger comfort. The autonomous driving system, through its decision-making and planning module, comprehensively considers various sensor information, prediction results, and vehicle dynamics constraints to calculate appropriate acceleration and steering wheel angles, enabling the planning and control of the future trajectory and thus smoothly completing the lane-changing operation from the ramp to the main lane.

[0011] 2) Conflicting vehicles on the target lane of the main road

[0012] Conflicting vehicles in the target lane of the main road refer to vehicles that have a potential risk of interaction and conflict with main vehicles on the ramp during the merging process. Typically, these vehicles travel at a predetermined speed in the main lane, but their behavior may change when faced with vehicles approaching from the ramp, such as actively slowing down, yielding in advance, or failing to adjust their speed in time, resulting in a cross-trajectory. The presence of conflicting vehicles increases the uncertainty of the merging process, and fluctuations in their motion states (such as speed, acceleration, and expected trajectory) are a key factor that autonomous driving systems must consider during decision-making.

[0013] 3) Vehicles traveling in the inner lane of the main road

[0014] Vehicles traveling in the inner lanes of the main road are typically located relatively far from the merging area, and their driving status and dynamics have a relatively small direct impact on merging vehicles. For the sake of model simplification and focus, this application does not consider the dynamic information of vehicles in the inner lanes of the main road when designing the weekly vehicle driving intention prediction model. Although these vehicles may affect the overall traffic flow equilibrium over a larger area, their direct interference effect on merging decisions is relatively limited; therefore, in this application, they are treated as background information and not as primary decision variables.

[0015] In the scenario of merging into a highway ramp, in addition to accurately perceiving its own status and the road environment, the lead vehicle must also make timely and accurate judgments on the potential behavior of other vehicles in the main lane to ensure the safety and efficiency of the lane merging operation. Figure 2The diagram illustrates a typical ramp merging scenario, where a vehicle is attempting to merge from the ramp into the target lane on the main road, facing the potential interaction of at least two conflicting vehicles (conflicting vehicle A and conflicting vehicle B) on the main road. Figure 2 The text indicates several possible vehicle decision-making methods. For example, the main vehicle can choose to "merge directly" or "wait to give way," while vehicles on the main road may choose to "keep going straight" or "change lanes to give way."

[0016] from Figure 2 It is evident that vehicle A in a conflict might change lanes to allow sufficient space for the main vehicle to merge, or it might choose to continue driving straight and occupy the target lane. Vehicle B, on the other hand, might maintain its speed within its own lane or proactively slow down to yield. Meanwhile, the main vehicle on the ramp can choose to merge directly or slow down and wait based on traffic flow conditions and risk assessment. Clearly, the combinations of these vehicle decisions are extremely diverse. If the autonomous driving system cannot predict the true intentions of surrounding vehicles in a timely manner, it can easily lead to collision risks or traffic flow instability. Only by accurately identifying these intentions can the autonomous driving system make targeted game-theoretic decisions and trajectory planning; currently, a precise and effective identification solution is lacking. Summary of the Invention

[0017] In order to at least partially solve one of the technical problems existing in the prior art, the purpose of this invention is to provide a method and related equipment for predicting the intention of a vehicle based on multimodal spatiotemporal fusion.

[0018] The first technical solution adopted in this invention is:

[0019] A method for predicting vehicle intentions based on multimodal spatiotemporal fusion includes the following steps:

[0020] Obtain historical trajectory data;

[0021] The obtained historical trajectory data is input into the trained vehicle intention prediction model, and the prediction result is output.

[0022] The vehicle intent prediction model includes two parallel feature extraction branches: one is a temporal feature extraction branch, which uses a long short-term memory network to encode the historical trajectory of each vehicle and dynamically captures the temporal features of the trajectory through an attention mechanism; the other is a spatial feature extraction branch, which uses a GraphSAGE network to model the spatial interaction relationship between multiple vehicles, and then combines the attention mechanism of a graph attention network to assign adaptive weights to different vehicles, so that the model can highlight the importance of key interfering vehicles.

[0023] The spatiotemporal features extracted from the two branches are input into the model's fusion module. The fusion module uses a gated fusion mechanism to adaptively weight and combine the features from the temporal feature extraction branch and the spatial feature extraction branch, and further improves the stability of network training through batch normalization. The fused feature information is finally input into the decoder module, and the Softmax function is used to output the probability distribution of various driving intentions of the conflicting vehicle at future time moments.

[0024] Furthermore, the temporal feature extraction branch includes an LSTM feature extraction module, a multi-head attention module, and a Dropout module;

[0025] In the LSTM feature extraction module, an independent LSTM encoder is constructed for each vehicle trajectory. The input to each LSTM encoder is one-dimensional time-series trajectory data. After LSTM encoding, the motion pattern features h of the vehicle trajectory are obtained. i ;

[0026] In the multi-head attention module, the motion pattern features h of each vehicle are... i As the attention query vector, the LSTM output features of all vehicles are used together as the key and value to calculate the contextual features c of the temporal interactions between vehicles. i ;

[0027] Features processed by the multi-head attention module are then processed by the Dropout module to prevent overfitting and improve the robustness of temporal features.

[0028] Furthermore, in the spatial feature extraction branch, a graph structure is used to represent the interaction relationships between vehicles. Vehicles are considered as nodes, and the interaction relationships between vehicles are considered as edges connecting these nodes, i.e., G = (V, E), where V = {n1, n2, ..., n}. i} is a set of vehicle nodes. It is the set of edges that define the relationships between vehicles;

[0029] Each node n in the node set i Corresponding to a feature vector X i The feature vector includes the vehicle's relative position, velocity, and acceleration at the current moment; the initial feature matrix H0 of the node is calculated as follows:

[0030]

[0031] Furthermore, the GraphSAGE network operates as follows:

[0032] For node v i Its neighbor node set N(v i ) represents all expressions related to v iConnected nodes, a K-layer GraphSAGE network is constructed at each node, and features of the k-layer node v are computed. First, the features of neighboring nodes sampled at layer k-1 are aggregated using an aggregation function. Perform aggregation to obtain the aggregation result. Then, the vectors of this feature and the features of node v at layer k-1 are concatenated, and finally, a nonlinear mapping is performed to generate the feature matrix of node v at layer k.

[0033] Furthermore, the graph attention network operates as follows:

[0034] An adjacency matrix is ​​constructed based on the relative relationships between nodes in a spatiotemporal trajectory, and its elements are denoted as... Represents node v j (t) and node v j (τ) Does an interaction relationship exist? Let τ satisfy any of the following conditions. Otherwise, set it to 0:

[0035] 1) Spatial proximity at the same time: If the Euclidean distance between two different nodes at the same time is less than the preset threshold d0, then the two objects are considered to be close to each other at the same time and there is a possibility of direct interaction.

[0036] 2) Continuous trajectory of the same object at adjacent time steps: If the node index i is the same and the timestamps differ by 1, then these two nodes are regarded as the trajectory extension of the same physical object in consecutive time steps, and the node interaction relationship is obtained. In this way, the connection relationship between nodes will change with time and the relationship between nodes, and the adjacency matrix A of the edge is obtained accordingly.

[0037] The feature matrix H of the nodes and the adjacency matrix A of the edges are used as inputs to the graph attention network. The attention coefficients between nodes are calculated using the feature matrix H and the adjacency matrix A. After processing with the LeakyReLU activation function, the attention coefficients of each pair of neighboring nodes are normalized using the softmax function. The normalized attention coefficient matrix is ​​then used to perform a weighted summation on the feature matrix H of the nodes to obtain the updated feature representation of the nodes. Finally, the node feature matrix after the attention mechanism is obtained.

[0038] Furthermore, the input to the fusion module includes: the feature vector H extracted by the temporal feature extraction branch. Temporal The feature vector H extracted by the spatial feature extraction branch Spatial and the eigenvector H generated by the residual network Residual ;

[0039] The three feature vectors are concatenated to obtain the original feature representation before fusion.

[0040]

[0041] The gating factor is obtained through a fully connected layer and a Softmax activation function:

[0042]

[0043] In the formula, W g b g These are the weights and bias parameters to be learned in the network, and the Softmax operation guarantees... The obtained gating factor is used to perform weighted fusion of the three features to obtain the fused features:

[0044]

[0045] The fused features are then normalized using a batch normalization layer.

[0046]

[0047] In the formula, E[·] represents the mean of the feature, Var[·] represents the variance of the feature, ε is a decimal to prevent division by zero, and γ and β are learnable scaling and translation parameters.

[0048] Furthermore, the decoder module consists of fully connected layers. Through layer-by-layer linear mapping and non-linear activation function processing, it gradually abstracts the high-dimensional feature vector into a low-dimensional feature representation corresponding to the number of driving intention categories. The calculation method for each fully connected layer is as follows:

[0049] h l+1 =ReLU(W l h l +b l )

[0050] After processing through multiple fully connected layers, the feature dimension is gradually reduced and further abstracted, ultimately outputting a vector with the same dimension as the number of intent categories:

[0051] z = [z LC ,z Dec ,z Keep ,z Acc ]

[0052] Each dimension corresponds to the probability of the vehicle's possible driving intentions during the week: z LC Indicates changing lanes to give way, z Dec Indicates deceleration and straight-line movement, z Keep Indicates uniform straight-line motion, z Acc Indicates accelerating straight ahead;

[0053] The above output is normalized using the Softmax function:

[0054]

[0055] In the formula, P i z represents the probability that the conflicting vehicle may adopt the i-th driving intention in the future. i This represents the network's original output value for the i-th intent category.

[0056] Furthermore, the historical trajectory data includes the vehicle position, speed, and acceleration of the main vehicle and the conflicting vehicle;

[0057] The acquisition of historical trajectory data includes:

[0058] Historical vehicle trajectories over a predetermined time period are selected as the model's original input data X. i :

[0059] X i ={(x i,t ,y i,t ,v i,t ,a i,t )}

[0060] In the formula, x i,t ,y i,t Let v represent the horizontal and vertical coordinates of vehicle i at time t, respectively. i,t and a i,t Let $\frac{ ...

[0061] The second technical solution adopted in this invention is:

[0062] An electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a multimodal spatiotemporal fusion-based vehicle intention prediction method as described above.

[0063] The third technical solution adopted in this invention is:

[0064] A computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a multimodal spatiotemporal fusion-based vehicle intention prediction method as described above.

[0065] The fourth technical solution adopted in this invention is:

[0066] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method for predicting vehicle intentions based on multimodal spatiotemporal fusion.

[0067] The beneficial effects of this invention are: by closely coordinating the temporal feature extraction branch and the spatial feature extraction branch, this invention fully integrates spatiotemporal information, which can accurately predict the driving intentions of conflicting vehicles, providing an efficient and robust solution for the research of autonomous driving technology in multi-vehicle interaction scenarios, and providing solid data and theoretical support for the decision-making and planning of autonomous driving systems in complex scenarios such as merging at highway ramps. Attached Figure Description

[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following description is provided with accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0069] Figure 1 This is a schematic diagram of a highway ramp merging;

[0070] Figure 2 This is a diagram showing the position and function of each vehicle when merging into a highway ramp.

[0071] Figure 3 This is an overall framework diagram of the vehicle intent prediction model in this embodiment of the invention;

[0072] Figure 4 This is a schematic diagram of the temporal feature extraction branch in an embodiment of the present invention;

[0073] Figure 5 This is a schematic diagram of the spatial feature extraction branch in an embodiment of the present invention;

[0074] Figure 6 This is a schematic diagram of the fusion module in an embodiment of the present invention;

[0075] Figure 7 This is a schematic diagram of the decoder module in an embodiment of the present invention;

[0076] Figure 8 This is a radar chart comparing the performance of various models in the embodiments of the present invention;

[0077] Figure 9This is a schematic diagram of confusion matrix comparison in an embodiment of the present invention;

[0078] Figure 10 This is a comparison chart of the F1 fractional box plots in the embodiments of the present invention;

[0079] Figure 11 This is a schematic diagram of the overall pseudocode of the vehicle intention prediction model in an embodiment of the present invention. Detailed Implementation

[0080] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0081] The terminology used in the embodiments of this application is for the purpose of describing specific embodiments only and is not intended to limit the embodiments of this application. The singular forms "a," "described," and "the" used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. Furthermore, unless otherwise expressly limited, terms such as "set," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0082] In the description of this application, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0083] In the description of this application, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0084] In the description of this application, "and / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the related objects before and after it are in an "or" relationship.

[0085] Terminology Explanation:

[0086] GraphSAGE: an abbreviation for Graph Sample and Aggregated, is a graph neural network model used for learning graph node embeddings;

[0087] Example 1

[0088] like Figure 3 As shown, this embodiment provides a method for predicting vehicle intentions based on multimodal spatiotemporal fusion, including the following steps:

[0089] S1. Obtain historical trajectory data;

[0090] S2. Input the obtained historical trajectory data into the trained vehicle intention prediction model and output the prediction result;

[0091] The vehicle intent prediction model comprises two parallel feature extraction branches: a temporal feature extraction branch, which uses a long short-term memory network to encode the historical trajectories of each vehicle and dynamically captures the temporal features of the trajectories through an attention mechanism; and a spatial feature extraction branch, which uses a GraphSAGE network to model the spatial interaction relationships between multiple vehicles, and then combines the attention mechanism of a graph attention network to assign adaptive weights to different vehicles, enabling the model to highlight the importance of key interfering vehicles. The model's fusion module uses a gating fusion mechanism to adaptively weight and combine the features from the temporal and spatial feature extraction branches. The fused feature information is finally input into the decoder module, which outputs the probability distribution of various driving intentions of the conflicting vehicles at future times.

[0092] This embodiment proposes a driving intention prediction network (i.e., a surrounding vehicle intention prediction model) that combines temporal features and spatial interaction information to effectively identify the potential driving intentions of surrounding vehicles in the scenario of merging onto a highway ramp. Figure 3The overall framework of the network is presented, which mainly consists of two core branches: a temporal extraction branch and a spatial extraction branch. In the temporal extraction branch, a Long Short-Term Memory (LSTM) network combined with an attention mechanism is employed. This not only captures the dynamic changes in vehicle trajectories but also weights the importance of each time step through self-attention, highlighting the features of key moments. This design effectively solves the problems of information redundancy and gradients in long-sequence data. Simultaneously, the spatial extraction branch utilizes GraphSAGE and a Graph Attention Network (GAT) to construct an interaction graph between vehicles. Through neighborhood sampling and adaptive weight adjustment, it accurately characterizes the relative positions, speed differences, and lane relationships between vehicles, reflecting the potential interference of conflicting vehicles on the main vehicle. Subsequently, the model employs a gated fusion mechanism to adaptively integrate temporal and spatial features, further supplemented by a residual network and batch normalization to ensure stable transmission of fused features within the deep network. Finally, after fully connected layers and Softmax activation, the model outputs the probability distribution of each driving intent.

[0093] The following detailed explanation of the vehicle intention prediction model of this embodiment, in conjunction with the accompanying drawings and specific implementation methods, will be provided.

[0094] (1) Model Input

[0095] The input module is the front end of the network, responsible for data acquisition and processing. This study sets the target scenario as a highway ramp merging scenario, using the historical trajectory data of the lead vehicle and surrounding vehicles as the network input. Considering that vehicle behavior may reveal more interactive information over a longer historical window, this network selects the vehicle historical trajectories of the past 4 seconds as the raw input data.

[0096] Each vehicle trajectory data includes its position (horizontal and vertical coordinates), velocity, acceleration, and heading angle in the road coordinate system. Assuming a sampling frequency of 10Hz, each vehicle generates state data for 40 time points within 4 seconds. This trajectory data forms a time series input. Assuming there are N vehicles in the scene (including the main vehicle and surrounding conflicting vehicles), the trajectory sequence of each vehicle within the historical time window is defined by equation (3-1), where x... i,t ,y i,t Let v represent the horizontal and vertical coordinates of vehicle i at time t, respectively. i,t and a i,t Let $\frac{ ...

[0097] X i ={(x i,t ,y i,t ,v i,t ,a i,t )},t∈[-4s,0],i=1,2,...,N (1)

[0098] (2) Temporal Feature Extraction Branch

[0099] In autonomous driving scenarios, accurately identifying the intentions of surrounding vehicles often requires reasoning based on their historical trajectories over a period of time. Therefore, effectively extracting temporal features from trajectory data is crucial for achieving high-precision predictions. While traditional Long Short-Term Memory (LSTM) networks have certain advantages in modeling temporal dependencies, they still have many shortcomings when dealing with long-term dependencies.

[0100] First, LSTM has limitations in highlighting key information. Although it uses gating mechanisms to control the flow of information, when the time series is long, the model often struggles to effectively capture key features at specific moments, relying more on the most recent input information. Second, LSTM's memory capacity remains insufficient for particularly long trajectory data. While memory units extend the information storage time, implicit long-term dependencies are still easily diluted or lost in extremely long time series, affecting the model's grasp of global information. Finally, LSTM struggles to achieve global attention across the entire time series. It propagates hidden states recursively, relying primarily on information from the previous time step for updates, lacking the ability to comprehensively process the global historical trajectory. Therefore, when the trajectory of a train spans a long period, important clues may be overlooked in repeated local updates.

[0101] To address the aforementioned issues, this embodiment introduces a multi-head attention (MHA) mechanism on top of the LSTM structure. This allows the model to dynamically adjust its focus on the input time series, enabling it to capture global information over long time spans and highlight key information, thereby improving the accuracy of driving intent recognition. Figure 4 As shown, this branch mainly consists of three modules: LSTM feature extraction module, multi-head attention module, and Dropout module.

[0102] 2.1) LSTM Feature Extraction Module

[0103] This module constructs an independent LSTM encoder for each vehicle's trajectory i. The input of each LSTM encoder at time t is... After LSTM encoding, the network obtains the hidden state vector of the vehicle trajectory. These hidden states represent the motion pattern characteristics of the vehicle trajectory within the historical window. The LSTM hidden layer dimension is set to 128. The key computational process of LSTM encoding includes the forget gate, input gate, candidate state, unit state update, and recursive calculation of the output gate. The specific formula is shown in (2):

[0104]

[0105] In the formula, ht-1 C t-1 These represent the hidden state and the cell state from the previous time step, respectively; f t i t ,o t These represent the gate values ​​for the forget gate, input gate, and output gate, respectively; σ is the sigmoid function; ⊙ represents element-wise multiplication; W f W i W C W o and b f ,b i ,b C ,b o These are the learnable parameter matrices and bias vectors for training the LSTM network. For each vehicle trajectory sequence, the hidden state sequence is finally obtained through the above recursive operations. That is, the output of the LSTM timing coding module.

[0106] 2.2) Multi-head attention module

[0107] While LSTM modules effectively capture the temporal features of a single vehicle, they cannot explicitly model the interactions between vehicles. To further enhance the differences in contributions from different historical moments in the vehicle trajectory, a multi-head attention layer based on a self-attention mechanism is added after LSTM encoding. For each target vehicle's final hidden feature h... i As the attention query vector; h is the LSTM output feature of all other vehicles. j Together, they serve as the key vector and value vector. By calculating the similarity between the target vehicle's query vector Q and the key vectors K of all other vehicles, the model can determine the attention heads of each vehicle and the target vehicle. Then, the attention heads are weighted and summed with their corresponding value vectors V to obtain the weighted features of the target vehicle.

[0108] The calculation of each attention head is shown in Equation (3), and the final fused interaction features H of each vehicle are obtained. i Attention See equation (4). In the equation, W O Let h be the learnable weight matrix, h be the number of attention heads, and Concat(·) denote feature concatenation. In this study, the number of heads in the multi-head attention mechanism is set to 3, and the feature dimension corresponding to each head is 32. The final output dimension is the same as that of the LSTM output, which facilitates subsequent feature fusion.

[0109]

[0110]

[0111] 2.3) Dropout module

[0112] Finally, the interaction features processed by the attention mechanism are passed through the Dropout module (with the Dropout probability set to 0.3) to prevent network overfitting, and the output is... (This also serves as the output of the entire temporal feature extraction branch) ).

[0113] The above network structure design effectively captures the temporal information of the vehicle's own historical trajectory through the LSTM network, explicitly models the interaction relationship between vehicles through the multi-head attention mechanism, dynamically assigns different weights to data at specific moments in the historical trajectory, and further improves the accuracy and generalization ability of feature representation by using the Dropout module.

[0114] (3) Spatial Feature Extraction Branch

[0115] During merging on highway ramps, driving behavior is influenced not only by the vehicle's own historical trajectory but also by the behavior of surrounding vehicles. Complex spatial interactions exist between the merging vehicle and the conflicting vehicle, manifested in their interactions, relative positions, and state changes. Traditional methods typically analyze the trajectory of individual vehicles in isolation, neglecting the interactions between them, making it difficult to comprehensively and accurately depict the complexity of traffic scenarios. To address this issue, this embodiment proposes a spatial feature extraction module based on graph neural networks, specifically using GraphSAGE and GAT to effectively model and extract spatial interaction features between vehicles.

[0116] like Figure 5 As shown, the main function of this module is to use a graph structure to express the interaction relationships between vehicles. In this embodiment, vehicles are considered as "nodes," and the interaction relationships between vehicles are considered as "edges" connecting these nodes, thus constructing a spatiotemporal relational graph of multi-vehicle interaction relationships. The dynamic interaction characteristics of multiple vehicles are represented by an undirected graph network structure, i.e., G = (V, E), where V = {n1, n2, ..., n}. i Let i = 1, 2, ..., N be the set of vehicle "nodes". It is the set of "edges" that connect vehicles. Each node n in the node set... i Corresponding to a feature vector X i This refers to the model input. The initial feature matrix H0 of the nodes is calculated as follows:

[0117]

[0118] 3.1) Obtain the node feature matrix H using Graph SAGE G-SAGE

[0119] For node n iIts neighbor node set N(n) i ) represents all pairs of n i The connected nodes form a k-layer Graph SAGE network at each node, where... Represents node n i Features at layer k. The Graph SAGE algorithm aggregates the features of each node with the mean of the features of all its neighbors, allowing each node to contain partial information from its neighbors in its initial state. This method captures more local information when sampling the same node, thus enhancing the model's expressive power.

[0120] In calculating node n at the k-th layer i Features At that time, the pooling aggregation function AGGREGATE is first used. pool Neighbor node features sampled from layer k-1 Perform aggregation to obtain the aggregation result. Then combine it with node n i Features in the (k-1)th layer The vector concatenation operation is performed, and finally, node n is generated through nonlinear mapping. i Features at the kth layer As shown in equation (3-6).

[0121] In the formula, W is the parameter matrix and b is the bias matrix.

[0122]

[0123] 3.2) Obtain the adjacency matrix A

[0124] In this embodiment, an adjacency matrix is ​​constructed based on the relative relationships between spatiotemporal trajectory nodes, and its elements are denoted as... Represents node n i (t) and node n j (τ) Does an interaction relationship exist? To uniformly characterize the relationship between two nodes, the following cases are proposed. If any one of these conditions is met, then let... Otherwise, set it to 0.

[0125] 1) Spatial proximity at the same time

[0126] If the Euclidean distance between two different nodes at the same time is less than a given threshold d0 (50m), then the two objects are considered to be close to each other at the same time and there is a possibility of direct interaction.

[0127] 2) Continuous trajectories of the same object at adjacent time points

[0128] If the node indices i are the same and the timestamps differ by 1, then these two nodes are considered as the trajectory extensions of the same physical object in consecutive time steps. In this case, it is also assumed that the two objects may interact directly. The node interaction relationship is obtained as shown in equation (7). Thus, the connection relationship between nodes changes with time and the relationship between nodes. Based on this, we obtain the adjacency matrix A of the "edges," as shown in equation (8):

[0129]

[0130] Based on the improvements to node feature processing and edges mentioned above, this study will use the node feature matrix H G-SAGE The adjacency matrix A of the edges is used as input to GAT to capture the mutual influence between vehicles.

[0131] 3.3) Attention coefficient normalization

[0132] First, through the node feature matrix H G-SAGE The attention coefficient e between nodes is calculated from the adjacency matrix A. ij After processing with the LeakyReLU activation function, the attention coefficients of each pair of neighboring nodes are normalized using the softmax function, thus obtaining the normalized attention coefficients α. ij See equation (9), where the indicator function (usually expressed as 1) {Aij=1} The purpose of this is to ensure that, when calculating softmax normalization, only nodes that actually have an adjacency relationship are calculated.

[0133]

[0134] Next, use α ij For the node characteristic matrix H G-SAGE Weighted summation is performed to obtain the updated feature representation of the nodes, as shown in equation (10). The final H is obtained... GAT This is the node feature matrix after GAT processing, and also the output H of the spatial extraction branch. Spatial .

[0135]

[0136] Overall, the spatial feature extraction module deeply captures the spatial interaction features between vehicles through the multi-layered aggregation structure of the Graph SAGE network, and further refines and enhances the expression of features using GAT, ultimately obtaining high-quality feature vectors representing the spatial interaction relationships between vehicles. Compared with traditional methods, the structural design and feature aggregation mechanism of this branch can more accurately reflect the details of vehicle interaction in complex traffic environments.

[0137] (4) Feature fusion optimization

[0138] In driving intention prediction tasks, a vehicle's driving behavior is often influenced by both its own historical trajectory (temporal features) and the interaction behavior of surrounding vehicles (spatial features). To fully utilize the multi-source feature information extracted from these two branches, this study designs a feature fusion module that efficiently fuses the features from the temporal branch, spatial branch, and residual network to improve the accuracy of the final driving intention prediction.

[0139] For example Figure 6 This module receives feature inputs from three different branches, namely the feature vector H extracted from the temporal feature branch. Temporal Feature vector H extracted by spatial feature branch Spatial and the eigenvector H generated by the residual network Residual The residual network is used to directly pass features from lower or earlier time steps to the fusion layer, enhancing the stability of feature transfer and avoiding information attenuation caused by depth stacking. This residual connection structure can be represented as:

[0140] H Residual =X i +MLP(X i (11)

[0141] Among them, X i For shallow or historical features, MLP(·) represents a multilayer perceptron structure used for nonlinear mapping and reconstruction of input features. These three sets of features respectively characterize the dynamic change trend of the vehicle's historical trajectory, the spatial interaction between the vehicle and its surrounding environment, and low-order fine-grained spatial feature information.

[0142] To adaptively integrate these heterogeneous features, this module employs a gated fusion mechanism. The core idea is to assign learnable weights to the temporal and spatial branches, thereby dynamically adjusting the importance of the two types of features. First, the three feature vectors are concatenated to obtain the original feature representation before fusion, as shown in equation (12):

[0143]

[0144] Subsequently, the gating factor is obtained through a fully connected layer and a Softmax activation function, and the calculation formula is shown in (13), where W g b g These are the weights and bias parameters to be learned in the network, and the Softmax operation guarantees... Finally, the obtained gating factor is used to perform weighted fusion of the three features to obtain the fused features, as shown in Equation (14):

[0145]

[0146] The fused features are then normalized by a batch normalization layer, as shown in equation (15), H Fusion This is the final output of the feature fusion module. Here, E[·] represents the mean of the features, Var[·] represents the variance of the features, ∈ is a decimal to prevent division by zero, and γ and β are learnable scaling and translation parameters.

[0147]

[0148] By combining the above-mentioned gating fusion mechanism with batch normalization, the fusion module can effectively aggregate feature information from different sources, suppress redundancy and noise, enhance feature representation capabilities, and provide high-quality feature support for the subsequent decoder module to accurately predict the driver's intentions of the vehicle.

[0149] (5) Decoder module and intent classification output

[0150] After feature fusion, the model needs to map the high-dimensional fused vector to specific intent categories. The decoder is the key module for accomplishing this function. Its core idea is to use fully connected layers to reduce the dimensionality of the fused vector or perform a linear transformation, and then use an activation function to map the result to a set of interpretable outputs, such as the different driving intentions a vehicle might take. The decoder module consists of several fully connected network layers, and finally connects to a softmax function to obtain the probability distribution of each intent, thus completing the final classification output.

[0151] like Figure 7 As shown, the core of this module is a multi-layer fully connected network. Each layer of the fully connected network further abstracts high-dimensional features through linear mapping and non-linear activation functions to extract high-order semantic information. Specifically, this decoder module consists of a multi-layer fully connected neural network. Through layer-by-layer linear mapping and non-linear activation function processing, it gradually abstracts high-dimensional feature vectors into low-dimensional feature expressions corresponding to the number of driving intention categories. The calculation method of each fully connected layer can be expressed by equation (16):

[0152] h l+1 =ReLU(W l h l +b l (16)

[0153] After processing through multiple fully connected layers, the network progressively reduces and further abstracts the feature dimensions, ultimately outputting a vector z = [z...]. LC ,z Dec ,z Keep ,z Acc ], where each dimension corresponds to the probability of the vehicle's possible driving intentions per week: z LC Indicates changing lanes to give way, zDec Indicates deceleration and straight-line movement, z Keep Indicates uniform straight-line motion, z Acc This indicates accelerating straight ahead. To convert the output feature vector into the probability of each driving intention category, this module further uses the Softmax function to normalize the above output, as shown in equation (17):

[0154]

[0155] To transform the output feature vector into probabilities for each driving intention category, this module further uses the Softmax function to normalize the output, as shown in equation (17). Where P i z represents the probability that the conflicting vehicle may adopt the i-th driving intention in the future. i This represents the network's original output value for the i-th intention category. The Softmax function normalizes the output, ensuring that the sum of the probabilities of the four driving intentions is 1. The pseudocode is as follows: Figure 11 As shown.

[0156] (6) Model training and result analysis

[0157] 6.1) Experimental Platform Setup and Training Details

[0158] The experiments in this embodiment were conducted in a high-performance computing environment to ensure efficient training and inference of the dual-branch vehicle intent recognition model based on multimodal spatiotemporal fusion. The experiments used an Intel Xeon Gold 6226R@2.90GHz CPU, an NVIDIA RTX 3090 (24GB VRAM) GPU, and 128GB of DDR4 memory, and trained the model using Python 3.8, PyTorch 1.10.1, and CUDA 11.3.

[0159] After data preparation, this study input the pre-defined training, validation, and test sets into the bi-branch spatiotemporal model for training, and set key training parameters to ensure efficient convergence and generalization performance. Specifically, the training process used a batch size of 128, which fully utilizes the parallel computing power of the GPU and helps stabilize gradient updates; the optimizer used was Adam, with an initial learning rate of 1×10⁻⁶. -31 β1 = 0.9 and β2 = 0.999 are used to achieve fast convergence in non-convex depth networks.

[0160] In addition, to avoid overfitting and getting stuck in local optima, the learning rate will be automatically reduced to 0.1 times its original value when the performance on the validation set does not improve significantly within 5 consecutive training epochs. The entire training process will run for a maximum of 50 epochs and will employ an early stopping strategy, terminating training when the model performance tends to converge on the validation set, while retaining the optimal model parameters.

[0161] 6.2) Baseline Model Selection and Evaluation Indicators

[0162] 6.2.1) Baseline Model Selection

[0163] The bi-branch spatio-temporal fusion-based intent recognition model proposed in this embodiment is hereinafter referred to as Bi-STFM. To verify the effectiveness of the proposed model, this study conducts comparative experiments on the following three baseline models:

[0164] 1. LSTM: Only LSTM is used to process time series information, without considering spatial interactions between vehicles, to measure the effectiveness of pure time series modeling.

[0165] 2. GCN model: It can capture spatial relationships in graph structure data, but its ability to capture temporal features is relatively limited. It is suitable as a benchmark for modeling individual spatial features.

[0166] 3. ST-GCN Model: The Spatio-Temporal Graph Convolutional Network (ST-GCN) combines spatial graph convolution with temporal convolution operations and is widely used in prediction tasks in dynamic interactive environments.

[0167] The models described above encompass various feature modeling mechanisms, including pure temporal feature modeling, pure spatial feature modeling, and a simple fusion of spatiotemporal features. The structural differences between these models and the Bi-STFM model proposed in this paper fully demonstrate the contribution of model structural innovation to performance improvement.

[0168] 6.2.2) Model Evaluation Indicators

[0169] In multi-class intent recognition scenarios, to comprehensively evaluate the performance of each model, this study uses the following evaluation metrics: accuracy, recall, and F1 score. For ease of explanation, we assume there are k intent labels, denoted as TP. k For the number of true positive samples in class k (i.e., the true class is k and the predicted class is also k), FP k FN represents the number of false positive samples (the true class is not k but the predicted class is k). kThis represents the number of false negative samples (the true class is k but the predicted class is another class), ignoring TN. k (Weakly correlated with the definition of each category in a multi-classification scenario).

[0170] 1) Accuracy: In multi-class classification, accuracy can be defined as the ratio of correctly predicted samples to the total number of samples. If N is the total sample size, then the more intents the model correctly identifies, the higher the accuracy. See formula (18):

[0171]

[0172] 2) Recall: For a specific category k, recall represents the proportion of samples correctly predicted as k among all samples with the true label k. Its physical meaning is: for a certain type of driving intention, such as speeding up or cutting in front of another vehicle, whether the model can capture as many positive examples as possible from all samples where that intention actually exists. See formula (19):

[0173]

[0174] 3) Precision: For class k, precision represents the percentage of samples predicted as k that are actually positive examples belonging to class k. A higher precision means that the model rarely misclassifies a particular intention. See formula (20):

[0175]

[0176] 4) F1 coefficient: F1 is the harmonic mean between precision and recall, used to balance the two to obtain the overall F1 performance, and is used to measure the balance of the model's recognition under different driving intentions. See formula (21):

[0177]

[0178] 5) Security: Confusion matrix M∈R 4×4 element M ij This represents the probability that the true intention is category i but the prediction is category j. Under this definition, we consider predictions along the diagonal and below (i.e., i≥j) as safe predictions, while predictions above the diagonal (i.e., i<j) are considered unsafe. This indicator design can intuitively reflect that when the model makes a conservative prediction, the overall safety value is higher. Therefore, it can be defined as the safety probability, as shown in formula (22):

[0179]

[0180] 6.3) Comparative Analysis of Intent Recognition Results

[0181] This embodiment selected a 15-second time window for vehicle intent recognition, with a recognition frequency of 0.1Hz per instance, and could predict the intent of approximately 10 vehicles per instance. Table 1 shows the collected and organized data, which analyzes the performance differences of each model in detail and explains the reasons for the performance improvement in conjunction with differences in model structure.

[0182] 1) Overall performance comparison and analysis

[0183] Firstly, as shown in Table 1 regarding overall performance metrics, the Bi-STFM model outperforms the LSTM, ST-GCN, and GCN baseline models in recall, precision, F1 score, accuracy, and security. For example, in recognizing the intent to "accelerate and cut in line," the Bi-STFM model achieved an F1 score of 92.98%, higher than LSTM's 86.82%, ST-GCN's 91.34%, and GCN's 88.63%. This difference demonstrates the overall performance advantage of the Bi-STFM model in intent prediction. To more intuitively illustrate the performance metrics of each model, the data in the table has been processed and converted into a radar chart for a clearer comparison of the performance of each model across different evaluation metrics. See [link to table]. Figure 8 As shown.

[0184] Table 1 Comparison of Algorithm Model Performance

[0185]

[0186]

[0187] from Figure 8 The radar chart comparison shows that spatiotemporal fusion methods (Bi-STFM and ST-GCN) generally exhibit higher prediction accuracy compared to models that use temporal or spatial features alone. In particular, the Bi-STFM model demonstrates higher prediction accuracy and stronger safety by fusing temporal and spatial features. Compared to traditional LSTM and GCN models, Bi-STFM performs better in complex environments. This aligns with the multimodal spatiotemporal feature fusion idea proposed in Section 3.3 of this paper: combining LSTM and attention mechanisms to extract temporal features, supplementing with GraphSAGE and GAT to construct refined spatial interaction features, and adaptively integrating the two types of features through a gating fusion strategy, enabling the model to capture dynamic driving intentions in complex environments more comprehensively and meticulously.

[0188] 2) Confusion matrix analysis for driver intent recognition

[0189] The confusion matrix visually displays the prediction outcomes between different categories. By observing the degree of confusion between different categories, we can gain a deeper understanding of the model's ability to discriminate between different driving intentions. Figure 9As can be seen, there are significant differences in classification errors among the models: the diagonal value of the Bi-STFM model is significantly higher than that of other models (e.g., Figure 9 As shown in (a), the values ​​of the non-diagonal elements are extremely low, indicating that the model effectively reduces the misclassification rate between categories, especially in important behaviors such as accelerating to cut in and changing lanes to avoid obstacles. LSTM (as shown in (a)) Figure 9 As shown in (b)) and GCN (as shown in...) Figure 9 The model shown in (d) exhibits significant confusion when predicting "lane-changing avoidance" and "accelerating to overtake" intentions. This reflects the inability of single-feature modeling methods (temporal or spatial only) to fully capture the complexity of vehicle intentions.

[0190] This significant difference further validates the effectiveness of the spatial feature dynamic adaptive attention mechanism and the temporal feature multi-head self-attention mechanism adopted by Bi-STFM: this structure can capture richer vehicle interaction and motion features, and improve the ability to distinguish different driving intentions.

[0191] 3) F1 fractional box plot analysis

[0192] Figure 10 Box plots of the F1 scores for each model in multiple experiments are shown. It can be seen that, due to the lack of complementary multidimensional information, the distribution of the F1 score for the single model fluctuates significantly, with a large gap between the upper and lower quartiles, indicating poor performance stability under different operating conditions. In contrast, the Bi-STFM model exhibits a higher F1 score in multiple experiments, with a narrower range between the upper and lower quartiles, indicating that this model has higher stability and robustness.

[0193] In summary, through the selection of baseline models, the construction of evaluation indicators, and the comprehensive comparative analysis of experimental results, the superior performance of the Bi-STFM model proposed in this study in the driving intention recognition task has been fully verified, demonstrating the innovation and effectiveness of the model structure design proposed in Section 3.3.

[0194] (1) Single temporal or spatial models have certain limitations due to their single source of information, making it difficult to cope with the uncertainties brought about by multi-vehicle interactions in complex traffic scenarios.

[0195] (2) The proposed dual-branch prediction network based on multimodal spatiotemporal fusion effectively integrates temporal dynamics and spatial interaction features, which greatly improves the accuracy and robustness of driving intention prediction.

[0196] (3) The synergistic effect of the gating fusion mechanism and the overall network architecture is the key to achieving performance breakthroughs. This is not only reflected in the confusion matrix and F1 score box plot, but also provides a direction for improvement for subsequent model optimization.

[0197] Example 2

[0198] This invention also provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set, or an instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to achieve the following: Figure 1 This paper presents a method for predicting the intention of vehicles around the road based on multimodal spatiotemporal fusion.

[0199] It is understood that the memory may include random access memory (RAM) or read-only memory. Optionally, the memory may include non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, code, code sets, or instruction sets. The memory may include a stored program area and a stored data area, wherein the stored program area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the stored data area may store data created according to the use of the server, etc.

[0200] A processor may include one or more processing cores. The processor connects to various parts of the server via various interfaces and lines, executing instructions, programs, code sets, or instruction sets stored in memory, and accessing data stored in memory to perform various server functions and process data. Optionally, the processor may be implemented using at least one of the following hardware forms: Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor may integrate one or more of the following: Central Processing Unit (CPU) and Modem. The CPU primarily handles the operating system and applications; the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor.

[0201] Since this electronic device is an electronic device corresponding to the multimodal spatiotemporal fusion-based vehicle intention prediction method in this embodiment of the invention, and the principle of solving the problem by this electronic device is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0202] Example 3

[0203] This invention also provides a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to achieve the following: Figure 1 This paper presents a method for predicting the intention of vehicles around the road based on multimodal spatiotemporal fusion.

[0204] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0205] Since the storage medium is the storage medium corresponding to the multimodal spatiotemporal fusion-based vehicle intention prediction method in this embodiment of the invention, and the principle of the storage medium in solving the problem is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.

[0206] Example 4

[0207] In some possible implementations, various aspects of the methods of the embodiments of the present invention can also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps of a multimodal spatiotemporal fusion-based vehicle intent prediction method according to various exemplary embodiments of this application as described above. The executable computer program code or "code" for performing the various embodiments can be written in high-level programming languages ​​such as C, C++, Python, Smalltalk, Java, JavaScript, Visual Basic, Structured Query Language (e.g., Transact-SQL), Perl, or in various other programming languages.

[0208] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0209] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0210] The above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made based on the essence of the content of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A method for predicting vehicle intentions based on multimodal spatiotemporal fusion, characterized in that, Includes the following steps: Obtain historical trajectory data; The obtained historical trajectory data is input into the trained vehicle intention prediction model, and the prediction result is output. The vehicle intent prediction model includes two parallel feature extraction branches: one is a temporal feature extraction branch, which uses a long short-term memory network to encode the historical trajectory of each vehicle and dynamically captures the temporal features of the trajectory through an attention mechanism; the other is a spatial feature extraction branch, which uses a GraphSAGE network to model the spatial interaction relationship between multiple vehicles, and then combines the attention mechanism of a graph attention network to assign adaptive weights to different vehicles, so that the model can highlight the importance of key interfering vehicles. The model's fusion module employs a gated fusion mechanism to adaptively weight and fuse features from the temporal feature extraction branch and the spatial feature extraction branch; the fused feature information is then input into the decoder module, which outputs the probability distribution of various driving intentions of the conflicting vehicles at future time points. The input to the fusion module includes: feature vectors extracted by the temporal feature extraction branch. Feature vectors extracted by the spatial feature extraction branch and the feature vectors generated by the residual network. ; The three feature vectors are concatenated to obtain the original feature representation before fusion. : The gating factor is obtained through a fully connected layer and a Softmax activation function: In the formula, , These are the weights and bias parameters to be learned in the network, and the Softmax operation guarantees... The obtained gating factor is used to weight and fuse the three features to obtain the fused features: The fused features are then normalized using a batch normalization layer. In the formula, E[ ] represents the mean of the feature, Var[ ] represents the variance of the feature. To prevent decimals from being divided by zero, , These are learnable scaling and translation parameters.

2. The method for predicting vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that, The temporal feature extraction branch includes an LSTM feature extraction module, a multi-head attention module, and a Dropout module; In the LSTM feature extraction module, an independent LSTM encoder is built for the trajectory of each vehicle, and the input of each LSTM encoder is one-dimensional time-series trajectory data. After LSTM encoding, the motion pattern features of the vehicle trajectory are obtained. ; In the multi-head attention module, the motion pattern features of each vehicle are... As the attention query vector, the LSTM output features of all vehicles are used together as keys and values ​​to calculate the contextual features of temporal interactions between vehicles. ; Features processed by the multi-head attention module are then processed by the Dropout module to prevent overfitting of the network.

3. The method for predicting vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that, In the spatial feature extraction branch, a graph structure is used to represent the interaction relationships between vehicles. Vehicles are considered nodes, and the interaction relationships between vehicles are represented as edges connecting these nodes. ,in It is a set of vehicle nodes. It is the set of edges that define the relationships between vehicles; Each node in the node set Corresponding to a feature vector The feature vector includes the vehicle's relative position, velocity, and acceleration at the current moment; The initial feature matrix of the node is calculated. : 。 4. The method for predicting vehicle intention based on multimodal spatiotemporal fusion according to claim 3, characterized in that, The GraphSAGE network operates as follows: For nodes Its set of neighboring nodes Indicates all with Connected nodes, a K-layer GraphSAGE network is built on each node, and the computation of the k-layer nodes is performed. v Features First, the features of neighboring nodes sampled at layer k-1 are aggregated using an aggregation function. Perform aggregation to obtain the aggregation result. Then combine it with the node v The features at the (k-1)th layer are concatenated into vectors, and finally nodes are generated through nonlinear mapping. v The feature matrix of the k-th layer .

5. The method for predicting vehicle intention based on multimodal spatiotemporal fusion according to claim 3, characterized in that, The graph attention network works as follows: An adjacency matrix is ​​constructed based on the relative relationships between nodes in a spatiotemporal trajectory, and its elements are denoted as... Represents a node With nodes Does an interaction relationship exist? Let if any of the following conditions are met, Otherwise, set it to 0: 1) Spatial proximity at the same time: If the Euclidean distance between two different nodes at the same time is less than a preset threshold If so, it is considered that the two objects are close to each other at the same point in time and there is a possibility of direct interaction; 2) Continuous trajectories of the same object at adjacent time points: if the node index If the timestamps are identical and differ by 1, then these two nodes are considered as extensions of the trajectory of the same physical object across consecutive time steps, thus revealing the node interaction relationships. In this way, the connection relationships between nodes change over time and with each other, allowing us to obtain the adjacency matrix of the edges. A ; The feature matrix of the node H Adjacency matrix of edges A As input to the graph attention network, through the feature matrix H and adjacency matrix A The attention coefficients between nodes are calculated, and after processing by the LeakyReLU activation function, the attention coefficients of each pair of neighboring nodes are normalized by the softmax function. The node's feature matrix is ​​analyzed using the normalized attention coefficient matrix. H We perform weighted summation to obtain the updated feature representation of the nodes, and finally obtain the node feature matrix after attention mechanism processing.

6. The method for predicting vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that, The decoder module consists of fully connected layers. Through layer-by-layer linear mapping and non-linear activation function processing, it gradually abstracts high-dimensional feature vectors into low-dimensional feature representations corresponding to the number of driving intention categories. The calculation method for each fully connected layer is as follows: After processing through multiple fully connected layers, the feature dimension is gradually reduced and further abstracted, ultimately outputting a vector with the same dimension as the number of intent categories: Each dimension corresponds to the probability of the vehicle's possible driving intentions during the week: This indicates that you should change lanes to give way. Indicates slowing down and proceeding straight. Indicates uniform speed in a straight line. Indicates accelerating straight ahead; The above output is normalized using the Softmax function: In the formula, The vehicles representing the conflict may take the first step in the future. i The probability of a certain driving intention. For the network to the first i The raw output value of the intent category.

7. The method for predicting vehicle intention based on multimodal spatiotemporal fusion according to claim 1, characterized in that, The historical trajectory data includes the vehicle position, speed, and acceleration of the main vehicle and the vehicle involved in the collision; The acquisition of historical trajectory data includes: Historical vehicle trajectories over a preset time period are selected as the model's raw input data. : In the formula, , Representing vehicles i At any moment t The horizontal and vertical coordinate positions, and The vehicles at the time t The velocity and acceleration.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one instruction, at least one program, a code set, or an instruction set, the at least one instruction, the at least one program, the code set, or the instruction set being loaded and executed by the processor to implement the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the method as described in any one of claims 1 to 7.