A microscopic simulation system of interweaving area based on space-time graph Transformer
Through the interleaving area microscopic simulation system based on the space-time graph Transformer, an interleaving scene diagram is constructed and a graph attention neural network is used to predict driving intentions and vehicle trajectories, which solves the problem of failure to fully consider the road environment, driver awareness and vehicle interaction in the prior art, and achieves a more accurate and migable vehicle trajectory prediction.
Patent Information
- Application Number
- CN202411361792.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-09-27
AI Technical Summary
Existing research fails to fully consider the coupling effects of road environment, driver awareness and interaction with surrounding vehicles when predicting the trajectory of autonomous vehicles in interleaved segments, and has significant limitations in situational perception and interpretability.
Using an interleaving area microscopic simulation system based on the space-time graph Transformer, the interleaving scene graph is constructed, the road environment, driver awareness and vehicle motion characteristics are integrated, and the graph attention neural network and multi-head self-attention mechanism are used to predict driving intentions and vehicle trajectories.
It realizes the dynamic changes of interleaved scenes more accurately, and improves the refinement of traffic scene modeling and the accuracy and migability of vehicle trajectory prediction.
Smart Images

Figure CN119312673B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to an interweaving area microscopic simulation system based on a space-time graph Transformer. Background Art
[0002] Safety assurance in complex scenarios has always been a top priority in the development of autonomous driving technology. Especially in weaving sections, vehicles enter or leave the highway through adjacent entrance ramps and exit ramps, with large speed changes and frequent lane changes in short distances, leading to serious traffic conflicts and making weaving sections a bottleneck in the highway system. Considering the inherent complexity of weaving scenarios, autonomous vehicles must fully understand the behavior of other traffic participants in order to make wise driving decisions. Therefore, accurately predicting vehicle trajectories in weaving sections is of great significance to improving road safety and advancing autonomous driving technology.
[0003] In the weaving section, the road environment, driver awareness, and interaction with surrounding vehicles will affect the vehicle trajectory, making the vehicle driving scene more complicated. On the one hand, after entering the weaving section, the vehicle usually changes lanes immediately. When the traffic density is low, the lane change positions are more distributed upstream of the weaving section. In a shorter weaving section, the limited length will increase the density of lane changes, making the merging and diverging tasks of vehicles in the weaving section more concentrated. The number of lanes in the weaving section will affect the space for weaving vehicles and non-weaving vehicles, thereby inhibiting the selective lane changes of all vehicles. The restrictions on driving tasks and road section length may produce more diverse driving behaviors, such as free driving, following vehicles, selective lane changes, forced lane changes, and active lane changes. Frequent lane changes of weaving vehicles need to adapt to the interaction with surrounding vehicles by adjusting speed and lane position to ensure reasonable spacing. In addition, bad weather, poor light, and slippery roads may reduce the driver's visibility and the vehicle's braking performance, leading to more cautious driving behavior. It can be seen that the road-driver-vehicle system is complex and dynamic, and each component may affect and interact with other components. However, existing research mainly focuses on the impact of a single factor and ignores the dynamic impact and relationship between factors.
[0004] On the other hand, the vehicle trajectory will also change significantly after the driver decides to execute his driving intention. At present, many methods have been used to predict driving intention, such as physical rule models, probabilistic graphical models, and deep learning models. Some models regard the driver's decision as a deterministic process and infer the driving intention by introducing the probability of lane change. The adaptive multi-dimensional continuous Gaussian mixture hidden Markov model can use steering wheel angle, steering wheel angular velocity and lateral acceleration to identify dynamic driving intention. In addition, some deep learning methods are also widely used to predict driving intention. For example, a residual neural network-based model is established to identify driving intention by using road slope, curvature and vehicle dynamics characteristics, and a dual converter network is applied to predict driving intention by considering the dependency of the historical trajectories of surrounding vehicles in mixed traffic environments. There are also studies that use a two-way trajectory contrast learning model to learn a generalized trajectory representation to improve the performance of driving intention prediction. However, intention refers to the driver's thoughts before taking action. Existing studies have not fully studied how drivers perceive and make decisions in driving scenarios, and the decision-making process of these models lacks interpretability.
[0005] In addition, most studies use historical position information to predict future vehicle trajectories through data-driven models, but they rarely consider the significant impact of situational awareness on vehicle trajectories. Although some methods attempt to use generative models to generate multimodal vehicle trajectories, generative models rely on short-term historical position information, which may lead to significant randomness and uncertainty in the learned long-term trajectory distribution. Another approach is to use reinforcement learning techniques to generate driving strategies by minimizing the feature expectation between expert demonstrations and predicted trajectories, but there is no vehicle trajectory prediction model suitable for weaving segments. Situational awareness refers to the perception and understanding of elements in the driving scene (i.e., roads, drivers, and vehicles) in a specific time and space, including driving intentions and spatiotemporal information related to the driving scene. Drivers need to quickly understand the rapid changes in the driving scene to make correct decisions and take appropriate actions. However, relying solely on historical vehicle trajectory data may not fully capture all situational awareness information, resulting in limited application of existing models in weaving segments.
[0006] In summary, most existing studies focus on analyzing location information based on historical vehicle trajectories, but fail to fully consider the coupling effects of road environment, driver awareness, and interaction with surrounding vehicles; moreover, in interweaving scenarios, existing studies have significant limitations in situational awareness and interpretability. Therefore, there is an urgent need for a model that can use real-world trajectory data to predict driving intentions and vehicle trajectories, thereby improving the accuracy and transferability of predictions. Summary of the invention
[0007] The purpose of the present invention is to provide a weaving area micro-simulation system based on a spatiotemporal graph Transformer, which can accurately characterize the dynamic changes of weaving scenes, realize refined traffic scene modeling, and enhance the accuracy and transferability of vehicle trajectory prediction.
[0008] To achieve the above object, the present invention provides a microscopic simulation system of an interlaced area based on a spatiotemporal graph Transformer, comprising the following steps:
[0009] The interweaving scene graph construction module constructs an interweaving scene graph according to the road environment characteristics, driver awareness characteristics, vehicle motion characteristics and vehicle interaction characteristics;
[0010] The driving intention prediction module obtains the situational awareness features of driving intention based on the interwoven scene graph through information transfer in the graph attention neural network and the multi-head self-attention mechanism;
[0011] The vehicle trajectory prediction module predicts the vehicle trajectory through the spatiotemporal graph Transformer network based on comprehensive factors such as road environment, driving intention and interaction with surrounding vehicles.
[0012] Preferably, road environment characteristics include weather, lane, length, speed limit and trajectory restriction; driver awareness characteristics include motivation and intention; vehicle motion characteristics include timestamp, vehicle code, vehicle model, lane, remaining distance, speed, acceleration and position; vehicle interaction characteristics include speed difference and spacing.
[0013] Preferably, the interwoven scene graph is a non-Euclidean data structure consisting of nodes and edges, wherein nodes represent real-world entities, edges represent relationships between entities, and the features of each node and edge describe their respective properties.
[0014] Preferably, entities of the interwoven scene graph include scenes, roads, vehicles and drivers, and relationships include belonging to, located in, interacting with and influencing.
[0015] Preferably, the interleaved scene graph is represented by an undirected graph, which is described as G = (V, E, A);
[0016] Where V = {v 1 ,…,v m} is a node set, E={e 1 ,…,e n} is the edge set and A is the adjacency matrix.
[0017] Preferably, the situational awareness features include the driver's understanding and cognition of the road environment, vehicle movement, interaction with surrounding vehicles, and the resulting driving intention decisions.
[0018] Preferably, the driving intention includes lane keeping, lane changing left and lane changing right.
[0019] Therefore, the present invention adopts the above-mentioned interlaced area microscopic simulation system based on the spatiotemporal graph Transformer, which has the following technical effects:
[0020] (1) By constructing an interweaving scene graph to integrate and quantify the driving environment and driver awareness, it is possible to more accurately characterize the dynamic changes of the interweaving scene and achieve refined traffic scene modeling.
[0021] (2) By introducing a context-aware enhanced spatiotemporal graph Transformer network, it is possible to generate vehicle trajectory predictions and traffic in weaving sections and simulate the driver's cognitive and decision-making process in complex driving scenarios, thereby improving the accuracy and transferability of predictions.
[0022] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a framework diagram of the interweaving zone micro-simulation system based on the space-time graph Transformer;
[0024] Figure 2 The figure is a schematic diagram of a weaving scene on a highway in an embodiment of a weaving area microscopic simulation system based on a spatiotemporal graph Transformer, wherein: Figure 2 (a) is an interlaced scene task. Figure 2 (b) is a straight-through motion. Figure 2 (c) is lane change;
[0025] Figure 3 The present invention is a schematic diagram of constructing an interweaving scene graph in an embodiment of an interweaving area microscopic simulation system based on a spatiotemporal graph Transformer, wherein: Figure 3 (a) is an interlaced scene graph. Figure 3 (b) is the bird’s-eye view from the first perspective of the ego vehicle;
[0026] Figure 4 It is a top view of a weaving section of expressway A in the embodiment of the weaving area microscopic simulation system based on the space-time graph Transformer. Figure 4 (a) is the data collected by the drone. Figure 4 (b) is the extracted vehicle schematic diagram. Figure 4 Middle (c) is an enlarged view of the interweaving area;
[0027] Figure 5 is a typical lane keeping and lane changing process in an embodiment of a weaving area microscopic simulation system based on a spatiotemporal graph Transformer, wherein: Figure 5 (a) is the position diagram of lane change. Figure 5 (b) is the lateral speed of lane change, Figure 5 (c) is the position diagram of lane keeping. Figure 5 (d) is the longitudinal speed of lane keeping;
[0028] Figure 6 is a ROC curve for driver intention prediction in an embodiment of a weaving area micro-simulation system based on a spatiotemporal graph Transformer, where: Figure 6 (a) is the ROC curve of each task type predicted by the CNN model. Figure 6 (b) is the ROC curve of each task type predicted by the LSTM model. Figure 6 (c) is the ROC curve of each task type predicted by the Transformer model. Figure 6 (d) is the ROC curve of each task type predicted by the SR-LSTM model. Figure 6 (e) is the ROC curve of each task type predicted by the STAR-Transformer model. Figure 6 (f) is the ROC curve of each task type predicted by the SASTGT model;
[0029] Figure 7 The figure is a vehicle trajectory prediction result diagram of the weaving section A of the expressway in the embodiment of the weaving area microscopic simulation system based on the spatiotemporal graph Transformer, wherein: Figure 7 (a) is free driving, Figure 7 (b) is following a car. Figure 7 (c) is to slow down and prepare for lane change. Figure 7 (d) is a random lane change. Figure 7 (e) in the middle is active lane change and merging. Figure 7 (f) is a forced lane change;
[0030] Figure 8 It is a vehicle trajectory prediction result under various traffic conditions on Guangdong Shuiguan Expressway in the embodiment of the weaving area micro-simulation system based on the spatiotemporal graph Transformer, in which: Figure 8 (a) is the vehicle trajectory prediction result during the morning peak period. Figure 8 (b) shows the vehicle trajectory prediction results during free flow. DETAILED DESCRIPTION
[0031] The present invention can be explained in more detail by the following examples. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following examples.
[0032] like Figure 1 As shown in FIG. 1 , the present invention provides a weaving area microscopic simulation system based on a spatiotemporal graph Transformer, which includes three modules: weaving scene graph construction (perception), driving intention prediction (decision-making), and vehicle trajectory prediction (action). Figure 2 As shown in , in weaving sections, vehicles usually need to change lanes to complete their driving tasks. Figure 2 As shown in (a), in the weaving link, the main driving tasks of the vehicle include through movement, that is, the vehicle maintains a straight path on the main line; merging, that is, the vehicle enters the main line from the ramp and merges with the vehicles on the main line; diverging, the vehicle leaves the main line and enters the ramp. Once the driving tasks are determined, the driver can generate consistent driving motivation. Figure 2 (b) and Figure 2 As shown in (c), according to the change of vehicle position, the driving motivation is discretized into specific driving intentions, including lane keeping, lane change left (lane change L) and lane change right (lane change R), and the local coordinate system of the weaving segment is also defined, where the X-axis represents the longitudinal direction along the road and the Y-axis represents the normal direction (lateral direction) along the road. The length of the weaving segment is defined as the distance from the end of the ramp acceleration lane to the start of the ramp deceleration lane.
[0033] (1) In order to construct a chart that can better describe the elements of the weaving scene and their interactions, this embodiment uses input variables from four sources, including road environment characteristics (such as weather, weaving section length, speed limit, etc.), driver awareness characteristics (such as driving motivation, driving intention), vehicle motion characteristics (such as timestamp, vehicle type, position, speed, acceleration, etc.), and interaction characteristics with surrounding vehicles (for example, gap, speed difference), as shown in Table 1. In addition, this embodiment also uses the minimum-maximum normalization method to normalize continuous input variables to eliminate differences in numerical ranges.
[0034] Table 1 Summary description of input variables
[0035]
[0036] Since the input variables contain different types of information (for example, some are related to road environment features and others are related to vehicle features), this information may come from different vehicles. In an interweaving scenario, it is difficult to form an effective connection between these heterogeneous features. To solve this problem, in an interweaving scenario, this embodiment organizes all features into a graph structure and adopts a top-down approach to effectively integrate context information to describe the relationship between input variables, such as Figure 3 shown.
[0037] The interweaving scene graph is a non-Euclidean data structure composed of nodes and edges, where nodes represent entities in the real world (e.g., roads, vehicles, and drivers), and edges represent the relationships between these entities. The relationships established by nodes and edges enable effective modeling of heterogeneous data in interweaving scenes. As shown in Table 2, the interweaving scene graph in this embodiment includes four entity types and relationships, each of which has certain attribute characteristics, and the characteristics of each node and edge describe their respective attributes. By utilizing the connectivity and scalability of the graph topology, the dynamics and complexity of the interweaving scene can be effectively described, including the changing states and interactions of vehicles, roads, and drivers.
[0038] Table 2 Entities and relations in the interleaved scene graph
[0039]
[0040] According to the direction of the edge, the graph structure can be divided into directed graph and undirected graph. Considering the interaction between roads, drivers and vehicles, this embodiment uses an undirected graph to represent the interweaving scene. The interweaving scene graph can be described as G = (V, E, A), where V = {v 1 ,…,v m} is a node set, E={e 1 ,…,e n} is an edge set. The node attributes are expressed as Where m is the number of nodes and d is the dimension of the node features. Similarly, the edge attributes are expressed as Where q is the number of edges and s is the dimension of the edge element. In the interwoven scene graph, the topological relationship between entities (i.e., roads, vehicles, and drivers) is represented by the adjacency matrix A = [a ij ] m×q It indicates that,
[0041]
[0042] Where i and j are the indices of the nodes.
[0043] (2) In order to improve the model's cognitive and decision-making capabilities in complex weaving sections, it is crucial to understand the driver's situational awareness and predict their driving intentions. Therefore, this embodiment uses an interpretable graph attention neural network (GAT) to establish a driving intention prediction module in weaving scenarios, which can simulate the driver's cognitive and decision-making process and enhance the situational awareness capabilities in complex driving scenarios.
[0044] The independent variable of the driving intention prediction module is the interwoven scenario graph, while the dependent variable is the situational awareness features including driving intention (i.e., lane keeping, lane change L, and lane change R). By utilizing the information transfer and multi-head self-attention mechanism in GAT, it can better simulate the driver's cognitive and decision-making process and improve the driving intention prediction performance.
[0045] The driving intention prediction module includes two GAT layers and a fully connected layer. The GAT layer uses a multi-head self-attention mechanism to capture the dynamic interactions between road, driver, and vehicle nodes. The fully connected layer calculates the probabilities of all possible driving intentions and outputs higher-level situational awareness features. Situational awareness features include the driver's understanding and cognition of the road environment, vehicle motion, interaction with surrounding vehicles, and the resulting driving intention decisions.
[0046] (3) Vehicle trajectory is essentially sequence data containing temporal and spatial position information. After the driving intention is predicted, in order to better consider the coupling effects of the road environment, driving intention, and interaction with surrounding vehicles, in the vehicle trajectory prediction module, this embodiment uses graph embedding and multi-head self-attention mechanism in the spatiotemporal graph transformer network to construct a spatiotemporal graph transformer network to predict the vehicle trajectory. The situational awareness function is introduced in this module, and the interweaving scene graph is regarded as an independent variable, and the future position of the vehicles involved in the interweaving scene is taken as the dependent variable, which can better capture the spatiotemporal dependence of the vehicle trajectory on situational awareness and improve the accuracy of vehicle trajectory prediction.
[0047] The spatiotemporal graph transformer network consists of four parts, including a graph embedding layer, a graph transformer layer, an external graph memory layer, and a fully connected layer. The graph embedding layer integrates the position features in the input interleaved graph with the situational awareness features and converts them into a graph embedding. In the graph transformer layer, the spatial transformer network extracts the spatial interactions between surrounding vehicles, while the temporal transformer network independently captures the temporal dependencies of individual vehicles. The external graph memory layer smoothes vehicle trajectories by matching historical position embedding information. The fully connected layer outputs the predicted positions of all vehicles in the future time interval Δt.
[0048] Experimental verification
[0049] In order to train and evaluate the weaving area micro-simulation model based on the spatiotemporal graph Transformer, this embodiment uses the "Highway-A" dataset in Chengdu, China, which includes three basic lanes, auxiliary lanes, entrance and exit ramps in each direction, the speed limits of the main line and ramps are 100 km / h and 40 km / h respectively, the length of the weaving section is 183.633m, and the width of each lane is 3.75m. Figure 4 shown.
[0050] Since the data of Highway A was collected using two drones simultaneously, covering a 250-meter basic highway section, a total of 22,306 vehicle trajectories were extracted from 315.2 minutes of drone videos, which were taken at different time periods and in both directions, and the weather during the trajectory collection was also recorded, including sunny and rainy days. In order to reduce the impact of data noise, the moving average method was used to smooth the trajectory and speed, and the filter was set to 15 frames, corresponding to a duration of 0.5 seconds.
[0051] After acquiring the data, lane keeping and lane changing events are extracted. Figure 5 (a) shows a typical lane change process and sets three time points to locate lane change events, namely lane change decision, lane change execution and lane change completion. Lane change decision is defined as the time point when the vehicle's lateral speed exceeds 0.1m / s (or -0.1m / s) for the first time; lane change execution refers to the time when the vehicle's lateral speed begins to increase continuously; lane change completion refers to the time when the vehicle changes lane position, that is, the lateral speed value becomes 0. Compared with lane change events, lane keeping events are characterized by minimal lateral movement, such as Figure 5 (b) After the lane change event is extracted, other trajectories with a duration of more than 3.0s are marked as lane keeping events. At the same time, the trajectory data of vehicles around the lane change and lane keeping vehicles are also extracted.
[0052] Existing studies have shown that a duration of 1.0s is a suitable duration for the driver to make a new decision, indicating that during this time interval, the driver's intention and the vehicle trajectory show consistency. In order to balance trajectory feature learning and driver decision time, in this embodiment, the input time window and prediction time window are set to 60 time steps (2.0s) and 30 time steps (1.0s), respectively, and the time step of the model is set to 1 / 30 second, allowing for the capture of finer-grained motion features. Moreover, in order to ensure continuous prediction, this embodiment uses a sliding window mechanism combined with a forward propagation method. Specifically, the prediction results are connected to the historical trajectory data, and the input window is updated by sliding forward 30 time steps.
[0053] In order to quantitatively evaluate the performance of the system, this embodiment uses the area under the curve (AUC) and accuracy (ACC) indicators to evaluate the driving intention prediction, and uses the average displacement error (ADE) and final displacement error (FDE) to evaluate the vehicle trajectory prediction.
[0054] Among them, the AUC indicator evaluates the model's ability to distinguish different categories by measuring the area under the receiver operating characteristic (ROC) curve; ADE is used to measure the mean square error of the overall estimated position in the predicted trajectory and the ground truth trajectory; and FDE quantifies the final position difference of these trajectories. The calculation formulas for each indicator are as follows:
[0055]
[0056] Where TP, FN, FP, and TN represent the number of true positives, false negatives, false positives, and true negatives, respectively, and nt is the number of time steps. is the predicted position at time step t, is the true position at time step t, is the predicted position at the last time step, is the true position of the last time step, and ‖·‖ represents the Euclidean distance.
[0057] This embodiment predicts the driving intention (classification task) and vehicle trajectory (regression task) of a certain interweaving section of Highway A by integrating the graph attention neural network and the spatiotemporal graph Transformer framework (i.e., the SASTGT model), thereby simulating the driver's perception, decision-making, and action, and compares them with existing SOTA models (such as CNN, LSTM, Transformer, SR-LSTM, and STAR-Transformer).
[0058] The SASTGT model uses interlaced scene graphs as input, and other SOTA models use variable sequences as input, and divide the vehicle trajectory samples into training sets (80%) and test sets (20%). The test results are shown in Table 3. It can be seen that the driving intention accuracy of the SOTA model on the test data set is 91.5%, 89.9%, 91.2%, 92.4% and 93.5%, respectively, which are much lower than the 97.4% of the SASTGT model. Furthermore, in order to more comprehensively analyze the performance of each model, this embodiment also compares the operating characteristic (ROC) curves of the intention prediction of each category of subjects, as shown in Table 3. Figure 6 As shown. According to the ROC curves of each category, the AUC values of the SASTGT model in lane keeping, lane change L, and lane change R are 0.96, 0.98, and 0.99, respectively. Overall, the SASTGT model outperforms other SOTA models in predicting driving intentions, which shows that in the graph attention network, the driver node transmits messages with adjacent nodes (i.e., road and vehicle nodes), and the multi-head self-attention mechanism simulates the driver's cognition of the global road environment and local vehicle interactions, which enhances the driving intention prediction performance.
[0059] Table 3 Performance comparison of each model
[0060] Model ACC ADE / m FDE / m CNN 0.915 1.462 3.131 LSTM 0.899 0.996 1.853 Transformer 0.912 0.722 1.693 SR-LSTM 0.924 0.682 1.412 STAR-Transformer 0.935 0.491 0.931 SASTGT 0.974 0.128 0.248
[0061] The vehicle trajectory prediction results on the test dataset are as follows: Figure 7 As shown in the figure, it can be seen that the SASTGT model is significantly better in predicting vehicle trajectories, with the lowest ADE and FDE values of 0.128 and 0.248 respectively, which are much lower than other SOTA models. This shows that the spatiotemporal graph transformer network can effectively capture the spatiotemporal dependency of vehicle trajectories on situational awareness.
[0062] On the other hand, in order to prove the prediction results of each vehicle, this embodiment selects five typical interweaving scenarios from the test data set and makes a one-to-one comparison of vehicle trajectories, such as Figure 7 shown. Figure 7 (a) is free driving, Figure 7 (b) is following a car. Figure 7 (c) is to slow down and prepare for lane change. Figure 7 (d) is a random lane change. Figure 7 (e) in the middle is active lane change and merging. Figure 7 (f) is a forced lane change. The results show that in these typical weaving scenarios, the predicted trajectories of the ego vehicle and surrounding vehicles are highly consistent with the true trajectories and can capture the longitudinal and lateral motion of the vehicles, further demonstrating the superior performance of the SASTGT model in weaving segments.
[0063] In addition, in order to verify the transferability of the SASTGT model in this embodiment, the model trained with the above dataset is used to verify its applicability on other highways.
[0064] In this embodiment, the transferability verification is performed using the field observation data collected by drones during two specific time periods: the morning peak (8:45 to 9:00 a.m.) and free flow (3:30 to 3:45 p.m.) of Shuiguan Expressway in Guangdong, China in 2023. Figure 8 The length of the weaving section of the highway is 683 meters, the main line speed limit is 100 km / h, and the ramp speed limit is 40 km / h. A total of 464 vehicle trajectories were obtained, including 203 lane change L events, 97 lane change R events, and 214 lane keeping events. The prediction results are shown in Table 4.
[0065] Table 4 Prediction performance of SASTGT model under different traffic conditions
[0066] Dataset ACC ADE(m) FDE(m) During the morning rush hour 0.924 0.286 0.531 Free movement period 0.947 0.199 0.358
[0067] It can be seen from Table 4 that the SASTGT model has higher driving intention prediction accuracy and lower vehicle trajectory error during both the morning peak and free flow periods. Among them, during the morning peak, the ACC is 92.4%, the ADE is 0.286m, and the FDE is 0.531m; while during the free flow period, the ACC is 94.7%, the ADE is 0.199m, and the FDE is 0.358m. This difference may be attributed to the higher interaction density with surrounding vehicles, which reduces the driver's situational awareness and increases the uncertainty of the vehicle's trajectory. Moreover, during the morning peak and free flow periods, the trajectories continuously predicted by the SASTGT model have higher consistency with the field observations, such as Figure 8 Therefore, the SASTGT model shows excellent performance in terms of reliability and transferability.
[0068] Therefore, the present invention adopts the above-mentioned weaving area micro-simulation system based on the spatiotemporal graph Transformer, which can effectively capture the global road environment characteristics, local interactions with surrounding vehicles, and the long-term temporal dependence of driving intentions and vehicle trajectories, and realize trajectory prediction under different traffic densities and road configurations.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A microscopic simulation system of interlaced areas based on spatiotemporal graph Transformer, characterized in that: include: The interweaving scene graph construction module constructs an interweaving scene graph according to the road environment characteristics, driver awareness characteristics, vehicle motion characteristics and vehicle interaction characteristics; The entities of the interwoven scene graph include scene, road, vehicle and driver, and the relationships include belong to, located in, interact and influence; the attributes of the interactive relationship are speed, acceleration, speed difference and gap distance; The driving intention prediction module obtains the situational awareness features of driving intentions based on the interwoven scene graph through information transfer in the graph attention neural network and the multi-head self-attention mechanism; driving intentions include lane keeping, lane change left, and lane change right; The vehicle trajectory prediction module predicts the vehicle trajectory through the spatiotemporal graph Transformer network based on the comprehensive factors of road environment, driving intention and interaction with surrounding vehicles; the spatiotemporal graph Transformer network consists of four parts, including graph embedding layer, graph Transformer layer, external graph storage layer and fully connected layer; the graph embedding layer integrates the position features and situation awareness features in the input interleaved graph and converts them into graph embeddings; in the graph transformer layer, the spatial Transformer network extracts the spatial interactions between surrounding vehicles, while the temporal Transformer network independently captures the temporal dependencies of individual vehicles; the external graph memory layer smoothes the vehicle trajectory by matching the historical position embedding information; the fully connected layer outputs the predicted positions of all vehicles within the future time interval Δt.
2. According to claim 1, a microscopic simulation system of an interlaced area based on a spatiotemporal graph Transformer is characterized in that: Road environment characteristics include weather, lane, length, speed limit and trajectory restriction; driver awareness characteristics include motivation and intention; vehicle motion characteristics include timestamp, vehicle code, vehicle model, lane, remaining distance, speed, acceleration and position; vehicle interaction characteristics include speed difference and spacing.
3. The interlaced area microscopic simulation system based on spatiotemporal graph Transformer according to claim 1 is characterized in that: The interwoven scene graph is a non-Euclidean data structure consisting of nodes and edges, where nodes represent real-world entities, edges represent relationships between entities, and the features of each node and edge describe their respective properties.
4. The interlaced area microscopic simulation system based on spatiotemporal graph Transformer according to claim 1 is characterized in that: The interleaved scene graph is represented by an undirected graph, described as G = (V, E, A); Where V = {v1,…,v m } is a node set, E={e1,…,e n } is the edge set and A is the adjacency matrix.
5. The interlaced area microscopic simulation system based on spatiotemporal graph Transformer according to claim 1 is characterized in that: Situational awareness characteristics include the driver's understanding and cognition of the road environment, vehicle movement, interaction with surrounding vehicles, and the resulting driving intention decisions.