Text event and graph attention network fusion-based traffic flow generation method

By combining traffic feature data with text event information, the weighted fusion of AutoencoderKL and BERT models and graph attention network is used to solve the traffic flow prediction problems under long-term prediction and emergencies, achieving higher accuracy and stable traffic state prediction.

CN120259487APending Publication Date: 2025-07-04XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510219176.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing traffic flow prediction methods have limitations in long-term prediction and emergency perception, and it is difficult to accurately predict future traffic conditions, especially in emergencies such as traffic accidents, road construction and extreme weather.

Method used

Combining traffic feature data and external text event information, the diffusion model is trained and weighted fusion is performed to generate a feature map that conforms to the actual traffic state through AutoencoderKL codec, BERT model and graph attention network.

Benefits of technology

It improves the accuracy and stability of medium- and long-term traffic flow prediction, can dynamically adapt to emergencies, predict traffic congestion level with an error of 0.25 (MAE) and speed prediction with 4.8km/h (MAE) to provide more reliable traffic management support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259487A_ABST
    Figure CN120259487A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic flow generation method based on fusion of a text event and a graph attention network. The method comprises the following steps: carrying out data preprocessing and feature dimension remodeling on traffic feature data in a data set; traffic characteristic data are compressed and coded through a pre-trained Autoencode KL codec, and diffusion and de-noising are carried out in a potential space; converting the text event description into condition input for generating traffic characteristics by utilizing a BERT model, and gradually denoising and generating a characteristic image conforming to an actual traffic state by combining the condition text input when a diffusion model is trained; training a graph attention network model to capture spatial correlation of traffic characteristics; carrying out weighted fusion on the result generated by the diffusion model and the output of the graph attention network, and finally decoding the fused traffic characteristic data; in the model reasoning stage, a corresponding road traffic feature map is generated by inputting event texts at future moments. The method can dynamically adapt to the influence of emergencies, and is excellent in performance in long-term traffic task prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a traffic flow generation and prediction method in an intelligent transportation system (ITS), and particularly to a traffic flow generation method based on the fusion of text events and graph attention networks. Background Art

[0002] Traffic prediction is an important basic task in an intelligent transportation system (ITS), and its main goal is to predict future traffic conditions based on historical data. In recent years, deep learning methods have become the mainstream research direction in the field of traffic flow prediction and are widely used to model the spatio-temporal characteristics of traffic flow. Currently, there have been many studies on the traffic flow prediction task in the prior art.

[0003] Zhang Xu et al. proposed a short-term urban traffic flow prediction method based on hybrid convolutional LSTM in the invention patent "A Short-Term Urban Traffic Flow Prediction Method and System Based on Hybrid Convolutional LSTM" with the application number CN202111508145.9. The method includes a feature fusion module, a hybrid convolutional module, and a spatial-aware multi-attention module. The feature fusion module fuses the traffic flow map with external factors to generate a feature fusion map, and the hybrid convolutional module generates and fuses a hidden map. After upsampling, the map is processed by the spatial-aware multi-attention module and finally convolved with the traffic flow map to generate a prediction map, and the traffic flow value is obtained by inverse normalization. This method effectively extracts time and space features and combines a multi-attention module to improve the prediction accuracy. Such methods perform well in prediction tasks with a short time span (such as within 30 minutes), but still have great limitations in dealing with complex long-term dependence problems and emergencies.

[0004] Yang Guoyan proposed an improved traffic prediction method in the invention patent "A Traffic Prediction Method Based on Graph Attention Neural Network and Spatio-Temporal Big Data" with the application number CN202210638919.8. This method establishes a road network topology, uses a graph attention network (GAT) to process the spatial features of historical traffic data, and combines a long short-term memory network (LSTM) to perform temporal modeling on the spatial feature information. This method optimizes the processing of temporal features by introducing an attention mechanism, thereby enhancing the analysis of spatial correlations in the traffic road network and improving the accuracy and stability of prediction. However, such methods have poor adaptability when sudden events (such as traffic accidents, road construction, extreme weather) occur and are difficult to adjust the prediction results in real time. Therefore, the existing methods based on spatio-temporal graph attention networks still have great limitations in predicting non-typical traffic patterns and emergencies.

[0005] Existing traffic prediction methods perform well in short-term prediction but face challenges in long-term prediction and the perception of emergencies. Current methods mainly rely on historical data for training and are difficult to handle emergencies such as traffic accidents, road construction, and extreme weather. As a result, when the traffic flow is disturbed, models based on historical patterns cannot accurately predict future conditions. Summary of the Invention

[0006] The present invention provides a traffic flow generation method based on the fusion of text events and graph attention networks, aiming to improve the prediction accuracy of traffic flow states by combining traffic feature data with external text event information, and having significant advantages especially in long-term prediction and the perception ability of emergencies.

[0007] To achieve the above object, the present invention is realized through the following technical solutions:

[0008] A traffic flow generation method based on the fusion of text events and graph attention networks, comprising the following steps:

[0009] S1: Perform data normalization preprocessing and feature dimension reshaping on the traffic feature data in the dataset to obtain the traffic feature data required for model training;

[0010] S2: Pre-train the AutoencoderKL encoder-decoder model and compress and encode the preprocessed traffic feature data to construct a latent space vector;

[0011] S3: Perform diffusion and denoising in the latent space, and use the BERT model to transform the text event data description into a conditional input for generating traffic features. During the training of the diffusion model, combine the conditional input, gradually denoise, and generate features that conform to the actual traffic state;

[0012] S4: Train the graph attention network model to capture the spatial correlation of traffic features; weight and fuse the results generated by the diffusion model in step S3 with the results of the trained graph attention network model to enhance the prediction effect, and finally decode the fused traffic feature data;

[0013] S5: Use the trained diffusion model and graph attention network model for inference. According to the input event text at a future moment, a corresponding road traffic feature map can be generated to assist traffic management and planning.

[0014] In an embodiment of the present invention, the specific implementation method of step S1 includes the following steps:

[0015] S1.1. For the traffic feature data in the dataset, adopt the maximum-minimum normalization method, linearly transform it to [0,1] through the min-max normalization method, and the normalization formula is:

[0016]

[0017] Among them, \(x\) represents the original traffic feature data, \(x'\) represents the normalized traffic feature data, \(x_{min}\) represents the minimum value of the traffic feature data in the dataset, and \(x_{max}\) represents the maximum value of the traffic feature data in the dataset;

[0018] S1.2. Reshape the traffic feature data to convert it into an image format that can be learned by the diffusion model. Define 36 virtual empty roads and fill them into the original traffic feature data matrix. Reshape the 1260 road features in the traffic feature data into a tensor of \(36\times36\times3\). Each pixel point represents the traffic feature of a specific road, and the three channels contain the same feature information.

[0019] In an embodiment of the present invention, the specific implementation method of step S2 includes the following steps:

[0020] S2.1. Pre-train the AutoencoderKL model for data encoding and decoding of the latent diffusion model;

[0021] S2.2. Use the pre-trained AutoencoderKL model to map the input traffic feature data into a latent space to obtain the latent representation \(z\) of the traffic feature data:

[0022]

[0023] Among them, \(x\) represents the input traffic feature data.

[0024] In an embodiment of the present invention, the specific implementation method of step S3 includes the following steps:

[0025] S3.1. In the forward diffusion stage, gradually add noise to the latent representation \(z\) to make it gradually approach pure Gaussian noise. The change of the noise intensity follows a linear scheduling strategy, increasing gradually from 0.00085 to 0.012. The diffusion process is set to 1000 time steps, and the diffusion process is expressed as:

[0026]

[0027] Among them, \(z_t\) represents the latent representation obtained at the \(t\)-th time step, and \(\alpha_t\) is a parameter controlling the noise intensity, usually adopting a linear scheduling strategy:

[0028] \(\alpha_t = 1 - \beta_t\), \(\beta_t\sim Linear(0.00085, 0.012)\);

[0029] S3.2. Use the BERT pre-trained language model to encode the input text, extract the semantic information of the text and convert it into a 768-dimensional embedding vector, that is, the text embedding vector, as the generation condition for the diffusion model training to ensure that the generation result of the diffusion model is consistent with the text description;

[0030] S3.3. In the reverse diffusion stage, the diffusion model gradually denoises from the pure Gaussian noise obtained in the forward diffusion process through the reverse diffusion process and gradually recovers the target data.

[0031] In an embodiment of the present invention, the specific implementation method of step S4 includes the following steps:

[0032] S4.1. Based on the BJTT dataset, by traversing the spatial connection relationship of road segments, an adjacency matrix containing 1260 road connection information is constructed;

[0033] S4.2. In the training of the graph attention network model, first calculate the attention coefficients between road nodes, assign different importance weights to different neighbor nodes based on the adaptive attention mechanism, and finally the attention weights between adjacent nodes are obtained through Softmax normalization:

[0034]

[0035] Among them, represents the neighbor set of node i, and αij is a learnable attention weight, representing the similarity between node i and node j;

[0036] S4.3. After calculating the attention weights, use the multi-head attention mechanism to capture the multi-dimensional relationships between nodes, calculate different weights through multiple independent attention heads, and splice them in the final stage. The formula is expressed as:

[0037]

[0038] Among them, each head k learns different attention weights and the linear transformation matrix W ( k ) , capturing node relationships from multiple perspectives;

[0039] S4.4. Through the gating mechanism, weightedly fuse the denoising result output by the diffusion model with the traffic structure information learned by the graph attention network to obtain the fused features;

[0040] S4.5. Map the fused features back to the latent space through the decoder in the pre-trained AutoencoderKL model to recover the final traffic state image.

[0041] In one embodiment of the present invention, the specific implementation method of step S5 includes the following steps:

[0042] S5.1. During the model inference process, the DDIM sampling algorithm is adopted to accelerate convergence, and the number of sampling steps is set to 200 to accelerate the generation process, and finally three-channel traffic feature data is obtained;

[0043] S5.2. The generated traffic feature data is converted into geographical coordinates through road topology mapping, and color coding is used to visually display the traffic states of different roads, and an interactive map is generated through Folium; finally, the prediction result is output in the form of a web map.

[0044] In one embodiment of the present invention, in step S3.3, the reverse denoising process is implemented through a U-Net network. The U-Net dynamically adjusts the latent representation through a cross-attention mechanism, and the formula is expressed as:

[0045]

[0046] where Q is the noise feature projection, and K and V are the conditional vector projections;

[0047] At each time step, the U-Net network processes features of different resolutions through a multi-scale attention mechanism, namely 1 / 4, 1 / 2, and full resolution, to improve the latent representation, and finally generates a traffic flow prediction result consistent with the text conditions.

[0048] In one embodiment of the present invention, in step S4.4, a gating mechanism is adopted to weight and fuse the output results of the graph attention network and the output results of the diffusion model. The following relationship exists in the process of fusing using the gating mechanism:

[0049] α = σ(fusion_conv(h))

[0050] where h is the denoised feature processed by the diffusion model, and σ is the Sigmoid activation function, ensuring that the obtained fusion weight α is within the range of [0, 1];

[0051] Using the calculated fusion weight α, the denoised feature h output by the diffusion model and the traffic structure information target learned by the graph attention network are weighted and summed:

[0052]

[0053] where represents the fused feature.

[0054] Advantages of the present invention:

[0055] The present invention proposes a traffic flow generation method based on the fusion of text events and graph attention networks, which is applicable to the medium- and long-term traffic prediction scenario under the condition of known future traffic event text information. Through experiments, the error in predicting the road traffic congestion level is 0.25 (MAE), and the error in predicting the road traffic speed is 4.8 km / h (MAE). The traffic flow generation method of the present invention can not only infer future trends based on historical data, but also generate a traffic map that conforms to the actual traffic state by combining the information of emergencies and the topological structure of the traffic network, providing more reliable prediction support for intelligent traffic management and urban planning. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 FIG. is a schematic flow chart of a traffic flow generation method based on the fusion of text events and graph attention networks provided by the present invention;

[0057] Figure 2 FIG. is a schematic diagram of model training for diffusion and denoising in the latent space;

[0058] Figure 3 FIG. is an intermediate process of denoising the road traffic congestion characteristics saved during the diffusion process training;

[0059] Figure 4 FIG. is a schematic diagram of weighted fusion of the results generated by the diffusion model and the results of the graph attention network model;

[0060] Figure 5 FIG. is a visualized traffic congestion situation map. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The schematic flow chart of the specific implementation steps of the traffic flow generation method based on the fusion of text events and graph attention networks provided by the embodiments of the present invention is as Figure 1 shown, and the traffic flow generation method specifically includes the following steps:

[0062] S1: Perform data normalization preprocessing and feature dimension reshaping on the traffic feature data in the dataset to obtain the traffic feature data required for model training.

[0063] Further, the specific implementation method of step S1 includes the following steps:

[0064] S1.1. Normalize the traffic feature data to ensure that different features are in the same numerical range, improving the stability and prediction accuracy of model training. The original traffic feature data includes congestion level and speed features. Among them, the congestion level is discretely labeled (1-5 levels), and the speed feature is between [0, 150].

[0065] In data preprocessing, first, the traffic feature data is extended to a three-channel format of 36×36×3, where the same feature information is stored in the three channels. Then, the traffic feature data is normalized. The congestion level is normalized to [0, 1], and the speed feature is normalized to [0, 1] to ensure that the model can stably learn the influence of different features and improve the reliability of traffic flow prediction.

[0066] S1.2. In the preprocessing of traffic feature data, in order to make the original traffic feature data adapt to the input requirements of the deep learning model, the original traffic feature data is rearranged and reshaped into an image format for use as the input of the diffusion model.

[0067] Specifically, 1260 roads in the dataset are reorganized so that they can be input into the deep learning network in the form of a three-dimensional tensor, where each road corresponds to a pixel point. To achieve this mapping, 36 virtual "empty roads" are defined and filled into the original data to expand the data scale to 1296 roads (36×36).

[0068] Since a standard RGB image consists of three channels, while the feature information of the BJTT dataset only includes two dimensions (i.e., traffic flow speed and traffic congestion level), therefore, in the present invention, each feature is replicated to three channels to ensure that the model can independently train these features.

[0069] S2: Pre-train the AutoencoderKL model and compress and encode the preprocessed traffic feature data to construct a latent space vector.

[0070] Further, the specific steps of step S2 are as follows:

[0071] S2.1. Pre-train the AutoencoderKL encoder. The training of this AutoencoderKL encoder uses the Adam optimizer, and the initial learning rate is set to 4.5e-6 to ensure stable convergence. The pre-trained AutoencoderKL encoder can efficiently extract the latent representation of traffic features and retain its key structural information in the low-dimensional space.

[0072] S2.2. Use the AutoencoderKL encoder pre-trained in step S2.1 to compress and encode the traffic feature data obtained in step S1 to construct a latent space vector.

[0073] First, the input traffic feature data (36×36×3, including multi-channel information such as congestion level or traffic flow speed) is compressed into a latent representation by the pre-trained AutoencoderKL encoder as follows, where the downsampling factor is 2, and the calculation expression is:

[0074]

[0075] Among them, x represents the input traffic feature data, z represents the compressed traffic feature data, i.e., the latent representation, Encoder represents the encoder, the resolution is downsampled from 36×36 to 18×18, and the number of channels remains 3. This operation effectively reduces the computational complexity while retaining more image details and semantic information of the input traffic feature data. At the same time, a scaling factor α = 0.18215 is applied to the latent representation to narrow the numerical range of the latent representation.

[0076] S3: By performing diffusion and denoising in the latent space and using the BERT model to transform the description of text event data (such as traffic accidents, construction, etc.) into the conditional input for generating traffic features, during the training of the diffusion model, combined with the conditional input, gradually denoise and generate features that conform to the actual traffic state. For the specific implementation process, please refer to Figure 2 as shown.

[0077] Furthermore, the specific implementation method of step S3 includes the following steps:

[0078] S3.1. In the forward diffusion stage, gradually add noise to the latent representation z so that it finally approaches isotropic Gaussian noise. The diffusion process can be expressed as:

[0079]

[0080] where zt represents the latent representation obtained at the t-th time step, and αt is a parameter controlling the noise intensity, usually adopting a linear scheduling strategy:

[0081] αt = 1 - βt, βt ~ Linear(0.00085, 0.012)

[0082] The intensity of the noise in this task gradually increases from 0.00085 to 0.012 by the linear noise scheduling strategy. As the time step t increases, the noise intensity gradually increases, and finally zt is transformed into pure noise. The time step t in this task is set to 1000 time steps, and t gradually increases from 0 until the model completely converts the latent representation (i.e., the latent space vector) into pure noise. Among them, the intermediate results of the diffusion process training are as Figure 3 shown.

[0083] S3.2. First, use the BERT pre-trained language model to encode the text event data, extract the semantics of the text and convert it into a 768-dimensional vector, that is, the text embedding vector as the generation condition for the diffusion model training. This high-dimensional representation can capture the detailed features of the text event data and enhance the correlation between the text event data and the traffic flow generation task.

[0084] The present invention uses a BERT model with a 32-layer Transformer structure, where each layer contains a self-attention mechanism and a feed-forward neural network, and can deeply model the semantic features of text event data.

[0085] The encoded text embedding vectors are input into the UNet network of the reverse diffusion process for conditional sampling, so that the generated data images can accurately reflect the time, location, and event types in the text information, thereby improving the controllability and prediction accuracy of the model.

[0086] S3.3. The goal of the reverse denoising process is to recover the latent representation from the pure Gaussian noise obtained in the forward diffusion process and generate the target traffic state in combination with the conditional input. In the reverse diffusion process, the U-Net network is the core process of denoising, and dynamically adjusts the latent representation through the cross-attention mechanism:

[0087]

[0088] Among them, Q is the noise feature projection, and K and V are the conditional vector projections. Specifically, at each time step, the U-Net network receives the latent representation of the current noise (with a size of 18×18×3) and the text embedding vectors from the conditional input, and the latter is fused with the latent representation through the cross-attention mechanism.

[0089] Among them, the configuration of the UNet network adopts a multi-scale attention mechanism, including feature extraction at different resolutions, namely 1 / 4, 1 / 2, and full resolution. This multi-scale modeling enhances the network's ability to capture global and local information, and finally gradually recovers the latent representation through step-by-step denoising operations.

[0090] In the training stage of the diffusion model, we use 1000 time steps to optimize the diffusion process.

[0091] S4: Train the graph attention network model to capture the spatial correlation of traffic features; weight and fuse the results generated in step S3 with the results of the trained graph attention network model to enhance the prediction effect, and finally decode the fused traffic feature data. For the specific implementation process, please refer to Figure 4 .

[0092] Furthermore, the specific implementation method of step S4 includes the following steps:

[0093] S4.1. Based on the BJTT dataset, by traversing the spatial connection relationships of road segments, an adjacency matrix containing 1260 road connection information is constructed.

[0094] The present invention is based on the BJTT dataset, which contains traffic information of 1260 roads. Each road consists of multiple road segments, and each road segment is defined by a set of latitude and longitude coordinates. The present invention constructs an adjacency matrix A in a connection-relationship-based manner, and the construction method can be described by a formula:

[0095]

[0096] Specifically, first traverse all the road segment coordinates to detect whether there are the same coordinate points; then judge the direct connection relationship between roads, that is, if two different roads share the same coordinate points on some road segments, it is considered that they have a direct connection; finally, construct the adjacency matrix. If there is a direct connection between road i and road j, assign 1 to the corresponding position Ai,j in the adjacency matrix A, otherwise assign 0. Finally, the dimension of the constructed adjacency matrix is 1260×1260, which completely characterizes the topological structure between roads.

[0097] S4.2. Perform a linear transformation on the traffic feature data of the road nodes input to the model, and use a learnable weight matrix W to map the original features of each node to a low-dimensional space, so that the features are more suitable for the calculation of the attention mechanism.

[0098] To ensure that the attention weights are calculated only between connected nodes, introduce the adjacency matrix Ai,j calculated in step S4.1, and set the attention scores of unconnected nodes to a minimum value (such as -1e5) to prevent affecting the results during Softmax normalization. Finally, the attention weights between all adjacent nodes are calculated through Softmax normalization:

[0099]

[0100] where αij represents the attention weight between adjacent nodes i and j, represents the neighbor set of node i. Softmax normalization ensures that the sum of the attention weights of all neighbor nodes is 1, making the attention distribution smooth and reasonable.

[0101] S4.3. After calculating the attention weights, GAT (Graph Attention Network) aggregates the features of neighbor nodes through weighted summation. The present invention adopts a multi-head attention mechanism, that is, calculates different weights through multiple independent attention heads and concatenates them at the final stage:

[0102]

[0103] where the overall weight matrix W is split into the weight matrices W(k) of multiple independent attention heads, and each head k learns different attention weights and the weight matrix W ( k) Capture node relationships from multiple perspectives. Compared with mean aggregation, the splicing method can retain features from different perspectives and improve the information expression ability.

[0104] S4.4. Jointly fuse the output results of the Graph Attention Network (GAT) and the output results of the diffusion model to obtain the fused features.

[0105] Since the diffusion model (UNet) may lose some traffic topological relationships during the denoising process in step S3.3, and GAT mainly relies on adjacency relationships for feature modeling, combining the two can improve the prediction effect.

[0106] In the fusion stage, a gating mechanism is used to control the combination method of the denoising output of UNet and the structural information of GAT. Specifically, a gating mechanism is used to weight and fuse the output results of the Graph Attention Network and the output results of the diffusion model. The following relationship exists during the fusion process using the gating mechanism:

[0107] α = σ(fusion_conv(h))

[0108] where h is the denoised feature processed by the diffusion model, and σ is the Sigmoid activation function, ensuring that the output weight α is in the range of [0, 1];

[0109] Calculate the fusion weight α through Sigmoid and perform weighted summation on the two parts of the features:

[0110]

[0111] where h is the denoised feature processed by the diffusion model, and target is the traffic structure information learned by the Graph Attention Network. This fusion strategy effectively improves the stability and accuracy of traffic state prediction.

[0112] S4.5. After the result fusion is completed in step S4.4, the generated fused features Need to be mapped through the decoder D of the trained AutoencoderKL in step S2.1 to restore the final target traffic state image.

[0113] The final generation result Is expressed as:

[0114]

[0115] This step ensures that while the computational cost of LDM is significantly reduced, it can still retain sufficient detailed information to guarantee the authenticity of the generated traffic state.

[0116] S5. Use the trained model for inference. According to the input event text at future time points, a corresponding visual road traffic feature map can be generated to assist traffic management and planning.

[0117] Furthermore, the specific implementation method of step S5 includes the following steps:

[0118] S5.1. In the inference stage, first load the model trained in step S4 and use the GPU for efficient computing. The model takes the traffic state described by text as the conditional input, and combines it with the road topology feature matrix to generate future traffic feature data through the diffusion sampling process.

[0119] During the inference process, the DDIM sampling algorithm is adopted, and the number of sampling steps is set to 200 to accelerate convergence, and finally three-channel traffic feature data is obtained.

[0120] To match the actual requirements, the model output is mapped to the original data distribution after inverse normalization processing and converted into the prediction result of the road feature state.

[0121] S5.2. Convert the generated traffic feature data into geographical coordinates through road topology mapping, and use color coding to visually display the traffic states of different roads, and generate an interactive map through Folium. Finally, the prediction result is output in the form of a picture and can be used for traffic management decisions. The specific effect is as Figure 5 shown.

[0122] The present invention proposes a traffic flow generation method based on the fusion of text events and graph attention networks, which is applicable to the medium- and long-term traffic prediction scenario under the condition of known future traffic event text information. Through experimental verification, the error in predicting the road traffic congestion level is 0.25 (MAE), and the error in predicting the road passing speed is 4.8 km / h (MAE). The traffic flow generation method of the present invention can not only infer future trends based on historical data, but also generate a traffic map that conforms to the actual traffic state by combining the information of unexpected events and the topological structure of the traffic network, providing more reliable prediction support for intelligent traffic management and urban planning.

[0123] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A traffic flow generation method based on the fusion of text events and graph attention network, characterized in that It includes the following steps: S1: Perform data normalization preprocessing and feature dimension reshaping on the traffic feature data in the dataset to obtain the traffic feature data required for model training; S2: Pre-train the AutoencoderKL codec model and compress and encode the preprocessed traffic feature data to construct a latent space vector; S3: Through diffusion and denoising in the latent space, and using the BERT model to convert the text event data description into a conditional input for generating traffic features. During the training of the diffusion model, combined with the conditional input, gradually denoise and generate features that conform to the actual traffic state; S4: Train the graph attention network model to capture the spatial correlation of traffic features; Weightedly fuse the results generated by the diffusion model in step S3 and the results of the trained graph attention network model to enhance the prediction effect, and finally decode the fused traffic feature data; S5: Use the trained diffusion model and graph attention network model for inference. According to the input event text at a future moment, the corresponding road traffic feature map can be generated to assist traffic management and planning.

2. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 1, wherein The specific implementation method of step S1 includes the following steps: S1.

1. For the traffic feature data in the dataset, adopt the min-max normalization method, and linearly transform it to [0,1] through the min-max normalization method. The normalization formula is: Among them, \(x\) represents the original traffic feature data, \(x'\) represents the normalized traffic feature data, \(x_{\min}\) min represents the minimum value of the traffic feature data in the dataset, and \(x_{\max}\) max represents the maximum value of the traffic feature data in the dataset; S1.

2. Reshape the traffic feature data, convert the traffic feature data into an image format that can be learned by the diffusion model, define 36 virtual empty roads and fill them into the original traffic feature data matrix, and reshape the 1260 road features in the traffic feature data into a tensor of 36×36×3. Each pixel represents the traffic feature of a specific road, and the three channels contain the same feature information.

3. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 1, wherein The specific implementation method of step S2 includes the following steps: S2.

1. Pre-train the AutoencoderKL model for data encoding and decoding of the latent diffusion model; S2.

2. Use the pre-trained AutoencoderKL model to map the input traffic feature data into a latent space to obtain the latent representation z of the traffic feature data: z = Encoder(x), where x represents the input traffic feature data.

4. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 1, wherein The specific implementation method of step S3 includes the following steps: S3.

1. In the forward diffusion stage, gradually add noise to the latent representation z to make it gradually approach pure Gaussian noise. The change of the noise intensity follows a linear scheduling strategy, increasing gradually from 0.00085 to 0.

012. The diffusion process is set to 1000 time steps, and the diffusion process is expressed as: where z t represents the latent representation obtained at the t-th time step, and α t is a parameter controlling the noise intensity, and usually adopts a linear scheduling strategy: α t = 1 - β t , β t ~ Linear(0.00085, 0.012); S3.

2. Use the BERT pre-trained language model to encode the input text, extract the semantic information of the text and convert it into a 768-dimensional embedding vector, that is, the text embedding vector, as the generation condition for the diffusion model training to ensure that the generation result of the diffusion model is consistent with the text description; S3.

3. In the reverse diffusion stage, the diffusion model gradually denoises from the pure Gaussian noise obtained in the forward diffusion process through the reverse diffusion process and gradually recovers the target data.

5. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 1, characterized in that The specific implementation method of step S4 includes the following steps: S4.

1. Based on the BJTT dataset, by traversing the spatial connection relationships of road segments, an adjacency matrix containing 1260 road connection information was constructed; S4.

2. During the training of the graph attention network model, first calculate the attention coefficients between road nodes, assign different importance weights to different neighbor nodes based on the adaptive attention mechanism, and finally the attention weights between adjacent nodes are obtained through Softmax normalization: Among them, represents the neighbor set of node i, and α ij is a learnable attention weight, representing the similarity between node i and node j; S4.

3. After calculating the attention weights, the multi-head attention mechanism is adopted to capture the multi-dimensional relationships between nodes, calculate different weights through multiple independent attention heads, and splice them at the final stage. The formula is expressed as: Among them, each head k learns different attention weights and the linear transformation matrix W (k) , capturing node relationships from multiple perspectives; S4.

4. Through the gating mechanism, the denoising result output by the diffusion model is weighted and fused with the traffic structure information learned by the graph attention network to obtain the fused features; S4.

5. The fused features are mapped back to the latent space through the decoder in the pre-trained AutoencoderKL model to restore the final traffic state image.

6. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 1, wherein The specific implementation method of step S5 includes the following steps: S5.

1. During the model inference process, the DDIM sampling algorithm is adopted to accelerate convergence, and the sampling step is set to 200 to accelerate the generation process, and finally three-channel traffic feature data is obtained; S5.

2. The generated traffic feature data is converted into geographical coordinates through road topology mapping, and color coding is used to visually display the traffic states of different roads, and an interactive map is generated through Folium; finally, the prediction results are output in the form of a web map.

7. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 4, wherein In step S3.3, the reverse denoising process is implemented through the U-Net network. The U-Net dynamically adjusts the latent representation through the cross-attention mechanism. The formula is expressed as: where Q is the noise feature projection, and K and V are the conditional vector projections; At each time step, the U-Net network processes features of different resolutions through the multi-scale attention mechanism, that is, 1 / 4, 1 / 2, and full resolution are used to improve the latent representation, and finally a traffic flow prediction result consistent with the text conditions is generated.

8. The traffic flow generation method based on the fusion of text events and graph attention network according to claim 5, wherein In step S4.4, the gating mechanism is used to weight and fuse the output result of the graph attention network and the output result of the diffusion model. The following relationship exists in the process of fusion using the gating mechanism: α = σ(fusion_conv(h)) where h is the denoised feature processed by the diffusion model, and σ is the Sigmoid activation function to ensure that the obtained fusion weight α is within the range of [0, 1]; Using the calculated fusion weight α, the denoised feature h output by the diffusion model and the traffic structure information target learned by the graph attention network are weighted and summed: Among them, represents the fused feature.

Citation Information

Patent Citations

  • Urban short-term traffic flow prediction method and system based on hybrid convolution LSTM

    CN114360242A

  • A traffic prediction method based on graph attention neural network and spatiotemporal big data

    CN114973678B