Deep Learning Traffic Flow Prediction Method Based on Meteorological Information Fusion
By introducing time alignment modules, enhancing spatiotemporal convolution networks and contrast learning modules into the traffic flow prediction model, the shortcomings of existing models in spatiotemporal dependency capture and external data fusion are solved, and the accuracy and robustness of traffic flow prediction are significantly improved.
Patent Information
- Application Number
- CN202411693544.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-11-25
AI Technical Summary
Existing traffic flow prediction models are difficult to fully capture the spatial and temporal dependence of traffic flow, especially when dealing with complex road network structures and external factors such as weather conditions, the accuracy of the prediction results is insufficient.
Deep learning traffic flow prediction method based on meteorological information fusion is adopted, including time alignment module, enhanced spatiotemporal convolution network and comparison learning module. Through these modules, traffic data and weather data are preprocessed, feature extraction and feature fusion, and predicted values of traffic flow are generated.
It effectively solves the problem of time dislocation between traffic data and weather data, improves the model's ability to capture space-time dependence relationships, and enhances the accuracy and robustness of prediction results.
Smart Images

Figure CN119181256B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information processing, and particularly relates to a deep learning traffic flow prediction method based on meteorological information fusion. Background Art
[0002] With the accelerated development of urbanization and the evolution of intelligent transportation systems, traffic flow prediction has become a core task in modern traffic management. However, existing prediction models show significant limitations when facing the following technical problems:
[0003] First of all, traffic flow has highly spatio-temporal dependence characteristics, but existing models often struggle to fully capture this complex dependence relationship. Especially when dealing with complex road network structures, traditional models have limited understanding of spatial relationships and are difficult to accurately capture the non-Euclidean spatial associations between road nodes.
[0004] Secondly, traffic flow is also affected by many external factors, especially weather conditions. Adverse weather such as rainfall, haze, etc. will significantly change the road traffic capacity and vehicle flow patterns, but traditional prediction models are difficult to integrate traffic data and weather data simultaneously, often ignoring the time misalignment relationship between the two, resulting in insufficient accuracy of prediction results.
[0005] Traffic flow prediction has always been a key task in intelligent transportation systems (ITS). Traditional prediction methods mainly rely on statistical models and time series analysis methods, such as autoregressive integrated moving average model (ARIMA) and support vector regression (SVR). These methods model through historical traffic data, which is simple and direct, but they show obvious limitations when facing complex spatio-temporal dependence relationships and non-linear data. In particular, traditional statistical models often struggle to provide sufficient prediction accuracy when dealing with long-term dependencies and variable traffic dynamics. In addition, these models usually cannot effectively utilize the potential impact of external data (such as weather, etc.) on traffic flow.
[0006] With the rise of deep learning technology, traffic flow prediction methods have been significantly improved. In recent years, deep learning-based models, such as recurrent neural network (RNN), long short-term memory network (LSTM), and temporal convolutional network (TCN), have gradually become mainstream. Models such as LSTM can better capture long-term dependence relationships by introducing memory units and perform better than traditional statistical models. However, although RNN and LSTM have obvious advantages in time series modeling, they are insufficient in dealing with the complex spatial relationships of traffic flow, especially when facing road networks with complex topological structures.
[0007] To address this issue, in recent years, graph neural networks (GNNs) and their variants have been introduced into the field of traffic flow prediction, achieving remarkable progress. The Diffusion Convolutional Recurrent Neural Network (DCRNN) is one of the representative works, which combines diffusion convolution and recurrent neural networks and can effectively model the spatio-temporal dependencies in traffic flow. Subsequently, Graph WaveNet further enhances the model's ability to model non-Euclidean spaces by introducing a learnable static adjacency matrix. Meanwhile, models such as STGCN (Spatio-Temporal Graph Convolutional Network) and ASTGCN (Attention-Based Spatio-Temporal Graph Convolutional Network) combine graph convolutional networks with time series analysis to capture the spatio-temporal characteristics of traffic flow through spatio-temporal convolution, providing new methods for traffic prediction.
[0008] Although these improved models have enhanced the ability to capture spatio-temporal dependencies, existing models still have deficiencies in dealing with long-term dependencies, external data fusion (such as weather), and data time misalignment. First, the time misalignment between traffic data and external factors such as weather is a long-standing problem, and existing models often struggle to handle these time differences, resulting in limited use of external data. Second, most models perform poorly in dealing with long-term sequence dependencies and are unable to fully capture the traffic dynamic changes over long time spans. Finally, existing multi-source data fusion methods lack effective feature fusion mechanisms and usually only use simple concatenation or weighting methods, failing to deeply explore the deep relationships between different modal data. Summary of the Invention
[0009] Object of the Invention: The technical problem to be solved by the present invention is to provide a deep learning traffic flow prediction method based on meteorological information fusion in view of the deficiencies of the prior art, including the following steps:
[0010] Step 1, preprocess traffic data and weather data;
[0011] Step 2, establish a time alignment module to solve the time misalignment problem between traffic data and weather data;
[0012] Step 3, perform embedding encoding on traffic data and weather data;
[0013] Step 4, establish an enhanced spatio-temporal convolutional network. Traffic data extracts spatio-temporal features through the enhanced spatio-temporal convolutional network, and meteorological data extracts temporal features through the enhanced spatio-temporal convolutional network;
[0014] Step 5, establish a contrastive learning module to perform feature fusion and contrastive learning on spatio-temporal features and temporal features;
[0015] Step 6, generate the predicted value of traffic flow.
[0016] Step 1 includes: performing time alignment on traffic data and weather data, processing missing values, and then performing normalization.
[0017] Step 2 includes: the time alignment module includes a cross - correlation estimator and a time - shift adjuster;
[0018] The cross - correlation estimator calculates the cross - correlation between traffic data features and weather data features through the following formula:
[0019] ,
[0020] where, represents the weather data feature and the traffic data feature at the time offset is the correlation coefficient, is the length of the sliding window, represents the time point, represents time is the weather data feature at time , which is a subtraction operation used to indicate moving forward (or backward) by δ time units based on time t. and are the mean and standard deviation of the weather data feature respectively; represents the value of the traffic data feature at time , and are the mean and standard deviation of the traffic data feature respectively.
[0021] In Step 2, the optimal time offset is found by maximizing .
[0022] In Step 2, the time - shift adjuster is used to adjust the time of the weather data, and the adjusted weather data feature is represented by the following formula:
[0023] ,
[0024] where, is the adjusted feature of the weather data feature , represents the original weather data feature , is the optimal time offset of the weather data feature ; Shift represents the time - offset operation.
[0025] Step 3 includes: for traffic data and weather data, perform embedding encoding through linear transformation and ReLU activation function respectively. The embedding process of traffic data is expressed as:
[0026] ,
[0027] where, represents the traffic data features after embedding, is the weight matrix of the embedding layer, is the original traffic data features, is the bias term for traffic data; ReLU is the activation function used to introduce non-linearity, and the formula is:
[0028] ,
[0029] where represents the variable;
[0030] The embedding process of weather data is expressed as:
[0031] ,
[0032] where, is the weather data features after embedding, is the weight matrix for weather data, represents the weather data after time alignment, is the bias term for weather data.
[0033] Step 4 includes: the enhanced spatio-temporal convolutional network first uses the enhanced temporal convolutional network ETCN to extract the temporal features in traffic data. The enhanced temporal convolutional network ETCN uses a combination of depthwise convolution DWConv and pointwise convolution PWConv. By using depthwise convolution DWConv, each channel will perform convolution operations independently, and the formula is:
[0034] ,
[0035] where, is the input feature, is the depthwise convolution kernel;
[0036] The pointwise convolution PWConv realizes cross-channel feature fusion through 1×1 convolution:
[0037] ,
[0038] ,
[0039] where, is the batch normalization operation, is an activation function, and ConvFFN1(x) represents a composite function that includes point convolution, batch normalization, and an activation function; PWConv1 represents a point convolution operation; represents another composite function that includes point convolution, batch normalization, and an activation function; PWConv2 represents another point convolution operation.
[0040] In step 4, the enhanced spatio-temporal convolutional network also introduces a graph attention network, which dynamically assigns connection weights to different nodes in the traffic network through an attention mechanism. The specific formula is:
[0041] ,
[0042] ,
[0043] ,
[0044] ,
[0045] Among them, is the feature vector of node p, representing the attributes of this node p; is the feature vector of node q, representing the attributes of another node q adjacent to node p; represents calculating the relevance for all neighbors y of node p; represents that point y is a node adjacent to node p. The above formula is to calculate the relevance of all nodes adjacent to node p. The calculated feature reflects the information aggregation of each node by GAT in this step, making the feature of each node not only contain its own information but also fuse the information of neighbor nodes, forming a spatial dependence representation of the relationship between nodes in the entire traffic network.
[0046] LeakyReLU is a leaky rectified linear unit, which is used to introduce non-linearity to improve the model's ability to express features; is the feature after the linear transformation of node . This linear transformation can project the feature into a new space for better subsequent similarity calculation and attention weight assignment; is the attention coefficient between node and , is the attention coefficient between node and k; represents the weight matrix for the linear transformation of node features; is the weight vector in the attention mechanism; is the neighbor set of node , is the attention weight after softmax normalization; exp is the natural exponential function; T represents matrix transpose. represents the new feature representation of node p, which is the final output of GAT, representing node features, that is, the spatial information of traffic data, will be combined with the temporal information of traffic, and finally the output of the final GETCN is obtained, which is the spatio-temporal feature of traffic data, that is, traffic feature.
[0047] The traffic features obtained from GETCN and the weather features obtained from ETCN need to be fused to obtain a unified feature representation for contrast learning. The fusion method is concatenation. After concatenation, sample pairs are constructed.
[0048] Positive sample pairs: represent similar samples. For example, the traffic conditions of two different road sections under similar weather conditions within the same time period.
[0049] Negative sample pairs: represent dissimilar samples. For example, traffic conditions at different times or under different weather conditions.
[0050] These positive and negative sample pairs are used to train the model so that it can learn that the features of similar samples are closer and the features of dissimilar samples are farther apart.
[0051] Step 5 includes: The loss function L of the contrast learning is:
[0052] ,
[0053] where, represents the m-th feature vector and the n-th feature vector of the cosine similarity, V is the number of sample pairs (that is, the positive and negative sample pairs constructed earlier). is the temperature parameter for controlling the feature separation scale.
[0054] Step 6 includes: converting the multi-dimensional feature vector into a specific output value through a fully connected layer:
[0055] ,
[0056] where, is the predicted traffic flow value, is the fused feature output by the contrast learning module, and They are the weights and biases of the fully connected layer respectively. The fused feature is obtained by concatenating and fusing the weather feature and the traffic feature in contrastive learning. The loss function L of the contrastive learning module is mainly used to optimize these fused features, making the features of similar samples closer and the features of dissimilar samples farther away. Therefore, the role of the loss function L is to enhance the discriminability of features by optimizing the similarity and difference of features, rather than directly fusing features.
[0057] The present invention proposes a multi-source spatio-temporal hybrid network (MSTHN, Multi-source Spatio-Temporal Hybrid Network). This network can effectively fuse traffic data and weather data from different sources. By introducing a temporal alignment module (Temporal Alignment Module, TAM), the problem of temporal misalignment between weather data and traffic data is solved, ensuring precise alignment of data before feature extraction. At the same time, MSTHN uses an enhanced temporal convolutional network (Enhanced Temporal Convolutional Network, ETCN) and a graph attention network (Graph Attention Network, GAT) to effectively capture the spatio-temporal dependencies in traffic data. To further improve the prediction accuracy of the model, the system also optimizes the fused multi-source features through a contrastive learning module, enabling the model to show higher accuracy and robustness in the face of complex traffic flow prediction tasks.
[0058] The present invention addresses the limitations of the prior art through the following innovative points:
[0059] Temporal alignment module (TAM): This module can solve the problem of temporal misalignment between traffic data and external data (such as weather data), ensuring precise alignment of data before feature extraction, thereby effectively capturing the temporal correlation in multi-source data and improving the prediction accuracy.
[0060] Enhanced temporal convolutional network (ETCN): Through improved temporal convolutional operations, ETCN can show stronger capture ability in the modeling of long-term dependencies, ensuring that the model can stably and effectively model complex dynamic relationships in the face of long-term sequence prediction tasks.
[0061] Contrastive learning module (Contrastive Learning Module): This module is used to optimize the feature fusion of multi-source data, especially in the combination of traffic data and weather data. By contrastive learning, it can better capture the deep relationships between different data modalities and improve the robustness and generalization ability of the fused features.
[0062] Beneficial effects: The present invention effectively solves the deficiencies of existing traffic flow prediction methods in aspects such as multi-source data integration, long-term dependence capture, and feature fusion, providing a new technical approach for improving the accuracy of traffic flow prediction and the universality of practical applications. Brief Description of the Drawings
[0063] Figure 1 It is the method architecture diagram of the present invention. Detailed Embodiments
[0064] The following further specifically describes the present invention in conjunction with the drawings and detailed embodiments, and the above and / or other advantages of the present invention will become clearer.
[0065] This embodiment provides a deep learning traffic flow prediction method based on meteorological information fusion. The method integrates a Time Alignment Module (TAM) to solve the problem of time asynchrony between traffic data and meteorological data, ensuring accurate synchronization before feature extraction. Subsequently, an Enhanced Temporal Convolutional Network (ETCN) is used to capture complex temporal patterns in traffic data, while a Graph Attention Network (GAT) models the inherent spatial dependencies in the traffic network. In addition, a contrast learning module is adopted to fuse the features of multi-source data, thereby improving the accuracy of traffic flow prediction. As Figure 1 shown, the method specifically includes the following steps:
[0066] Step 1, data input and preprocessing: Traffic data and weather data are the two core data sources of the system. The former comes from road surface sensors, and the latter comes from meteorological monitoring stations. To ensure that the system can eliminate the differences in data sources during processing, it is necessary to perform time alignment and normalization processing on the data first.
[0067] Figure 1 In, P1 represents a positive sample pair, which is usually a sample under similar conditions, such as traffic characteristics during the same time period or under similar weather conditions.
[0068] N1 represents a negative sample pair, that is, a sample pair under dissimilar conditions.
[0069] Feature comparison ( ): The representation of the input features is extracted through the function , and the similarity between the positive and negative samples is calculated. The goal of contrast learning is to make the distance between positive sample pairs closer and the distance between negative sample pairs farther.
[0070] , , represents the attention weight, and represents the attention coefficient between nodes.
[0071] : Represents N layers.
[0072] In the data input stage, the traffic data received by the system includes indicators such as vehicle speed and lane occupancy at each time point, and the time granularity is usually 5 minutes. The weather data, on the other hand, includes environmental factors such as temperature, humidity, and precipitation. The time granularity of these data is inconsistent with that of the traffic data, and there may be problems of lag or out-of-sync. To solve this problem, it is first necessary to align the time of traffic and weather data and divide them into a unified time step so that the model can process data from different time steps.
[0073] In addition, the data preprocessing stage also includes the handling of missing values, such as using interpolation methods or data filling methods to ensure data integrity. After that, the data is standardized using the following normalization formula:
[0074] (1)
[0075] Where, is the original input data, such as the original value of traffic flow or weather data characteristics, represents the mean of the original input data, is the standard deviation of the original input data, represents the normalized data. By subtracting the mean from the original data and dividing by the standard deviation, normalization ensures that the values of different data characteristics are scaled to the same scale, thus avoiding some characteristics having too much influence on model training.
[0076] Step 2, establish a Temporal Alignment Module (TAM). The change of traffic flow is closely related to weather conditions, but the impact of weather often has a certain time lag effect. For example, rainfall may not immediately affect traffic flow in the short term, but as time goes by, the wet road surface will cause the vehicle speed to decrease. Therefore, this embodiment proposes a Temporal Alignment Module (TAM), and its main technical solution is to solve the time misalignment problem between traffic data and weather data through a Cross-Correlation Estimator (CCE) and a Shift Adjuster (SA). After completing data preprocessing, the model solves the time offset problem between traffic data and weather data through the Temporal Alignment Module (TAM). The Cross-Correlation Estimator (CCE) is used to calculate the correlation between the two types of data, and then determine the optimal time offset.
[0077] The working principle of the Cross-Correlation Estimator is to calculate the cross-correlation between traffic data characteristics and weather data characteristics through the following formula:
[0078] (2)
[0079] Where, Represents weather data characteristics and traffic data characteristics at a time offset when the correlation coefficient is the length of the sliding window, used to calculate the correlation within a certain period of time represents a time point represents time when the weather data characteristics , and are respectively the mean and standard deviation of the weather data characteristics . Represents the value of the traffic data characteristics at time when and are respectively the mean and standard deviation of the traffic data characteristics. By maximizing , the optimal time offset is found to achieve the time synchronization of the weather data and the traffic data
[0080] After determining the offset, the system adjusts the weather data accordingly through a shift adjuster (SA) to ensure its synchronization with the traffic data. If there is a negative correlation between some weather data characteristics and traffic data characteristics, the "negative correlation" here can be determined by calculating the correlation coefficient, rather than relying entirely on human judgment. Specifically, statistical methods can be used to determine the relationship between the two sets of data. For example, during the data alignment process, the system can calculate the correlation coefficient between the traffic data and the weather data: R is the correlation coefficient. If the correlation coefficient is negative (i.e., -1 < r < 0), it indicates that there is a negative correlation between the two characteristics, which means that when one characteristic increases, the other characteristic tends to decrease. In this case, the system will flip the values of the negatively correlated weather characteristics to ensure that their relationship can correctly match the input characteristics of the traffic flow prediction model
[0081] The shift adjuster will flip the values of the relevant characteristics to ensure data consistency. The adjusted weather data characteristics are represented by the following formula
[0082] (3)
[0083] where is the adjusted characteristic of the weather data characteristic , represents the original weather data characteristic , is the weather data characteristic The optimal time offset. The Chinese meaning of "Shift" is "time offset" or "time shift".
[0084] Specifically, "Shift" means shifting the original weather data features in the time domain. This shift is performed by the optimal time offset to align the weather data with the traffic data in time.
[0085] Step 3: Embedding encoding. After time alignment, the next step is to perform embedding encoding on the traffic data and weather data. The role of the embedding layer is to convert high-dimensional data into low-dimensional vector representations to reduce computational complexity and ensure that the associations between different features can be represented more effectively.
[0086] For the traffic data and weather data, embedding encoding is performed through linear transformation and the ReLU activation function respectively. The embedding process of the traffic data is expressed as:
[0087] (4)
[0088] where represents the traffic data features after embedding, is the weight matrix of the embedding layer, representing the trainable parameters that map the traffic data features from high-dimensional space to low-dimensional space, is the original traffic data features, is the bias term for the traffic data. ReLU is a commonly used activation function for introducing non-linearity, and the formula is: (5)
[0089] where represents the variable;
[0090] The embedding process of the weather data is expressed as:
[0091] (6)
[0092] where is the weather data features after embedding, is the weight matrix for the weather data, represents the weather data after time alignment, is the bias term for the weather data. Embedding encoding maps high-dimensional data into low-dimensional feature representations through linear transformation and activation functions, facilitating subsequent processing.
[0093] Step 4, establish a Graph-Enhanced Temporal Convolutional Network (GETCN).
[0094] After the embedding encoding is completed, the Graph-Enhanced Temporal Convolutional Network (GETCN) is used to capture the temporal and spatial dependencies in the data. Although the traditional Temporal Convolutional Network (TCN) can capture temporal dependencies, it ignores the spatial dependencies in the traffic network. The traffic network naturally has a graph structure, and there are complex spatial associations between road segments. Therefore, the present invention proposes a Graph-Enhanced Spatiotemporal Convolutional Network (GETCN). GETCN realizes the spatiotemporal joint modeling of the traffic network by combining temporal convolution and graph attention network.
[0095] The Graph-Enhanced Spatiotemporal Convolutional Network first uses an Enhanced Temporal Convolutional Network (ETCN) to extract the temporal features in the traffic data. Although the traditional Temporal Convolutional Network (TCN) can effectively model time series, it has a large computational overhead when dealing with complex multi-channel data. Therefore, in this embodiment, a combination of depthwise convolution (DWConv) and pointwise convolution (PWConv) is used in the ETCN to solve this problem.
[0096] DWConv is used to independently perform convolution operations on each channel. In a convolutional neural network, "channel" generally refers to different feature dimensions in the input feature map. In this embodiment, "channel" can be understood as the various features of the traffic data, such as vehicle speed, traffic flow, lane occupancy, etc. Each channel represents a dimension of the traffic data. For example, vehicle speed is one channel, and lane occupancy is another channel. By using depthwise convolution (DWConv), each channel will independently perform convolution operations, and the formula is:
[0097] (7)
[0098] Where, is the input feature, and this input feature is obtained from the previous preprocessing steps and the Time Alignment Module (TAM). Specifically:
[0099] The traffic data is preprocessed and then aligned with the weather data through the Time Alignment Module to ensure temporal synchronization.
[0100] After these processes, the traffic data is passed into the Enhanced Temporal Convolutional Network (GETCN) for further feature extraction.
[0101] Before entering the GETCN, the traffic data has undergone embedding encoding and becomes the input feature that can be directly used by the GETCN.
[0102] is a depth convolution kernel. DWConv effectively reduces the computational complexity of the model and is suitable for the characteristics of a large number of sensor channels in traffic data.
[0103] PWConv realizes cross-channel feature fusion through 1×1 convolution:
[0104] (8)
[0105] (9)
[0106] Among them, is the batch normalization operation, is the activation function, providing non-linear modeling ability. By combining DWConv and PWConv, the model can not only capture long-term dependencies but also perform effective inter-channel interactions. PWConv1: represents the point convolution operation. Point convolution is a 1×1 convolution used for cross-channel feature fusion and is usually used to reduce the dimension of the feature map.
[0107] What this formula outputs is the spatio-temporal feature and the temporal feature, which are the features that enter the contrastive learning later.
[0108] ConvFFN1(x): represents a composite function containing point convolution, batch normalization, and activation function. The composite function here is used to enhance the non-linear modeling ability of the input features.
[0109] PWConv2: represents another point convolution operation. Similar to PWConv1, it is used to further fuse features across channels.
[0110] PWConv2: represents another point convolution operation, with the same structure as PWConv1 but used to further fuse features across channels.
[0111] ConvFFN2(x): has the exact same structure as ConvFFN1(x), used to further enhance the non-linear modeling ability of the features and has independent weight parameters.
[0112] To further enhance the modeling of the spatial features of traffic data, a graph attention network (GAT) is introduced in GETCN. GAT dynamically assigns connection weights to different nodes in the traffic network through the attention mechanism, and the specific formula is:
[0113] (10)
[0114] (11)
[0115] (12)
[0116] (13)
[0117] Among them, is the feature after the linear transformation of node ; is the node and The attention coefficient between them is calculated through the LeakyReLU activation function and is used to quantify the mutual relationship between the two nodes. represents the weight matrix for the linear transformation of node features. It linearly transforms the original node features to meet the calculation requirements of the graph attention network (GAT). is the weight vector in the attention mechanism and is usually used to calculate the correlation between nodes; is the node 's neighbor set, is the attention weight after softmax normalization; exp is the natural exponential function; T represents matrix transpose; compared with the traditional graph convolutional network (GCN), GAT can better handle dynamic spatial relationships, enabling the model to adaptively assign more weights to important nodes, thereby improving the prediction performance.
[0118] There are significant differences in the characteristics between traffic data and meteorological data. Traffic data has complex spatial dependencies, so GETCN is used to process traffic data. For meteorological data, since it mainly shows dynamic changes in time and lacks significant spatial dependencies, ETCN is used to extract its temporal features.
[0119] Step 5, Feature fusion and contrast learning: After the traffic data extracts spatio-temporal features through GETCN and the meteorological data extracts temporal features through ETCN, this embodiment designs a contrast learning module to fuse these multi-source features. In the process of multi-source data feature fusion, traditional methods often struggle to fully utilize the mutual relationship between traffic and meteorological data. Therefore, a contrast learning module is introduced to further optimize the fused feature space. Contrast learning improves the discrimination of feature representations by bringing similar data features closer and different data features farther apart. The loss function L of contrast learning is:
[0120] (14)
[0121] Among them, represents the cosine similarity between the m-th feature vector and the n-th feature vector . These feature vectors are the outputs of traffic and weather data through the enhanced graph convolutional temporal convolutional network (GETCN) and the enhanced temporal convolutional network (ETCN). V is the number of sample pairs. (In the summation, this is equivalent to ) is to ensure that negative samples do not contain themselves, thus ensuring the rationality and effectiveness of the contrastive learning loss function. The temperature parameter is used to control the scale of feature separation. Through contrastive learning, it can ensure that traffic and meteorological features from similar environments are more closely integrated, while features with large differences are distinguished, thereby enhancing the model's predictive ability for complex scenarios.
[0122] Step 6, Generate Prediction: After feature extraction and fusion, the model maps these features to the predicted value of traffic flow through the fully connected layer. The function of the fully connected layer is to convert the multi-dimensional feature vector into a specific output value, that is, the final traffic flow prediction. The formula of the fully connected layer is:
[0123] (15)
[0124] in, is the predicted traffic flow value, is the fusion feature output by the contrast learning module. and are the weights and biases of the fully connected layer respectively.
[0125] Through the linear transformation of the fully connected layer, the fused features are mapped to the output space. This step can effectively integrate the spatiotemporal features extracted from different modules to generate accurate traffic flow prediction values. The design of the fully connected layer has low computational complexity while ensuring prediction accuracy, allowing the model to efficiently perform real-time predictions in practical applications.
[0126] In the process of solving the challenge of multi-source data fusion in traffic flow prediction, this embodiment constructs a traffic prediction model based on a multi-source spatiotemporal hybrid network (MSTHN). This model is specifically optimized for the accuracy and robustness issues when integrating and analyzing the spatiotemporal features of multi-source data such as traffic and meteorology. By combining deep learning technology and spatiotemporal feature fusion, this embodiment proposes a novel method to capture and understand the complex dynamics in the traffic system, thereby significantly improving the accuracy and effectiveness of traffic flow prediction.
[0127] First, this embodiment designs a time alignment module (TAM) for synchronous processing of traffic data and meteorological data. Since there are time differences between traffic and meteorological data, for example, the impact of meteorological conditions on traffic may be delayed, this embodiment realizes the time alignment processing of traffic and meteorological data through the combination of cross-correlation estimator and time step adjuster, so that the two data can reach the optimal synchronization state before feature extraction, providing an accurate basis for subsequent spatiotemporal feature analysis.
[0128] Secondly, in this embodiment, an enhanced graph convolutional temporal convolutional network (GETCN) is introduced into MSTHN to extract complex spatio-temporal features in traffic data. The enhanced temporal convolutional network (ETCN) is used to extract temporal features from meteorological data. GETCN combines the enhanced temporal convolutional network and the graph attention network (GAT) to model the long-term and spatial dependencies in traffic data. In particular, ETCN effectively captures the long-term characteristics of traffic and meteorological data through the combination of depth convolution and point convolution, while GAT deeply models the spatial dependence of traffic flow by dynamically assigning weights to different traffic nodes. This structure enables the comprehensive extraction of the features of traffic and meteorological data and enhances the sensitivity and adaptability of the model to data changes.
[0129] To further fuse the multi-source features from traffic and meteorological data, a contrastive learning module is introduced into the model. Through the method of contrastive learning, this module brings the feature vectors in similar environments closer and separates the dissimilar features, thus effectively distinguishing and fusing the features of multi-source data. This innovative feature fusion method enables the model to make full use of the complementary information between the two types of data in the complex multi-source data background, enhancing the accuracy and robustness of traffic flow prediction.
[0130] Finally, the multi-source data that has undergone feature extraction and fusion is mapped through a fully connected layer to generate the final traffic flow prediction result. This step enables the model to provide specific prediction outputs on the basis of efficiently utilizing the extracted spatio-temporal features. Combining the advantages of multi-source data, our MSTHN model significantly improves the prediction performance of traffic flow through precise time alignment, effective spatio-temporal feature extraction, and innovative multi-source feature fusion.
[0131] Through this technical solution, not only the performance of the traffic flow prediction model in multi-source data processing is improved, but also strong technical support is provided for intelligent traffic management, contributing to the accurate prediction of traffic flow and the efficient management of urban traffic systems.
[0132] The meanings of some terms in the embodiments of the present invention are as follows:
[0133] Traffic data: refers to the original traffic information data collected through traffic sensors, surveillance cameras, etc., including basic indicators reflecting traffic flow such as vehicle speed, lane occupancy, and traffic volume.
[0134] Traffic data features: refers to the processed feature values extracted from the original traffic data. These features are used as the input of the model after preprocessing, normalization, etc. For example, the vehicle speed and traffic volume data after time alignment and normalization are traffic data features.
[0135] Weather data: Refers to the original data collected by weather stations that reflect environmental conditions, including temperature, humidity, precipitation, etc.
[0136] Weather data features: Refers to the preprocessed weather data features used for fusing with traffic data to improve the model's prediction ability for traffic flow.
[0137] Feature vector: Refers to the input vector formed by integrating traffic data features, weather data features, and other processed features for the neural network model. This is the final input feature of the model.
[0138] In one embodiment of the present invention, the traffic flow of a main road section in Nanjing and local meteorological data are used to verify the effectiveness of the method proposed in this embodiment. The traffic data is sourced from road surface sensors and includes information such as vehicle speed and traffic volume, with a time granularity of once every 5 minutes; the meteorological data includes information such as temperature, rainfall, and humidity, with a time granularity of once every 10 minutes. The purpose of this embodiment is to improve the accuracy of traffic flow prediction through multi-source data fusion, and specifically includes the following steps:
[0139] Step 1: Data input and preprocessing;
[0140] Traffic data:
[0141] The original data contains indicators such as vehicle speed (unit: km / h) and traffic volume (unit: vehicles / hour). The time period is from July 1st to July 31st, and it is recorded once every 5 minutes.
[0142] For the vehicle speed data, its mean is 60 km / h and the standard deviation is 15 km / h.
[0143] Normalize the data using the formula: . For example, for a vehicle speed value of 75 km / h, its normalized value is: .
[0144] Weather data:
[0145] The original data includes temperature (unit: °C), rainfall (unit: mm), humidity (unit: %), and the time period is the same as the traffic data, recorded once every 10 minutes.
[0146] To make the traffic data and weather data have the same time granularity, first perform linear interpolation on the weather data to adjust its time granularity to once every 5 minutes.
[0147] Time alignment:
[0148] The Cross-Correlation Estimator (CCE) calculates the correlation between traffic flow and rainfall and determines that the optimal time offset for rainfall is 15 minutes.
[0149] The Shift Adjuster (SA) shifts the rainfall data forward by 15 minutes to align the temporal relationship between traffic flow and the impact of rainfall.
[0150] Step 2: Application of the Time Alignment Module (TAM);
[0151] The TAM is used to align traffic data and weather data. For example, during a certain period on July 10th, there is a lag in the impact of the original rainfall on traffic flow. Through the Time Alignment Module, the rainfall data is advanced by 15 minutes to ensure that traffic and weather data are synchronized before model feature extraction.
[0152] Result: The aligned dataset includes synchronized traffic features and weather features, both with a time granularity of 5 minutes.
[0153] Step 3: Weather data enters the Enhanced Temporal Convolutional Network (ETCN);
[0154] Since weather data mainly exhibits dynamic changes over time, it is only processed through the Enhanced Temporal Convolutional Network (ETCN).
[0155] The ETCN uses a combination of Depthwise Convolution (DWConv) and Pointwise Convolution (PWConv) to extract weather features.
[0156] The weather data features after being processed by the ETCN are . (Temporal features)
[0157] Step 4: Traffic data enters the Graph-Enhanced Temporal Convolutional Network (GETCN);
[0158] For traffic data, since traffic flow not only has temporal dependence but also significant spatial dependence, it is processed using the Graph-Enhanced Temporal Convolutional Network (GETCN).
[0159] The ETCN extracts the temporal features in traffic data and reduces the computational complexity through DWConv and PWConv.
[0160] The Graph Attention Network (GAT) further models the spatial relationships of the traffic network by assigning weights to important nodes through the attention mechanism.
[0161] Finally, the spatio-temporal features after being processed by the GETCN are .
[0162] Step 5: The Contrastive Learning Module (CL) performs feature fusion;
[0163] The contrast learning module ensures that the similar parts of the two types of features are close and the dissimilar parts are far apart by fusing the weather features and traffic features . Namely , which is extracted by the enhanced graph convolutional temporal convolutional network (GETCN) and is used to extract spatio-temporal features from traffic data. In GETCN, the GAT outputs node features, which are concatenated with the traffic data features output by depthwise convolution (DWConv) and pointwise convolution (PWConv) to obtain the final spatio-temporal features.
[0164] Namely , which is extracted by the enhanced temporal convolutional network (ETCN) and is used to extract temporal features from weather data because weather data mainly shows dynamic changes over time. In ETCN, the results output by depthwise convolution (DWConv) and pointwise convolution (PWConv) are this weather feature, that is, the temporal feature.
[0165] Calculation of the loss function: For example, for two similar traffic and weather conditions, the cosine similarity of the fused feature vectors is 0.85, while the similarity is only 0.25 under different traffic conditions.
[0166] Result: Through contrast learning, a fused feature is generated, and its discrimination and robustness are significantly improved.
[0167] Step 6: Generate predictions;
[0168] Input the fused feature into the fully connected layer to generate the predicted value of traffic flow.
[0169] Result: For example, during the morning rush hour on July 15th, the predicted traffic flow is 500 vehicles per hour, the actual observed value is 520 vehicles per hour, and the prediction error is 3.8%.
[0170] Experimental results and technical advantages:
[0171] Prediction accuracy: The deep learning traffic flow prediction method based on multi-source data fusion adopted in this embodiment reduces the average prediction error by 15.2% compared with the traditional ARIMA model, and reduces the error by 3.4% compared with the model that only uses traffic data.
[0172] Technical problems to be solved: Through multi-source data fusion, the present invention solves the problem that weather data cannot be effectively utilized in traditional traffic flow prediction; through the time alignment module, the problem of data time misalignment is solved; through enhanced temporal convolutional and graph attention networks, the spatio-temporal dependencies in traffic flow are effectively captured, significantly improving the prediction accuracy in complex traffic scenarios.
[0173] Through this specific embodiment, it fully demonstrates the effectiveness of each module of the present invention in actual traffic flow prediction tasks, as well as its technological innovation in solving multi-source data fusion and complex spatio-temporal dependency modeling problems.
[0174] The present invention provides a deep learning traffic flow prediction method based on meteorological information fusion. There are many methods and ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A deep learning traffic flow prediction method based on meteorological information fusion, characterized in that: The following steps are involved: Step 1, preprocessing traffic data and weather data; Step 2: Establish a time alignment module to solve the time misalignment problem between traffic data and weather data; Step 3, embedding and encoding traffic data and weather data; Step 4: Establish an enhanced spatiotemporal convolutional network. The traffic data is extracted from the spatiotemporal features through the enhanced spatiotemporal convolutional network, and the meteorological data is extracted from the time series features through the enhanced spatiotemporal convolutional network. Step 5: Establish a contrastive learning module to perform feature fusion and contrastive learning on spatiotemporal features and temporal features; Step 6, generating a predicted value of traffic flow; Step 1 includes: time alignment of traffic data and weather data, processing of missing values, and then normalization; Step 2 includes: the time alignment module includes a cross-correlation estimator and a time shift adjuster; The cross-correlation estimator calculates the cross-correlation between traffic data features and weather data features by the following formula: Among them, R i,j (δ) represents the correlation coefficient between weather data feature i and traffic data feature j at time offset δ, L is the length of the sliding window, t represents the time point, X i,t-δ Represents the weather data feature i at time t-δ, μ i and σ i are the mean and standard deviation of weather data feature i; X j,t represents the value of traffic data feature j at time t, μ j and σ j are the mean and standard deviation of the traffic data characteristics, respectively; In step 2, by maximizing |R i,j (δ)|, find the optimal time offset; In step 2, the time shift adjuster is used to adjust the time of the weather data. The characteristics of the adjusted weather data are expressed by the following formula: X aligned,i =Shift(X weather,i ,δ i ), Among them, X aligned,i is the adjusted feature of weather data feature i, X weather,i Represents the original weather data feature i, δ i is the optimal time offset of weather data feature i; Shift represents the time offset operation; Step 3 includes: embedding and encoding traffic data and weather data through linear transformation and ReLU activation function respectively. The embedding process of traffic data is expressed as: E traffic JReLU(W traffic ·V traffic +b traffic ), Among them, E traffic represents the embedded traffic data features, W traffic is the weight matrix of the embedding layer, V traffic is the original traffic data feature, b traffic is a bias term for traffic data; ReLU is an activation function used to introduce nonlinearity, and the formula is: ReLU(x1)=max(0,x1), Where x1 represents the variable; The embedding process of weather data is expressed as: E weather JReLU(W weather ·V weather_aligned +b weather ), Among them, E weather is the embedded weather data feature, W weather is the weight matrix for weather data, V weather_aligned represents the weather data after time alignment, b weather is the bias term for weather data; Step 4 includes: the enhanced spatiotemporal convolutional network first uses the enhanced temporal convolutional network ETCN to extract the temporal features in the traffic data. The enhanced temporal convolutional network ETCN uses a combination of deep convolution DWConv and point convolution PWConv. By using deep convolution DWConv, each channel will perform convolution operations independently. The formula is: DWConv(x)=x*W DW , Among them, x is the input feature, W DW is the depth convolution kernel; The point convolution PWConv achieves cross-channel feature fusion through 1×1 convolution: ConvFFN1(x)=PWConv1(GELU(BN(PWConv1(x)))), ConvFFN2(x)=PWConv2(GELU(BN(PWConv2(x)))), Among them, BN is a batch normalization operation, GELU is an activation function, ConvFFN1(x) represents a composite function containing point convolution, batch normalization and activation function; PWConv1 represents a point convolution operation; ConvFFN2(x) represents another composite function containing point convolution, batch normalization and activation function; PWConv2 represents another point convolution operation; In step 4, the enhanced spatiotemporal convolutional network also introduces a graph attention network, which dynamically allocates the connection weights of different nodes in the traffic network through the attention mechanism. The specific formula is: h' p =Wh p , e pq =LeakyReLU(a T [Wh p ||Wh q ]), Among them, h p is the feature vector of node p, representing the attribute of node p; h q is the feature vector of node q, representing the attribute of another node q adjacent to node p; h y Indicates the calculation of the relevance of all neighbors y of node p; LeakyReLU is a leaky linear rectifier unit; h' p is the linearly transformed feature of node p; e pq is the attention coefficient between nodes p and q, e pk is the attention coefficient between nodes p and k; W represents the linear transformation weight matrix for node features; a is the weight vector in the attention mechanism; N(p) is the set of neighbors of node p, α pq is the attention weight after softmax normalization; exp is the natural exponential function; T represents the matrix transpose; Represents the new feature representation of node p; Step 5 includes: the loss function L of the contrastive learning is: Among them, sim(z m ,z n ) represents the mth eigenvector z m and the nth eigenvector z n The cosine similarity of , V is the number of sample pairs, and τ is the temperature parameter that controls the feature separation scale; Step 6 includes: converting the multidimensional feature vector into a specific output value through a fully connected layer: y pred =W fc ·F+b fc , Among them, y pred is the predicted traffic flow value, F is the fusion feature output by the contrastive learning module, and W fc and b fc are the weights and biases of the fully connected layer respectively.
Citation Information
Patent Citations
Traffic flow prediction method based on time and space and related equipment
CN114360254A
Traffic volume prediction method under adverse weather conditions
CN118015832A