Traffic flow prediction method and system based on feature fusion and frequency enhancement
By constructing a dynamic spatiotemporal map and introducing spatiotemporal attention convolution, combining external meteorological data and a hierarchical encoder-decoder model, the problem of insufficient spatial topological relationship capture and incomplete fusion of external factors in the existing traffic flow prediction methods is solved, and the accuracy and robustness of traffic flow prediction are improved.
Patent Information
- Application Number
- CN202510626621.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-04
AI Technical Summary
Existing traffic flow prediction methods cannot effectively capture the topological relationships of road space, lack of robust noise reduction mechanisms, failure to effectively fusion of external factors, gradient vanishing problems in long-sequence prediction, and poor adaptability to noise and external factors.
By constructing a dynamic spatiotemporal map, integrating geographical neighbor maps with functional similar maps, introducing spatiotemporal attention convolution and graph convolution, combining external meteorological data, long-term short-term memory networks and hierarchical encoder-decoder models are used to predict traffic flow.
It improves the accuracy and robustness of traffic flow prediction, significantly improves the adaptability to complex traffic scenarios, captures more potential spatial and temporal correlations and external factors, and improves long-term prediction performance.
Smart Images

Figure CN120260301A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic flow prediction, and particularly relates to a traffic flow prediction method and system based on feature fusion and frequency enhancement. Background Art
[0002] With the expansion and development of urban scale and infrastructure, urban roads are becoming increasingly complex. The limited urban traffic resources can no longer meet the growing traffic demand. How to improve road traffic efficiency and relieve traffic congestion has become an urgent problem to be solved. Therefore, the Intelligent Traffic System (ITS) has emerged, and traffic flow prediction has become an important research hotspot in the key technologies of intelligent transportation systems. Traffic flow prediction is to analyze historical traffic data (such as traffic volume, speed, density) and external environmental information (such as weather, time period), and use mathematical models or machine learning techniques to predict the traffic status (such as congestion level, traffic efficiency) of key nodes in the road network in a specific future time period. Its core goal is to optimize traffic management strategies, improve the utilization rate of road resources, and provide real-time route suggestions for travelers, so as to relieve traffic congestion, reduce carbon emissions, and promote the development of smart cities. This prediction process needs to comprehensively consider spatio-temporal features and capture the periodicity (such as morning and evening rush hours) and sudden changes (such as the impact of accidents) of traffic flow.
[0003] Generally, traditional methods are mainly divided into two categories: statistical parameter models and machine learning-based non-parametric models. The former is represented by ARIMA (Autoregressive Integrated Moving Average), which predicts future trends through linear combination of historical data, but relies on the assumption of data stationarity; the latter such as LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit) use recurrent neural networks to capture the non-linear relationships in time series. In addition, the VAR (Vector Autoregression) model attempts to model the linear dependencies between multiple nodes.
[0004] The Chinese patent application with the publication number CN118762513A discloses a traffic flow time series prediction method based on dual-domain normalization. By normalizing in both the time domain and the frequency domain simultaneously, it dynamically captures the distribution changes of traffic flow data, eliminates non-stationary factors in the time series data, and then uses a distribution prediction model for prediction. After that, the denormalization process is carried out to reconstruct its non-stationary information, ensuring that the prediction results can accurately reflect the non-stationary characteristics of the original data, thus guaranteeing the reliability and robustness of the prediction results and significantly improving the accuracy and stability of traffic flow time series prediction. Specifically, frequency domain normalization decomposes the time series into high-frequency and low-frequency components to capture fast-changing and sudden change information; time domain normalization calculates local statistics such as the mean and standard deviation to dynamically reflect the fast changes in the time series.
[0005] However, methods such as ARIMA and LSTM only focus on the time dimension and do not effectively integrate the road space topology relationship, resulting in the inability to capture dynamic spatial correlations. Random fluctuations (such as sudden accidents) in traffic data are likely to interfere with the prediction results. Traditional methods lack a robust noise reduction mechanism, and the impacts of meteorological conditions (such as rainfall, wind speed) and air quality (such as PM2.5) on traffic flow are not systematically integrated. RNN (Recurrent Neural Networks) - like models are prone to gradient vanishing in long sequence prediction, making it difficult to model long-term dependencies such as weekly cycles, and the prediction errors accumulate rapidly with the sequence length. In addition, although these methods can solve some problems, they generally ignore the topological structure of the road network and have poor adaptability to noise and external factors (such as weather). Summary of the Invention
[0006] The present invention proposes a traffic flow prediction method and system based on feature fusion and frequency enhancement, aiming to solve the problems existing in the prior art in the field of traffic flow prediction, such as the inability to capture the road space topology relationship, the lack of a robust noise reduction mechanism, the ineffective integration of external factors, the gradient vanishing problem in long sequence prediction, and the poor adaptability to noise and external factors.
[0007] To solve the above technical problems, the traffic flow prediction method proposed by the present invention includes the following steps:
[0008] Obtain traffic flow data and external meteorological data from traffic monitors and perform preprocessing;
[0009] Construct a dynamic spatio-temporal graph generation module. The dynamic spatio-temporal graph generation module constructs a dynamic spatio-temporal graph that integrates a geographical neighbor graph and a functional similarity graph based on the preprocessed traffic flow data. The geographical neighbor graph is an inherent spatial attribute of traffic monitors and the traffic network, and the functional similarity graph is generated by multi-level wavelet decomposition and dynamic time warping of the traffic flow data in the time series.
[0010] Input the dynamic spatio-temporal graph into the spatio-temporal attention convolutional layer to obtain the spatio-temporal matrix;
[0011] Fuse the spatio-temporal matrix with the external meteorological data processed by the long short-term memory network one or more times, and use the output after passing through a fully connected layer of the fused data as the traffic flow prediction data; or,
[0012] Perform adjacent time slice processing on the fused data, construct a hierarchical encoder-decoder model based on Transformer, and generate traffic flow prediction data after the sliced data passes through the model.
[0013] Preferably, the preprocessing includes data cleaning, normalization, time alignment and denoising processing;
[0014] Among them, the time alignment aligns the traffic flow data at a preset time interval with the external meteorological data to generate a matrix including weather conditions and air quality data;
[0015] The denoising performs multi-level wavelet decomposition on the traffic flow data and removes high-frequency noise through soft / hard thresholding.
[0016] Preferably, the generation method of the geographical neighbor graph is specifically as follows:
[0017] Based on the spatial data of the traffic flow, define the initial adjacency matrix of the traffic network according to the road topology structure. For two nodes in the traffic network, if they are adjacent, the matrix element is 1, otherwise it is 0. The processed matrix is the geographical neighbor graph;
[0018] The generation method of the functional similarity graph is specifically as follows:
[0019] Based on the traffic flow of the time series, perform multi-level wavelet decomposition and denoising processing on each traffic flow time series by using wavelet transform; calculate the similarity of the denoised sequences between different nodes through dynamic time warping, select the first several node pairs with the smallest dynamic time warping distance, and construct a functional similarity graph in matrix form.
[0020] Preferably, the dynamic spatio-temporal graph is also divided into time granularities before being input into the spatio-temporal convolutional layer, and is divided into recent time segments, daily periodic time segments and weekly periodic time segments;
[0021] The recent time segment continuously intercepts a data sequence with a length of T along the time axis at the current time t0, h forming a recent time segment for capturing the short-term dynamic change trend of the traffic flow;
[0022] The daily periodic time segment is based on the sampling rate q, and intercepts a data segment with a length of T within the corresponding time window before each day, d for capturing the daily periodic change rule;
[0023] The weekly periodic time segment is based on the same time node of each week, and a historical data segment with a length of T w is extracted for modeling the regular changes of the fixed pattern every week.
[0024] Preferably, the spatio-temporal attention convolutional layer is one or a plurality of serially arranged ones. Each spatio-temporal attention convolutional layer includes a time attention layer, a spatial attention layer, a graph convolutional layer, a time-axis convolutional layer, and a residual convolutional layer arranged in sequence.
[0025] Preferably, the adjacent time slices are specifically: the time dimension information of a fixed length of each monitor is sliced into segments, and then connections are established between the segments; the length of the segments is a hyperparameter; the specific slicing method is:
[0026] seg i,d ={x t,d |(i - 1)×L < t ≤ i×L}
[0027] In the formula, seg i,d represents the finally segmented segment, i represents the segment index of the time series, x t,d represents the value of the d-th dimension at the t-th moment in the original time series, and L represents the length of the segment.
[0028] Preferably, the model includes an encoder, a slice attention layer, a frequency enhancement channel attention layer, a decoder, and an activation function arranged in sequence;
[0029] The initial layer of the encoder is the output of the adjacent time slice algorithm. After that, in each layer, the adjacent two slices of the previous layer are merged in the same dimension, and then the merged result is multiplied by a learnable moment to obtain the final result of this layer. Finally, the final result is input into the slice attention to capture the cross-time and cross-dimensional dependencies to obtain the output of the encoder of this layer;
[0030] The number of layers of the decoder is the same as that of the encoder. Each layer of the decoder takes the feature matrix output by the corresponding layer of the encoder as the input, and then outputs a decoded two-dimensional matrix.
[0031] Preferably, the slice attention layer includes a time-axis slice attention and a feature-axis slice attention;
[0032] After the time-axis slice attention uses the multi-head attention of the Transformer to capture the dependencies between time periods in the same dimension, it is used as the input of the feature-axis slice attention;
[0033] The feature axis slice attention sets a fixed number of learnable vectors for each time step as the intermediate layer; the intermediate layer first aggregates messages from all dimensions through the self-attention mechanism and uses the vectors of all dimensions as keys and values; then the intermediate layer distributes the received messages between dimensions by using the dimension vectors and the aggregated messages as keys and values to establish a full-to-full connection between all dimensions.
[0034] Preferably, the frequency enhancement channel attention layer extracts the frequency domain information of the input features by using the discrete cosine transform. First, the features are divided into several subgroups according to the input dimension, and each subgroup is processed by the corresponding discrete cosine transform components from low frequency to high frequency; then a gating mechanism with sigmoid activation is selected to capture the channel dependencies and comprehensively extract the time information from the frequency domain.
[0035] Correspondingly, the present invention also proposes a traffic flow prediction system based on feature fusion and frequency enhancement. The system is used to implement the above traffic flow prediction method and includes: a data preprocessing module, a dynamic spatio-temporal graph generation module, and a hierarchical encoder-decoder model;
[0036] The data preprocessing module is used to preprocess the input traffic flow data and external meteorological data, including data cleaning, normalization, time alignment, and denoising processing;
[0037] The dynamic spatio-temporal graph generation module constructs a dynamic spatio-temporal graph that fuses a geographical neighbor graph and a functional similarity graph based on the preprocessed traffic flow data. The geographical neighbor graph is the inherent spatial attribute of traffic monitors and traffic networks, and the functional similarity graph is generated by multi-level wavelet decomposition and dynamic time warping of the time series traffic flow data;
[0038] The hierarchical encoder-decoder model is constructed based on Transformer, receives the sliced data, and outputs traffic flow prediction data after passing through the activation function;
[0039] Among them, the dynamic spatio-temporal graph generation module further includes a wavelet analysis sub-module, a dynamic time warping sub-module, a multi-time granularity division sub-module, and a spatio-temporal attention convolution module;
[0040] The hierarchical encoder-decoder model further includes an encoder sub-module, a slice attention sub-module, a frequency enhancement channel attention sub-module, a decoder sub-module, and an activation function sub-module arranged in sequence.
[0041] Compared with the prior art, the present invention has the following technical effects:
[0042] 1. The traffic flow prediction method proposed by the present invention generates a dynamic spatio-temporal graph by fusing a geographical neighbor graph and a functional similarity graph, introduces spatio-temporal attention and graph convolution to process the dynamic spatio-temporal graph, and is based on a graph convolutional network supplemented by a variety of different attention mechanisms to model node representations and capture cross-time and cross-dimensional dependencies to improve the performance of traffic prediction and generate traffic flow prediction results. By mining potential spatio-temporal correlations and the influence of external factors, the performance of the model in predicting traffic flow is effectively improved.
[0043] 2. The traffic flow prediction method proposed by the present invention introduces adjacent time slice processing and a hierarchical encoder-decoder model. By adding a slicing algorithm and a Transformer model, cross-time and cross-dimensional dependencies are captured, and the prediction performance of the model can be improved.
[0044] 3. The traffic flow prediction method proposed by the present invention significantly improves the accuracy and robustness of the prediction model by introducing external meteorological data during the traffic flow prediction process. Specifically, the external meteorological data includes weather conditions and air quality data, both of which participate in the prediction modeling as time series features. Compared with the existing technical solutions that only consider weather conditions, the present invention further incorporates air quality factors into consideration, fully reflecting the influence of the external environment on the dynamic changes of traffic flow, thereby effectively improving the adaptability of the model to complex traffic scenarios.
[0045] 4. The traffic flow prediction method proposed by the present invention divides the original traffic flow data into time granularities, including recent, daily cycle, and weekly cycle time series data, predicts the results of various time granularities respectively, and finally fuses the output results of the three components based on a parameter matrix to obtain the final prediction result. More potential time features in traffic data can be captured, thereby improving the performance of the network model.
[0046] 5. In the hierarchical encoder-decoder model proposed by the present invention, a slicing attention and a frequency enhancement module are introduced, enabling the model to have better prediction effects and performance in long-term prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a schematic flowchart of the traffic flow prediction method described in the present invention;
[0048] Figure 2 is a schematic diagram of the dynamic spatio-temporal graph generation module described in an embodiment of the present invention;
[0049] Figure 3 is a schematic diagram of the multi-level discrete wavelet decomposition structure described in an embodiment of the present invention;
[0050] Figure 4 is a schematic diagram of a visualized wavelet decomposition result described in an embodiment of the present invention;
[0051] Figure 5 It is a schematic diagram of traffic flow in different time periods on different working days described in the embodiments of the present invention;
[0052] Figure 6 It is a schematic diagram of traffic flow on the same day of different weeks described in the embodiments of the present invention;
[0053] Figure 7 It is a schematic diagram of the fully connected layer for verifying mixed data in the embodiments of the present invention;
[0054] Figure 8 It is a schematic diagram of the hierarchical encoder-decoder model described in the embodiments of the present invention;
[0055] Figure 9 It is the evaluation result using the PeMSD4 dataset on different prediction intervals;
[0056] Figure 10 It is the evaluation result using the PeMSD8 dataset on different prediction intervals. Detailed implementation manners
[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present application and with reference to the accompanying drawings.
[0058] Embodiment 1
[0059] This embodiment is a traffic flow prediction method based on feature fusion and frequency enhancement. As Figure 1 shown, it includes the following steps 1 to 5:
[0060] Step 1: Obtain traffic flow data and external meteorological data of traffic monitors and perform preprocessing.
[0061] For external meteorological data, since various characteristics of weather condition data are also time series data at each time step, some researchers have already considered the influence of weather conditions on traffic flow prediction based on spatio-temporal fusion graphs, such as STFGCN. However, at the same time, air quality is also time series data at each time step and will also have a certain impact on traffic flow prediction. Therefore, in this embodiment, both weather conditions and air quality are included in the category of external meteorological data.
[0062] The dataset used in this embodiment contains two subsets: weather conditions and air quality. Weather conditions include real-time temperature (°C), dew point temperature (°C), relative humidity (%), and wind speed (m / s), and all these data are recorded every half hour, as shown in Table 1.
[0063]
[0064] Table 1 Example of weather condition data
[0065] Air quality, including real-time Air Quality Index (AQI) and concentrations of atmospheric particulate matter (PM2.5 and PM10), SO2, NO2, CO, and O3 (μg / m 3 ), all of which are recorded hourly as shown in Table 2.
[0066]
[0067] Table 2 Example of air quality data
[0068] Since the detectors in the dataset of this embodiment collect data every 5 minutes, which is different from the sampling granularity of weather conditions and air quality data, a data sharing method is adopted for data filling. Specifically, the weather condition data from 06:00 to 06:05 shares the recorded data from 06:00 to 06:30, as shown in the first row of Table 1. Similarly, the corresponding air quality data from 06:00 to 06:05 shares the recorded data from 06:00 to 07:00, as shown in the first row of Table 2.
[0069] The preprocessing measures taken in this embodiment include data cleaning, normalization, time alignment, and denoising;
[0070] Among them, the time alignment aligns traffic flow data at a preset time interval with external meteorological data to generate a matrix including weather conditions and air quality data.
[0071] The denoising performs multi-level wavelet decomposition on traffic flow data and removes high-frequency noise through soft / hard thresholding.
[0072] After the above processing, a matrix including weather conditions and air quality can be obtained.
[0073] Step 2: Construct a dynamic spatio-temporal graph generation module. The dynamic spatio-temporal graph generation module constructs a dynamic spatio-temporal graph that fuses a geographical neighbor graph and a functional similarity graph based on the preprocessed traffic flow data. The geographical neighbor graph is the inherent spatial attribute of traffic monitors and traffic networks, and the functional similarity graph is generated by multi-level wavelet decomposition and dynamic time warping of traffic flow data based on time series.
[0074] As Figure 2 shown, this embodiment proposes a dynamic spatio-temporal graph generation module, which uses the wavelet analysis method to encode the spatial correlation between different monitors into two graphs through a graph convolutional network: a geographical neighbor graph and a functional similarity graph.
[0075] The generation method of the geographical neighbor graph is as follows:
[0076] Based on the spatial data of traffic flow, an initial adjacency matrix of the traffic network is defined according to the road topology structure. For two nodes in the traffic network, if they are adjacent, the matrix element is 1, otherwise it is 0. The processed matrix is the geographical neighbor graph.
[0077] The method for generating the functional similarity graph is as follows:
[0078] Based on the traffic flow of the time series, multi-level wavelet decomposition and denoising processing are performed on each traffic flow time series by using wavelet transform; the similarity of the denoised sequences between different nodes is calculated through dynamic time warping, and the top several node pairs with the smallest dynamic time warping distance are selected to construct a functional similarity graph in matrix form.
[0079] Specifically, the implementation process of the dynamic spatio-temporal graph generation module is as follows:
[0080] S21: Define the adjacency matrix based on the road topology structure. The regional static attributes are invariant over time and are essentially the inherent attributes of each traffic monitor and the traffic network, that is, the original topological information of the graph, which is defined as the matrix This matrix is an adjacency matrix with a dimension of n×n, which records the geographical connection relationships between all nodes in the traffic network and is called the geographical neighbor graph, as shown in the following formula:
[0081]
[0082] In the formula, represents the adjacency matrix of node i and node j.
[0083] For the functional similarity graph, the regional dynamic attributes should be related to the historical traffic flow sequences. In this embodiment, a traffic flow spatial correlation measurement method based on multi-level wavelet decomposition and dynamic time warping is proposed to calculate the spatial similarity matrix This matrix is the functional similarity graph.
[0084] S22: For two different time-slice traffic flow data L = {l1, l2, …, l n} and K = {k1, k2, …, k m}, first use wavelet transform to denoise the data to provide more and more accurate information for the subsequent DTW (Dynamic Time Warping).
[0085] Wavelet decomposition is a technique that allows the analysis of time series data at different scales and frequencies. By applying wavelet decomposition to time series data, local and global patterns in the time series can be captured. The discrete wavelet transform (DWT) is used to decompose the input time series into a set of approximation and detail coefficients. The process of multi-level discrete wavelet decomposition is asFigure 3 as shown
[0086] Initially, the original signal x[n] undergoes a first-level decomposition of L and H, and downsampling ↓2 respectively, to obtain the high-frequency component XH(1) and the low-frequency component XL(1) with low resolution. Next, the same process is applied to the low-frequency component, that is, through a low-pass filter and a high-pass filter, repeating the previous steps until the specified decomposition level is reached. Finally, multiple sub-signals are obtained, and each sub-signal represents the component of the original signal at different frequencies.
[0087] The noise is hidden in the high-frequency components after wavelet decomposition. The noise reduction process is to threshold the high-frequency vectors and finally reconstruct them to be close to the original traffic data. In this embodiment, different wavelet basis functions are selected for denoising, and the data is aggregated after different denoising methods, which enables the prediction model to learn the information after multiple denoising methods. Based on previous wavelet transform research, since different wavelet basis functions are suitable for predictions at different times, this embodiment selects a combination of wavelet basis functions, sets the decomposition level, and selects a threshold reconstruction method.
[0088] The threshold reconstruction method is mainly divided into soft threshold and hard threshold. The basic idea of the soft threshold is that the data values with absolute values less than the threshold in a sequence are replaced with substitute values; the data values with absolute values greater than or equal to the threshold are shrunk towards zero by value. In hard thresholding, the data values with absolute values less than the threshold are replaced with substitute values, and the data values with absolute values greater than or equal to the threshold remain unchanged.
[0089] As an example, the soft threshold method is adopted, sym5 is selected as the wavelet basis function, through five-level wavelet decomposition, the threshold coefficient is set to 5, and the visualization result of wavelet analysis is as Figure 4 shown. In Figure 4 , A n (n = 1, 2, 3, 4, 5) represents the low-frequency component after the nth filtering, and D n (n = 1, 2, 3, 4, 5) represents the high-frequency component after the nth filtering. The thresholding is only for all high-frequency components. It can be observed that as the wavelet decomposition sequence increases, the frequency of the low-frequency component becomes lower, and it can better represent the general law of traffic flow fluctuations. The frequency of the high-frequency component becomes lower, and the randomness of the traffic flow gradually weakens. The increase of the sequence can more clearly characterize the general law of traffic flow fluctuations. However, the degree of change in the scale space and the wavelet space becomes smaller, and the workload increases exponentially. Therefore, a reasonable number of decomposition sequences should be selected.
[0090] S23: Fuse the geographical neighbor graph A adj with the spatial similarity matrix A s to obtain the final adjacency matrix A adj-s of the traffic network graph, as the dynamic spatio-temporal graph:
[0091]
[0092] In the formula, σ represents the normalization function, η represents the correlation coefficient, represents the magnitude of the influence of the spatial similarity matrix on the geographical neighbor graph, and k is the number of wavelet basis functions, which is used to transform A s to generate multiple spatial similarity matrices A s1 , A s2 , … A sk , and finally the generated adjacency matrix A adj-s The values in are all between 0 and 1, representing the magnitude of the correlation between each node in the traffic data network. This matrix is used as the input for the subsequent GCN (Graph Convolutional Network).
[0093] Step 3: Input the dynamic spatio-temporal graph into the spatio-temporal attention convolutional layer to obtain the spatio-temporal matrix.
[0094] To more comprehensively capture the spatio-temporal information in the traffic road network graph, this embodiment proposes a spatio-temporal attention convolutional layer. As Figure 2 shown, the spatio-temporal attention convolutional layer is one or a plurality of serially arranged ones. Each spatio-temporal attention convolutional layer includes a time attention layer, a spatial attention layer, a graph convolutional layer, a time-axis convolutional layer, and a residual convolutional layer arranged in sequence. Among them, the graph convolutional layer is a k-order graph convolutional network.
[0095] The adjacency matrix A adj-s is first input into a spatio-temporal attention convolutional layer. In the spatio-temporal attention convolutional layer, it sequentially passes through the time attention layer, the spatial attention layer, the graph convolutional layer, the time-axis convolutional layer, and the residual convolutional layer. The output of each layer is used as the input of the next layer. In the case of multiple spatio-temporal attention convolutional layers, the output of each spatio-temporal attention convolutional layer is used as the input of the next-level spatio-temporal attention convolutional layer. The output of the last spatio-temporal attention convolutional layer is the spatio-temporal matrix.
[0096] For the spatial attention layer, in the spatial dimension, the traffic conditions at different positions affect each other, and the mutual influence is highly dynamic. That is to say, the same node at different times will be affected by neighbor nodes to different degrees as the time changes. Therefore, this embodiment uses the attention mechanism to construct the spatial attention layer, adaptively capturing the dynamic correlation between nodes in the spatial dimension, so that when performing graph convolution, the adjacency matrix is fused with the spatial attention matrix to dynamically adjust the influence weight between nodes.
[0097] Similarly, for the temporal attention layer, similar to the case of the spatial dimension, in the temporal dimension, there is a correlation between the traffic conditions of different time slices, and the correlation varies in different situations. In this embodiment, the attention mechanism is also used to adaptively capture the dynamic correlation in the temporal dimension, and finally, the input is dynamically adjusted by fusing the temporal correlation information.
[0098] For the graph convolutional layer among them, graph convolution is performed in both the spatial dimension and the temporal dimension. The spatial graph convolution first only considers the spatial graph on a certain time slice to study the modeling method of spatial features. In this embodiment, the spectral graph method is used to generalize the convolution operation to graph-structured data, and the features of each node can be regarded as signals on the graph. Therefore, in order to make full use of the topological characteristics of the traffic network, on each time slice, this embodiment directly processes the signals using graph convolution based on spectral graph theory, and utilizes the signal correlation of the traffic network in the spatial dimension to capture meaningful patterns and features in the space.
[0099] In spectral graph theory, a graph can be represented by its corresponding Laplacian matrix. By analyzing the Laplacian matrix and its eigenvalues, the properties of the graph structure can be obtained. The Laplacian matrix of the graph is defined as L = D - A, and its normalized form is where A is the adjacency matrix, L N is the identity matrix, and the degree matrix is a diagonal matrix composed of node degrees. The eigenvalue decomposition of the Laplacian matrix is L = U∧U T , U is the Fourier basis of graph G, and Λ is the eigenvalue diagonal matrix. Taking the traffic flow data at time t as an example, the graph signal is The Fourier transform of the graph signal can be expressed as According to the properties of the Laplacian matrix, it can be known that U is an orthogonal matrix, so the inverse Fourier transform is obtained Graph convolution is to equivalently replace the classical convolution operator with a linear operator diagonalized in the Fourier domain. The implemented convolution operation uses the convolution kernel g θ Perform a convolution operation on graph G. Transforming the graph to the spectral domain to implement the convolution operation on the graph is graph convolution. However, when the scale of the graph is large, directly performing eigenvalue decomposition on the Laplacian matrix is costly. Therefore, in this paper, the Chebyshev polynomial approximation expansion is used to solve it, which is equivalent to using the convolution kernel to extract the information of the 0th to K - 1th order neighbors centered on each node in the graph. The graph convolution module uses the rectified linear unit (ReLU) as the activation function.
[0100] For temporal graph convolution, after the graph convolution operation captures the adjacent information of each node on the graph in the spatial dimension, standard convolutional layers in the temporal dimension are further stacked to update the node signals by integrating information from adjacent time slices. After performing a single temporal convolution operation, the feature information of each node is fused and updated with the features of its adjacent time points. During this process, the data of the node and its adjacent time points have integrated the features of their adjacent nodes at the same time point through the graph convolutional network. Therefore, after being processed by a single-layer spatio-temporal convolution, the model can capture the features of the data in the temporal and spatial dimensions, as well as the correlations between them. To further extract more extensive spatio-temporal features, a network is composed by stacking multiple spatio-temporal convolution modules. Finally, these features are mapped to the dimension of the prediction target through a fully connected layer, and a rectified linear unit is used as the activation function in the fully connected layer.
[0101] It should be noted that due to the complex spatio-temporal dependence and the uncertainty of traffic data itself, predicting traffic flow becomes challenging. Limited by hardware conditions such as acquisition devices, the original traffic flow data often consists of a large number of single-structured and single time-series data. Using such data directly for high-precision traffic flow prediction often fails to meet expectations. Therefore, in-depth data mining of existing data and refining the information hidden in single traffic flow data become important methods to improve prediction accuracy. As a data mining method that enhances the temporal dimension attributes, the multi-time granularity fusion algorithm has been widely used in the field of time series prediction. Therefore, in some embodiments of the present invention, the dynamic spatio-temporal graph is also divided into time granules before inputting into the spatio-temporal convolutional layer, including recent time segments, daily cycle time segments, and weekly cycle time segments.
[0102] To more clearly observe the periodicity of traffic flow in a short period of time, as an example of the monitor's observation, Figure 5 a time-distributed traffic flow graph is shown. It can be seen from the graph that the traffic flow trends from Monday to Friday on weekdays have obvious daily similarities. Overall, the traffic flow value remains very low before 4 am, starts to show an upward trend during the 6-7 am period, reaches the morning peak at 8-9 am, then starts to decline, and keeps oscillating until 1 pm. After 1 pm, the traffic flow significantly decreases, starts to rise after 5 pm, reaches the evening peak at around 7 pm, and then gradually decreases.
[0103] Select and Figure 5 the observation data of a monitor at the same time period on different weeks, such as the traffic flow data on February 3, 10, and 17, 2018. Its traffic flow curve is as Figure 6As shown. It can be found that the traffic flow changes in these three days are highly similar. Except for the large fluctuation around 10 o'clock, the traffic flow changes in other time periods are basically consistent with the overall changes and tend to be stable. Therefore, this embodiment fully considers the time cycle characteristics of traffic flow, divides the data into recent, daily and weekly cycles, and then comprehensively considers to mine the implicit information of the data.
[0104] This embodiment uses the above three time series, namely recent, daily and weekly time series data, to predict various time granularity results respectively, and finally fuses the output results of the three components based on the parameter matrix to obtain the final prediction result. This structure captures more potential time features in traffic data, thereby improving the performance of the network model.
[0105] Specifically, the recent time segment is continuously intercepted along the time axis at the current time t0 for a length of T h The data sequence is used to form a recent time segment to capture the short-term dynamic change trend of traffic flow. The daily cycle time segment is based on the sampling rate q, and the length of the corresponding time window before each day is intercepted to T d The data segments are used to capture the daily periodic changes. The weekly periodic time segments are based on the same time node every week, and the extraction length is T w A piece of historical data is used to model regular changes in a weekly fixed pattern.
[0106] For the above three time granularities, assuming that the characteristic length of each node data at the same time is L, represents the value of all features of node i at time t, represents the value of all features of all nodes in time slice t, where N represents the number of nodes. Assume that the sampling frequency is q times per day. Assume that the current time is t0 and the prediction window size is T p . The length of the intercept along the time axis is T h , T d , T w The three time series segments of are used as the input of the recent, daily cycle and weekly cycle components, respectively, where T h , T d , T w All T p The detailed information of the three time series segments are as follows:
[0107] Recent clips:
[0108]
[0109] The traffic flow at each moment cannot be formed suddenly. The congestion on the road is also aggregated from the traffic flow data of the previous time slice. That is to say, in terms of time, the traffic flow at the latter moment will inevitably be affected by the traffic flow at the previous moment.
[0110] Daily cycle segment:
[0111]
[0112] The purpose of introducing the daily cycle segment is to find the regular changes in the daily traffic volume, and then assist in traffic flow prediction.
[0113] Weekly cycle segment:
[0114]
[0115] It consists of the historical time series parts of the recent few weeks, and these parts have the same weekly attributes and time intervals as the prediction period. Just like people's regular activities every day, people's activities at fixed times every week also show regular changes. Introducing the weekly cycle segment can not only capture the regular changes in such weekly activities, but also make up for the problem of reduced prediction accuracy caused by the different regularities between weekdays and weekends in the daily cycle.
[0116] Each time slice shares the same network model structure, and the data of each time slice is input into the network for training. Finally, the output results of each time slice are merged through a parameter matrix, and the output result of this module is finally obtained. Specifically, if the output results of three time slices after passing through the model are The three learnable parameter weight matrices set are Then the final result obtained by weighted fusion is As follows:
[0117]
[0118] In the formula, represents the Hadamard product of the corresponding elements of the matrix, and W h , W d , W w are learning parameters, which reflect the influence degrees of the three time dimension characteristics of the recent period, daily cycle, and weekly cycle on the prediction target.
[0119] Step 4: Fuse the spatio-temporal matrix with the external meteorological data processed by the long short-term memory (LSTM) network once or multiple times, and use the output after passing through a fully connected layer of the fused data as the traffic flow prediction data; or use the fused data for Step 5, and output the traffic flow prediction data in Step 5.
[0120] As an example, this embodiment adopts a two-layer stacked LSTM structure. The first layer of LSTM has 128 neurons, and the second layer has 276 neurons.
[0121] To verify the usability of the output data in Steps 3 and 4, this embodiment uses the PeMS (Performance Measurement System) of the California Department of Transportation in the United States, which is a publicly available highway database, to train and experimentally verify the model. The system includes traffic data collected in real time from more than 39,000 individual detectors in California. These detectors span the highway systems of all major metropolitan areas in California. In this experiment, two datasets are selected from the PeMS website. One is PeMSD4, which covers the San Francisco Bay Area; the other is PeMSD8, which covers San Bernardino County. The external meteorological data comes from the repository of the National Oceanic and Atmospheric Administration (NOAA) of the United States.
[0122] The traffic flow data in the PeMS system is recorded at 5-minute intervals. There are 12 consecutive records in one hour. 16,992 records in the PeMSD4 dataset and 17,568 records in the PeMSD8 dataset are used. Standard normalization is used to process the data. The dataset is divided into a training set, a validation set, and a test set in a ratio of 7:1:2 in chronological order.
[0123] All the experiments conducted in this experiment are run on a server based on the Linux operating system. The specific hardware environment configuration for the experimental part is shown in Table 3.
[0124]
[0125] Table 3 Experimental Hardware Environment Configuration The specific software environment configuration for the experimental part is shown in Table 4.
[0126]
[0127] Table 4 Experimental Software Environment Configuration The hyperparameter settings for the model network part in the experiment are shown in Table 5.
[0128]
[0129]
[0130] Table 5 Hyperparameter Settings for the Network Model
[0131] To verify the effectiveness of the model, the experiments in this step use the Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE) to quantify the performance and prediction effect of the model. Among them, RMSE reflects the dispersion degree of the deviation distribution of the prediction results, MAE reflects the absolute error between the predicted value and the true value, and MAPE represents the relative error between the predicted value and the true value in the form of a percentage. The smaller their values are, the better the performance of the model.
[0132] The model design of this embodiment is based on Figure 2 , and after mixing the data, a fully connected layer output is added. The structure of the fully connected layer is as Figure 7 shown.
[0133] The following benchmark models are selected for comparison, including traditional methods, machine learning-based methods, and deep graph neural network-based models.
[0134] HA: Refers to the historical average model, which simply averages the historical traffic conditions over a period of time (such as a week) to predict the traffic flow in the next time period. It is a very primitive method.
[0135] ARIMA: A classic time series prediction model that uses autoregressive (AR) and moving average for traffic prediction.
[0136] VAR: Refers to the vector autoregressive model, which is usually used to model stochastic processes. It is a more advanced time series model that can capture the pairwise relationships between all traffic flow sequences. However, due to the large number of parameters, the time efficiency is not high.
[0137] LSTM: Long Short-Term Memory network, a variant of the recurrent neural network. The coordinated work of three gating mechanisms is used to maintain and transmit historical data.
[0138] GRU: Gated Recurrent Unit network, a special RNN model with fewer parameters than LSTM. It can capture long-term and short-term dependencies faster and perform traffic prediction.
[0139] DCRNN: One of the most classic road traffic prediction models. It combines the graph convolutional network with the RNN in an encoder-decoder manner. It uses bidirectional random walks on the graph to capture spatial correlations and uses the seq2seq architecture and pre-sampling to capture temporal correlations.
[0140] STGCN: Spatio-Temporal Graph Convolutional Network, which uses ChebNet graph convolution in the spatial dimension and a 2D convolutional network in the temporal dimension to model the correlations in spatio-temporal graph data.
[0141] This embodiment was compared with eight baseline methods on PeMSD4 and PeMSD8. Table 6 shows the average results of the future one-hour traffic flow prediction performance on the PeMSD4 and PeMSD8 datasets.
[0142]
[0143] Table 6 Comparison of Average Performance of Different Models Ⅰ
[0144] As can be seen from Table 6, compared with traditional models such as ARIMA and VAR, the prediction errors of deep learning models such as LSTM and GRU are much smaller than those of traditional models. Among them, the RMSE, MAE, and MAPE values of the LSTM model are reduced by 20.02, 2.66, and 1.85 respectively compared with the ARIMA model, and by 8.62, 4.31, and 3.32 respectively compared with the VAR model; the RMSE, MAE, and MAPE values of the GRU model are reduced by 20.31, 3.46, and 1.98 respectively compared with the ARIMA model, and by 8.91, 5.11, and 3.45 respectively compared with the VAR model. It is found that traditional machine learning methods fail to show acceptable results because they cannot model the effective non-linear spatio-temporal correlations between traffic flows. At the same time, DCRNN and STGCN have about 8%, 17%, and 16% improvements respectively compared with LSTM and GRU in the three metrics. This is because compared with LSTM and GRU that only model time features, these two models also capture the spatial features in traffic road network information, which shows the importance of combining the spatial correlations of sensors for traffic prediction. Among these methods, the model used in this embodiment also obtains the best performance, with improvements of 7.7%, 11.9%, and 12.4% respectively compared with the typical DCRNN model, and 8.8%, 11.2%, and 11.8% respectively compared with the STGCN model, which also proves the effectiveness of this embodiment in mining potential spatio-temporal correlations and the influence of external factors.
[0145] Step five, perform adjacent time slice processing on the fused data, construct a hierarchical encoder-decoder model based on Transformer, and the sliced data generates traffic flow prediction data after passing through the model.
[0146] The adjacent time slice is specifically: the time dimension information of a fixed length for each monitor is sliced into segments, and then connections are established between the segments; the length of the segment is a hyperparameter; the specific slicing method is:
[0147] segi,d = {x t,d | (i - 1)×L < t ≤ i×L}
[0148] wherein, seg i,d represents the finally segmented segment, i represents the segment index of the time series, x t,d represents the value of the d-th dimension at time t in the original time series, and L represents the length of the segment.
[0149] Meanwhile, the segments of all dimensions within a period of time are aggregated and represented as S m:n , which is specifically represented in the following form:
[0150]
[0151] wherein, D represents all dimensions in the time series, m represents the start time, and n represents the end time.
[0152] As Figure 8 shown, the model includes an encoder, a slice attention layer, a frequency enhancement channel attention layer, a decoder, and an activation function arranged in sequence;
[0153] The initial layer of the encoder is the output of the adjacent time slice algorithm. After that, in each layer on the same dimension, the adjacent two slices of the previous layer are merged, and the merged result is multiplied by a learnable matrix to obtain the final result of this layer. Finally, the final result is input into the slice attention to capture cross-time and cross-dimensional dependencies to obtain the output of the encoder of this layer;
[0154] The number of layers of the decoder is the same as that of the encoder. Each layer of the decoder takes the feature matrix output by the corresponding layer of the encoder as input, and then outputs a decoded two-dimensional matrix.
[0155] Specifically, the slice attention layer includes time-axis slice attention and feature-axis slice attention. The time-axis slice attention uses the multi-head attention of Transformer to capture the dependencies between time periods on the same dimension, and then serves as the input of the feature-axis slice attention.
[0156] The feature-axis slice attention sets a fixed number c of learnable vectors for each time step as the intermediate layer, where c is much smaller than D. The intermediate layer first aggregates the messages from all dimensions through the self-attention mechanism, and uses the vectors of all dimensions as keys and values; then the intermediate layer distributes the received messages between dimensions by using the dimension vectors and the aggregated messages as keys and values, and establishes a full-to-full connection between all dimensions.
[0157] Frequency is a natural auxiliary means for time series analysis. Introducing frequency information into a time series model is important and intuitive. To capture this information, the usual method is to use the Fourier transform to extract frequency information from the time series. However, if the values at both ends of the sequence differ greatly, the Fourier transform will introduce high-frequency noise, which will cause errors in the boundary information, known as the Gibbs phenomenon. Moreover, traffic flow data itself has strong randomness, and it may also have large fluctuations in a short period of time. If the Fourier transform is used to extract frequency domain features, a lot of noise will be introduced at the same time. Therefore, in this embodiment, a frequency enhancement channel attention layer is set up. The frequency enhancement channel attention layer uses the discrete cosine transform to extract the frequency domain information of the input features. First, the features are divided into several subgroups according to the input dimension, and each subgroup is processed by the corresponding discrete cosine transform components from low frequency to high frequency; then a gated mechanism with sigmoid activation is selected to capture the channel dependencies and comprehensively extract time information from the frequency domain.
[0158] In this step, an experiment is designed to evaluate the performance of the hierarchical encoder-decoder model. In this experiment, new hyperparameters are introduced as shown in Table 7.
[0159]
[0160] Table 7 Hyperparameters of the hierarchical encoder-decoder model
[0161] To verify the effectiveness of the model, the root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE) are also used in this experiment to quantify the performance and prediction effect of the model.
[0162] To verify the prediction ability of the hierarchical encoder-decoder model, several well-known deep learning-based prediction models in traffic flow prediction are selected for comparison. Their detailed information is as follows:
[0163] Graph WaveNet: A graph neural network that uses pre-defined and adaptive adjacency matrices for diffusion convolution to capture spatial dependency graphs, and applies one-dimensional dilated causal convolution to capture temporal dependencies.
[0164] GMAN: Graph multi-attention network, which adopts an encoder-decoder framework. It designs various spatio-temporal attention mechanisms in the encoder and decoder to simulate spatio-temporal correlations, and designs a transformation attention mechanism to transfer information from the encoder to the decoder.
[0165] ASTGCN: Attention-based spatio-temporal graph convolutional network, which designs spatial and temporal attention mechanisms to capture spatial and temporal patterns respectively.
[0166] STSGCN: Spatio-Temporal Synchronized Graph Convolutional Network, which designs a spatio-temporal synchronization modeling mechanism to capture local spatio-temporal correlations.
[0167] AGCRN: Adaptive Graph Convolutional Recurrent Network, which learns a data-adaptive adjacent matrix for graph convolution to model spatial correlations and uses a gated recurrent unit (GRU) to model temporal correlations.
[0168] The model in this step was compared with five baseline methods on PeMSD4 and PeMSD8, as well as the models proposed in Steps 3 and 4. Table 8 shows the average results of the future one-hour traffic flow prediction performance on the PeMSD4 and PeMSD8 datasets.
[0169]
[0170] Table 8 Comparison of Average Performance of Different Models II
[0171] The performance on both the PeMSD4 and PeMSD8 datasets can prove that by adding the slicing algorithm and the Transformer model in this chapter to capture cross-time and cross-dimensional dependencies, the prediction performance of the model can be improved.
[0172] To more intuitively and clearly see the model results, each of the above models was tested at 5-minute intervals in the prediction interval from 5 minutes to 1 hour. As the prediction interval increases, the performance changes of different methods are as Figure 9 、 10 shown. Figure 9 -(1) is the comparison of RMSE of different models on the PeMSD4 dataset, Figure 9 -(2) is the comparison of MAE of different models on the PeMSD4 dataset, Figure 9 -(3) is the comparison of MAPE of different models on the PeMSD4 dataset; Figure 10 -(1) is the comparison of RMSE of different models on the PeMSD8 dataset, Figure 10 -(2) is the comparison of MAE of different models on the PeMSD8 dataset, Figure 10 -(3) is the comparison of MAPE of different models on the PeMSD8 dataset.
[0173] Figure 9 、 10Shows the changes in the prediction performance of various methods as the prediction interval increases. It can be clearly seen from the figure that in short-term prediction, the hierarchical encoder-decoder model in this step always outperforms all other models on the two datasets, demonstrating its effectiveness in learning the spatio-temporal features of traffic condition prediction. In long-term prediction, the model in this step has a more obvious improvement compared to the models used in Steps 3 and 4 in short-term prediction. At the same time, the slope in the line chart becomes smaller, indicating that the decline rate of its prediction accuracy is more stable, which proves the effectiveness of the slice attention and frequency enhancement modules proposed in this step. Overall, the model proposed in this chapter has the best overall effect among all models, followed by AGCRN, the models used in Steps 3 and 4, STSGCN, and ASTGCN. The last two models, WaveNet and GMAN, perform the worst. The model in this step only performs slightly worse than AGCRN in some metrics on some datasets, and it is all in the long-term prediction part, but the gap is within 1%. At the same time, the model in this chapter has achieved a relatively large improvement in long-term prediction accuracy compared to the model in the previous chapter, which can also prove from the side that the model in this step not only has an advantage in short-term prediction, but also the efforts made in long-term prediction are effective and it does not lag behind the current advanced methods.
[0174] Example 2
[0175] This embodiment is a traffic flow prediction system based on feature fusion and frequency enhancement. The system is used to implement the traffic flow prediction method as described in Example 1, including: a data preprocessing module, a dynamic spatio-temporal graph generation module, and a hierarchical encoder-decoder model;
[0176] The data preprocessing module is used to preprocess the input traffic flow data and external meteorological data, including data cleaning, normalization, time alignment, and denoising processing;
[0177] The dynamic spatio-temporal graph generation module constructs a dynamic spatio-temporal graph that fuses a geographical neighbor graph and a functional similarity graph based on the preprocessed traffic flow data. The geographical neighbor graph is the inherent spatial attribute of traffic monitors and traffic networks, and the functional similarity graph is generated by multi-level wavelet decomposition and dynamic time warping of time series traffic flow data;
[0178] The hierarchical encoder-decoder model is constructed based on Transformer, receives the sliced data, and outputs traffic flow prediction data after passing through the activation function;
[0179] Among them, the dynamic spatio-temporal graph generation module further includes a wavelet analysis sub-module, a dynamic time warping sub-module, a multi-time granularity division sub-module, and a spatio-temporal attention convolution module;
[0180] The hierarchical encoder-decoder model further includes an encoder sub-module, a slice attention sub-module, a frequency enhancement channel attention sub-module, a decoder sub-module, and an activation function sub-module that are sequentially arranged.
[0181] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A traffic flow prediction method based on feature fusion and frequency enhancement, characterized in that It includes the following steps: Obtain traffic flow data and external meteorological data of traffic monitors and perform preprocessing; Construct a dynamic spatio-temporal graph generation module, which constructs a dynamic spatio-temporal graph integrating a geographical neighbor graph and a functional similarity graph based on the preprocessed traffic flow data. The geographical neighbor graph is the inherent spatial attribute of traffic monitors and traffic networks, and the functional similarity graph is generated by multi-level wavelet decomposition and dynamic time warping of traffic flow data based on time series; Input the dynamic spatio-temporal graph into the spatio-temporal attention convolutional layer to obtain a spatio-temporal matrix; Fuse the spatio-temporal matrix with the external meteorological data processed by the long short-term memory network one or more times, and use the output after passing through a fully connected layer of the fused data as traffic flow prediction data; or, Perform adjacent time slice processing on the fused data, construct a hierarchical encoder-decoder model based on Transformer, and generate traffic flow prediction data through the model for the data after slice processing.
2. The method according to claim 1, characterized in that The preprocessing includes data cleaning, normalization, time alignment, and denoising processing; Among them, the time alignment aligns traffic flow data and external meteorological data at a preset time interval to generate a matrix including weather conditions and air quality data; The denoising performs multi-level wavelet decomposition on traffic flow data and removes high-frequency noise through soft / hard thresholding.
3. The method according to claim 1, characterized in that, The specific method for generating the geographical neighbor graph is as follows: Based on the spatial data of traffic flow, define the initial adjacency matrix of the traffic network according to the road topology structure. For two nodes in the traffic network, if they are adjacent, the matrix element is 1, otherwise it is 0. The processed matrix is the geographical neighbor graph; The specific method for generating the functional similarity graph is as follows: Based on the traffic flow in time series, perform multi-level wavelet decomposition and denoising processing on each traffic flow time series by using wavelet transform; calculate the similarity of the denoised sequences between different nodes through dynamic time warping, select the first several node pairs with the smallest dynamic time warping distance, and construct a functional similarity graph in matrix form.
4. The method according to claim 1, wherein Before the dynamic spatio-temporal graph is input into the spatio-temporal convolutional layer, time granularity division is also performed, which is divided into recent time segments, daily cycle time segments, and weekly cycle time segments; The recent time segment is continuously intercepted along the time axis at the current time t0 with a length of T h to form a data sequence, which is used to capture the short-term dynamic change trend of traffic flow; The daily cycle time segment is based on the sampling rate q and intercepts a data segment with a length of T within the corresponding time window before each day to capture the daily periodic change pattern; d ; The weekly cycle time segment is based on the same time node every week and extracts a historical data segment with a length of T w for modeling the regular changes in the fixed pattern every week.
5. The method according to claim 1, wherein The spatio-temporal attention convolutional layer is one or multiple serially arranged. Each spatio-temporal attention convolutional layer includes a time attention layer, a spatial attention layer, a graph convolutional layer, a time axis convolutional layer, and a residual convolutional layer arranged in sequence.
6. The method according to claim 1, wherein The adjacent time slice is specifically: the time dimension information of a fixed length of each monitor is sliced into segments, and then connections are established between the segments; the length of the segment is a hyperparameter; the specific slicing method is: seg i,d = {x t,d | (i - 1)×L < t ≤ i×L} where, seg i,d represents the finally segmented segment, i represents the segment index of the time series, and x t,d represents the value of the d-th dimension at time t in the original time series, and L represents the length of the segment.
7. The method according to claim 1, wherein The model includes an encoder, a slice attention layer, a frequency enhancement channel attention layer, a decoder, and an activation function arranged in sequence; The initial layer of the encoder is the output of the adjacent time slice algorithm. After that, in each layer, the adjacent two slices in the same dimension of the previous layer are merged, and then the merged result is multiplied by a learnable matrix to obtain the final result of this layer. Finally, the final result is input into the slice attention to capture cross-time and cross-dimensional dependencies to obtain the output of the encoder of this layer; The number of layers of the decoder is the same as that of the encoder. Each decoder layer takes the feature matrix output by the corresponding encoder layer as input and then outputs a decoded two-dimensional matrix.
8. The method according to claim 7, characterized in that, The slice attention layer includes temporal slice attention and feature-axis slice attention; The temporal slice attention uses the multi-head attention of the Transformer to capture the dependencies between time periods on the same dimension and serves as the input to the feature-axis slice attention; The feature-axis slice attention sets a fixed number of learnable vectors for each time step as the intermediate layer; The intermediate layer first aggregates messages from all dimensions through the self-attention mechanism and uses the vectors of all dimensions as keys and values; Then, the intermediate layer distributes the received messages between dimensions by using the dimension vectors and the aggregated messages as keys and values to establish a full-to-full connection between all dimensions.
9. The method according to claim 7, wherein The frequency-enhanced channel attention layer uses the discrete cosine transform to extract the frequency-domain information of the input features. First, the features are divided into several subgroups according to the input dimension, and each subgroup is processed by the corresponding discrete cosine transform components from low frequency to high frequency; then, a gated mechanism with sigmoid activation is selected to capture the channel dependencies and comprehensively extract the temporal information from the frequency domain.
10. A traffic flow prediction system based on feature fusion and frequency enhancement, characterized in that, The system is used to implement the traffic flow prediction method according to any one of claims 1-9, and includes: a data preprocessing module, a dynamic spatio-temporal graph generation module, and a hierarchical encoder-decoder model; The data preprocessing module is used to preprocess the input traffic flow data and external meteorological data, including data cleaning, normalization, time alignment, and denoising processing; The dynamic spatio-temporal graph generation module constructs a dynamic spatio-temporal graph that fuses the geographical neighbor graph and the functional similarity graph based on the preprocessed traffic flow data. The geographical neighbor graph is the inherent spatial attribute of traffic monitors and traffic networks, and the functional similarity graph is generated by multi-level wavelet decomposition and dynamic time warping of the time-series traffic flow data; The hierarchical encoder-decoder model is constructed based on the Transformer, receives the sliced data, and outputs the traffic flow prediction data after passing through the activation function; Among them, the dynamic spatio-temporal graph generation module further includes a wavelet analysis sub-module, a dynamic time warping sub-module, a multi-time granularity division sub-module, and a spatio-temporal attention convolution sub-module; The hierarchical encoder-decoder model further includes an encoder sub-module, a slice attention sub-module, a frequency-enhanced channel attention sub-module, a decoder sub-module, and an activation function sub-module arranged in sequence.
Citation Information
Patent Citations
Traffic flow time sequence prediction method based on double-domain normalization
CN118762513A
Cited By
Metro station passenger flow prediction method and device based on multi-relation self-attention time-space diagram neural network, and electronic equipment
CN121903093A
Freight vehicle flow prediction method based on GPS track data
CN122416738A
Determination of an optimal vehicle maneuvering plan in a traffic congestion situation
US12662149B2
Determination of an optimal vehicle maneuvering plan in a traffic congestion situation
US20250333071A1