Traffic flow prediction method and device based on dynamic comparative learning and multi-scale 3D convolution
Through the combination of dynamic contrast learning and multi-scale 3D convolution, the existing traffic flow prediction model is solved in capturing nonlinear space-time dependence and dynamic emergencies, and accurate modeling and rapid response to traffic flow are achieved, which improves prediction accuracy and real-timeness.
Patent Information
- Application Number
- CN202510742205.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Existing traffic flow prediction models are difficult to capture the synergy between nonlinear space-time dependence, dynamic emergencies and environmental factors, and traditional methods have shortcomings in real-time and accuracy.
The combination of dynamic contrast learning and multi-scale 3D convolution is adopted, and the global and local spatiotemporal characteristics of traffic flow data are extracted through the multi-scale 3D convolution module, and the positive and negative sample strategies are automatically adjusted in combination with dynamic contrast learning, dynamic temperature coefficients are introduced, and the environmental feature gating fusion module is used for deep fusion to build a joint loss function optimization model.
It significantly improves the prediction accuracy and real-time response capabilities for emergencies and complex environmental conditions, and can more sensitively capture the dynamic changes in traffic flow, meeting the needs of smart traffic management and autonomous driving.
Smart Images

Figure CN120260295A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of urban traffic flow prediction, and particularly relates to a traffic flow prediction method and device based on dynamic contrast learning and multi-scale 3D convolution. Background Art
[0002] With the continuous deepening of urbanization and the construction of smart cities, modern traffic management is undergoing a transformation from traditional regulation to intelligent prediction and decision-making. Real-time traffic data is massively recorded through sensors, cameras, and other data collection devices distributed throughout the city, providing rich information resources for traffic flow prediction. Accurate traffic flow prediction can provide timely regulation basis for traffic management departments, optimize signal light configuration and road scheduling. At the same time, it can also assist autonomous driving vehicles in selecting the optimal path, reducing accident risks, reducing environmental pollution, and providing support for logistics distribution and public travel planning. Currently, traditional time series analysis and machine learning methods, such as ARIMA, SVM, random forest, etc., can achieve certain effects in some scenarios, but they often rely on preset linear relationships and fixed spatio-temporal dependence structures, and it is difficult to capture the non-linear and dynamic change characteristics in traffic flow. In addition, these methods have limited efficiency in processing large-scale data and are difficult to meet the requirements of real-time prediction.
[0003] In recent years, the application of deep learning technology has gradually become a trend. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) and their variants (such as LSTMs, GRUs) have made significant progress in modeling the spatio-temporal dependence of traffic data. For example, spatio-temporal graph convolutional networks (STGCNs) effectively extract the topological structure and dynamic features of traffic networks by combining graph convolution and convolutional operations. For instance, Chinese patent document CN115482666A discloses a multi-graph convolutional neural network traffic prediction method based on data fusion, which uses 3D convolution to form a 3D RepVGG component to extract the spatio-temporal features of Euclidean traffic flow data and uses them as the features of each node in the traffic topology graph; designs a periodic gating logic unit to process non-Euclidean traffic flow data, improves the model's ability to extract time features, and uses a clustering algorithm to identify and quantify the regional traffic conditions, improving the model's generalization ability and reducing the number of parameters; constructs multiple traffic topology graphs based on regional traffic states, non-Euclidean traffic flow data, and node feature data to enhance the model's ability to extract long-distance spatial features and further improve the prediction accuracy.
[0004] Chinese Patent Document CN115601960A discloses a multi-modal traffic flow prediction method and system based on graph contrast learning, which establishes local and global traffic flow heterogeneous graphs based on historical traffic flow data; encodes the global and local traffic flow heterogeneous graphs to obtain corresponding heterogeneous graph traffic flow features; calculates the mutual information of the local traffic flow heterogeneous graph traffic flow features to optimize the local traffic flow heterogeneous graph traffic flow features; multiple local traffic flow heterogeneous graph traffic flow features are fused into a global traffic flow feature through an attention mechanism, and graph contrast learning is performed with the global traffic flow heterogeneous graph traffic flow features to optimize the global traffic flow heterogeneous graph traffic flow features; the optimized local and global traffic flow heterogeneous graph traffic flow features are input into a spatial graph convolutional neural network to predict multi-modal traffic flows respectively.
[0005] However, in current traffic flow prediction models, deep learning methods have two limitations: on the one hand, not all deep learning models are convolution-based. Although some models use a convolutional kernel structure with a fixed scale to extract spatio-temporal features, this only represents the design idea of convolutional neural networks. Other models, such as those based on recurrent neural networks or self-attention mechanisms, also have similar problems in spatio-temporal feature extraction, that is, it is difficult to adaptively balance between capturing global trends and local details. On the other hand, traditional contrast learning methods usually rely on preset fixed time windows to generate samples. That is to say, during the data sampling process, positive and negative samples are extracted from a fixed-length time period. This fixed strategy cannot be flexibly adjusted according to the drastic changes in real-time traffic conditions, so it may not be able to capture sudden traffic events or subtle traffic fluctuations in a timely manner. Again, the interaction between environmental features and traffic flow features is usually only processed through simple splicing, and the dynamic synergy between the two in different scenarios is not fully reflected.
[0006] Based on this, the present invention proposes a traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution to solve the problems existing in the above-mentioned prior art. Summary of the Invention
[0007] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides a traffic flow prediction method and device based on dynamic contrast learning and multi-scale 3D convolution. The core idea is to achieve accurate modeling and rapid response to traffic flow changes through an end-to-end processing flow, from raw data collection to final prediction output; this method aims to solve the deficiencies of traditional prediction methods in capturing non-linear spatio-temporal dependencies, dynamic emergency response, and collaborative modeling of environmental factors.
[0008] The detailed technical solution of the present invention is as follows: A traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution, the method includes: S1. Obtain traffic trajectory data and environmental data, and perform preprocessing operations on them respectively to obtain a traffic feature matrix and an environmental feature vector. Map the environmental feature vector to the spatial dimension of the traffic grid to obtain an environmental feature matrix, and splice the traffic feature matrix and the environmental feature matrix to obtain a target spatio-temporal matrix as the model input; S2. Construct a traffic flow prediction model including a multi-scale 3D convolution module, a dynamic contrast learning module, an environmental feature gated fusion module, and a spatio-temporal deconvolution prediction module. Use the target spatio-temporal matrix as input data, and construct a joint loss function based on the dynamic contrast learning loss, the gated fusion loss, and the prediction loss to optimize the traffic flow prediction model; Among them, the multi-scale 3D convolution module is used to extract different-scale feature information of the traffic feature matrix in the target spatio-temporal matrix and fuse them to generate a traffic fusion feature; the dynamic contrast learning module is used to calculate the traffic difference between two consecutive time steps in the traffic fusion feature to obtain a real-time traffic change rate, set a positive and negative sample generation strategy based on the real-time traffic change rate, and obtain positive and negative sample pairs based on this strategy. Use the positive and negative sample pairs as training data, and introduce a dynamic temperature coefficient to construct a dynamic contrast loss to optimize the traffic flow prediction model; the environmental feature gated fusion module is used to perform a non-linear transformation on the environmental feature matrix in the target spatio-temporal matrix to obtain a high-dimensional environmental embedding representation. Based on the traffic fusion feature and the high-dimensional environmental embedding representation, calculate a dynamic gating weight, and use the dynamic gating weight to calculate a target high-dimensional fusion feature. Construct a gated fusion loss based on the traffic fusion feature and the target high-dimensional fusion feature to train the traffic flow prediction model; the spatio-temporal deconvolution prediction module is used to perform an upsampling operation on the target high-dimensional fusion feature to generate a prediction map of future traffic flow; and construct a prediction loss based on the prediction result to optimize the traffic flow prediction model.
[0009] Preferably according to the present invention, in S1, the preprocessing operation on the traffic trajectory data specifically includes: Perform data cleaning on the traffic trajectory data, and perform a spatial discretization operation on the cleaned traffic trajectory data to obtain the traffic flow values of each grid in all time windows, and generate a corresponding single-frame two-dimensional flow map ; Stack the single-frame two-dimensional flow maps generated by each grid in all time windows in chronological order to form a trajectory spatio-temporal matrix ; Dynamically fill the missing values in the trajectory spatio-temporal matrix by using a spatio-temporal weighted interpolation method, combining historical time series and adjacent spatial information. The specific formula is: (2); (2); In Equation (2): represents the time step and the grid position is traffic flow value; represents the time step size, ranging from 1 to ; represents the time step and the grid position is traffic flow value; represents the position of the adjacent grid; represents the time step and the adjacent grid position traffic flow value; is the weight coefficient in the time dimension; is the weight coefficient in the space dimension; M(i,j) represents the neighborhood of the current grid (i,j), including the local neighborhood range formed by the adjacent grids in its up, down, left, right, and diagonal directions; For each grid in the interpolated trajectory spatio-temporal matrix perform max-min normalization independently: (3); In Equation (3): represents the grid position normalized traffic flow value, with a value range of [0, 1]; represents the grid position original traffic flow value; , are respectively the maximum and minimum values of the grid position ;
[0010] According to the preference of the present invention, in S1, the environmental data includes weather type, holiday identifier, and temperature data. Preprocessing operations are performed on the environmental data, specifically including: Perform data cleaning on the environmental data, and fill in missing values for the cleaned environmental data, including: interpolate and fill in the missing weather type according to the weather type in adjacent time periods; use linear interpolation method to fill in the missing temperature data, and its formula is: , where represents the original temperature value at time step ; represents the original temperature value at time step ; represents the original temperature value at time step ; Perform encoding operations on the filled environmental data, including: perform one-hot encoding on the weather type to generate a dimension of vector; encoding the holiday identifier using binary encoding to generate a vector with a dimension of ; converting the temperature data into a scalar value through Z-score normalization, with its dimension being = 1, that is: (4); In formula (4): represents the normalized temperature value, and its value range is usually [-1, 1] or [0, 1], specifically depending on the distribution of the original temperature value distribution; represents the mean of the temperature data in the training set; represents the standard deviation of the temperature data in the training set; Concatenating the encoding results of the weather type, holiday identifier, and temperature data to form an environmental feature vector .
[0011] According to the preference of the present invention, in S1, taking the vectors of the encoded weather type and holiday identifier in the environmental feature vector as global features and performing a feature mapping operation to generate a global feature matrix, and its formula is: (5); In formula (5): represents the time step and the grid position is global feature vector; represents the global feature vector at time step ; Taking the normalized temperature value as the spatial distribution feature and performing a feature mapping operation to generate a spatial distribution feature matrix: (6); In formula (6): represents the spatial distribution feature vector at time step and the grid position is ; represents the th sensor's observation value at time step ; represents the number of sensors; represents the grid position to the th sensor's Euclidean distance; is a smoothing constant, such as = 1e−6, used to prevent the denominator from being zero; Concatenating the generated global feature matrix and spatial distribution feature matrix to form a complete environmental feature matrix : (7); Concatenate the traffic feature matrix and the environmental feature matrix in the embedding layer to obtain the target spatio-temporal matrix : (8).
[0012] Preferably according to the present invention, in S2, the multi-scale 3D convolution module includes parallel coarse-grained branches, medium-grained branches, and fine-grained branches; wherein, the convolution kernel of the coarse-grained branch is , and is equipped with a stride , and its expression is: (9); The convolution kernel of the medium-grained branch , and the stride is set to 1, and its expression is: (10); The fine-grained branch uses a 1×1×1 convolution kernel, and its expression is: (11); In the above formula: represents the output feature tensor after being processed by the coarse-grained branch; represents the output feature tensor after being processed by the medium-grained branch; represents the output feature tensor after being processed by the fine-grained branch; Fuse the feature tensors generated by each branch in the multi-scale 3D convolution module to form a fused feature tensor ; and use the channel attention mechanism to integrate the feature tensors generated by each branch in the multi-scale 3D convolution module: (12); In formula (12): represents any one of the coarse-grained coarse, medium-grained mid, or fine-grained fine branches; represents the weight of each branch in the multi-scale 3D convolution module; is a learnable parameter; represents the th branch Perform global average pooling operation on the feature map to obtain a channel descriptor; Based on the above, finally obtain the traffic fusion feature as: (13).
[0013] Preferably according to the present invention, in S2, the calculation formula of the real-time traffic change rate is: (14); In formula (14): represents the real-time traffic change rate; represents the time step of the fused feature tensor, represents the time step of the fused feature tensor; Based on the real-time traffic change rate, a positive and negative sample generation strategy is set, including: if the real-time traffic change rate is higher than the set threshold, then for positive samples select the fused feature tensors of the current time step and its adjacent time steps and ; for negative samples sample from the area that is at least L grid cells away from the current grid position in space; if the real-time traffic change rate is lower than the set threshold, then use the historical data of the same period as positive samples , and sample negative samples in the same area ; Project the traffic fused features into a low-dimensional embedding space through a linear mapping , and measure the similarity between positive and negative sample pairs through cosine similarity: (15); And use an improved InfoNCE loss function to perform contrastive training on positive and negative sample pairs, and its formula is: (16); In formula (16): represents the InfoNCE loss function; represents the similarity between positive sample pairs; represents the similarity between a positive sample and the th negative sample; represents the number of negative samples; is a dynamic temperature coefficient, used to adjust the similarity distribution between positive and negative samples; and, the dynamic temperature coefficient is: (17); In formula (17): , are both adjustment coefficients; is an exponential function.
[0014] According to the preference of the present invention, in S2, for the target spatio-temporal matrix Perform a non-linear transformation on the environmental feature matrix E to obtain the high-dimensional embedded representation of the environment as follows: (18); In formula (18): represents the environmental feature matrix; represents the fully connected layer; represents the high-dimensional embedded representation of the environment; Calculate the attention weights between traffic features and environmental features through the multi-head attention mechanism, including: converting the traffic fusion feature into query Q, mapping the high-dimensional embedded representation of the environment to key and value V, that is: (19); In formula (19): represents the query vector, which is calculated by the traffic fusion feature through the query fully connected layer ; represents the key vector, which is calculated by the high-dimensional embedded representation of the environment through the key fully connected layer ; represents the value vector, which is calculated by the high-dimensional embedded representation of the environment through the value fully connected layer ; then there is: (20); In formula (20): represents the dimension of the key vector; represents the transpose operation on the key vector matrix ; Calculate the dynamic gating weight : (21); In formula (21): represents the Sigmoid activation function; represents the gating fully connected layer; Through adaptive fusion with the gating weight, obtain the target high-dimensional fusion feature : (22); Construct the loss of the environmental feature gating fusion module : (23); In formula (23): represents the gating fusion loss; represents the Traffic fusion features of a sample; Indicates the Target high-dimensional fusion features of the sample; Indicates the number of samples in the batch.
[0015] According to a preferred embodiment of the present invention, in S2, the target high-dimensional fusion features Are subjected to an upsampling operation to generate a prediction map of future traffic flow, including: The target high-dimensional fusion features Are subjected to a first-layer 3D transposed convolution operation, and its expression is: (24); Subsequently, a second-layer 3D transposed convolution operation is performed, and its expression is: (25); Subsequently, a prediction output is generated to generate a prediction map of future traffic flow: (26); In formula (26): Indicates the prediction map of future traffic flow, Is the number of future prediction time steps; Finally, a prediction loss is constructed based on the obtained prediction result : (27); In formula (27): Indicates the true traffic matrix of the i-th sample; Indicates the prediction result of the i-th sample; Is the total number of samples in the training set.
[0016] According to a preferred embodiment of the present invention, in S2, a joint loss function is constructed as: (28); In formula (28): α, ξ, and η are preset weight coefficients for balancing the prediction loss , the dynamic contrast learning loss And the gated fusion loss Of.
[0017] In another aspect of the present invention, there is provided an apparatus for implementing a traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution, the apparatus including: A data acquisition module, configured to acquire traffic trajectory data and environmental data and perform preprocessing operations on them respectively to obtain a traffic feature matrix and an environmental feature vector, map the environmental feature vector to the spatial dimension of a traffic grid to obtain an environmental feature matrix, and splice the traffic feature matrix and the environmental feature matrix to obtain a target spatio-temporal matrix as the input of the model; A model construction module, configured to construct a traffic flow prediction model including a multi-scale 3D convolution module, a dynamic contrast learning module, an environmental feature gated fusion module, and a spatio-temporal deconvolution prediction module, use the target spatio-temporal matrix as input data, and construct a joint loss function based on a dynamic contrast learning loss, a gated fusion loss, and a prediction loss to optimize the traffic flow prediction model; Among them, the multi-scale 3D convolution module is configured to extract feature information of different scales of the traffic feature matrix in the target spatio-temporal matrix and fuse them to generate a traffic fusion feature; The dynamic contrast learning module is configured to calculate the traffic difference between two consecutive time steps in the traffic fusion feature to obtain a real-time traffic change rate, set a positive and negative sample generation strategy based on the real-time traffic change rate, obtain positive and negative sample pairs based on this strategy, use the positive and negative sample pairs as training data, and introduce a dynamic temperature coefficient to construct a dynamic contrast loss to optimize the traffic flow prediction model; The environmental feature gated fusion module is configured to perform a non-linear transformation on the environmental feature matrix in the target spatio-temporal matrix to obtain a high-dimensional environmental embedding representation, calculate a dynamic gated weight based on the traffic fusion feature and the high-dimensional environmental embedding representation, calculate a target high-dimensional fusion feature using the dynamic gated weight, and construct a gated fusion loss based on the traffic fusion feature and the target high-dimensional fusion feature to train the traffic flow prediction model; The spatio-temporal deconvolution prediction module is configured to perform an upsampling operation on the target high-dimensional fusion feature to generate a prediction map of future traffic flow; and construct a prediction loss based on the prediction result to optimize the traffic flow prediction model.
[0018] Compared with the prior art, the beneficial effects of the present invention are: (1) Through the deep integration of dynamic contrast learning and multi-scale 3D convolution technology, the present invention has achieved a significant performance improvement in the field of traffic flow prediction, especially showing excellent performance in dealing with sudden traffic events and complex environmental conditions.
[0019] (2) The present invention adopts multi-scale 3D convolution, and synchronously extracts the global and local spatio-temporal features of traffic flow data through convolution kernels of coarse, medium, and fine scales, effectively solving the deficiencies of existing methods in capturing the non-linear and dynamic change characteristics of traffic flow, and providing a comprehensive and detailed description of traffic states.
[0020] (3) The present invention combines dynamic contrast learning, automatically adjusts the selection strategy of positive and negative samples according to the real-time traffic change rate, and introduces a dynamic temperature coefficient to adaptively smooth the similarity distribution in the feature space, making the model more sensitive to sudden traffic events, thereby significantly improving the response ability to emergencies.
[0021] (4) The present invention utilizes a gated attention mechanism to dynamically fuse the embedded environmental features and traffic flow features, fully reflecting the dynamic collaborative effect of environmental factors and traffic flow features in different scenarios, and enhancing the adaptability and robustness of the model to external condition changes.
[0022] (5) In terms of real-time performance, the present invention realizes real-time response to drastic fluctuations in traffic flow through spatio-temporal deconvolution prediction combined with a sliding window and dynamic step size adjustment strategy. The sliding window strategy can dynamically adjust the prediction step size according to the traffic change rate, shortening the step size when the traffic fluctuates violently, ensuring that the model can quickly respond to traffic dynamic changes, meet the requirements of real-time prediction, and provide a wide range of application prospects for fields such as intelligent traffic management, autonomous driving assistance, and logistics distribution optimization. Description of the Drawings
[0023] Figure 1 is a flowchart of the traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to the present invention.
[0024] Figure 2 is a flowchart of the preprocessing of traffic trajectory data and environmental data in Embodiment 1 of the present invention.
[0025] Figure 3 is a schematic flowchart of constructing a trajectory spatio-temporal matrix in Embodiment 1 of the present invention.
[0026] Figure 4 is a network structure diagram of the traffic flow prediction model constructed in Embodiment 1 of the present invention.
[0027] Figure 5 is a schematic flowchart of feature extraction and fusion through multi-scale 3D convolution in Embodiment 1 of the present invention.
[0028] Figure 6 is a schematic flowchart of training the model through dynamic contrast learning in Embodiment 1 of the present invention.
[0029] Figure 7 is a schematic flowchart of fusing traffic features and environmental features through environmental feature gating in Embodiment 1 of the present invention.
[0030] Figure 8 is a schematic flowchart of deconvolution prediction based on target high-dimensional fusion features in Embodiment 1 of the present invention. Detailed implementation manners
[0031] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0032] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners of the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0033] In the case of no conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0034] Aiming at the deficiencies of the prior art, the present invention provides a traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution. First, through the multi-scale 3D convolution module, the global and local spatio-temporal features of traffic flow data are synchronously extracted by using coarse-grained, medium-scale, and fine-grained convolution kernels, so as to realize a comprehensive description of the traffic state. Subsequently, the dynamic contrast learning module is adopted to automatically adjust the selection strategy of positive and negative samples according to the traffic flow change rate, and a dynamic temperature coefficient is introduced to adaptively smooth the similarity distribution in the feature space, making the model more sensitive to sudden traffic events. Finally, the present invention uses the gated attention mechanism to deeply and dynamically fuse the embedded environmental features and traffic flow features to achieve adaptive cooperation between the two, thereby improving the overall prediction accuracy and real-time performance.
[0035] The traffic flow prediction method and device based on dynamic contrast learning and multi-scale 3D convolution of the present invention will be further described below in conjunction with specific embodiments.
[0036] Embodiment 1 This embodiment provides a traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution, aiming to achieve accurate modeling and rapid response to traffic flow changes. Refer Figure 1 , the method includes: S1. Obtain traffic trajectory data and environmental data and perform preprocessing operations on them respectively to obtain a traffic feature matrix and an environmental feature vector , map the environmental feature vector to the spatial dimension of the traffic grid to obtain an environmental feature matrix , splice the traffic feature matrix with the environmental feature matrix to obtain a target spatio-temporal matrix as the model input .
[0037] In this embodiment, the input data mainly includes two categories: traffic trajectory data and environmental data. Among them, the traffic trajectory data records the GPS information of vehicles. These data reflect the geographical location and driving speed of each vehicle at a specific moment, and are an important basis for subsequent traffic flow modeling. In addition to traffic trajectory data, external environmental data closely related to traffic conditions is also collected, mainly including weather types, such as sunny, rainy, and snowy; holiday flags, where 0 represents non-holiday and 1 represents holiday; and temperature data, whose unit is °C. These environmental data will be used as auxiliary inputs in the subsequent model to help explain traffic flow fluctuations caused by weather changes or special dates.
[0038] After obtaining the above traffic trajectory data and environmental data, first preprocess them, as shown in Figure 2 . Specifically: In this embodiment, first perform a preprocessing operation of data cleaning on the obtained traffic trajectory data and environmental data. Data cleaning is a key step to ensure the quality of the model input data. The data cleaning of traffic trajectory data and environmental data mainly includes three steps: removing duplicate data, correcting data formats, and removing noise data: 1) Removing duplicate data: Check the GPS records of vehicles and de-duplicate them based on timestamps and geographical coordinates. If it is found that there are multiple GPS records of a vehicle at the same timestamp, only keep one of them. Check the environmental data records and de-duplicate them based on the observation time. For example, if there are duplicate weather data records at the same timestamp, only keep one. 2) Correcting data formats: Unify the timestamp format to ensure that the time representation of all traffic trajectory data is consistent; perform format checks and corrections on the longitude and latitude data among them, and remove outliers that are significantly inconsistent with geographical coordinates. Ensure that the time representation of environmental data is consistent with the trajectory data; perform format checks on data such as weather types and temperatures to ensure the correctness of the data. 3) Removing noise data: For traffic trajectory data, remove abnormal points based on the reasonableness of vehicle speeds. For example, if a record shows that the vehicle speed exceeds twice the speed limit of the section, it is determined as noise data and removed. For environmental data, remove obviously incorrect weather state records, and check and correct abnormal fluctuations in temperature data.
[0039] In addition, the construction of the trajectory spatio-temporal matrix is the core foundation for the efficient operation of the model. For the preprocessing of traffic trajectory data, it also includes converting the original traffic trajectory data into a three-dimensional tensor suitable for the input of the model's 3D convolution through strict mathematical modeling and spatio-temporal alignment strategies. Specifically, as shown in Figure 3 , this process mainly includes two steps: spatial discretization and time serialization: 1) Perform a spatial discretization operation on the traffic trajectory data, that is, map from GPS coordinates to a regular grid to obtain the traffic flow value of each grid in all time windows. within.
[0040] The original traffic trajectory data (such as taxi GPS points, shared bicycle riding records, etc.) contains longitude, latitude, timestamp and speed information. To achieve spatial discretization, first divide the urban area into a regular grid of H×W, such as 500m×500m, and use the Mercator projection to convert the longitude and latitude coordinates of the trajectory points therein into plane coordinates. The formula is as follows: (1); In formula (1): represents the longitude value; represents the latitude value; represents the converted plane coordinate value.
[0041] This projection ensures the uniform distribution of the geographical coordinates of the trajectory points on the plane and avoids deformation errors in high-latitude regions.
[0042] The traffic flow value of each grid is determined by counting the number of trajectory points falling into the grid within a fixed time window (such as 5 minutes), and a single-frame two-dimensional flow map is generated . For example, the flow value of a certain grid may reach 300 vehicles / 5 minutes during the morning rush hour and drop below 50 at night. This step not only converts unstructured traffic trajectory data into a structured matrix, but also provides a basis for spatial alignment for subsequent multi-scale 3D convolution to extract local and global features.
[0043] 2) Perform a time serialization operation on the traffic flow values of each grid in all time windows within.
[0044] To capture the temporal evolution law of traffic flow, continuous time needs to be divided into equal-length time windows (such as 5 minutes), and the single-frame two-dimensional flow maps generated by each grid in all time windows are stacked in chronological order , forming a three-dimensional spatio-temporal tensor, that is, a trajectory spatio-temporal matrix , represents the time dimension. Taking the prediction of traffic flow in the next 12 hours as an example, the time dimension T of the input three-dimensional spatio-temporal tensor is T = 144 = 12 hours × 12 frames / hour. This three-dimensional structure directly adapts to the input requirements of 3D convolution: the convolutional kernel slides along the time axis in the time dimension to capture the periodicity (such as morning and evening rush hours) and suddenness (such as traffic accidents) of traffic flow changes; in the spatial dimension, the convolutional kernel extracts local spatial patterns (such as the spread of congestion at intersections) and global trends (such as the migration of vehicle flows between regions) on the H×W grid.
[0045] Based on the constructed three-dimensional spatio-temporal tensor above, in order to further improve the effect of feature extraction from the three-dimensional spatio-temporal tensor, the method of this embodiment designs a key adaptation scheme: in the multi-scale 3D convolution branch, the coarse-grained branch uses a 5×5×5 convolution kernel with a large stride, effectively capturing long-term trends across regions while downsampling; the medium-grained branch uses a 3×3×3 convolution kernel to maintain a high spatio-temporal resolution, being able to extract dynamic changes within local regions and connect global and local features as a whole; the fine-grained branch uses a 1×1×1 convolution kernel to maintain the original resolution and focuses on capturing instantaneous local traffic fluctuations caused by factors such as signal light switching. Through this multi-scale adaptation design, the model can comprehensively extract and fuse traffic flow change information in both the time and space dimensions, thus providing a solid feature basis for subsequent traffic flow prediction.
[0046] Furthermore, in practical applications, some grids may have missing data due to sensor failures or sparse trajectories. Therefore, it is also necessary to process the obtained trajectory spatio-temporal matrix for missing values.
[0047] 3) By using the spatio-temporal weighted interpolation method, dynamically fill in the missing values by combining historical time series and adjacent spatial information, that is: (2); In formula (2): represents the traffic flow value at time step and the grid position is ; represents the time step length, ranging from 1 to ; represents the traffic flow value at time step and the grid position is ; represents the position of the adjacent grid; represents the traffic flow value at time step and the adjacent grid position ; the weight coefficient in the time dimension assigns higher weights to recent data (for example, data points within the past 10 minutes); the weight coefficient in the space dimension is based on distance attenuation of the grid position; M(i,j) is defined as the neighborhood of the current grid (i,j), including its adjacent grids in the up, down, left, right, and diagonal directions, forming a local neighborhood range.
[0048] This formula (2) dynamically fills in the missing values by combining historical time series data and adjacent spatial information, ensuring data continuity and avoiding noise introduced by zero filling or simple mean interpolation.
[0049] Furthermore, to improve the stability of model training, the filled trajectory spatio-temporal matrix is subjected to data normalization operation.
[0050] 4) Perform min-max normalization independently for each grid: (3); In formula (3): represents the grid position The normalized traffic flow value, with a value range of [0, 1]; represents the grid position The original traffic flow value; and are respectively the maximum and minimum values of the grid position This operation scales the traffic flow value of each grid to the interval [0, 1], while retaining the relative differences between grids. For example, the maximum traffic flow in the city center grid may be 5 times that in the suburbs. The normalization parameters
[0051] are statistically based on the training set to avoid information leakage in the test set. and Based on the above operations, the normalized traffic flow values of all grid positions in the urban area can be obtained and constructed into a traffic feature matrix
[0052] For the preprocessing of environmental data, it also includes missing value filling and encoding operations, etc. The specific implementation process is as follows: .
[0053] For the preprocessing of environmental data, it also includes missing value filling and encoding operations, etc. The specific implementation process is as follows: 1) Fill in the missing values of the environmental data. For weather data, if the weather record of a certain time period is missing, it is interpolated and filled according to the weather conditions of adjacent time periods. For example, if the weather data at 10:00 am on the same day is missing, the weather patterns at 9:55 am and 10:05 am on the same day are weighted and averaged, and the result is used as the weather data at 10:00 am on the same day. For temperature data, the linear interpolation method is used to fill in the missing values, that is: . Among them, represents the time step The original temperature value; represents the time step The original temperature value; represents the time step The original temperature value. For extreme outliers, such as -50°C or 50°C, based on the historical temperature range in the same period, that is, the mean ± 3 , they are excluded or corrected. Among them, represents the standard deviation of the temperature data in the training set.
[0054] It should be understood that for holiday markings, they are directly marked according to the list of national legal holidays without interpolation.
[0055] 2) Perform an encoding operation on the filled environmental data. For weather data, according to the included weather types, one-hot encoding is used to generate a vector with a dimension of . For example, if the weather types include sunny, rainy, and snowy, then the encoding dimension = 3, sunny is represented as , rainy is represented as , and snowy is represented as . Through one-hot encoding, the model can better learn the impact of different weather types on traffic flow. For holidays, binary encoding 0 / 1 is used, and the dimension = 1. For example, weekdays are represented as 0, and holidays are represented as 1. Holidays usually lead to changes in traffic flow, and through binary marking, it is convenient for the model to distinguish between weekday and holiday traffic flow patterns. For temperature data, it is converted into a scalar value through Z-score standardization, and the dimension = 1, and the formula is: (4); In formula (4): represents the standardized temperature value, and its value range usually is [-1, 1] or [0, 1], specifically depending on the distribution of the original temperature value ; represents the mean of the temperature data in the training set; represents the standard deviation of the temperature data in the training set.
[0056] By calculating the mean and standard deviation of the temperature data, it is converted into a standard normal distribution so that the model can better handle the impact of temperature changes on traffic flow. For example, assuming the mean of the temperature data is 20°C and the standard deviation is 5°C, after Z-score standardization, the temperature data will be converted into a standard normal distribution, thereby reducing the impact of temperature changes on the model.
[0057] Concatenate the above encoding results in the feature dimension to form a unified environmental feature vector , where d = . Then map the obtained environmental feature vector to the spatial dimension of the traffic grid to obtain the environmental feature matrix .
[0058] Specifically, after completing the encoding of the environmental data, it is necessary to map the obtained environmental feature vector to the spatial dimension of the traffic grid to ensure the spatial alignment of the environmental features and traffic features. In this step, the environmental feature vector The vector of the encoded weather type and holiday identifier in the middle is used as the global feature, and the standardized temperature value is used as the spatial distribution feature, and the feature mapping operations are performed separately. The specific steps are as follows: 1) The global feature is a feature related to the entire transportation network, which usually does not change with the change of spatial position, or its change can be ignored. Therefore, it is necessary to copy these features to all grid positions to generate the global feature matrix , represents the dimension of the global feature vector. That is: (5); In formula (5): represents the time step and the grid position is of the global feature vector; represents the time step of the global feature vector.
[0059] 2) For the spatial distribution feature, the influence of this feature varies with different spatial positions. It is necessary to allocate it to each grid position according to the temperature data and grid position information. The inverse distance weighting (IDW) method is used to interpolate the temperature data into the transportation grid to generate the spatial distribution feature matrix , represents the dimension of the spatial distribution feature vector. That is: (6); In formula (6): represents the time step and the grid position is of the spatial distribution feature vector; represents the th sensor's observation value at time step ; represents the number of sensors; represents the grid position to the th sensor's Euclidean distance; is a smoothing constant, such as = 1e−6, which is used to prevent the denominator from being zero.
[0060] Finally, the generated global feature matrix and the spatial distribution feature matrix are concatenated to form the complete environmental feature matrix , that is: (7).
[0061] This step ensures the alignment of environmental features in the spatial dimension, providing a basis for subsequent feature fusion.
[0062] Next, the traffic feature matrix is concatenated with the complete environmental feature matrix in the embedding layer to obtain the target spatio-temporal matrix as the input of the model. Specifically, after the spatial mapping of the environmental feature vector , the normalized traffic feature matrix and the complete environmental feature matrix containing global features and spatial distribution features are concatenated in the channel dimension to form a unified input representation , where C is the number of channels of traffic features; d is the total dimension of environmental features, and its expression is: (8).
[0063] This step integrates the two types of features into the model input for the subsequent multi-scale 3D convolution module to extract spatio-temporal features.
[0064] S2. Construct a traffic flow prediction model including a multi-scale 3D convolution module, a dynamic contrast learning module, an environmental feature gated fusion module, and a spatio-temporal deconvolution prediction module. Use the target spatio-temporal matrix as input data, and construct a joint loss function based on the dynamic contrast learning loss, the gated fusion loss, and the prediction loss to optimize the traffic flow prediction model.
[0065] As shown in Figure 4 , this step specifically includes the following operations.
[0066] The multi-scale 3D convolution module is used to extract feature information of different scales from the traffic feature matrix in the target spatio-temporal matrix and fuse them to generate traffic fusion features . Specifically, the preprocessed target spatio-temporal matrix is used as input, and the traffic feature matrix in it contains feature information such as traffic flow and is subjected to feature extraction by the multi-scale 3D convolution module. This module adopts a parallel branch structure to extract spatio-temporal features from different scales and forms a comprehensive feature representation through gradual fusion, providing rich semantic information for subsequent modules.
[0067] As shown in Figure 5 , the multi-scale 3D convolution module contains three parallel branches: the first branch is the coarse-grained branch, which uses a convolutional kernel with a larger size , for example, 5×5×5, with a relatively large stride s1. Its main function is to downsample the original input, thereby capturing the overall traffic trends across regions and over long time periods. For example, during peak hours, the trend of vehicle flow spreading from the outskirts of the city to the city center will be clearly reflected in this branch. Its expression is: (9); In formula (9): represents the output feature tensor after being processed by the coarse-grained branch, which contains the downsampled global trend information. Here, the BatchNorm operation is used to stabilize the training process, and the ReLU activation function introduces non-linearity, making the feature response more obvious. The feature tensor generated by the coarse-grained branch is downsampled both in time and space, retaining the global trend information, but the details may be blurred.
[0068] To balance the global trend and local details, the second branch is the medium-grained branch, which uses a medium-sized convolutional kernel , for example, 3×3×3, and the stride is set to 1 to maintain the original resolution of the input. This can not only capture the dynamic characteristics of traffic propagation between adjacent grids but also model the detailed information of local areas. The expression of this branch is: (10); In formula (10): represents the output feature tensor after being processed by the medium-grained branch, which retains the original resolution of the input and can capture the detailed information of local areas. Since the resolution is not changed, the output features retain the spatio-temporal detailed information and have strong expressive ability for the spread of local congestion at intersections and the short-term fluctuations of the mutual influence between vehicles.
[0069] The third branch is the fine-grained branch. For extremely subtle local changes, such as the instantaneous traffic fluctuations caused by traffic accidents or signal light changes, this branch uses a 1×1×1 convolutional kernel to perform a single-point non-linear transformation on each spatio-temporal unit. Its expression is: (11); In formula (11): represents the output feature tensor after being processed by the fine-grained branch, which can highlight the sharp local changes. Although this branch does not change the spatial structure, it can highlight the sharp local changes through non-linear mapping, making the model more sensitive in detecting emergencies.
[0070] During the multi-scale feature extraction process, branches of different scales may produce feature maps with inconsistent sizes. To ensure the effective alignment of information from each branch during subsequent feature fusion, a bilinear interpolation method is adopted after the output of each branch to unify them to the same spatial resolution. The interpolated feature maps have the same number of channels as before interpolation, ensuring that no additional channel alignment operations are required during subsequent concatenation or channel attention fusion. This operation is different from the interpolation of the original data during the data preprocessing stage and is an alignment process for feature maps to ensure that the coarse, medium, and fine-scale information can be seamlessly fused in a unified spatial dimension to form a fused feature tensor: , represents the sum of the output channel numbers of the three branches of the multi-scale 3D convolution module.
[0071] To further integrate features at each scale, a channel attention mechanism is adopted. The specific approach is to first perform global average pooling on each branch to obtain the channel descriptors of each branch, and then calculate the weight of each branch through the learnable parameter as follows: (12); In Equation (12): represents any one of the branches of coarse, mid, or fine; represents the weight of each branch in the multi-scale 3D convolution module; is a learnable parameter; represents the global average pooling operation on the feature map of the th branch to obtain the channel descriptor.
[0072] Based on the above, the finally obtained traffic fusion feature is: (13).
[0073] The traffic fusion feature here contains rich spatio-temporal information from global trends to local details, providing a solid foundation for subsequent dynamic contrast learning and environment fusion. Among them, represents the number of channels of the traffic fusion feature , which is the total number of channels after weighted fusion of the channel numbers of the feature tensors of the coarse, medium, and fine branches.
[0074] The dynamic contrast learning module is used to calculate the traffic flow difference between two consecutive time steps in the traffic fusion feature to obtain the real-time traffic flow change rate , and based on the real-time traffic flow change rate Set the positive and negative sample generation strategy, obtain positive and negative sample pairs based on this strategy, use the positive and negative sample pairs as training data, and introduce a dynamic temperature coefficient , construct a dynamic contrast loss to optimize the traffic flow prediction model.
[0075] Specifically, in order to enhance the model's response ability to sudden changes in traffic flow, the dynamic contrast learning module further optimizes the traffic fusion features . Specifically, the parameters Figure 6 , including the following steps: First is the calculation of the traffic change rate , calculate the traffic difference between two consecutive time steps in the time dimension, and measure it by the Euclidean distance: (14); In formula (14): represents the real-time traffic change rate; represents the fusion feature tensor at time step , represents the fusion feature tensor at time step , both contain rich spatio-temporal information from global trends to local details.
[0076] The real-time traffic change rate reflects the severity of the current traffic state. If exceeds the preset threshold, it is regarded as a sudden scenario. And according to the value of , different positive and negative sample generation strategies are adopted: 1) In this embodiment, if is relatively high, indicating a sudden scenario, the positive sample selects the fusion feature tensors of the current time step and its adjacent time steps and , and the positive sample is represented as ; while the negative sample is sampled from the area that is at least L grid cells away from the current grid position in space, where L is a threshold predefined according to the specific traffic network layout. For example, in the urban traffic grid, L can be defined as 20% of the total number of grids or a fixed value such as 5 grid cells, and the negative sample is represented as . By setting a clear distance threshold, it is ensured that the selected negative sample has a significant distance difference from the current grid in space, thereby highlighting the obvious traffic differences between different regions and enhancing the model's discrimination ability for instantaneous dynamic changes. 2) If is relatively low, indicating a stable scenario and the overall features are relatively stable. At this time, in order to further capture local details and minor changes, historical data of the same period can be used as positive samples , such as data of the same period in the previous 24 hours, and negative samples are sampled within the same area , so that the model can pay more attention to the subtle traffic changes within the same geographical location, rather than the overall traffic differences caused by geographical location differences. In the dynamic contrast learning module, cosine similarity is used as the core metric to measure the similarity between positive and negative sample pairs. Specifically, the traffic fusion features are projected into a low-dimensional embedding space through linear mapping , which facilitates the calculation of similarity. For the positive sample representation and the negative sample representation in the embedding space, the cosine similarity is calculated by the following formula: (15); In formula (15): and represent the feature vectors of the positive and negative samples in the low-dimensional embedding space respectively. This similarity function quantifies the similarity between samples by calculating the dot product of two vectors and normalizing to obtain the cosine value of the cosine angle between them.
[0077] Based on the above strategy, positive and negative samples are obtained. Further, in this embodiment, an improved InfoNCE loss function is used to perform contrast training on the positive and negative sample pairs. The improved InfoNCE loss function is: (16); In formula (16): represents the InfoNCE loss function; represents the similarity between positive sample pairs, using cosine similarity; represents the similarity between the positive sample and the th negative sample; is the dynamic temperature coefficient, used to adjust the similarity distribution between positive and negative samples; represents the number of negative samples, that is, the number of negative samples used for contrast learning with the positive sample. Among them, the dynamic temperature coefficient is: (17); In formula (17): , are both adjustment coefficients; is the exponential function. The introduction of the dynamic temperature coefficient enables the model to adaptively smooth the similarity distribution in the feature space, thereby improving the ability to distinguish positive samples and reject negative samples in sudden scenarios.
[0078] By optimizing this contrastive loss function, the similarity of positive sample pairs in the embedding space is increased, while the similarity of negative sample pairs is decreased. This optimization strategy helps the model better capture the dynamic change characteristics of traffic flow, especially in sudden scenarios, thereby significantly enhancing the model's ability to distinguish sudden changes in traffic flow. Specifically, in sudden scenarios, positive sample pairs , such as the current frame and its adjacent frames, have their similarity maximized; while negative sample pairs (such as sampling from regions that are at least L grid cells away from the current grid position in space) have their similarity minimized. This optimization strategy enables the model to more acutely identify and distinguish different traffic states, thereby enhancing the model's adaptability and prediction accuracy in complex traffic scenarios. In this way, the dynamic contrastive learning module can effectively improve the model's response ability to sudden changes in traffic flow, providing a more accurate basis for traffic flow prediction.
[0079] The environmental feature gated fusion module is used to perform a non-linear transformation on the environmental feature matrix E in the target spatio-temporal matrix to obtain a high-dimensional environmental embedding representation . Based on the traffic fusion feature and the high-dimensional environmental embedding representation , the dynamic gating weight is calculated, and the dynamic gating weight is used to calculate the target high-dimensional fusion feature . Based on the traffic fusion feature and the target high-dimensional fusion feature , a gated fusion loss is constructed to train the traffic flow prediction model.
[0080] Specifically, to make full use of the impact of external environmental data on traffic flow, after completing basic numerical encoding in the preprocessing stage, the environmental feature gated fusion module further performs deep feature learning and mapping on environmental data to obtain a more discriminative environmental representation. Thus, environmental information can be better integrated with traffic features in depth.
[0081] Refer Figure 7 , this process includes the following steps: 1) Deep feature learning, that is, by performing a non-linear transformation on the environmental feature matrix E in the target spatio-temporal matrix to obtain a high-dimensional environmental embedding representation .
[0082] In the preprocessing stage, after the environmental features are spatially mapped, they have been preliminarily integrated with the traffic features to form a complete environmental feature matrix E that includes global features and spatial distribution features, and are concatenated with the traffic features in the channel dimension to form a unified input representation. However, this preliminarily integrated feature representation may not be sufficient to capture the complex interaction relationship between the environment and traffic. To further enhance the discriminability of the features, the environmental feature gated fusion module performs a non-linear transformation on the environmental feature matrix E through a fully connected layer to generate a deep environmental embedding representation, thereby achieving more effective traffic flow feature fusion. Its expression is: (18); In Equation (18): represents the environmental feature matrix; represents the fully connected layer; represents the high-dimensional environmental embedding representation for subsequent multi-head attention interaction and dynamic gated fusion.
[0083] 2) Multi-head attention interaction, that is, calculating the attention weights between the traffic features and the environmental features through the multi-head attention mechanism.
[0084] To capture the complex interaction between the environmental data and the traffic features, first convert the traffic fusion feature into the query Q; at the same time, map the high-dimensional environmental embedding representation to the key and the value V: (19); In Equation (19): represents the query vector, which is calculated by the traffic fusion feature through the query fully connected layer It represents the query representation of the traffic features at different positions and is used to query the keys and values related to the environmental features in the attention mechanism; represents the key vector, which is calculated by the deep high-dimensional environmental embedding representation through the key fully connected layer It represents the key representation of the environmental features at different positions and is used to match with the query vector to calculate the attention weights; represents the value vector, which is calculated by the deep high-dimensional environmental embedding representation through the value fully connected layer It represents the value representation of the environmental features at different positions and is used to output the weighted environmental feature values according to the calculated attention weights.
[0085] Through the multi-head attention mechanism, calculate the traffic fusion feature and the high-dimensional environmental embedding representation The attention weights between them reflect the influence degree of environmental information on each spatio-temporal position, and its expression is: (20); In formula (20): represents the dimension of the key vector, which is used to scale the dot product result to prevent the gradient of the softmax function from vanishing or exploding due to an overly large dot product result; represents the transpose operation on the key vector matrix to facilitate matrix multiplication calculation; the transposed matrix is multiplied by the query vector matrix Q to obtain the dot product between the query vector and the key vector, reflecting their correlation.
[0086] 3) Dynamic gating mechanism: By concatenating the traffic fusion feature with the attention output, and passing through a linear mapping and Sigmoid activation, the dynamic gating weight is calculated: (21); In formula (21): represents the Sigmoid activation function; represents the gating fully connected layer, which is used to generate the gating weight to control the fusion ratio of environmental features and traffic features.
[0087] Then, through adaptive fusion with the gating weight, the target high-dimensional fusion feature is obtained: (22).
[0088] With this design, the environmental feature gating fusion module can not only make full use of the embedding layer concatenation result in the preprocessing stage, but also further enhance the interaction between environmental features and traffic features, thereby improving the model's understanding ability and prediction accuracy for complex traffic scenarios. Finally, based on the traffic fusion feature and the target high-dimensional fusion feature the loss of the environmental feature gating fusion module is constructed to measure the fusion effect of environmental features and traffic features. The specific calculation method is as follows: (23); In formula (23): represents the gating fusion loss; represents the traffic fusion feature of the th sample; represents the target high-dimensional fusion feature of the th sample; represents the number of samples in the batch.
[0089] The spatio-temporal deconvolution prediction module is used to perform upsampling on the target high-dimensional fusion feature to generate a prediction map of future traffic flow; and construct a prediction loss based on the prediction result to optimize the traffic flow prediction model.
[0090] Based on the target high-dimensional fusion feature obtained from the above steps , where is the size of the time dimension, that is, the length of the time series, H×W is the spatial size of the urban area, and C is the number of channels, which has integrated the multi-scale spatio-temporal features of traffic flow and external environment information. Next, the goal of the spatio-temporal deconvolution prediction module is to upsample this high-dimensional feature, restore it to the original spatio-temporal resolution, and generate a prediction map of future traffic flow.
[0091] First, since the multi-scale 3D convolution module and the fusion process may cause downsampling in the time and space dimensions, in order to ensure that the prediction result has the same resolution as the original data, 3D deconvolution needs to be used to gradually restore the downsampled features. The complete deconvolution prediction process is as Figure 8 shown, and its process includes the following steps: 1) Preliminary upsampling, that is, the target high-dimensional fusion feature goes through the first layer of 3D deconvolution operation.
[0092] This layer uses a relatively large-sized convolution kernel , such as 4×4×4, and the corresponding stride , such as 2×2×2, to achieve fast upsampling. The deconvolution operation not only increases the time and space dimensions but also fuses and smooths the features to a certain extent. Its expression is: (24); In formula (24): is the output of the first layer of 3D deconvolution. At this time, the size of the output feature is significantly enlarged compared to the target high-dimensional fusion feature , laying a foundation for subsequent refined processing.
[0093] 2) Layer-by-layer refined upsampling, that is, perform the second layer of 3D deconvolution to further restore details and the complete spatio-temporal structure.
[0094] This layer uses a medium-sized convolution kernel (such as 3×3×3) and a smaller stride (such as 1×1×1) to refine the rough feature obtained in the previous step, and at the same time further restore the original spatial and time dimensions; its expression is: (25); In Equation (25): is the output of the second-layer 3D transposed convolution. At this stage, the model not only upsamples the data but also reconstructs the features through the transposed convolution operation to ensure that the global trends and local details extracted by the previous multi-scale 3D convolution can be fully restored during the upsampling process.
[0095] 3) Final prediction output, that is, generating a prediction map of future traffic flow.
[0096] The last transposed convolution uses a 1×1×1 convolution kernel and a 1×1×1 stride to compress the number of channels and map the features to the dimension of the target output. At the same time, the output is constrained to an appropriate numerical range through the Sigmoid activation function; its expression is:[[]] (26); In Equation (26): represents the prediction map of future traffic flow, where is the number of time steps for future prediction, which is consistent with the previous upsampling target.
[0097] During the entire spatio-temporal transposed convolution prediction process, the model gradually restores the time and space resolutions reduced by the convolution operation through multiple transposed convolutions, while ensuring the continuous transmission of feature information. The upsampling operation is not only used to restore the data size, but more importantly, to "decode" the complex spatio-temporal features extracted by the previous modules through the activation function and batch normalization, and convert them into interpretable traffic flow values. This process forms an end-to-end closed loop with the previous multi-scale 3D convolution feature extraction, making each step of data change from the input spatio-temporal matrix to the final prediction output have clear mathematical definitions and physical meanings.
[0098] Finally, a prediction loss is constructed based on the obtained prediction results :[[]] (27); In Equation (27): represents the true flow matrix of the i-th sample (extracting samples from the target spatio-temporal matrix to form , and each sample corresponds to the flow data within a time window); represents the prediction result of the i-th sample; is the total number of samples in the training set.
[0099] Furthermore, in this embodiment, during the training process of the traffic flow prediction model, it comprehensively considers the prediction loss, the InfoNCE loss of the dynamic contrast learning module, and the loss of the environmental feature gating fusion module. These loss functions are weighted and summed through preset weights α, ξ, and η to form a joint loss function, which jointly guides the optimization process of the model.
[0100] The joint loss function adopts the form of weighted summation, that is: (28); In Equation (28): α, ξ, and η are preset weight coefficients used to balance the prediction loss , the dynamic contrast learning loss , and the gating fusion loss . By adjusting the values of α, ξ, and η, the relative importance of the three losses in the total loss can be controlled, thereby achieving fine control of the model training process.
[0101] In addition, to achieve real-time performance, this method adopts a sliding window strategy and dynamically adjusts the window step size according to the current traffic flow change rate during the prediction stage. When a significant traffic fluctuation is detected, such as when the real-time traffic flow change rate Rt is high, the model will automatically shorten the step size of the prediction window to more quickly respond to traffic dynamic changes. Through backpropagation, the prediction error (such as the mean square error MSE) is fed back to each deconvolution layer and the previous modules, continuously optimizing the model parameters, and finally improving the overall prediction accuracy and real-time response ability. This spatio-temporal deconvolution prediction module is closely connected to the front-end multi-scale feature extraction and dynamic contrast learning module, ensuring a seamless conversion from high-dimensional fusion features to prediction outputs, providing an accurate and real-time solution for urban traffic flow prediction.
[0102] The following further illustrates the training, validation, and testing processes of the model in this method in combination with specific application scenarios.
[0103] In the model training stage, two publicly available datasets, TaxiBJ and BikeNYC, are used as the original data sources. Among them, TaxiBJ contains the GPS trajectory data of taxis in Beijing, and BikeNYC covers the riding records of public bicycles in New York City. Based on the aforementioned data preprocessing method in this embodiment, the data in these two datasets are preprocessed. After the data preprocessing is completed, the data is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2 to ensure time continuity. Among them, the input format of the training set is a three-dimensional spatio-temporal matrix of 12 consecutive hours, and the label is the traffic flow matrix for the next hour. To further improve the adaptability of the model to complex traffic scenarios, external environmental data (such as weather, temperature, holidays) are also preprocessed, encoded, and standardized, and are adaptively fused with traffic flow data through an environmental feature gating fusion module to form a high-dimensional fusion feature that simultaneously contains spatio-temporal traffic features and environmental impacts.
[0104] In terms of traffic trajectory data preprocessing, first, the original traffic trajectory data is divided into fixed grids according to urban areas. For example, TaxiBJ is divided into 32×32 grids, and BikeNYC is divided into 16×8 grids. The traffic flow data in each grid is counted within a fixed time window (such as every 30 minutes or every 5 minutes) to generate a single-frame two-dimensional traffic flow map. , and then stacked into a three-dimensional spatio-temporal tensor in chronological order. To ensure consistent numerical scales, the maximum-minimum normalization method is used for each grid to scale the traffic flow values to the range [0, 1], and its normalization parameters are all statistically based on the training set. At the same time, the environmental data preprocessing includes data collection, alignment in time and space, and encoding. Specifically, the weather type uses one-hot encoding, such as sunny: ; rainy: ; snowy: ); the holidays use binary encoding, where 0 represents weekdays and 1 represents holidays; the temperature data is processed using Z-score standardization. The processed environmental data is assigned to the corresponding traffic grids through a spatial mapping method, and then preliminarily concatenated with the traffic flow tensor in the embedding layer, and an environmental feature gating fusion module is used to achieve adaptive fusion based on multi-head attention and a dynamic gating mechanism to generate a fusion feature representation that simultaneously contains spatio-temporal and environmental information.
[0105] During the model training process, the AdamW optimizer is selected, and the initial learning rate is set to 0.001. The learning rate is dynamically adjusted through the cosine annealing strategy, and the minimum learning rate is limited to 1e-5 to prevent oscillation. During training, 32 samples are input in each batch, and the maximum number of training epochs is set to 100. At the same time, an early stopping mechanism is introduced, and the patience value is set to 10 epochs to monitor the loss of the validation set and prevent overfitting. To enhance the generalization ability of the model, a Dropout layer (ratio 0.2) is added after the convolutional layer. At the same time, L2 weight decay is set, and the coefficient is set to 1e-4. Gradient clipping is adopted, and the threshold is set to 5.0 to avoid gradient explosion. After the relevant parameter configuration is completed, the constructed joint loss function is used to optimize and train the model. In addition, to evaluate the consistency of the embedding space, the following formula is used to calculate the similarity difference between positive and negative sample pairs: (29).
[0106] The prediction accuracy is quantified by the mean absolute error MAE and the root mean square error RMSE. The formulas are as follows: (30); (31).
[0107] If the mean absolute error MAE of the validation set does not decrease for 3 consecutive epochs, the learning rate will be halved; if the similarity difference between positive and negative samples is insufficient, the contrast loss weight will be adjusted, for example, increased to 0.5. During the training process, the environmental data adaptively adjusts its contribution at different spatio-temporal positions through the environmental feature gating fusion module, which generates gating weights , determining the proportion of environmental information in feature fusion. This mechanism enables factors such as weather, temperature, and holidays to have an effective impact on the model parameters during the backpropagation process, thereby improving the prediction accuracy of the model under abnormal environmental conditions.
[0108] In the model validation stage, the validation set is independently divided from the training data. Its core goal is to monitor the generalization ability of the model and the optimization of hyperparameters. During the validation process, in addition to evaluating the consistency of the embedding space by calculating the similarity difference between positive and negative sample pairs using formula (29), the overall prediction accuracy is also evaluated by the mean absolute error MAE and the root mean square error RMSE metrics. If the validation set metrics do not improve for multiple consecutive epochs, the learning rate and the contrast loss weight are dynamically adjusted to further improve the model performance.
[0109] During the model testing phase, the test set is strictly independent of the training and validation data. For example, the test data of TaxiBJ is selected from the last two months of 2019. The sliding window strategy is adopted in the testing phase, with a fixed window length of 12 hours and an initial sliding step of 1 hour. When the flow rate change rate Rt > 0.3 is detected, the step is dynamically shortened to 0.5 hours to respond more quickly to sudden traffic events. In addition to the global MAE and RMSE, local errors are calculated for traffic hotspots (such as business district grids), and the single-frame prediction delay is recorded, with the target being less than 30 milliseconds. Furthermore, cross-dataset tests are designed, such as TaxiBJ→BikeNYC and BikeNYC→TaxiBJ, to evaluate the model generalization ability by comparing the prediction effects under different traffic modes.
[0110] In summary, throughout the entire process, the TaxiBJ and BikeNYC datasets run through the training, validation, and testing phases, combining the preprocessing and deep fusion of environmental data, and using dynamic learning strategies, multi-objective joint loss functions, and strict evaluation mechanisms to systematically improve the accuracy and real-time performance of traffic flow prediction, providing a solid data support and model foundation for intelligent decision-making in complex traffic scenarios. Further, through the deep fusion of dynamic contrast learning and multi-scale 3D convolution technology, the present invention has achieved significant performance improvement in the field of traffic flow prediction, especially performing excellently in dealing with sudden traffic events and complex environmental conditions.
[0111] Embodiment 2 This embodiment provides a device for implementing a traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution. The device includes: A data acquisition module, configured to acquire traffic trajectory data and environmental data and perform preprocessing operations respectively to obtain a traffic feature matrix and an environmental feature vector, map the environmental feature vector to the spatial dimension of the traffic grid to obtain an environmental feature matrix, and splice the traffic feature matrix and the environmental feature matrix to obtain a target spatio-temporal matrix as the model input; A model construction module, configured to construct a traffic flow prediction model including a multi-scale 3D convolution module, a dynamic contrast learning module, an environmental feature gated fusion module, and a spatio-temporal deconvolution prediction module, use the target spatio-temporal matrix as input data, and construct a joint loss function based on the dynamic contrast learning loss, the gated fusion loss, and the prediction loss to optimize the traffic flow prediction model; Wherein, the multi-scale 3D convolution module is configured to extract different-scale feature information of the traffic feature matrix in the target spatio-temporal matrix and perform fusion to generate a traffic fusion feature; The dynamic contrast learning module is used to calculate the traffic difference between two consecutive time steps in the traffic fusion feature to obtain the real-time traffic change rate. Based on the real-time traffic change rate, a positive and negative sample generation strategy is set, and positive and negative sample pairs are obtained based on this strategy. The positive and negative sample pairs are used as training data, and a dynamic temperature coefficient is introduced to construct a dynamic contrast loss for optimizing the traffic flow prediction model; The environmental feature gated fusion module is used to perform a non-linear transformation on the environmental feature matrix in the target spatio-temporal matrix to obtain an environmental high-dimensional embedding representation. Based on the traffic fusion feature and the environmental high-dimensional embedding representation, a dynamic gating weight is calculated, and the target high-dimensional fusion feature is calculated using the dynamic gating weight. A gated fusion loss is constructed based on the traffic fusion feature and the target high-dimensional fusion feature for training the traffic flow prediction model; The spatio-temporal deconvolution prediction module is used to perform an upsampling operation on the target high-dimensional fusion feature to generate a prediction map of future traffic flow; and a prediction loss is constructed based on the prediction result for optimizing the traffic flow prediction model.
[0112] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific embodiments of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution, characterized in that, The method includes: S1. Obtain traffic trajectory data and environmental data and perform preprocessing operations on them respectively to obtain a traffic feature matrix and an environmental feature vector. Map the environmental feature vector to the spatial dimension of the traffic grid to obtain an environmental feature matrix, and splice the traffic feature matrix and the environmental feature matrix to obtain a target spatio-temporal matrix as the input of the model; S2. Construct a traffic flow prediction model including a multi-scale 3D convolution module, a dynamic contrast learning module, an environmental feature gated fusion module, and a spatio-temporal deconvolution prediction module. Use the target spatio-temporal matrix as input data, and construct a joint loss function based on the dynamic contrast learning loss, the gated fusion loss, and the prediction loss to optimize the traffic flow prediction model; Among them, the multi-scale 3D convolution module is used to extract feature information of different scales of the traffic feature matrix in the target spatio-temporal matrix and fuse them to generate a traffic fusion feature; The dynamic contrast learning module is used to calculate the traffic difference between two consecutive time steps in the traffic fusion feature to obtain a real-time traffic change rate. Set a positive and negative sample generation strategy based on the real-time traffic change rate, and obtain positive and negative sample pairs based on this strategy. Use the positive and negative sample pairs as training data, and introduce a dynamic temperature coefficient to construct a dynamic contrast loss to optimize the traffic flow prediction model; The environmental feature gated fusion module is used to perform a non-linear transformation on the environmental feature matrix in the target spatio-temporal matrix to obtain a high-dimensional environmental embedding representation. Calculate a dynamic gated weight based on the traffic fusion feature and the high-dimensional environmental embedding representation, and use the dynamic gated weight to calculate a target high-dimensional fusion feature. Construct a gated fusion loss based on the traffic fusion feature and the target high-dimensional fusion feature to train the traffic flow prediction model; The spatio-temporal deconvolution prediction module is used to perform an upsampling operation on the target high-dimensional fusion feature to generate a prediction map of future traffic flow; and construct a prediction loss based on the prediction result to optimize the traffic flow prediction model.
2. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, wherein In S1, the preprocessing operation on the traffic trajectory data specifically includes: Clean the traffic trajectory data, and perform spatial discretization on the cleaned traffic trajectory data to obtain the traffic flow values of each grid within all time windows, and generate corresponding single-frame two-dimensional flow maps , where represents the size of the regular grid in the urban area where the traffic trajectory is located; Stack the single-frame two-dimensional flow maps generated by each grid in all time windows in chronological order to form a trajectory spatio-temporal matrix , representing the time dimension; Dynamically fill the missing values in the trajectory spatio-temporal matrix through spatio-temporal weighted interpolation method, combining historical time series and adjacent spatial information. The formula is as follows: in which: (2); In Equation (2): represents the time step and the traffic flow value at the grid position ; represents the time step size, ranging from 1 to ; represents the time step and the traffic flow value at the grid position ; represents the position of the neighboring grid; represents the time step and the traffic flow value at the neighboring grid position ; is the weight coefficient in the time dimension; is the weight coefficient in the space dimension; M(i, j) represents the neighborhood of the current grid (i, j), including the local neighborhood range formed by the adjacent grids in its up, down, left, right, and diagonal directions; Perform min-max normalization independently for each grid in the interpolated trajectory spatio-temporal matrix : (3); In formula (3): represents the grid position The normalized traffic flow value, with a value range of [0, 1]; represents the grid position of the original traffic flow value; and are respectively the maximum and minimum values of the grid position 3. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, wherein In S1, the environmental data includes weather type, holiday flag, and temperature data. The preprocessing operation on the environmental data specifically includes: Perform data cleaning on the environmental data, and fill in missing values for the cleaned environmental data, including: Interpolate and fill in the missing weather type according to the weather type in adjacent time periods; The missing temperature data is filled by using the linear interpolation method, and its formula is: , where represents the original temperature value at time step ; represents the original temperature value at time step ; represents the original temperature value at time step ; And perform an encoding operation on the filled environmental data, including: Perform one-hot encoding on the weather type to generate a vector with a dimension of ; Perform an encoding operation on the holiday identifier using binary encoding to generate a vector with a dimension of ; The temperature data is converted into a scalar value through Z-score normalization, and its dimension is = 1, and the formula is: (4); In formula (4): represents the standardized temperature value; represents the mean of the temperature data in the training set; represents the standard deviation of the temperature data in the training set; Concatenate the encoding results of the weather type, holiday identifier, and temperature data in the feature dimension to form a unified environmental feature vector , where d = .
4. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 3, characterized in that, In S1, use the vectors of the encoded weather type and holiday flag in the environmental feature vector as global features, and perform a feature mapping operation to generate a global feature matrix, and its formula is: (5); In formula (5): represents the time step and the global feature vector at the grid position is ; represents the time step ; And use the standardized temperature value as a spatial distribution feature, and perform a feature mapping operation to generate a spatial distribution feature matrix, and its formula is: (6); In Equation (6): represents the time step and the spatial distribution feature vector at the grid position ; represents the observation value of the th sensor at the time step ; represents the number of sensors; represents the Euclidean distance from the grid position to the th sensor; is a smoothing constant used to prevent the denominator from being zero; Splice the generated global feature matrix and spatial distribution feature matrix to form an environmental feature matrix, and its formula is: (7); In formula (7): represents the environmental feature matrix; Concatenate the traffic feature matrix and the environmental feature matrix in the embedding layer to obtain the target spatio-temporal matrix as follows: (8); In formula (8): represents the target spatio-temporal matrix.
5. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, wherein, In S2, the multi-scale 3D convolution module includes a parallel coarse-grained branch, a medium-grained branch, and a fine-grained branch; among them, the convolution kernel of the coarse-grained branch is , and is paired with a stride , that is: (9); In formula (9): represents the target spatio-temporal matrix; represents the output feature tensor after the coarse-grained branch processing; The convolution kernel of the medium-grained branch , and the stride is set to 1, and its expression is: (10); In formula (10): represents the output feature tensor after medium-grained branch processing; The fine-grained branch uses a 1×1×1 convolutional kernel, and its expression is: (11); In formula (11): represents the output feature tensor after fine-grained branch processing; Fuse the feature tensors generated by each branch in the multi-scale 3D convolution module to form a fused feature tensor , represents the sum of the output channel numbers of all branches of the multi-scale 3D convolution module; And a channel attention mechanism is used to integrate the feature tensors generated by each branch in the multi-scale 3D convolutional module, and its expression is: (12); In formula (12): represents any one of the coarse, mid, or fine branches; represents the weight of each branch in the multi-scale 3D convolution module; is a learnable parameter; represents the feature map of the th branch is subjected to global average pooling operation to obtain a channel descriptor; Based on the above, the final traffic integration features are obtained as follows: (13)。 6. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, characterized in that In S2, the calculation formula of the real-time traffic change rate is: (14); In formula (14): represents the real-time flow rate change; represents the time step of the fused feature tensor, represents the time step of the fused feature tensor; Based on the real-time traffic change rate, a positive and negative sample generation strategy is set, including: If the real-time traffic change rate is higher than the set threshold, indicating a burst scenario, then the positive sample selects the current time step and its adjacent time steps and the fused feature tensor , and the positive sample is represented as ; the negative sample is sampled from the area that is at least L grid cells away from the current grid position in space, and the negative sample is represented as ; If the real-time traffic change rate is lower than the set threshold, indicating a stable scenario, historical data for the same period is used as positive samples , and negative samples are sampled within the same area ; Project the traffic integration feature onto a low-dimensional embedding space through linear mapping , and measure the similarity between positive and negative sample pairs through cosine similarity. The formula is as follows: (15); In formula (15): and respectively represent the feature vectors of the positive sample and the negative sample in the low-dimensional embedding space; And the improved InfoNCE loss function is used to perform contrastive training on the positive and negative sample pairs, and the improved InfoNCE loss function is: (16); In Equation (16): represents the InfoNCE loss function; represents the similarity between positive sample pairs; represents the similarity between a positive sample and the th negative sample; is the dynamic temperature coefficient, used to adjust the similarity distribution between positive and negative samples; represents the number of negative samples; Among them, the dynamic temperature coefficient is as follows: (17); In formula (17): and are both adjustment coefficients; is an exponential function.
7. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, characterized in that In S2, perform a non-linear transformation on the environmental feature matrix in the target spatio-temporal matrix to obtain the environmental high-dimensional embedding representation as: (18); In formula (18): represents the environmental feature matrix; represents the fully connected layer; represents the high-dimensional embedded representation of the environment; Calculate the attention weights between traffic features and environmental features through the multi-head attention mechanism, including: Convert the traffic fusion feature to query Q, and map the high-dimensional environmental embedding representation to key and value V, i.e.: (19); In formula (19): represents the query vector, which is obtained from the traffic fusion feature by calculating through the query fully-connected layer ; represents the key vector, which is obtained from the high-dimensional environmental embedding representation by calculating through the key fully-connected layer ; represents the value vector, which is obtained from the high-dimensional environmental embedding representation by calculating through the value fully-connected layer ; Then there is: (20); In formula (20): represents the dimension of the key vector; represents the transpose operation on the key vector matrix ; Calculating dynamic gating weights : (21); In formula (21): represents the Sigmoid activation function; represents the gated fully connected layer; Obtain the target high-dimensional fusion feature through gated weight adaptive fusion : (22); Loss of the constructed environmental feature gating fusion module : (23); In formula (23): represents the gated fusion loss; represents the traffic fusion feature of the th sample; represents the target high-dimensional fusion feature of the th sample; represents the number of samples in the batch.
8. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, characterized in that In S2, perform an upsampling operation on the target high-dimensional fusion feature to generate a prediction map of future traffic flow, including: Perform the first 3D transposed convolution operation on the target high-dimensional fusion feature, and its expression is: (24); In formula (24): represents the target high-dimensional fusion feature; is the output of the first layer of 3D transposed convolution; Subsequently, perform the second 3D transposed convolution operation, and its expression is: (25); In formula (25): is the output of the second layer of 3D deconvolution; Subsequently, predict the output to generate a prediction map of future traffic flow: (26); In formula (26): represents the predicted graph of future traffic flow,[[]] is the number of time steps predicted for the future.[[]] Finally, a prediction loss is constructed based on the obtained prediction results : (27); In formula (27): represents the true traffic matrix of the i-th sample; represents the prediction result of the i-th sample; is the total number of samples in the training set.
9. The traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution according to claim 1, characterized in that, In S2, construct the joint loss function as: (28); In Equation (28): α, ξ, and η are preset weight coefficients for balancing the prediction loss , the dynamic contrast learning loss , and the gated fusion loss .
10. An apparatus for implementing a traffic flow prediction method based on dynamic contrast learning and multi-scale 3D convolution, characterized in that, The device includes: A data acquisition module, which is used to acquire traffic trajectory data and environmental data and perform preprocessing operations respectively to obtain a traffic feature matrix and an environmental feature vector, map the environmental feature vector to the spatial dimension of the traffic grid to obtain an environmental feature matrix, and concatenate the traffic feature matrix and the environmental feature matrix to obtain the target spatio-temporal matrix as the model input; A model construction module, which is used to construct a traffic flow prediction model including a multi-scale 3D convolutional module, a dynamic contrast learning module, an environmental feature gated fusion module, and a spatio-temporal transposed convolution prediction module, use the target spatio-temporal matrix as input data, and construct a joint loss function based on the dynamic contrast learning loss, the gated fusion loss, and the prediction loss to optimize the traffic flow prediction model; Among them, the multi-scale 3D convolutional module is used to extract different-scale feature information of the traffic feature matrix in the target spatio-temporal matrix and fuse them to generate a traffic fusion feature; The dynamic contrast learning module is used to calculate the traffic difference between two consecutive time steps in the traffic fusion feature to obtain the real-time traffic change rate, set a positive and negative sample generation strategy based on the real-time traffic change rate, and obtain positive and negative sample pairs based on this strategy, use the positive and negative sample pairs as training data, and introduce a dynamic temperature coefficient to construct a dynamic contrast loss to optimize the traffic flow prediction model; The environmental feature gated fusion module is used to perform a non-linear transformation on the environmental feature matrix in the target spatio-temporal matrix to obtain the environmental high-dimensional embedding representation, calculate the dynamic gated weight based on the traffic fusion feature and the environmental high-dimensional embedding representation, and use the dynamic gated weight to calculate the target high-dimensional fusion feature, and construct a gated fusion loss based on the traffic fusion feature and the target high-dimensional fusion feature to train the traffic flow prediction model; The spatio-temporal deconvolution prediction module is used to perform an upsampling operation on the target high-dimensional fusion feature to generate a prediction map of future traffic flow; and construct a prediction loss based on the prediction result to optimize the traffic flow prediction model.
Citation Information
Patent Citations
Traffic flow prediction method based on time-varying fusion graph convolutional network
CN118262517A
Traffic flow prediction system based on deep learning and dynamic network analysis and application method thereof
CN118675324A
Multi-task learning-based multi-mode trip flow collaborative prediction method, system and device
CN119862994A
Short-term traffic flow prediction method based on causal gated-low-pass graph convolutional network
US20240029556A1
Road traffic speed prediction method fusing multi-feature neural network
WO2024244300A1
Cited By
Binaryzation neural network model training method and device and computer equipment
CN120579592A
Traffic flow prediction method based on self-supervised spatio-temporal representation and scene adaptation
CN121708754A
Traffic flow prediction method based on self-supervised spatio-temporal representation and scene adaptation
CN121708754B
Traffic flow prediction method under high data missing rate condition and related equipment
CN121963490A