A traffic feature data interpolation method
By constructing a traffic feature data interpolation network and enhancing the ability to capture spatiotemporal features, the problem of insufficient feature capture in the spatiotemporal dimensions of existing models is solved, and more accurate traffic data interpolation and analysis are achieved.
Patent Information
- Application Number
- CN202510883929.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing traffic data interpolation model based on diffusion probability has insufficient feature capture capabilities in the spatiotemporal dimensions, which affects the accuracy of data analysis.
A traffic feature data interpolation network is constructed, including a conditional information construction module, a conditional feature extraction module, a noise estimation unit and a convolution layer. The spatiotemporal feature capture capability is enhanced through multi-scale trend information extraction, graph attention and graph convolution modules. The spatiotemporal collaborative dependency is mined using a multi-layer perceptron for noise estimation and interpolation.
The ability to capture the spatiotemporal characteristics of traffic data interpolation has been improved, and the predicted interpolated values are closer to the true values, which improves the accuracy and reliability of data analysis.
Smart Images

Figure CN120429556B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of traffic data processing, and particularly relates to a traffic feature data interpolation method. BACKGROUND
[0002] In the intelligent transportation system (ITS), traffic data interpolation is a key task to support core functions such as traffic prediction and congestion analysis. However, traffic data collected by sensor networks is often missing due to faults or weather factors, which seriously affects the accuracy of data analysis. In recent years, the interpolation model based on diffusion probability has shown application potential in the field of spatio-temporal data interpolation. This model extracts features by modeling conditional information and converts Gaussian noise into missing value estimates using inverse processes. Its flexible neural network architecture combined with attention mechanisms effectively avoids the error accumulation problem of recurrent neural networks. However, the existing interpolation model based on diffusion probability still has the problem of insufficient feature capture ability in the spatio-temporal dimension in traffic data interpolation. To solve the above problems, the present application proposes a traffic feature data interpolation method that can effectively enhance the feature capture ability in the spatio-temporal dimension. SUMMARY
[0003] In view of the defects and deficiencies in the prior art, the present application proposes a traffic feature data interpolation method.
[0004] To achieve the above-mentioned application purposes, the present application adopts the following technical solutions:
[0005] A traffic feature data interpolation method, comprising the following steps: obtaining a mask vector, an adjacency matrix, a Gaussian noise sample, and time embedding information based on the traffic feature data to be interpolated; inputting the traffic feature data to be interpolated, the mask vector, the Gaussian noise sample, and the time embedding information into a traffic feature data interpolation network, so as to obtain the estimated noise predicted.
[0006] The traffic feature data interpolation network comprises a conditional information construction module, a conditional feature extraction module, a noise estimation unit, a first Add layer, and a convolution layer connected in sequence; the output end of the second Add layer is connected to the input end of the noise estimation unit.
[0007] The conditional information construction module obtains traffic data estimation information based on the traffic feature data to be interpolated and the mask vector, and performs multi-scale extraction on the traffic data estimation information to obtain multi-scale information. The traffic data estimation information and the multi-scale information constitute the conditional information.
[0008] The conditional feature extraction module is used to capture the spatio-temporal features in the conditional information and the adjacency matrix, and obtain the conditional weight containing static and dynamic spatial correlation information and temporal correlation information.
[0009] The second Add layer is configured to add the traffic data estimation information, the Gaussian noise sample and the time embedding information to obtain noise information.
[0010] The noise estimation unit estimates noise based on the conditional weight, the adjacency matrix and the noise information.
[0011] The first Add layer is configured to add the noise estimation values output by the noise estimation unit.
[0012] The convolution layer is configured to perform convolution operation on the output result of the first Add layer to output estimated noise with rich spatiotemporal correlation information.
[0013] Preferably, the manner of obtaining the mask vector based on the traffic feature data to be imputed is that the mask vector is obtained by preprocessing the traffic feature data to be imputed, including the following steps: the values at the missing positions in the traffic feature data to be imputed are marked with 0, and the values at the non-missing positions in the traffic feature data to be imputed are marked with 1, so as to obtain the mask vector.
[0014] Preferably, the manner of obtaining the adjacency matrix based on the traffic feature data to be imputed is the same as the calculation manner of the weighted adjacency matrix W disclosed in “The emerging field of signal processing on graphs: Extending high-dimensional data analysis to networks and other irregular domains”.
[0015] Preferably, the manner of obtaining the Gaussian noise sample based on the traffic feature data to be imputed includes the following steps: the missing positions in the traffic feature data to be imputed are filled with standard Gaussian noise.
[0016] Preferably, the manner of obtaining the time embedding information based on the traffic feature data to be imputed is consistent with the manner of obtaining the Diffusion time embedding disclosed in “PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation”.
[0017] Preferably, the condition information construction module comprises a variational autoencoder and a multi-scale trend information extraction module; the variational autoencoder performs initial imputation on the traffic feature data with missing values to obtain initial imputation information; then, the initial imputation information and the mask vector are filled with missing values to obtain traffic data estimation information; then, the multi-scale trend information extraction module performs multi-scale extraction on the trend information in the traffic data estimation information to obtain multi-scale trend information; and the traffic data estimation information and the multi-scale trend information constitute the condition information.
[0018] Preferably, the multi-scale trend information extraction module comprises n trend information extraction units connected in parallel, n is greater than or equal to 3, each trend information extraction unit comprises a convolution layer and an average pooling layer connected in sequence, and the convolution kernel sizes of the convolution layers in different trend information extraction units are different.
[0019] Preferably, the condition feature extraction module comprises a Mamba module and a graph convolution module connected in parallel with the Mamba module, an output end of the Mamba module is connected with a normalization layer, output ends of the normalization layer are respectively connected with a self-attention layer and a graph attention layer, and an output end of the self-attention layer is also connected with an input end of the graph attention layer; the graph attention layer and the graph convolution module are both connected with an Add layer, and the Add layer is connected with a multi-layer perceptron.
[0020] Preferably, the noise estimation unit comprises first, second, third and fourth noise estimation modules connected in sequence, output ends of the first, second, third and fourth noise estimation modules are also connected with an input end of a first Add layer, an output end of a second Add layer is connected with an input end of the first noise estimation module in the noise estimation unit, and the input of the first noise estimation module in the noise estimation unit comprises a condition weight, an adjacency matrix and noise information.
[0021] Preferably, before the traffic feature data to be imputed, the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information are input into the traffic feature data imputation network to obtain the predicted estimation noise, the traffic feature data imputation network is trained to obtain a traffic feature data imputation network model.
[0022] Preferably, the traffic feature data imputation network is trained to obtain a traffic feature data imputation network model, and the training specifically comprises the following steps:
[0023] 1) Obtain a training set and a test set;
[0024] 2) Under the guidance of the total optimization loss, train the traffic feature data imputation network by using the training set to obtain a traffic feature data imputation network model; and the training specifically comprises the following steps:
[0025] The traffic feature data with missing values, the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information in the training set are input into the traffic feature data imputation network, the total optimization loss of the traffic feature data imputation network is calculated, and the total optimization loss is guided to perform back propagation, the weight parameters of the traffic feature data imputation network are updated, the training process of one epoch is completed, and the training process of the traffic feature data imputation network is completed after 100 epoch training processes, and the traffic feature data imputation network model is obtained.
[0026] Preferably, step 1) specifically comprises the following steps: dividing the traffic feature data with missing values and its corresponding mask vector, Gaussian noise sample and time embedding information in a time sequence in a ratio of 7:3 to obtain a sub-training set and a sub-test set, and then adding the adjacency matrix to the sub-training set and the sub-test set to obtain the training set and the test set; wherein the traffic feature data with missing values, the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information are obtained based on the traffic feature data of the existing traffic feature data set.
[0027] Preferably, the traffic feature data with missing values is obtained based on the traffic feature data of the existing traffic feature data set, comprising the following steps: sequentially performing normalization processing and mask processing on the traffic feature data in the existing traffic feature data set to obtain the traffic feature data with missing values.
[0028] Preferably, the way of obtaining the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information based on the traffic feature data of the existing traffic feature data set is the same as the way of obtaining the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information based on the traffic feature data to be imputed.
[0029] Preferably, the normalization processing on the traffic feature data of the existing traffic feature data set comprises the following steps: calculating the mean value μ and the standard deviation σ of all traffic feature data in the existing traffic feature data set along the spatial dimension; then, based on the mean value μ and the standard deviation σ, each traffic feature data is normalized to obtain the normalized traffic feature data.
[0030] Compared with the prior art, the application has the beneficial technical effects that:
[0031] The traffic feature data interpolation network constructed in the application not only has strong feature capturing capability in the time dimension, but also has strong feature capturing capability in the space dimension. In the application, the time features output by the normalization layer, the dynamic space features output by the graph attention layer, and the static space features output by the graph convolution module are added by the Add layer and then nonlinearly transformed by the multilayer perceptron (MLP) to mine the spatiotemporal collaborative dependency relationship between the time features, the dynamic space features, and the static space features, obtain the conditional weight containing the static and dynamic spatial correlation information and the time correlation information, and estimate the noise based on the conditional weight, the adjacency matrix, and the noise information to obtain the first noise estimation value to the fourth noise estimation value. The first Add layer adds the first noise estimation value to the fourth noise estimation value, and then performs convolution operation by the convolution layer to obtain the predicted estimated noise with rich spatiotemporal correlation information. It can be known through testing that the traffic feature data interpolation value predicted by the traffic feature data interpolation method described in the application is closer to the true value at the position corresponding to the missing value of the traffic feature data. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 A schematic diagram for traffic feature data mask processing;
[0033] Figure 2 A schematic diagram for traffic feature data preprocessing with missing values;
[0034] Figure 3 A general flowchart of the application;
[0035] Figure 4 A schematic diagram of the conditional information construction module;
[0036] Figure 5 A multi-scale trend information extraction module;
[0037] Figure 6 A schematic diagram of the conditional feature extraction module. DETAILED DESCRIPTION
[0038] A traffic feature data interpolation method, specifically comprising the following steps:
[0039] S1, based on the existing traffic feature data set, traffic feature data with missing values, mask vectors, adjacency matrices, Gaussian noise samples, and time embedding information are obtained; the existing traffic feature data set used in the embodiment is the PEMS08 data set; specifically:
[0040] The traffic feature data in the PEMS08 data set is normalized to obtain normalized traffic feature data, and the normalized traffic feature data is masked to obtain traffic feature data with missing values, such asFigure 1 As shown; Then, the traffic feature data with missing values are preprocessed to obtain the mask vector, as shown Figure 2 As shown; the specific steps are as follows:
[0041] The traffic characteristic data in the PEMS08 dataset are normalized, including the following steps: calculating the mean μ and standard deviation σ of all traffic characteristic data in the existing traffic characteristic dataset along the spatial dimension; then, normalizing each traffic characteristic data based on the mean μ and standard deviation σ to obtain the normalized traffic characteristic data. , The calculation formula is shown in formula (1):
[0042] (1)
[0043] In formula (1), The existing traffic feature dataset characteristic data of traffic characteristic data;
[0044] Masking the normalized traffic characteristic data includes the following steps: masking the normalized traffic characteristic data using an artificial missing value injection strategy to obtain traffic characteristic data with missing values; wherein the method of masking the normalized traffic characteristic data using the artificial missing value injection strategy is consistent with the method of masking Block Missing settings disclosed in the paper "Filling theg_ap_s: Multivariate time series imputation by graph neural networks";
[0045] Traffic feature data with missing values are preprocessed to obtain a mask vector; specifically, the mask vector is obtained by preprocessing traffic feature data with missing values, which includes the following steps: marking the values at the missing positions in the traffic feature data with missing values with 0, and marking the values at the non-missing positions in the traffic feature data with missing values with 1.
[0046] The method of obtaining the adjacency matrix based on traffic characteristic data with missing values is the same as the calculation method of the weighted adjacency matrix W disclosed in "The emerging field of signal processing on graphs: Extending high-dimensional data analysis to networks and other irregular domains" (i.e., formula (1) disclosed in the paper).
[0047] The way of obtaining Gaussian noise samples based on traffic feature data with missing values includes the following steps: filling the missing positions in the traffic feature data to be interpolated with standard Gaussian noise.
[0048] The way of obtaining time embedding information based on traffic feature data with missing values is consistent with the Diffusion time embedding obtaining method disclosed in PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation.
[0049] S2, constructing a traffic feature data imputation network; the traffic feature data imputation network includes a conditional information construction module, a conditional feature extraction module, a noise estimation unit, a first Add layer and a convolutional layer connected in turn; the output end of the second Add layer is connected to the input end of the noise estimation unit, as shown in Figure 3
[0050] In this application, the conditional information construction module is used to obtain conditional information based on traffic feature data with missing values and a mask vector; wherein, the conditional information construction module, as shown in Figure 4 includes a variational autoencoder and a multi-scale trend information extraction module.
[0051] The working principle of the conditional information construction module is that the variational autoencoder performs initial interpolation on the traffic feature data with missing values to obtain initial interpolation information; the variational autoencoder used in this application is consistent with the Variational autoencoder structure disclosed in Uncertainty-Aware Variational-Recurrent Imputation Network for Clinical Time Series, and the functions are also consistent.
[0052] Then, the initial interpolation information and the mask vector are filled with missing values to obtain traffic data estimation information; in this application, the calculation method of filling the initial interpolation information and the mask vector with missing values is the same as Equation (8) disclosed in Bidirectional spatial-temporal traffic data imputation via graph attention recurrent neural network.
[0053] Then, the multi-scale trend information extraction module performs multi-scale extraction on the trend information in the traffic data estimation information to obtain multi-scale trend information; the traffic data estimation information and the multi-scale trend information constitute the conditional information.
[0054] The structure of the multi-scale trend information extraction module in the present application is shown in FIG. 2, which includes n trend information extraction units connected in parallel, n is greater than or equal to 3, and n is 3 in the embodiment; each trend information extraction unit includes a convolution layer and an average pooling layer connected in sequence, and the convolution kernel sizes of the convolution layers in different trend information extraction units are different; in the embodiment, the multi-scale trend information extraction module includes three trend information extraction units connected in parallel, and the convolution kernel sizes of the convolution layers in the three trend information extraction units are 2, 4 and 8 respectively; in the present application, the convolution layers in the trend information extraction units are used to extract trend information, and the average pooling layers are used to perform average pooling operations on the extracted trend information. Figure 5 In the present application, the conditional feature extraction module takes the conditional information and its corresponding adjacency matrix as input, and is used to capture the spatio-temporal features in the conditional information and the adjacency matrix, enhance the feature capturing ability of the traffic feature data imputation network in the spatio-temporal dimension, and obtain conditional weights containing static and dynamic spatial correlation information and temporal correlation information.
[0055] The structure of the conditional feature extraction module in the present application is shown in FIG. 3, which includes a Mamba module and a graph convolution module connected in parallel with the Mamba module, the output end of the Mamba module is connected with a normalization layer, the output end of the normalization layer is respectively connected with a self-attention layer and a graph attention layer, and the output end of the self-attention layer is also connected with the input end of the graph attention layer; the graph attention layer and the graph convolution module are both connected with an Add layer, and the Add layer is connected with a multi-layer perceptron.
[0056] Figure 6 The Mamba module takes the conditional information as input and extracts the temporal correlation information (the temporal correlation information contains time series dependency and dynamic change law) of the conditional information, thereby enhancing the feature capturing ability of the traffic feature data imputation network in the time dimension, so that the traffic feature data imputation network described in the present application has strong feature capturing ability in the time dimension and obtains the temporal correlation information.
[0057] The normalization layer performs normalization processing on the temporal correlation information to map the temporal correlation information output by the Mamba module to a specific interval, and outputs time features with time series dependency and dynamic change law.
[0058] The self-attention layer processes the time features through a self-attention mechanism and outputs a dynamic graph, and the dynamic graph can reflect the characteristics of the traffic data changing over time.
[0059] The self-attention layer processes the time features through a self-attention mechanism and outputs a dynamic graph, and the dynamic graph can reflect the characteristics of the traffic data changing over time.
[0060] The graph attention layer learns the dynamic change rule of the time feature over time and extracts dynamic spatial correlation information in the dynamic graph, and outputs dynamic spatial features with time-varying information.
[0061] The graph convolution module takes the conditional information and the adjacency matrix as input, and performs convolution operation on the conditional information and the adjacency matrix, extracts static spatial correlation information in the adjacency matrix, and outputs static spatial features with global structured information. In the present application, the setting of the graph attention layer and the graph convolution module can effectively enhance the feature capturing ability of the traffic feature data imputation network in the spatial dimension, so that the traffic feature data imputation network described in the present application has strong feature capturing ability in the spatial dimension.
[0062] The time feature output by the normalization layer, the dynamic spatial feature output by the graph attention layer, and the static spatial feature output by the graph convolution module are added by the Add layer, and then are nonlinearly transformed by the multi-layer perception MLP to mine the spatiotemporal collaborative dependency relationship between the time feature, the dynamic spatial feature, and the static spatial feature, so as to obtain the conditional weight; in the present application, the multi-layer perception MLP is consistent with the MLP structure disclosed in the existing paper “PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation”.
[0063] In the present application, the noise estimation unit includes four noise estimation modules, and the structures of the four noise estimation modules are the same, which are consistent with the structure of the Noise Estimation Module disclosed in the existing paper “PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation”, and the functions are also consistent.
[0064] The noise information is obtained by adding the traffic data estimation information, a Gaussian noise sample and time embedding information using a second Add layer; the Gaussian noise sample is obtained by filling the missing positions in the traffic feature data with missing values using a standard Gaussian noise, and the time embedding information is also obtained based on the traffic feature data with missing values, wherein the method of filling the missing positions in the traffic feature data with missing values using a standard Gaussian noise is consistent with the Noise sample acquisition method disclosed in PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation; the method of obtaining the time embedding information based on the traffic feature data with missing values is consistent with the Diffusion time embedding acquisition method disclosed in PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation;
[0065] In the present application, the four noise estimation modules are a first noise estimation module, a second noise estimation module, a third noise estimation module and a fourth noise estimation module; wherein the first noise estimation module, the second noise estimation module, the third noise estimation module and the fourth noise estimation module are connected in sequence, and the output ends of the first noise estimation module, the second noise estimation module, the third noise estimation module and the fourth noise estimation module are connected with the input end of the first Add layer; in the present application, the output end of the second Add layer is connected with the input end of the first noise estimation module in the noise estimation unit; the input of the first noise estimation module in the noise estimation unit includes an adjacency matrix, noise information and a conditional weight (the conditional weight contains static and dynamic spatial correlation information and time correlation information);
[0066] In the present application, the first noise estimation module performs noise estimation based on the noise information, the conditional weight and the adjacency matrix, and outputs a first noise estimation value; the second noise estimation module performs noise estimation based on the noise estimation value output by the first noise estimation module, and outputs a second noise estimation value; the third noise estimation module performs noise estimation based on the noise estimation value output by the second noise estimation module, and outputs a third noise estimation value; the fourth noise estimation module performs noise estimation based on the noise estimation value output by the third noise estimation module, and outputs a fourth noise estimation value; the first Add layer adds the first noise estimation value to the fourth noise estimation value, and then performs convolution operation using a convolution layer to obtain the predicted estimated noise with rich spatiotemporal correlation information; wherein the spatiotemporal correlation information contains static and dynamic spatial correlation information and time correlation information.
[0067] S3, dividing the traffic feature data with missing values obtained in step S1 into a training set and a test set, and training the traffic feature data imputation network by using the training set to obtain a traffic feature data imputation network model;
[0068] The step S3 specifically comprises the following steps:
[0069] 1) obtaining the training set and the test set; specifically, the traffic feature data with missing values and the corresponding mask vector, the Gaussian noise sample and the time embedding information are divided in time sequence according to a proportion of 7:3 to obtain a sub-training set and a sub-test set, and then the adjacency matrix is added to the sub-training set and the sub-test set respectively to obtain the training set and the test set;
[0070] 2) training the traffic feature data imputation network by using the training set under the guidance of the total optimization loss to obtain a traffic feature data imputation network model; specifically comprising the following steps:
[0071] The traffic feature data with missing values, the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information in the training set are input into the traffic feature data imputation network, the total optimization loss of the traffic feature data imputation network is calculated, and the back propagation is performed under the guidance of the total optimization loss, the weight parameters of the traffic feature data imputation network are updated, one epoch of training process is completed, and the training process of the traffic feature data imputation network is completed after 100 epochs of iterative training process, thereby obtaining the traffic feature data imputation network model. In the present application, the learning rate is set to 0.0001 during the training process of the traffic feature data imputation network to ensure that the model can converge stably during the training process; the batch size is set to 32, which ensures the training efficiency while taking into account the sufficient learning of the traffic feature data imputation network model to the data.
[0072] In the present application, the total loss of the network L includes an initial imputation loss L 1 and a noise estimation loss L 2, the total loss of the network L is calculated as shown in equation (2):
[0073] L =α L 1 +β L 2 (2)
[0074] In equation (2), α and β are balance parameters, α is 0.4, and β is 0.6;
[0075] The initial imputation loss is used to calculate the loss generated by the initial imputation process of the variational autoencoder, and the initial imputation loss L1 including reconstruction loss L a and KL divergence loss L b , initial imputation loss L 1 The calculation formula is shown in equation (3):
[0076] L 1 = L a + L b (3)
[0077] In equation (3), the reconstruction loss L a is calculated in the same way as the reconstruction loss term disclosed in Uncertainty-Aware Variational-Recurrent Imputation Network for Clinical Time Series, and the KL divergence loss L b is calculated in the same way as the Kullback-Leibler divergence term disclosed in the above-mentioned paper.
[0078] Noise estimation loss L 2 is calculated by calculating the mean square error between the estimated noise and the standard Gaussian noise (the standard Gaussian noise refers to the standard Gaussian noise in the Gaussian noise sample that fills in the missing position in the traffic feature data with missing values), so as to optimize the accuracy of the traffic feature data imputation network to obtain the estimated noise. The calculation method of the noise estimation loss L 2 is shown in equation (4):
[0079] (4)
[0080] In equation (4), i represents the ith traffic feature data with missing values, n represents the number of traffic feature data with missing values, represents the estimated noise corresponding to the ith traffic feature data with missing values, is the standard Gaussian noise corresponding to the ith traffic feature data with missing values.
[0081] S4, based on the traffic feature data to be interpolated, obtain a mask vector, an adjacency matrix, a Gaussian noise sample and time embedding information; input the traffic feature data to be interpolated, the mask vector, the Gaussian noise sample and the time embedding information into the trained traffic feature data interpolation network model, forward propagate once to obtain estimated noise; then, using the inverse process of the denoising probability diffusion model, restore the standard Gaussian noise to the traffic feature data interpolation value, which is the traffic feature data interpolation value predicted by the traffic feature data interpolation network. The way of using the inverse process of the denoising probability diffusion model to restore the standard Gaussian noise to the traffic feature data interpolation value is consistent with Algorithm2 disclosed in the paper <PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation>. The way of obtaining the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information based on the traffic feature data to be interpolated is the same as the way of obtaining the mask vector, the adjacency matrix, the Gaussian noise sample and the time embedding information based on the existing traffic feature data set.
[0082] Test:
[0083] In order to compare the superiority of the traffic feature data imputation method described in the present application, the GP-VAE imputation method (from the paper "GP-VAE: Deep Probabilistic Multivariate Time Series Imputation"), the CSDI imputation method (from the paper "CSDI: Conditional Score-based Diffusion Models for Probabilistic Time Series Imputation"), the PriSTI imputation method (from the paper "PriSTI: A Conditional Diffusion Framework for Spatiotemporal Imputation"), and the BayOTIDE imputation method (from the paper "BayOTIDE: Bayesian Online Multivariate Time Series Imputation with Functional Decomposition") and the traffic feature data imputation method described in the present application are trained based on the training set described in the present application and using the steps described in step S3-2 to obtain the data imputation models of the above five methods, and then the traffic feature data in the test set described in the present application is input into the data imputation models of the above five methods for forward propagation once for testing, and the test results are shown in Table 1:
[0084] Table 1 Test results of data imputation models of the above five methods
[0085]
[0086] In Table 1, Ours represents the traffic feature data imputation method described in the present application. In the present application, MAE, RMSE, and GRPS are used as evaluation indexes to compare the imputation effects of the network models in the above five methods.
[0087] The smaller the MAE, RMSE and GRPPS values are, the closer the traffic feature data imputation value is to the true value at the corresponding position of the missing traffic feature data, and the more accurate the imputed traffic feature data obtained by using the traffic feature data imputation value is; the MAE represents the average value of the absolute error between the traffic feature data imputation value and the true value at the corresponding position of the missing traffic feature data, which is called the mean absolute error, and the MAE is used to measure the average deviation degree of the traffic feature data imputation value and the true value at the corresponding position of the missing traffic feature data; the RMSE represents the root mean square error between the traffic feature data imputation value and the true value at the corresponding position of the missing traffic feature data, and the root mean square error is obtained by first calculating the average value of the error square between the true value at the corresponding position of the missing traffic feature data and the corresponding traffic feature data imputation value, and then calculating the square root, and the RMSE is more sensitive to large errors and can highlight the influence of the traffic feature data imputation value with large deviation; the CRPS is a continuous ranking probability score, which is used to measure the overall difference between the probability distribution of the traffic feature data imputation value and the true value at the corresponding position of the missing traffic feature data; the smaller the CRPS value is, the higher the matching degree of the predicted probability distribution of the traffic feature data imputation value and the true value at the corresponding position of the missing traffic feature data is, which means that the model estimates the probability distribution of the traffic data more accurately.
[0088] The traffic feature data imputation method described in the present application has better effects than the above four existing traffic data imputation methods in the three evaluation indexes of MAE, RMSE and GRPPS; since the CSDI imputation method performs best among the above four existing traffic data imputation methods, the test results of the traffic feature data imputation method described in the present application and the CSDI imputation method are compared and analyzed as follows:
[0089] MAE index comparison: the traffic feature data imputation method described in the present application reaches 10.97 in the MAE index, which is reduced by 5.1% compared with the CSDI imputation method; this shows that the average absolute error between the true value at the corresponding position of the missing traffic feature data and the predicted traffic feature data imputation value is smaller, i.e. the predicted traffic feature data imputation value is closer to the true value at the corresponding position of the missing traffic feature data;
[0090] RMSE index comparison: the traffic feature data imputation method described in the present application reaches 18.65 in the RMSE index, which is reduced by 11.9% compared with the CSDI imputation method; this shows that the traffic feature data imputation method described in the present application has stronger inhibition ability to large errors and can effectively reduce the imputation deviation in extreme cases;
[0091] GRPS indicator comparison: the traffic feature data imputation method described in the present application achieves 0.0364 on the GRPS indicator, which is reduced by 5.2% compared with the CSDI imputation method; this shows that the traffic feature data imputation value predicted by the traffic feature data imputation method described in the present application has a higher matching degree with the true value at the corresponding position of the traffic feature data missing value; this means that the traffic feature data imputation method described in the present application can not only predict traffic feature data imputation values closer to the true value at the corresponding position of the traffic feature data missing value, but also more accurately depict the uncertainty of the traffic feature data and accurately describe the probability distribution characteristics of the traffic feature data; in actual traffic scenarios, whether it is the daily fluctuations of traffic flow or data anomalies caused by sudden situations, the traffic data imputation method described in the present application can better capture the change rule of traffic feature data, provide more reference value probability distribution information for traffic planning, flow regulation and other decisions, and show stronger adaptability and reliability.
Claims
1. A traffic characteristic data interpolation method, characterized by: The following steps are involved: The mask vector, adjacency matrix, Gaussian noise sample and time embedding information respectively obtained based on the traffic characteristic data to be interpolated are input into the traffic characteristic data interpolation network to obtain the predicted estimated noise; The traffic feature data interpolation network includes a conditional information construction module, a conditional feature extraction module, a noise estimation unit, a first Add layer, and a convolutional layer, which are connected in sequence. The second Add layer is connected to the noise estimation unit. The conditional information construction module obtains traffic data estimation information based on the traffic feature data to be interpolated and the mask vector, and then performs multi-scale extraction to obtain multi-scale information. The traffic data estimation information and the multi-scale information constitute the conditional information. The conditional feature extraction module is used to capture the conditional information and the spatiotemporal features in the adjacency matrix, and obtain the conditional weights containing static and dynamic spatial correlation information and temporal correlation information; The second Add layer is used to add traffic data estimation information, Gaussian noise samples and time embedding information to obtain noise information; The noise estimation unit performs noise estimation based on the conditional weight, the adjacency matrix and the noise information; the first Add layer is used to add the noise estimation values output by the noise estimation unit; The convolution layer is used to perform a convolution operation on the output of the first Add layer, outputting estimated noise with rich spatiotemporal correlation information; Gaussian noise samples are obtained by filling the missing positions in the traffic feature data with missing values using standard Gaussian noise. Time embedding information is also obtained based on the traffic feature data with missing values. The conditional information construction module includes a variational autoencoder and a multi-scale trend information extraction module; the multi-scale trend information extraction module includes n trend information extraction units connected in parallel, where n is greater than or equal to 3. Each trend information extraction unit includes a convolutional layer and an average pooling layer connected in sequence. The convolution kernel sizes of the convolutional layers in different trend information extraction units are different. The convolutional layers in the trend information extraction units are all used to extract trend information, and the average pooling layers are all used to perform average pooling operations on the extracted trend information. The conditional feature extraction module includes a Mamba module and a graph convolution module connected in parallel with the Mamba module. The output of the Mamba module is connected to the normalization layer, and the output of the normalization layer is connected to the self-attention layer and the graph attention layer respectively. The output of the self-attention layer is also connected to the input of the graph attention layer. The graph attention layer and the graph convolution module are both connected to the Add layer, and the Add layer is connected to the multi-layer perceptron. The noise estimation unit includes four noise estimation modules, each having the same structure; the four noise estimation modules are respectively a first noise estimation module, a second noise estimation module, a third noise estimation module, and a fourth noise estimation module; wherein the first noise estimation module, the second noise estimation module, the third noise estimation module, and the fourth noise estimation module are connected in sequence, and the output ends of the first noise estimation module, the second noise estimation module, the third noise estimation module, and the fourth noise estimation module are all connected to the input end of the first Add layer; the output end of the second Add layer is connected to the input end of the first noise estimation module in the noise estimation unit; the input end of the first noise estimation module in the noise estimation unit includes an adjacency matrix, noise information, and conditional weights; Before obtaining the estimated noise, the traffic characteristic data interpolation network is trained to obtain a traffic characteristic data interpolation network model, which specifically includes the following steps: 1) Obtain training and test sets. The training and test sets are obtained based on traffic feature data with missing values and their corresponding mask vectors, Gaussian noise samples, time embedding information, and adjacency matrices. The traffic feature data with missing values and their corresponding mask vectors, Gaussian noise samples, time embedding information, and adjacency matrices are obtained based on the PEMS08 dataset. The traffic characteristic data interpolation method is applied to traffic planning and flow control.
2. The traffic characteristic data interpolation method according to claim 1, characterized in that: The variational autoencoder performs initial interpolation on traffic feature data with missing values to obtain initial interpolation information; Fill missing values with the initial interpolation information and mask vector to obtain traffic data estimation information; The multi-scale trend information extraction module extracts the trend information in the traffic data estimation information at multiple scales to obtain the multi-scale trend information; Traffic data estimation information and multi-scale trend information constitute conditional information.
3. The traffic characteristic data interpolation method according to claim 1, characterized in that: The traffic characteristic data interpolation network is trained to obtain a traffic characteristic data interpolation network model, which specifically includes the following steps: 2) Under the guidance of the total optimization loss, the traffic feature data interpolation network is trained using the training set to obtain a traffic feature data interpolation network model; specifically, the following steps are included: The traffic feature data with missing values, mask vectors, adjacency matrix, Gaussian noise samples and time embedding information in the training set are input into the traffic feature data interpolation network. The total optimization loss of the traffic feature data interpolation network is calculated. Under the guidance of the total optimization loss, backpropagation is performed to update the weight parameters of the traffic feature data interpolation network. The training process of one epoch is completed. After iterating the training process for 100 epochs, the training process of the traffic feature data interpolation network is completed, and the traffic feature data interpolation network model is obtained.
4. The traffic characteristic data interpolation method according to claim 3, characterized in that: Step 1) specifically includes the following steps: dividing the traffic feature data with missing values and its corresponding mask vectors, Gaussian noise samples, and time embedding information in chronological order at a ratio of 7:3 to obtain sub-training sets and sub-test sets, and then adding the adjacency matrix to the sub-training set and sub-test set respectively to obtain training sets and test sets; wherein, the traffic feature data with missing values, mask vectors, adjacency matrix, Gaussian noise samples, and time embedding information are all obtained based on the traffic feature data of the existing traffic feature dataset.
5. The traffic characteristic data interpolation method according to claim 4, characterized in that: Obtaining traffic feature data with missing values based on traffic feature data of an existing traffic feature data set includes the following steps: performing normalization processing and mask processing on the traffic feature data in the existing traffic feature data set in sequence to obtain traffic feature data with missing values.
Citation Information
Patent Citations
Traffic data interpolation method based on self-supervised learning algorithm
CN115618180A
Missing value interpolation method for single cell sequencing data
CN118335191A