Traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model
Through the traffic data interpolation method of time-frequency feature fusion and conditional diffusion model, the problems of error accumulation and insufficient uncertainty reflection in the existing technology are solved, high-quality traffic data interpolation is achieved, and the accuracy and robustness of the interpolation are improved.
Patent Information
- Application Number
- CN202510678768.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-16
AI Technical Summary
Existing traffic data interpolation methods are prone to introduce cumulative errors when dealing with missing data, cannot effectively reflect the uncertainty of missing data, and lack effective utilization of spatiotemporal dependency condition information.
A method based on time-frequency feature fusion and conditional diffusion model is adopted. By constructing a noise prediction network, the cross-attention mechanism is extracted using time domain and frequency domain features, and the conditional information and adjacency matrix are combined to gradually remove noise to generate high-quality interpolation data.
It effectively avoids error accumulation, improves the accuracy of interpolation results, can provide the probability distribution of missing data, truly reflects the uncertainty of missing data, and improves the robustness and accuracy of interpolation.
Smart Images

Figure CN120653898A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart transportation and deep learning technology, and specifically is a traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model. Background Art
[0002] With the rapid development of Intelligent Traffic Systems (ITS), traffic flow data collected by road sensors and other devices is playing an increasingly important role in the field of smart transportation. However, collecting real-world traffic data is often fraught with difficulties and uncertainties. Factors such as sensor failures and system instability can lead to missing values in the data collected by devices. Missing traffic data is insufficient to provide accurate information, which complicates downstream tasks.
[0003] Early traffic data interpolation methods leveraged statistical learning and traditional machine learning to interpolate data along temporal or spatial dimensions, including but not limited to linear interpolation, expectation-maximization algorithms, and nearest neighbor algorithms. However, these methods all assume strong probability distributions for spatiotemporal data to estimate missing data, which is not realistic. With the development of deep learning, prediction-based interpolation methods are proliferating. The most representative approach uses recurrent neural networks (RNNs) for interpolation, capturing dependencies in time series and predicting missing values through latent state information transfer. Due to the spatial nature of traffic data, graph-based modeling methods can be used to represent the characteristics of graphs, using graph neural networks to learn node features and reconstruct missing values. However, these methods inevitably introduce cumulative errors and rely heavily on historical data, resulting in poor interpolation results. Furthermore, these methods only predict certain missing values and fail to more accurately reflect the uncertainty associated with missing data.
[0004] Recently, the diffusion probability model (DDPM), as a powerful generative model, has demonstrated wide applicability in multiple fields and has been used for time series interpolation tasks. This model corrupts the training data by continuously adding Gaussian noise and then learns to recover the data through inverse denoising, thereby more finely controlling the generation process and generating higher-quality samples than VAE and GAN. When using the diffusion probability model for interpolation tasks, modeling the diffusion probability model and constructing conditional information are inevitable challenges. For traffic data interpolation tasks, the problems faced are: (1) avoiding cumulative errors; (2) constructing and utilizing spatiotemporal dependent conditional information. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the technical problem to be solved by the present invention is to provide a traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model.
[0006] The present invention solves the technical problem by adopting the following technical solutions:
[0007] A traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model includes the following steps:
[0008] Step 1: Obtain the data of each observation node and reconstruct it to obtain traffic data; perform missing data processing on the traffic data to obtain the missing data, and use the missing data as the interpolation target data; construct the adjacency matrix formed by all observation nodes;
[0009] Step 2: Based on the conditional diffusion model, a noise prediction network is constructed;
[0010] The interpolated target data is forward-noised and reverse-denoised using a noise prediction network. The noise prediction network includes a conditional information extraction module and a noise prediction module. In the conditional information extraction module, the missing traffic data is pre-interpolated to obtain pre-interpolated data. A 1×1 convolution operation is used to map the pre-interpolated data into a latent space to obtain a latent representation. The latent representation is then passed through a time domain feature extraction module and a frequency domain feature extraction module to obtain the time domain features and frequency domain features of the pre-interpolated data. The time domain features and frequency domain features of the pre-interpolated data are fused using a cross-attention mechanism to obtain the conditional information.
[0011] The pre-interpolation data is spliced with the interpolation target data after noise addition in the t-th diffusion step to obtain the noise data; the noise data, conditional information, adjacency matrix and diffusion step t are input into the noise prediction module for noise prediction; the noise prediction module is composed of multiple layers, each of which includes a temporal attention module, a spatial attention module and a gated activation unit; in the first layer, the noise data and the diffusion step t are respectively subjected to 1×1 convolution and input into the temporal attention module to extract the first layer of temporal features; the first layer of temporal features and the adjacency matrix are passed through the spatial attention module to obtain the first layer of spatial features; the first layer of spatial features are divided into the first layer of residual connection features and the first layer of skip connection features through the gated activation unit; the first layer of residual connection features and the noise data are added as the noise data input to the second layer, and the first layer of skip connection features are used as the output of the first layer; multiple layers are stacked in this way to obtain the skip connection features of each layer, and the skip connection features of each layer are added to obtain the aggregated features; the aggregated features are subjected to two one-dimensional convolution operations to obtain the predicted noise;
[0012] Step 3: Train the noise prediction network to obtain the trained noise prediction network;
[0013] Step 4: Perform pre-interpolation operation on the traffic data to be interpolated Q to obtain pre-interpolation data Q1; extract condition information C′ based on the pre-interpolation data Q1 con, the pre-interpolated data Q1 is spliced with pure noise that conforms to the Gaussian distribution to obtain the noise data X′ in ; The noise data X′ in , pre-interpolation data Q1, the adjacency matrix of the traffic data to be interpolated and the diffusion step T′ are input into the trained noise prediction network to predict the noise and obtain the interpolation data
[0014] According to the observation mask M′ of the traffic data to be interpolated, the interpolation data Combined with the traffic data to be interpolated, the complete traffic data is obtained as
[0015] Furthermore, the time domain feature extraction module includes a temporal convolutional network and a graph convolutional network. The temporal convolutional network is used to extract the temporal features of the pre-interpolated data. The extraction process is expressed as follows:
[0016]
[0017] Where, represents the time characteristics of the pre-interpolation data, H in represents potential representation, represents a two-dimensional dilated convolution operation with a kernel size of k and a dilation rate of 1, Chomp(·) represents a cropping operation, and Dropout(·) represents a regularization operation;
[0018] The graph convolutional network is used to extract the time domain features of the pre-interpolated data. The extraction process is expressed as:
[0019]
[0020] Where, represents the time domain characteristics of the pre-interpolation data, A gcn Represents the normalized adjacency matrix, o is the order of graph convolution, Conv2D(·) represents a 1×1 two-dimensional convolution operation, Concat(·) represents a concatenation operation, and ReLU represents the ReLU activation function.
[0021] Furthermore, the condition information is expressed as:
[0022]
[0023] Where C con represents conditional information, softmax represents the softmax function, Q STF , K STF 、V STF represents the query, key, and value vectors of the cross-attention mechanism, and d represents the key vector K STF Dimensions, represents the frequency domain characteristics of the pre-interpolated data, represents the weight matrix.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] Using the conditional diffusion model for traffic data interpolation can leverage its non-recursive nature to produce high-quality interpolation results without relying on potentially error-prone prior estimates. This effectively avoids error accumulation and improves the accuracy of the interpolation results. To extract conditional information, a pre-interpolation operation is first used to obtain rough information. Time and frequency domain features are then extracted from the pre-interpolated data in parallel. These features are then fused using a cross-attention mechanism to capture rapid changes and periodic patterns in the data, effectively extracting conditional information and achieving more robust interpolation. The interpolation results follow a probability distribution, providing a possible region or probability distribution for the missing value rather than a single, fixed value, more realistically reflecting the uncertainty brought about by missing data. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is the overall flow chart of the present invention;
[0027] Figure 2 This is a structural diagram of the noise prediction module of the present invention. DETAILED DESCRIPTION
[0028] Specific embodiments are given below in conjunction with the accompanying drawings. The specific embodiments are only used to introduce the technical solutions of the present invention in detail and are not intended to limit the scope of protection of the present application.
[0029] like Figure 1 As shown, the present invention provides a traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model, comprising the following steps:
[0030] Step 1: Obtain the original traffic data and reconstruct the data; obtain the traffic data, process the missing traffic data, and obtain the interpolation target data; construct the adjacency matrix based on the spatial position of the observation node;
[0031] The average speed of vehicles at each observation node is obtained as the original traffic data, and the original traffic data is reconstructed. That is, the original traffic data is integrated into a two-dimensional matrix with observation nodes as columns and time series as rows, thereby converting the unstructured data form of "detection point-speed-time" into a standard two-dimensional matrix form;
[0032] According to the reconstructed original traffic data, traffic data is obtained Among them, X l ∈R NRepresents the observation value of N observation nodes at time l, L represents the number of observation moments; at each moment, not all observation nodes have observation values, so a binary observation mask is used to indicate whether there is an observation value; the observation mask at time l is expressed as in, Indicates that there is an observation value at observation node i, Indicates that there is no observation value for observation node i;
[0033] Perform missing data processing on traffic data X, and use the missing data as interpolation target data for model training and verification use represents the interpolation target mask, and the traffic data after missing processing is recorded as X mask ∈R N×L ; Since there are fewer missing data in real data, the missing data are manually selected for traffic data; specifically, two methods, point missing and block missing, are used: (1) point missing, 25% of the observations are randomly masked; (2) block missing, 5% of the observations are randomly masked, and then the continuous observations of each observation node for 1 to 4 hours are masked with a probability of 0.15%.
[0034] The adjacency matrix is constructed based on the connectivity and distance parameters between N observation nodes in the traffic network. First, a graph G = (V, E, A) is constructed based on the observation nodes of the traffic network, where V is the set of observation nodes and E represents the set of adjacent relationships between observation nodes. If the observation node v i and v j adjacent, then e ij =1, otherwise e ij =0,v i 、v j ∈V,e ij ∈E;A∈R N×N is the adjacency matrix, element d ij is the observation node v i and v j The distance between them, and a ij =a ji .
[0035] Step 2: Based on the conditional diffusion model, a noise prediction network is constructed;
[0036] The conditional diffusion model is divided into two processes: forward denoising and reverse denoising; Forward noise addition: In the process of forward noise addition, the noise variance table {β1,...,β t ,...,β T}, after T diffusion steps of noise addition operation, the interpolated target data will become pure noise data that conforms to the standard normal distribution; the forward noise addition process is a Markov process, which is represented by the Markov chain:
[0037]
[0038] Where q(·) represents the probability distribution generated during the forward noise addition process, β t ∈[0,1] represents the noise variance of the t-th diffusion step, denote the interpolated target data after noise addition in the t-1th, tth, and Tth diffusion steps, respectively. N(·) denotes a normal distribution, and I denotes a variance of 1.
[0039] The probability distribution of the interpolated target data after noise addition in the t-th diffusion step is expressed as therefore, is represented as:
[0040]
[0041] Where, ∈ represents Gaussian noise.
[0042] Interpolation target data generated by forward noise addition Perform reverse denoising; reverse denoising is to train a noise prediction network ∈ θ , through the denoising operation of T diffusion steps, Restore and obtain interpolated data The goal is to make the denoising operation offset the corresponding noise addition operation, so that the interpolated data As close as possible The reverse denoising process is expressed as:
[0043]
[0044] Where p θ (·) represents the probability distribution generated during the reverse denoising process, Respectively represent the interpolated data after denoising in the t-1, t, and T diffusion steps, C con Indicates conditional information. represents the average value, represents the variance, represents the prediction noise generated by the t-th diffusion step, α t =1-β t .
[0045] The noise prediction network includes a condition information extraction module and a noise prediction module; the condition information extraction module is based on the traffic data X after missing processing. maskExtraction condition information C con , the noise prediction module is based on the noise data X in , Condition Information C con , adjacency matrix A and diffusion step t prediction noise.
[0046] 2-1) Extract condition information
[0047] Use the forward interpolation method to process the missing traffic data X mask Perform pre-interpolation, that is, copy the most recently known previous moment data to the missing position to obtain the pre-interpolation data X1∈R N×L ,Pre-interpolation is to construct rough information to facilitate obtaining ,effective conditional information;
[0048] Use 1×1 convolution operation to map the pre-interpolated data X1 into the latent space to obtain the latent representation H in ∈R d ×N×L ; d is the channel size; potential representation H in The time domain features and frequency domain features of the pre-interpolated data are obtained by the time domain feature extraction module and discrete cosine transform respectively;
[0049] The time domain feature extraction module is used to capture temporal and spatial dependencies, including the temporal convolutional network (TCN) and the graph convolutional network (GCN); TCN is used to capture temporal dependencies and obtain the temporal features of pre-interpolated data. Expressed as:
[0050]
[0051] Where, represents a two-dimensional dilated convolution operation with a kernel size of k and a dilation rate of 1, Chomp(·) represents a cropping operation, and Dropout(·) represents a regularization operation;
[0052] GCN is used to capture spatial dependencies. The temporal features of the pre-interpolated data are passed through GCN to obtain the temporal features of the pre-interpolated data. Expressed as:
[0053]
[0054] Where A gcn represents the normalized adjacency matrix of the graph G, o is the order of graph convolution, Conv2D(·) represents a 1×1 two-dimensional convolution operation, Concat(·) represents a concatenation operation, and ReLU represents the ReLU activation function;
[0055] The frequency domain feature extraction module is used to reveal the periodicity and trend of the pre-interpolation data. The discrete cosine transform (DCT) is used to obtain the frequency domain features of the pre-interpolation data. Expressed as:
[0056]
[0057] Where, represents the time step, represents the total number of time steps, cos[·] represents the cosine operation;
[0058] The temporal features of the pre-interpolated data are integrated through the cross-attention mechanism and frequency domain characteristics By integrating the short-term time domain features, we can determine which long-term frequency domain information to focus on, achieve the complementary fusion of short-term dynamics and long-term trends, and mine the periodicity and globality of the data while maintaining the local characteristics of the data, and obtain the conditional information C con , expressed as:
[0059]
[0060] In the formula, softmax represents the softmax function, Q STF , K STF 、V STF represents the query, key, and value vectors of the cross-attention mechanism, and d represents the key vector K STF Dimensions, represents the weight matrix;
[0061] 2-2) Prediction noise
[0062] The noise prediction module uses conditional information C con As a guide, capture the noisy data X in The temporal and spatial dependencies of the data are analyzed, and the attention mechanism is used to help alleviate the randomness of the injected Gaussian noise to the data distribution; the input of the noise prediction module includes the noise data Condition Information C con , adjacency matrix A and diffusion step t; where || represents the splicing operation;
[0063] like Figure 2 As shown in the figure, the noise prediction module consists of multiple layers, each layer has the same structure, including a temporal attention module, a spatial attention module and a gated activation unit; the input of the first layer is the noise data X in , Condition Information C con , adjacency matrix A and diffusion step t, noise data X in After 1×1 convolution and diffusion step t, they are input into the time attention module to extract the first layer of temporal features X tem1 ; First layer temporal feature X tem1 And the adjacency matrix A passes through the spatial attention module to obtain the first layer of spatial features X spa1 ; The first layer of spatial features Xspa1 Through the gated activation unit, it is divided into the first layer of residual connection features X residual1 and the first layer skip connection feature X skip1 ; The first layer residual connection feature X residual1 and noisy data X in After addition, the noise data is used as the input of the second layer, and the first layer jump connection feature X skip1 Aggregated output; stacking multiple layers in this way to obtain the skip connection features of each layer, adding the skip connection features of each layer to obtain the aggregated features; the aggregated features undergo two one-dimensional convolution operations to obtain the predicted noise, which is the output of the noise prediction module.
[0064] The temporal attention module extracts temporal features X tem The process is expressed as:
[0065]
[0066] Where, Attn tem (·) represents the temporal attention mechanism, Q tem , K tem 、V tem Denote the query, key, and value vectors of the temporal attention mechanism, respectively, and d1 denotes the key vector K tem Dimensions, represents the weight matrix;
[0067] The spatial attention module extracts spatial features X spa The process is expressed as:
[0068]
[0069] Where, Attn spa (·) represents the spatial attention mechanism, G(·) represents the graph convolution operation, Norm(·) represents the normalization operation, MLP(·) represents the fully connected operation, Q spa , K spa 、V spa Denote the query, key, and value vectors of the spatial attention mechanism, respectively, and d2 denotes the key vector K spa Dimensions, represents the weight matrix.
[0070] Step 3: Train the noise prediction network;
[0071] Perform random masking on the traffic data X to obtain the missing traffic data X mask ; Traffic data X after missing processing mask Perform pre-interpolation operation to obtain pre-interpolation data X1; Add noise to obtain the interpolated target data after noise addition The interpolated target data after noise addition The pre-interpolated data X1, the adjacency matrix A and the diffusion step t are input into the noise prediction network for training. As the training time increases, the gap between the real noise and the predicted noise becomes smaller and smaller until the error between the real noise and the predicted noise converges, and the trained noise prediction network is obtained.
[0072] Step 4: Perform pre-interpolation operation on the traffic data to be interpolated Q to obtain pre-interpolation data Q1; extract condition information C′ based on the pre-interpolation data Q1 con , the pre-interpolated data Q1 is combined with the pure noise that conforms to the Gaussian distribution Splicing to get the noise data X′ in ; The noise data X′ in , pre-interpolation data Q1, adjacency matrix A′ and diffusion step T′ of traffic data to be interpolated are input into the trained noise prediction network to predict noise, and interpolation data are obtained by step-by-step denoising. According to the observation mask M′ of the traffic data to be interpolated, the interpolation data Combined with the traffic data to be interpolated Q, complete traffic data is obtained
[0073] Example
[0074] This example uses traffic data recorded by highway sensors as an example. The public METR-LA dataset is selected. This dataset consists of traffic data collected by 207 sensors on Los Angeles freeways over a four-month period. A chronological partitioning strategy is used, allocating 70% of the data to the training set, 10% to the validation set, and the remainder to the test set. The training, validation, and test sets contain 23,967, 3,404, and 286 sample points, respectively, each containing 24 time steps of data. This example uses speed as a feature. The dataset information is shown in Table 1.
[0075] Table 1 Overview of the METR-LA dataset
[0076] Number of sensors Time step Total time steps 207 5 minutes 34272
[0077] The missing points in the METR-LA dataset were processed and interpolated using the proposed method and the baseline method. The interpolation results are shown in Table 2.
[0078] Table 2 Comparison of interpolation results of different methods on the METR-LA dataset
[0079]
[0080] Table 3 Comparison of interpolation results of different methods on the METR-LA dataset
[0081]
[0082] Table 2 shows the performance of different methods in terms of MAE and MSE. For methods that generate probability distributions for missing values, such as V-RIN, GP-VAE, CSDI, PriSTI, and our method, we further compared their imputed data (see Table 3). Specifically, given missing values, 100 samples were randomly generated to simulate the probability distribution, and the median of these samples was taken to determine the deterministic imputation result. The results in Tables 2 and 3 show that our method outperforms other baseline methods on the point-missing pattern in the METR-LA dataset.
[0083] Methods based on statistical learning and traditional machine learning usually make strict assumptions about the data and cannot obtain the complex spatiotemporal correlations of real data sets. Among deep learning methods, GRIN is a graph neural network method for multivariate time series interpolation based on bidirectional GRU. It is an autoregressive method and is better than BRITS based on bidirectional RNN because it considers spatial correlation more than BRITS. However, both methods have the problem of autoregressive error accumulation. The method of the present invention uses a conditional diffusion model and a built-in error correction mechanism through a step-by-step denoising process from "full noise" to "clear prediction". Each denoising step minimizes and corrects the error as much as possible, avoiding the problem of continuous accumulation and expansion of errors in traditional autoregressive methods, thereby performing more stable and reliable in prediction.
[0084] For generative methods, methods based on VAE (V-RIN, GP-VAE) and GAN (rGAIN) cannot outperform methods based on diffusion models (CSDI, PriSTI, and the present method). Diffusion models can more accurately capture and reconstruct data distributions and provide high-quality samples. The present method can effectively extract spatiotemporal features as conditional information, guiding the noise prediction network to more accurately predict noise. The construction of conditional information and spatiotemporal correlation can improve the interpolation ability of the diffusion model. Experimental results also verify the overall effectiveness of the present method.
[0085] References:
[0086] [1]TutzG,RamzanS.Improvedmethodsfortheimputationofmissingdatabynearest neighbormethods[J].ComputationalStatistics&DataAnalysis,2015,90:84-99.
[0087] [2]MulyadiAW,JunE,SukH-I.Uncertainty-awarevariational-recurrentimputationnetwork forclinicaltimeseries[J].IEEETransactionsonCybernetics,2021,52(9):9684-94.
[0088] [3]FortuinV,BaranchukD, G,etal.Gp-vae:Deepprobabilistictimeseries imputation[C]. / / Internationalconferenceonartificialintelligenceandstatistics.PMLR,2020:1651-61.
[0089] [4]YoonJ,JordonJ,SchaarM.Gain:Missingdataimputationusinggenerativeadversarial nets[C]. / / Internationalconferenceonmachinelearning.PMLR,2018:5689-98.
[0090] [5]CaoW,WangD,LiJ,etal.Brits:Bidirectionalrecurrentimputationfortimeseries[J].Advancesinneuralinformationprocessingsystems,2018,31.
[0091] [6]TashiroY,SongJ,SongY,etal.Csdi:Conditionalscore-baseddiffusionmodelsfor probabilistictimeseriesimputation[J].AdvancesinNeuralInformationProcessingSystems,2021,34:24804-16.
[0092] [7] LiuM, HuangH, FengH, etal.Pristi: Aconditionaldiffusionframeworkforspatiotemporalimputation[C]. / / 2023IEEE39thInternationalConferenceonDataEngineering(ICDE).IEEE,2023:1927-39.
[0093] Any matters not described in the present invention are applicable to the prior art.
Claims
1. A traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model, characterized in that: The following steps are involved: Step 1: Obtain the data of each observation node and reconstruct it to obtain traffic data; perform missing data processing on the traffic data to obtain the missing data, and use the missing data as the interpolation target data; construct the adjacency matrix formed by all observation nodes; Step 2: Based on the conditional diffusion model, a noise prediction network is constructed; The interpolation target data is forwardly denoised and reversely denoised using a noise prediction network; the noise prediction network includes a conditional information extraction module and a noise prediction module; In the conditional information extraction module, the missing traffic data is pre-interpolated to obtain pre-interpolated data. A 1×1 convolution operation is used to map the pre-interpolated data into a latent space to obtain a latent representation. The latent representation is then passed through the time domain feature extraction module and the frequency domain feature extraction module to obtain the time domain features and frequency domain features of the pre-interpolated data. The time domain features and frequency domain features of the pre-interpolated data are fused through the cross-attention mechanism to obtain conditional information; The pre-interpolated data is concatenated with the interpolated target data after noise addition at the t-th diffusion step to obtain the noise data. The noise data, conditional information, adjacency matrix, and diffusion step t are input into the noise prediction module for noise prediction. The noise prediction module is composed of multiple layers, each of which includes a temporal attention module, a spatial attention module, and a gated activation unit. In the first layer, the noise data and diffusion step t are each subjected to a 1×1 convolution and then input into the temporal attention module to extract the first-layer temporal features. The first layer of temporal features and adjacency matrix are passed through the spatial attention module to obtain the first layer of spatial features; The first-layer spatial features are divided into the first-layer residual connection features and the first-layer skip connection features through the gated activation unit; The residual connection features and noise data of the first layer are added together as the noise data of the second layer input, and the skip connection features of the first layer are used as the output of the first layer. Multiple layers are stacked in this way to obtain the skip connection features of each layer, and the skip connection features of each layer are added together to obtain the aggregated features. The aggregated features undergo two one-dimensional convolution operations to obtain the prediction noise; Step 3: Train the noise prediction network to obtain the trained noise prediction network; Step 4: Perform pre-interpolation operation on the traffic data to be interpolated Q to obtain pre-interpolation data Q1; extract condition information C′ based on the pre-interpolation data Q1 con , the pre-interpolated data Q1 is spliced with pure noise that conforms to the Gaussian distribution to obtain the noise data X′ in ; The noise data X′ in , pre-interpolation data Q1, the adjacency matrix of the traffic data to be interpolated and the diffusion step T′ are input into the trained noise prediction network to predict the noise and obtain the interpolation data According to the observation mask M′ of the traffic data to be interpolated, the interpolation data Combined with the traffic data to be interpolated to obtain complete traffic data 2. The traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model according to claim 1 is characterized in that: The time domain feature extraction module includes a temporal convolutional network and a graph convolutional network. The temporal convolutional network is used to extract the temporal features of the pre-interpolated data. The extraction process is expressed as follows: Where, represents the time characteristics of the pre-interpolation data, H in represents potential representation, represents a two-dimensional dilated convolution operation with a kernel size of k and a dilation rate of 1, Chomp(·) represents a cropping operation, and Dropout(·) represents a regularization operation; The graph convolutional network is used to extract the time domain features of the pre-interpolated data. The extraction process is expressed as: Where, represents the time domain characteristics of the pre-interpolation data, A gcn Represents the normalized adjacency matrix, o is the order of graph convolution, Conv2D(·) represents a 1×1 two-dimensional convolution operation, Concat(·) represents a concatenation operation, and ReLU represents the ReLU activation function.
3. The traffic data interpolation method based on time-frequency feature fusion and conditional diffusion model according to claim 1 or 2 is characterized in that: Condition information is represented as: Where C con represents conditional information, softmax represents the softmax function, Q STF , K STF 、V STF represents the query, key, and value vectors of the cross-attention mechanism, and d represents the key vector K STF Dimensions, represents the frequency domain characteristics of the pre-interpolated data, represents the weight matrix.
Citation Information
Cited By
Conditional diffusion model-based digital human posture action generation method
CN121304873A
Vibration signal space-time reconstruction method based on multi-modal condition diffusion model
CN121502240A
A vibration signal space-time reconstruction method based on a multi-modal conditional diffusion model
CN121502240B
Traffic flow data completion method and system based on frequency domain diffusion model
CN121682273A