A time series data missing value filling method based on diffusion model

By using a diffusion-based imputation method combined with Diffwave and Transformer/GRU modules, efficient and accurate imputation of time series data is achieved, solving the problems of feature correlation and error accumulation in existing technologies and improving data quality.

CN119202555BActive Publication Date: 2025-12-09CIVIL AVIATION UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411260499.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-12-09
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

Existing time series missing value imputation algorithms do not adequately consider feature correlation and temporal variability, and suffer from error accumulation and a lack of research on the patterns of missing data.

Method used

A diffusion-based filling method is adopted, which combines the Diffwave denoising diffusion module with the evaluation filling module of Transformer and GRU. Through self-supervised training, multi-level attention calculation and incremental filling algorithm, the attention capture ability and filling accuracy of time series are improved.

Benefits of technology

It effectively improves the accuracy and quality of filling missing values ​​in time series data, solves the problem of missing data in time series, and enhances the value and reliability of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202555B_ABST
    Figure CN119202555B_ABST
Patent Text Reader

Abstract

The application provides a time series data missing value filling method based on a diffusion model, which comprises the following steps: constructing a filling model, training the filling model, and filling data missing values by using the trained filling model; the filling model comprises a denoising diffusion module and an evaluation filling module; the denoising diffusion module comprises an input data processing module, a residual connection block and an output integration module; the evaluation filling module comprises an incremental filling module and a sampling result integration module; the application is based on the most advanced diffusion model architecture DiffWave, adopts a self-supervised training method, generates filling targets by masking non-missing data, and uses observable part information to assist the diffusion model in predicting noise, so that efficient and accurate training of the diffusion model is realized; the application can effectively improve the time series missing value filling effect, solve the problem of data missing in the time series, and improve the data quality and value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computers, and particularly relates to a time series data missing value filling method based on a diffusion model. BACKGROUND

[0002] In order to solve the problem of time series missing values, a large number of scholars have carried out in-depth research, and various filling algorithms have been proposed based on different technologies. The main filling algorithms are based on statistical learning, deep learning and generation model.

[0003] In the filling algorithm based on statistical learning, the early algorithm uses statistical values such as mean, median and mode to fill in the missing values, which lacks consideration of the relevance of features and the variability of time, and the filling effect is not good. The filling algorithm based on deep learning infers new filling values from historical filling values containing bias, which is inevitably affected by error accumulation. In the generation model, the standard attention mechanism is mostly used for non-missing data, which lacks the capture of the space-time features of time series, and the filling process also has the problem of directly using the picture generation and speech synthesis workflow, which lacks the study of data missing rules and the use of observable information. SUMMARY

[0004] Therefore, the application aims to overcome the above-mentioned problems in the prior art, and provides a time series data missing value filling method based on a diffusion model.

[0005] To achieve the above-mentioned purposes, the technical scheme of the application is as follows:

[0006] A time series data missing value filling method based on a diffusion model, comprising constructing a filling model, training the filling model, and filling data missing values by using the trained filling model.

[0007] The filling model comprises a denoising diffusion module composed of a classical speech generation model Diffwave and an evaluation filling module based on Transformer and GRU. The denoising diffusion module comprises an input data processing module, a residual connection block and an output integration module. The evaluation filling module comprises an incremental filling module and a sampling result integration module.

[0008] The input data processing module is used for artificially masking the input data containing missing values to generate a mask matrix and a condition information matrix, and randomly sampling a diffusion step t, and projecting and embedding the data, the diffusion step and the condition information through a convolution layer and an activation function.

[0009] The residual connection block is used for multi-level, multi-dimensional mapping and attention calculation and fusion processing of input data, including a time attention module TAM, a diffusion step embedding projection convolution layer, a conditional information embedding projection convolution layer and an intermediate expansion decomposition layer, the diffusion step embedding projection convolution layer is used for mapping input diffusion steps to the input dimension of the residual connection block through a convolution network, the conditional information embedding projection convolution layer is used for mapping input conditional information to the input dimension of the residual connection block through a convolution network, and the time attention module TAM is used for decomposing time attention into parallel in-channel static attention and inter-channel dynamic attention, and series of time attention along the time dimension and feature attention along the feature dimension; the intermediate expansion decomposition layer is used for decomposing data into two parts of equal size, one part as the input of the next layer residual block, and the other part as the output of the current residual block;

[0010] The output integration module is used for adding and integrating the inputs of multiple residual connection blocks, and mapping data mapped to a high dimension to the original dimension of the data again through a convolution layer and an activation function;

[0011] The incremental filling module reasonably evaluates observable data with high attention scores or adjacent to the data to be filled, and inputs the filling values meeting the constraints as observable information into the next round of the generation model, so as to realize incremental filling.

[0012] The sampling result integration module automatically weights and integrates the sampling results by using an unsupervised clustering algorithm.

[0013] Further, the time attention module TAM includes a time attention layer and a feature attention layer, the time attention layer takes the tensor of each feature as input, extracts the time dependence of data, and the feature attention layer takes the tensor of each time point as input, extracts the feature dependence of data.

[0014] Further, the time attention layer and the feature attention layer are both one-layer Transformer encoders.

[0015] Further, the rationality evaluation module TSRE adopts sequence evaluation from two angles of time and feature, the time rationality evaluation divides the data according to time points, all feature values of each time point are combined into a sequence, the attention of the observable value to the filling value is calculated by using the Transform, and then whether each time point meets the requirement of the original data distribution is output through a linear layer, the feature rationality evaluation divides the data according to the feature dimension, and the data in a period of time corresponding to each feature is combined into a sequence, the data in the sequence is input into the GRU network according to time sequence, the information of the whole sequence is condensed by means of the GRU, and then whether the data under the feature meets the requirement of the original data distribution is output after a linear layer, the results of the time rationality and the feature rationality are multiplied, and the evaluation result of the whole sequence is obtained.

[0016] Further, the sampling result integration module realizes the process as follows:

[0017] The results of multiple sampling of the model are input into a K-Means model, the sampling results are aggregated into two categories according to the distance between samples, denoted as set1 and set2, the number of samples in each category is counted, denoted as n1 and n2, and the final filling is shown in the following formula:

[0018]

[0019] P1 and p2 respectively represent the proportion of the number of samples in the opposite category to the total number of samples, and the logarithmic value of the proportion is used to determine the weight of each category, multiplied by the filling mean value of the respective category, as the final filling result, so as to realize the weighted integration of the sampling results.

[0020] Further, the training of the filling model includes the training of the denoising diffusion module and the evaluation and filling module based on the Transform and the GRU, wherein the training of the denoising diffusion module adopts a self-supervised training mode, a part of the values of the multivariate time series X containing missing values are deleted by using a mask strategy to obtain a new X and a corresponding mask matrix M, and the true value of the mask part is used to guide the training of the denoising diffusion module.

[0021] Compared with the prior art, the time series data missing value filling method based on the diffusion model has the following advantages:

[0022] 1. The present application is based on the most advanced diffusion model architecture DiffWave, adopts a self-supervised training mode, generates a filling target by masking the non-missing part of the data, and uses the observable part information to assist the diffusion model to predict noise, so as to realize efficient and accurate training of the diffusion model.

[0023] 2. In the aspect of feature extraction, the time sequence attention module for fusing calculation of multi-level and multi-angle attentions among channels, within channels, time dimensions and feature dimensions is designed, so as to improve the attention capturing ability for time series containing missing values;

[0024] 3. The incremental filling algorithm is designed, partial stage filling results are reserved as observable information of subsequent filling process, and the value of stage achievements generated by filling work is fully explored;

[0025] 4. The application can effectively improve the filling effect of time series missing values, solve the problem of data missing in time series, and improve the data quality and value. DETAILED DESCRIPTION

[0026] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0027] Figure 1 the filling model structure schematic diagram of the present application;

[0028] Figure 2 the time sequence attention module TAM structure schematic diagram of the present application;

[0029] Figure 3 the denoising diffusion module training flowchart of the present application;

[0030] Figure 4 the evaluation filling module flowchart of the present application. DETAILED DESCRIPTION

[0031] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.

[0032] In the description of the present application, it needs to be understood that the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation of the present application. In addition, the terms "first", "second" and the like are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first", "second" and the like can be explicitly or implicitly included one or more. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0033] In the description of the present application, it needs to be understood that the terms "installation", "connection", "connection" should be understood in a broad sense, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through intermediate medium, or the communication between two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood through specific circumstances.

[0034] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0035] The present application provides a time series data missing value filling method based on a diffusion model, which comprises constructing a filling model, training the filling model, and filling data missing values by using the trained filling model; the filling model is as shown in Figure 1 The filling model mainly consists of two parts, one is a denoising diffusion module based on a classic speech generation model Diffwave, and the other is an evaluation filling module based on Transformer and GRU. In the denoising diffusion module, it is mainly divided into an input data processing module, a residual connection block and an output integration module.

[0036] Specifically, the input data processing module is responsible for artificially masking the input data containing missing values to generate a mask matrix and a condition information matrix, and randomly sampling diffusion steps t. Then the data, diffusion steps and condition information are projected and embedded through convolution layer and activation function.

[0037] Specifically, the residual connection block is responsible for multi-level, multi-dimensional mapping and attention calculation and fusion processing of the input data. It includes a time attention module TAM, a diffusion step embedding projection convolution layer, a conditional information embedding projection convolution layer, and an intermediate expansion decomposition layer. The diffusion step embedding projection convolution is responsible for mapping the input diffusion step to the input dimension of the residual connection block through the convolution network. The conditional information embedding projection convolution layer is responsible for mapping the input conditional information to the input dimension of the residual connection block through the convolution network. The time attention module TAM decomposes the time attention into the parallel of intra-channel static attention and inter-channel dynamic attention, and the series of time attention along the time dimension and feature attention along the feature dimension. The intermediate expansion decomposition layer is used to decompose the data into two parts of the same size, one part as the input of the next layer residual block, and the other part as the output of the current residual block. The two parts of data are respectively mapped by the activation function, and are differentially nonlinearly transformed to meet different task requirements. In the present application, the diffusion step embedding projection convolution layer is responsible for mapping the input diffusion step to the input dimension of the residual connection block through the convolution network, then adding the input data, and finally inputting the result to the time attention module TAM. The conditional information embedding projection convolution layer is responsible for mapping the input conditional information to the input dimension of the residual connection block through the convolution network, then adding the calculation result of the time attention module TAM, and inputting to the intermediate expansion decomposition layer to obtain the input of the next layer residual connection block and the output of the current residual connection block.

[0038] Specifically, the model structure of the time attention module TAM is as shown in Figure 2 For intra-channel attention, the model uses depth separable convolution, dilated depth separable convolution and 1x1 convolution to simulate large kernel convolution, which decomposes the large kernel convolution into three parts, while saving the computational overhead of the model. For inter-channel attention, the model uses an average pooling layer to condense the information of each channel, then uses a linear layer and an activation function to capture the dependency between channels, and finally uses a product form to connect the intra-channel attention and the inter-channel attention. In addition, in view of the two characteristics of time order and multi-dimension in multi-element time series, the model designs a time attention layer and a feature attention layer, both of which are 1-layer Transformer encoders. The time attention layer takes the tensor of each feature as input to extract the time dependency of the data, and the feature attention layer takes the tensor of each time point as input to extract the feature dependency of the data. Through multi-level and multi-angle attention calculation from inter-channel, intra-channel, time dimension, feature dimension, etc., and with the non-sensitivity of missing values by the modules such as dilated convolution and average pooling, the influence of missing values on attention calculation is weakened, and the attention capturing ability of the model for time series containing missing values is improved.

[0039] Specifically, the output integration module is responsible for adding and integrating the inputs of multiple residual connection blocks, and then using convolutional layers and activation functions to remap the data mapped to the high-dimensional dimensions back to the original dimensions of the data.

[0040] The evaluation filling module mainly includes an incremental filling module and a sampling result integration module.

[0041] Specifically, the incremental imputation module addresses the lack of utilization of the phased results generated during the imputation process in existing time series imputation algorithms by designing an incremental imputation algorithm. Unlike time series prediction algorithms, where the true value to be predicted has not yet been generated at the time of prediction, making the predicted value unevaluable during runtime, the missing values ​​in time series imputation algorithms are data that has been generated but not recorded; in other words, the true value exists but is unobservable. Therefore, the reasonableness of imputation can be evaluated using observable data from adjacent data points or those with high attention scores. Based on the evaluation results, some imputation values ​​that meet the constraints are used as observable information and input into the next round of the generative model, achieving incremental imputation. Using entropy theory from thermodynamics, the entropy of a time series containing missing values ​​is highest initially. As the imputation process progresses, some imputation positions tend to stabilize or are replaced by imputation values ​​that meet the constraints, and the entropy of the time series gradually decreases. As the entropy gradually decreases, the imputation process becomes simpler, more stable, and more accurate. Therefore, incremental filling can make the filling process more stable and the filling results more accurate. This invention relates to a reasonableness evaluation module (TSRE), which evaluates sequences from both temporal and feature perspectives. Figure 1 As shown in the lower right part, the temporal rationality evaluation divides the data according to time points. All feature values ​​at each time point are combined into a sequence. A Transformer is then used to calculate the attention of observables to the imputed values. A linear layer then outputs whether each time point meets the requirements of the original data distribution. The feature rationality evaluation divides the data according to feature dimensions and combines the data within a time period corresponding to each feature into a sequence. The data within the sequence are then input into a GRU network in chronological order. The GRU condenses the information of the entire sequence, and after passing through a linear layer, it outputs whether the data under that feature meets the requirements of the original data distribution. Finally, the temporal rationality and feature rationality results are multiplied to obtain the evaluation result of the entire sequence. Imputation results evaluated as reasonable and stable serve as known information for the next round of imputation, enabling effective and stable training and adaptive transmission of necessary information, thereby improving the accuracy of the diffusion model-based time series missing value imputation algorithm.

[0042] Specifically, in order to further weaken the influence of the sampling results with large partial deviation generated by the probability filling model on the final filling results, the sampling result integration module designs a sampling result automatic weighting integration algorithm based on an unsupervised clustering algorithm. Specifically, the results of multiple sampling of the model are input into a K-Means model, and the sampling results are aggregated into two categories, set1 and set2, according to the distance between samples. The number of samples in each category is counted and denoted as n1 and n2. The final filling is shown in equations (1), (2) and (3):

[0043]

[0044]

[0045] p1 and p2 represent the proportion of the number of samples in the opposite category to the total number of samples, and the logarithmic value of the proportion is used to determine the weight of each category. Then, the filling mean value of each category is multiplied by the weight to obtain the final filling result, thereby realizing the weighted integration of the sampling results. Equation (1) mainly uses the function characteristics of the logarithmic function with a base greater than 0 when the independent variable is less than 1, i.e., amplifying the main category and suppressing the minor category, thereby weakening the sampling results with large partial deviation and improving the accuracy of the final filling results.

[0046] The model training and filling are also divided into two stages, including diffusion model training and incremental filling. For a multivariate time series X ∈ R K×L , where K is the number of features of the time series and L is the length of the time series. For the mask matrix M = {m 1:K,1:L} ∈ {0, 1} K×L , where

[0047]

[0048] Therefore, each time series can be represented as {X, M}. The time series missing value filling task is to estimate the distribution of the missing values of X using the observed values of X.

[0049] The diffusion model consists of two Markov chains. The Markov chain of the forward diffusion process converts data into noise, and the Markov chain of the reverse diffusion process converts noise into data. Both the forward diffusion and the reverse diffusion process consist of T diffusion steps. x0 represents the original data, x T represents random noise, which is usually a simple normal distribution. The transition kernel q(x t |x t-1) is usually pre-designed, aiming to add noise to the data step by step in T diffusion steps, and convert the original data distribution q(x0) into a simple prior distribution N(0, I).p θ (x t-1 |x t ) is the transition kernel of the reverse Markov chain, aiming to remove the noise added in the forward diffusion process step by step from the random noise x T step by step in T diffusion steps, so as to generate a sample x0 conforming to the original data distribution q(x0).

[0050] The flow of diffusion model training in the application is shown in Figure 3 For the missing value filling task, the model adopts a self-supervised training manner, deletes a part of the values of the non-missing parts in X through a mask strategy to obtain a new X and a corresponding mask matrix M, and then uses the true value of the masked part to guide the training of the diffusion model. The true value of the masked part is denoted as X m , which is also the filling target in the training stage, and the observable part of the data after masking is denoted as X s , which is also the conditional information in the training stage, so X = X m + X s . Therefore, in the application, the training target of the diffusion model is to approximate the real data distribution q(x θ |x m ) with the model distribution p s (x m |x s ), and the transition kernel of the reverse Markov chain of the diffusion model is expanded as shown in formula (4) and (5):

[0051]

[0052] At the same time, the final optimization target of the diffusion model is also expanded as shown in formula (6):

[0053]

[0054] In the formula, t represents the diffusion step, x0 is the data to be filled, the filling target and the observable information are obtained by masking x0. In the initial state, that is, when t = T, the data of the part to be filled is random noise of normal distribution, denotes the value of the part to be filled at the t-th diffusion step in the reverse diffusion denoising process. The filling task is to recover the value of the part to be filled based on the observable information through T diffusion steps, so that it is sufficiently close to the original value of the filling target

[0055] In formula (6), ∈ θThe diffusion model in the application is as follows represents the added noise in the tth diffusion step in the forward diffusion process, represents the noise value of the to-be-filled part of the tth diffusion step predicted by the model based on the observation information. Thus, the training target of the model is changed from predicting the data of each step to predicting the noise value added to the filling part of each step, and then removing the noise value step by step to recover the original data.

[0056] In each iteration of the model training stage, the to-be-filled data x0, the diffusion step t and the Gaussian noise ∈ are first sampled, and then x0 is masked to obtain the filling target and the observable information Finally, the model is calculated and gradient descent is performed based on this until the model converges.

[0057] The model filling process is as shown in Figure 4 The application uses the trained noise prediction model ∈ θ to fill, which is different from the filling algorithm based on the diffusion model in the prior art. The output of the previous round is directly used as the input of the next round in the prior art, and the observable information remains unchanged. The application adds a time series evaluation network to evaluate the denoising result , retain the data of the filling points that meet the distribution constraint, form new and as the input of the next round of diffusion model, so as to realize incremental filling.

[0058] In order to evaluate the performance of the model proposed in the application, the application is comprehensively compared with the classic model and the latest time series filling model. The benchmark model includes mean filling Median, filling model BRITS based on RNN, filling model GAIN based on GAN, filling model SAITS based on attention mechanism, missing value filling model CSDI based on diffusion model and SSSD.

[0059] The experiment is performed on three public data sets, air quality data set AQI, power transformer data set ETTh1 and weather data set Weather, for 10 times. The data missing mode adopts a completely random missing mode, and the data missing rate is set to 10%, 20%, 50% and 90% in turn. When evaluating the performance of the model under different missing rates, 10% to 90% of the data points are randomly masked to simulate different degrees of missing.

[0060] Table 1, Table 2 and Table 3 respectively show the MAE and RMSE index values of the I2TDM model and the baseline model on the AQI, ETTh1 and Weather data sets. On all experimental data sets, the simple statistical-based Meadian method performs poorly, and the deep learning-based method is better than the simple filling method. As can be seen from the table, the filling model GAIN designed for time series data performs poorly when the missing rate is less than 50%. The RNN-based filling model BRITS performs well when the missing rate is low, because RNN relies heavily on the output value of the previous time, and when the missing rate is high, the deviation of the filled value of the previous time is relatively large. From the historical values containing deviation, the error accumulation effect will continue to amplify, resulting in poor filling effect. Benefiting from the excellent performance of the self-attention mechanism, SAITS performs better than BRITS in filling effect on all data sets at all missing rates. However, compared with the CSDI based on the diffusion model, the data generated by the diffusion model is of higher quality and more diverse, and the CSDI again surpasses the SAITS. However, SSSD based on the diffusion model fails to surpass SAITS in filling effect due to not using a multi-channel input format. I2TDM uses a multi-channel input format, and designs a time series attention module for time series containing missing values, and an incremental filling scheme for missing values. Only in the case of 90% missing rate of the AQI data set and the ETTh1 data set, I2TDM is partially inferior to CSDI. The evaluation index of the filling result of I2TDM on almost all missing rates of the above data sets exceeds that of the above benchmark model. Compared with the current optimal result, I2TDM improves the MAE index by 3%-8% and the RMSE index by 4%-10%.

[0061] Table 1 is the missing value filling experiment result on the AQI data set

[0062]

[0063] Table 2 is the missing value filling experiment result on the ETT-h1 data set

[0064]

[0065] Table 3 is the missing value filling experiment result on the Weather data set

[0066]

[0067]

[0068] A large number of experiments prove that the present application can effectively improve the time series missing value filling effect, solve the problem of data missing in time series, and improve the data quality and value.

[0069] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for filling missing values in time series data based on a diffusion model, characterized in that: The method comprises constructing a filling model, training the filling model, and filling data missing values by using the trained filling model; The filling model comprises a denoising diffusion module composed of a classical speech generation model Diffwave and an evaluation filling module based on a Transformer and a GRU; the denoising diffusion module comprises an input data processing module, a residual connection block and an output integration module; the evaluation filling module comprises an incremental filling module and a sampling result integration module; The input data processing module is configured to artificially mask input data containing missing values to generate a mask matrix and a condition information matrix, and randomly sample a diffusion step t, and project and embed data, a diffusion step and condition information through a convolution layer and an activation function; The residual connection block is configured to perform multi-level and multi-dimensional mapping, attention calculation and fusion processing on the input data, and comprises a time sequence attention module TAM, a diffusion step embedding projection convolution layer, a condition information embedding projection convolution layer and an intermediate expansion decomposition layer; the diffusion step embedding projection convolution layer is configured to map the input diffusion step to the input dimension of the residual connection block through a convolution network; the condition information embedding projection convolution layer is configured to map the input condition information to the input dimension of the residual connection block through a convolution network; the time sequence attention module TAM is configured to decompose time sequence attention into parallel channel-in static attention and channel-inter dynamic attention, and series time sequence attention along a time dimension and feature attention along a feature dimension; the intermediate expansion decomposition layer decomposes data into two parts of equal size, one part as the input of the next layer residual block, and the other part as the output of the current residual block; The output integration module is configured to add and integrate the inputs of a plurality of residual connection blocks, and map data mapped to a high dimension to the original dimension of the data again through a convolution layer and an activation function; The incremental filling module performs rationality evaluation on observable data points adjacent to or having high attention scores of the data to be filled, and inputs the filling values meeting the constraints as observable information into the next generation model to realize incremental filling; the incremental filling module performs sequence evaluation from the time and feature angles; the time rationality evaluation divides data according to time points, combines all feature values at each time point into a sequence, calculates the attention of observable values on filling values by using a Transformer, and outputs whether each time point meets the requirements of the original data distribution through a linear layer; the feature rationality evaluation divides data according to feature dimensions, combines data in a period corresponding to each feature into a sequence, inputs data in the sequence according to time sequence into a GRU network, condenses information of the whole sequence by means of the GRU, and outputs whether the data under the feature meets the requirements of the original data distribution after a linear layer; and the results of the time rationality and the feature rationality are multiplied to obtain the evaluation result of the whole sequence; The sampling result integration module automatically weights and integrates the sampling results by using an unsupervised clustering algorithm. 2.The method of claim 1, wherein: The time attention module TAM comprises a time attention layer and a feature attention layer, the time attention layer takes a tensor of each feature as input to extract time dependency of data, and the feature attention layer takes a tensor of each time point as input to extract feature dependency of data.

3. The method of claim 2, wherein the method is characterized by: The time attention layer and the feature attention layer are both a layer of a Transformer encoder.

4. The method of claim 1, wherein the method is based on a diffusion model. The implementation process of the sampling result integration module is as follows: The results of multiple sampling of the model are input into a K-Means model, and the sampling results are aggregated into two categories according to the distance between samples, denoted as , , the number of samples in each category is counted, denoted as , , and the final filling is as shown in the following formula: ; ; ; and respectively represent the proportion of the number of samples of opposite categories in the total number of samples, and the logarithmic value of the proportion will be used to determine the weight of each category, multiplied by the filling mean value of the respective category, as the final filling result, so as to realize the weighted integration of the sampling result.

5. The method of claim 1, wherein the method is based on a diffusion model. The training of the filling model comprises training of the denoising diffusion module, and the trained denoising diffusion module is used for filling, wherein the training of the denoising diffusion module adopts a self-supervised training mode, a part of non-missing positions in the multivariate time series X containing missing values is deleted by a mask strategy to obtain a new X and a corresponding mask matrix M, and the true value of the mask position is used to guide the training of the denoising diffusion module.

Citation Information

Patent Citations

  • Conditional diffusion model-based time series data prediction method and system

    CN117076931A

  • Time series data generation method and system based on diffusion model

    CN118410289A