A method and system for filling missing values in power data based on time-aware diffusion

By building a TIED-GAN model, combining time-aware diffusion and generative adversarial network, the problem of failure to effectively consider timing and long-term dependence of the power data missing value completion model in the existing technology is solved, and the accurate completion of the power data missing value and data quality are achieved.

CN119884626BActive Publication Date: 2025-07-01SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510368683.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing diffusion-based power data missing value completion model fails to effectively consider the timing and long-term dependence of power data, especially in the case of continuous missing during power outages, making it difficult to achieve accurate missing value completion.

Method used

A method for completing the missing value of power data based on time-aware diffusion is proposed. By constructing a TIED-GAN model, using a noise predictor and a discriminator, combining the time-aware transformer layer and the feature transformer layer, capturing the missing time interval-related information, and improving the completion performance by generating an adversarial network structure.

Benefits of technology

Accurate completion of missing values ​​of power data is achieved, data quality is improved, downstream tasks are ensured, especially in terms of timing and long-term dependence of power data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884626B_ABST
    Figure CN119884626B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for completing missing values in power data based on time-aware diffusion, belonging to the technical field of power data governance. The method includes: acquiring power data and performing preprocessing, randomly masking a part of the power data to represent missing values; sampling noise data, adding the noise data to the masked part of the power data, and inputting it together with the unmasked original power data into a pre-constructed generative adversarial model based on time-aware diffusion for training to obtain predicted noise data; removing the predicted noise data from the masked part of the power data with the added noise data to obtain the completion result of the missing values in the power data. Based on the powerful generation effect of the score-based diffusion model, the present invention ensures the generation quality of power data; the proposed generative adversarial model based on time-aware diffusion takes the time interval as an additional input for the continuous missing characteristics of power data, improving the accuracy of completion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power data governance, and particularly relates to a method and system for filling missing values of power data based on time-aware diffusion. Background Art

[0002] The statements herein only provide background art related to the present invention and do not necessarily constitute prior art.

[0003] With the continuous development of technologies such as smart grids, renewable energy, and electrification, the power system is undergoing digital transformation. A large number of sensors and monitoring devices are deployed in the power system to collect various parameter data in real time, providing comprehensive monitoring and control capabilities for system operation. The data assets of power enterprises exhibit typical big data characteristics. These power data come from various links of power production and power consumption, including power generation, transmission, transformation, distribution, and dispatching, and include various types of data such as power grid operation, equipment management, marketing services, and enterprise management, containing rich information reflecting the production and operation of power enterprises and customer service status.

[0004] However, in the new power system, there are many types of equipment, wide distribution, and large differences, resulting in increased uncertainty in power grid measurement data collection and a relatively high random data missing rate. Power outages will cause the power supply in the entire area to be interrupted. Therefore, relevant power data cannot be recorded or collected during power outages, which will lead to data loss during power outages. The time series data at the last moment before the power outage is very likely to be closely related to the data at the first moment after restoration. For power data, most of its applications and algorithms require data without missing values. The quality of power data will also affect the efficiency of downstream tasks (such as data anomaly detection and data classification). Poor power data will bring biases to subsequent tasks and question the effectiveness of subsequent tasks. Therefore, it is urgent to mine potential missing laws based on massive power data, and then construct a power data missing value filling model to effectively guarantee the quality of power data and downstream tasks.

[0005] In recent years, with the advancement of the field of power data governance technology, diffusion models based on scores have stood out due to their excellent generation performance and completion effect. Diffusion models based on scores are a class of deep generative models that generate samples by gradually converting noise into credible data samples through denoising. Moreover, such models have achieved state-of-the-art sample quality in many tasks such as image generation and audio synthesis, outperforming other generative models including autoregressive models. Diffusion models based on scores can also be applied to the completion of missing data values. Thanks to the powerful generation effect of diffusion models, these models have the advantage of generating results with relatively small average errors in estimation work. However, in previous diffusion-based completion models, the sampled noise information was not well utilized, and the results could only be used to train noise predictors. In addition, the inventors found that due to the particularity of power data missing, such as the power outage phenomenon mentioned above, power data may have consecutive missing phenomena over a period of time, which requires the corresponding missing model to consider the missing time interval. However, previous diffusion-based completion models did not consider the time interval when applied to power data, which does not match the missing situation of power data, so it is difficult to achieve accurate completion of missing power data values. Summary of the Invention

[0006] The purpose of the present invention is to overcome the above-mentioned deficiencies existing in the prior art, and provide a method and system for completing missing power data values based on time-aware diffusion. According to the collected massive power data, a TIED-GAN model is constructed to complete the missing power data values.

[0007] To achieve the above purpose, the present invention is implemented through the following technical solutions:

[0008] On the one hand, the technical solution of the present invention provides a method for completing missing power data values based on time-aware diffusion, including:

[0009] Obtain power data and perform preprocessing, and randomly mask a part of the power data to represent missing values;

[0010] Sample noise data, add the noise data to the masked part of the power data, and input it together with the unmasked original power data into a pre-constructed time-aware diffusion-based generative adversarial model for training to obtain predicted noise data;

[0011] Remove the predicted noise data from the masked part of the power data with the added noise data to obtain the completion result of the missing power data values.

[0012] In at least one embodiment, the time-aware diffusion-based generative adversarial model includes a noise predictor and a discriminator; the noise predictor is used to output predicted noise data based on the power data of the masked part with added noise data and the unmasked original power data; the discriminator discriminates between real sampled noise and generated noise data based on the noise data output by the noise predictor and the sampled noise data, and assists in generating a better noise predictor.

[0013] In at least one embodiment, the noise predictor includes a time-aware transformer layer and a feature transformer layer; according to the missing features of the power data, the position encoding module of the transformer is rewritten in the time-aware transformer layer to capture information related to the missing time interval; the feature transformer layer takes the tensor at each time point as input and learns the time dependence.

[0014] In at least one embodiment, the discriminator includes a downsampling convolutional block, a self-attention convolutional block, and an upsampling convolutional block.

[0015] In at least one embodiment, during training, the loss function of the noise predictor consists of a reconstruction error and the discriminator's discrimination result.

[0016] In at least one embodiment, the loss function of the noise predictor is:

[0017] ;

[0018] where is the reconstruction error, is the discriminator's discrimination result.

[0019] In at least one embodiment, the reconstruction error is the reconstruction error between the noise result output by the noise predictor and the real sampled noise, and the calculation method is:

[0020] ;

[0021] where is the original data, is the real sampled noise data, is the predicted noise data generated by the noise predictor, t is the time step; is the noise target, is the given conditional observation data.

[0022] In at least one embodiment, the discriminator's discrimination result The calculation formula is:

[0023] ;

[0024] where is the noise data of real sampling, and is the predicted noise data generated by the noise predictor.

[0025] In at least one embodiment, the power data includes multivariate power data of a substation and daily power consumption data of power users; the preprocessing includes data cleaning, data definition and storage; a part of information is randomly masked according to the missing rate to represent the missing data; the sampled noise data is Gaussian noise.

[0026] On the other hand, the technical solution of the present invention also provides a power data missing value completion system based on time-aware diffusion, including:

[0027] A data acquisition module, configured to: acquire power data and perform preprocessing, and randomly mask a part of the power data to represent the missing value;

[0028] A noise data prediction module, configured to: sample noise data, add the noise data to the masked part of the power data, and input it together with the unmasked original power data into a pre-constructed generative adversarial model based on time-aware diffusion for training to obtain the predicted noise data;

[0029] A data completion module, configured to: remove the predicted noise data from the power data with the masked part added with the noise data to obtain the completion result of the missing value of the power data.

[0030] The beneficial effects of the above technical solution of the present invention are as follows:

[0031] (1) According to the collected massive power data, the present invention completes the missing values of the power data by constructing a TIED-GAN model, which not only provides complete and accurate power data, but also ensures the data quality of downstream tasks (such as power data anomaly detection and power data classification, etc.).

[0032] (2) Aiming at the time series and long-term dependence of power data, the present invention uses a score-based diffusion model to complete the missing values of power data. On the one hand, the powerful generation effect of the score-based diffusion model ensures the generation quality of power data. On the other hand, the proposed TIED-GAN model, aiming at the continuous missing characteristics of power data, takes the time interval as an additional input of the time encoder, improving the accuracy of completion.

[0033] (3) The present invention uses the structure of the generative adversarial network, adds a discriminator after the noise predictor to assist in generating a better noise predictor, thereby improving the completion performance of the model.

[0034] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings

[0035] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not unduly limit the present invention.

[0036] Figure 1 is a schematic diagram of the overall process of a method for filling missing values in power data based on time-aware diffusion provided by an embodiment of the present invention;

[0037] Figure 2 is a schematic diagram of the structure of each module of a method for filling missing values in power data based on time-aware diffusion provided by an embodiment of the present invention;

[0038] Figure 3 is the overall architecture diagram of the TIED-GAN model provided by an embodiment of the present invention;

[0039] Figure 4 is the structure diagram of the noise predictor of the TIED-GAN model provided by an embodiment of the present invention;

[0040] Figure 5 is the structure diagram of the discriminator of the TIED-GAN model provided by an embodiment of the present invention. Detailed Embodiments

[0041] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0042] Term Explanation:

[0043] Score-based Diffusion Model: The score-based diffusion model is a mathematical model that describes how information or influence spreads in a complex network. Different from traditional integer-order diffusion models, the score-based diffusion model uses the concept of fractional calculus to describe the non-local and non-Markov properties in the propagation process.

[0044] The score-based diffusion model usually consists of fractional differential equations or fractional difference equations, where the fractional derivative describes the memory effect and long-range dependence in the propagation process. The characteristic of this model is that it can better capture the non-linear and non-local interaction characteristics existing in the actual system.

[0045] Generative Adversarial Networks (GAN): Generative Adversarial Networks (GAN) is a deep learning model proposed by Ian Goodfellow et al. in 2014. GAN consists of two neural networks: a Generator and a Discriminator, which are trained competitively and adversarially to generate realistic samples.

[0046] Generator: The Generator takes a random noise vector as input and attempts to generate fake samples similar to real samples. It generates new samples by learning the characteristics of the data distribution, with the goal of deceiving the Discriminator as much as possible so that it cannot distinguish between the generated fake samples and real samples.

[0047] Discriminator: The Discriminator takes the fake samples generated by the Generator and real samples and attempts to distinguish between them. It is similar to a binary classifier, with the goal of distinguishing real samples from the generated fake samples. The task of the Discriminator is to determine whether the input sample is a real sample or a fake sample generated by the Generator.

[0048] As introduced in the background art, the purpose of the present invention is to overcome the deficiencies in the above-mentioned existing technologies, and provide a method and system for filling missing values in power data based on time-aware diffusion. According to the collected massive power data, by constructing a TIED-GAN model, the missing values in the power data are filled.

[0049] Embodiment 1

[0050] In a typical implementation manner of the present invention, as Figure 1 and Figure 2 shown, this embodiment discloses a method for filling missing values in power data based on time-aware diffusion, including:

[0051] S100. Obtain power data and perform preprocessing, and randomly mask a part of the power data to represent missing values;

[0052] S200. Sample noise data, add the noise data to the masked part of the power data, and input it together with the unmasked original power data into a pre-constructed generative adversarial model based on time-aware diffusion for training to obtain predicted noise data;

[0053] S300. Remove the predicted noise data from the masked part of the power data with added noise data to obtain the filling result of the missing values in the power data.

[0054] The following will be described in detail in conjunction with the attached Figures 1 - 5 drawings.

[0055] S100. Obtain power data and preprocess it, randomly masking a part of the power data to represent missing values.

[0056] In this embodiment, taking the data of various databases obtained based on an online platform as an application example, a large amount of different types of power data related to electricity is collected and preprocessed, and then a part of the information is randomly masked according to the missing rate to represent missing data.

[0057] Specifically, S101: Power data acquisition, including multivariate power data (ETT dataset) of substations and daily electricity consumption data (Electricity dataset).

[0058] The ETT dataset comes from two substations and covers two years of data. The ETT dataset provides seven variable data including oil temperature, high-voltage side characteristic load, high-voltage side low load, medium-voltage side characteristic load, medium-voltage side low load, low-voltage side characteristic load, and low-voltage side low load. For ETTm1, "m" means the data is recorded every 15 minutes, with a total of 69,680 time points. On the other hand, ETTh1 is equivalent to the hourly data of ETTm1, with each hour containing 17,420 time steps. The ETT dataset helps to understand the operating status of power transformers and provides important information for the monitoring and maintenance of transformers, as well as for the monitoring and maintenance of the power system.

[0059] The Electricity dataset records the electricity consumption of 321 users. These data are recorded every hour and cover two years, with a total of 26,304 time steps. The Electricity dataset provides a detailed understanding of electricity consumption and helps to analyze and predict users' electricity consumption behavior and electricity consumption situation.

[0060] S102: Data preprocessing: Preprocess the large amount of power data obtained, including data cleaning, data definition, and storage. Data cleaning eliminates outliers based on statistical methods to improve data integrity and consistency. Data definition enhances the expression of time series patterns by unifying units and aligning timestamps. Data storage is implemented in a distributed system (HDFS / S3) to achieve efficient access, supplemented by time partitioning and indexing to accelerate queries.

[0061] S103: Random masking: Randomly mask a part of the information of the preprocessed power data according to the missing rate to represent missing data.

[0062] Suppose is a time series power data sample, whose shape is , where is the time series length, is the number of data features. In this embodiment, the time length Set them equal. Assume ( ) is an observation mask sample. If , it represents missing. If , it represents not missing. During the training process, multiple mask matrices with the same shape can be used to complete the purpose of self-supervised training through masking.

[0063] S200. Sample noise data, add the noise data to the masked part of the power data, and input it together with the unmasked original power data into a pre-constructed generative adversarial model based on time-aware diffusion for training to obtain predicted noise data.

[0064] S201: Sample noise. In this embodiment, Gaussian noise data is sampled as the input and output for training the generative adversarial model based on time-aware diffusion (TIED-GAN model).

[0065] S202: Construct a TIED-GAN model.

[0066] The overall structure of the TIED-GAN model is as Figure 3 shown. It includes a noise predictor and a discriminator. The noise predictor is trained with the masked part of the power data with added noise data, the unmasked original power data, and some additional parameters (e.g., the number of time embedding layers and the number of feature embedding layers) as inputs, and outputs the predicted noise data. The discriminator is trained based on the noise data output by the noise predictor and the sampled noise data. After training, the discriminator can distinguish between real sampled noise and generated noise data, and use the discrimination result as part of the loss function of the noise predictor to assist in generating a better noise predictor.

[0067] Among them, the noise predictor is a score-based diffusion model. This type of model is a latent variable model composed of a forward process and a reverse process. Both the forward process and the reverse process are parameterized Markov chains.

[0068] Among them, the forward process refers to the process of gradually adding Gaussian noise to the data until the data becomes random noise. For the original data , each step of the total diffusion process of T steps adds Gaussian noise to the data obtained in the previous step in the following way:

[0069] (1)

[0070] Here The variance adopted for each step, which ranges from 0 to 1. Usually, the later steps will adopt a larger variance, that is, satisfying . Under a well-designed variance strategy, if the number of diffusion steps T is large enough, then the finally obtained will completely lose the original data and become a random noise. Each step of the forward process generates a noisy data , and the entire forward process is also a Markov chain:

[0071] (2)

[0072] The reverse process is a denoising process. If the true distribution of each step of the reverse process is known, then from a random noise these distributions can be estimated using a neural network. The reverse process is also defined as a Markov chain, but it is composed of a series of Gaussian distributions parameterized by neural network parameters. Gradually denoising can generate a real sample, so the reverse process is also the process of generating data, as shown below:

[0073] (3)

[0074] (4)

[0075] Here , and are parameterized Gaussian distributions, and their means and variances are given by the trained networks and . is the covariance matrix of the conditional distribution in the reverse process. In DDPM, is simplified to the square of a variance vector . In fact, the diffusion model is to obtain these trained networks because they constitute the final generative model:

[0076] (5)

[0077] The Denoise Diffusion Probabilistic Model (DDPM), which considers the forward process with the specific parameterization as in formula (5):

[0078] where is a trainable denoising function, represents the variance (i.e., noise intensity) of the Gaussian noise added at the t -th step in the forward process, denotes the previoust The proportion of the original data retained step by step; In the reverse process, from Denoising generates The mean value of; Is the variance vector of the reverse process and is fixed as a parameter, Is ; Is to adjust the variance of the conditional distribution in the reverse process to ensure matching with the forward process. This means replacing the original predicted mean with the predicted noise, that is, training a good noise predictor.

[0079] Under this parameterization, the reverse process can be trained by solving the following optimization problem:

[0080] (6)

[0081] Denoising function Estimate the noise vector It is added to the noisy input On, this training objective is also regarded as a weighted combination of denoising score matching for training a score-based generative model. After training is completed, Can be sampled from formula (3).

[0082] The reverse process of the diffusion model gradually denoises through the noise predictor trained by the forward process, so the reverse process is a denoising process and also a process of generating data. To have excellent generated data quality or excellent missing value completion results, first, a good noise predictor needs to be trained . For , the DDPM training method is mainly adopted. Specifically, given the conditional observed data And the missing value completion target , sample the noise target For subsequent training.

[0083] In this embodiment, a discriminator Is added after the noise predictor To assist d Generate noise results close to the sampling. The discriminator d Distinguish the real sampling noise , from the generated noise . For the trained , if Is closer to the real sampling noise , Is closer to 1. The discrimination result will be used as part of the training of the noise predictor .

[0084] ​Training the noise predictor The loss function of consists of two parts: the reconstruction error and the discriminator's discrimination result.

[0085] Among them, the reconstruction error is the reconstruction error between the noise result output by the noise predictor and the real sampled noise, and the calculation method is:

[0086] (7)

[0087] The calculation formula for the discriminator's discrimination result is:

[0088] (8)

[0089] Based on this, the loss function of the noise predictor is:

[0090] (9)

[0091] As Figure 4 shown, in terms of structure, the noise predictor introduces two single-layer Transformer encoders, namely the time-aware Transformer layer and the feature Transformer layer: According to the missing features of the power data, in the time-aware Transformer layer, the position encoding module of the Transformer is rewritten so that it can capture information related to the missing time intervals. The feature Transformer layer takes the tensor at each time point as input and learns the time dependence.

[0092] When training this generator, first process the input data ( ) through Conv1×1 and ReLU. The number of diffusion embedding steps is converted through a fully connected layer (FC) and the SiLU activation function, connected with additional information and processed through Conv1×1. In residual layer 0, the time-aware Transformer layer and the feature Transformer layer capture the time and feature relationships, the gated unit regulates the information flow, and Conv1×1 completes the information fusion and transformation. The processing results interact with subsequent layers through skip connections and enter residual layer 1 until residual layer N −1. The information processed by each layer goes through operations such as Conv1×1, ReLU, etc., and finally outputs. During training, the parameters are continuously adjusted to make the generated data close to the real data distribution. Transformer is a model architecture that avoids repetition and completely relies on the attention mechanism to draw the global dependencies between the input and output. When constructing the model input data, data sequence information is added by adding position encoding to the input data. The dimension of the position encoding information is the same as the feature dimension of the input data at each time point, and the concatenated data and the encoded position are used as the input of the model. The position encoding formula of Transformer is as follows:

[0093] (10)

[0094] In the formula, t is the timestamp of the input data in a day, i is the dimension, K is the sampling frequency of the data. TIED-GAN uses time-aware positional encoding to replace the positional encoding of Transformer. Specifically:

[0095] (11)

[0096] Based on this, the variables i and j are used to represent the position index and the time-aware positional embedding dimension respectively. TPos ( i ) The function is used to merge the information of the position index i and the time interval, helping the model to obtain time-aware position information through the trainable parameters , and b so that the model can learn flexible functions.

[0097] As Figure 5 shown, the discriminator includes a downsampling convolutional block, a self-attention convolutional block, and an upsampling convolutional block.

[0098] Inside the downsampling convolutional block, the Leaky-ReLU activation function and the 3×3 convolutional layer appear alternately. Leaky-ReLU can effectively avoid the problem of gradient disappearance and provide a more stable gradient flow for network training; while the 3×3 convolutional layer focuses on extracting local features in the input data and captures the patterns and structures of the data in a small range through convolutional operations. Subsequently, the pooling operation reduces the dimension of the data, reduces the subsequent calculation amount, expands the receptive field, and provides a more representative low-dimensional feature representation for subsequent feature processing.

[0099] The self-attention convolutional block is the key part of the whole structure, which is dedicated to exploring the global dependency relationships between data features. First, the input data is linearly transformed through multiple 1×1 convolutional layers to adjust the channel dimension of the data. Then, the downsampling operation further reduces the data dimension. Then, the Softmax function is used to calculate the attention weights between different positions, enabling the model to pay attention to the long-range information interaction in the data. Finally, the attention weights are weighted and fused with the original data to enhance the model's understanding of the overall structure and semantics of the data.

[0100] The role of the upsampling convolutional block is to restore the low-dimensional features processed by the self-attention convolutional block to an appropriate dimension, while further extracting and optimizing features. Among them, the application of spectral normalization technology can effectively prevent the model from experiencing gradient explosion during training, ensuring the stability of training. The ReLU activation function introduces non-linearity into the model, enhancing the model's expressive ability and enabling it to learn more complex functional relationships. The upsampling layer increases the data dimension through methods such as interpolation, while the 3×3 convolutional layer extracts features from the data after dimension restoration again, further refining and integrating feature information.

[0101] During the discrimination process, the input data flows through the downsampling convolutional block, the self-attention convolutional block, and the upsampling convolutional block in sequence. During this process, the data is continuously extracted for features, dependencies are mined, and the dimension is adjusted. Finally, operations such as global pooling are used to summarize the processed data and output the discrimination result, thereby determining whether the input data comes from the true data distribution or the data generated by the generative model.

[0102] During the training process, real data and generated data are respectively input into this discriminator network. The discriminator discriminates the input data according to the above processing flow and outputs a judgment result. Subsequently, this judgment result is compared with the true label of the data (the true data label is marked as 1, and the generated data label is marked as 0), and formula (8) is used as the loss function to calculate the difference between the two. Based on the calculated loss value, the gradient is transmitted to each parameter of the network (such as the weight parameters of the convolutional layer) through the backpropagation algorithm, thereby updating these parameters, continuously adjusting the internal structure and parameter values of the discriminator, and gradually enhancing the discriminator's ability to distinguish real data and generated data, enabling it to make more accurate judgments in subsequent discrimination tasks.

[0103] Introducing a discriminator as the opponent of the generator in the GAN structure can discriminate real sampling noise and generated noise data, and assist in generating a better noise predictor. The noise closer to the real sampling , , the closer it is to 1. Training the discriminator d 's loss function is as follows:

[0104] (12)

[0105] Among them, is the noise data of real sampling, is the predicted noise data generated by the noise predictor.

[0106] S300. Remove the predicted noise data from the power data in the masked part where noise data is added to obtain the completion result of the missing values in the power data.

[0107] In the power data in the masked part where noise data is added, remove the noise data generated by the trained noise predictor. The obtained result is the completion result of the missing values in the power data.

[0108] Divide the power data set into a training set, a validation set, and a test set according to the ratio of 8:1:1. The training set is used for model training, the validation set is used for determining hyperparameters, and the test set is used for validating the model performance. Tables 1 and 2 show the comparison of the experimental results between the method of the present invention and the existing methods. Among them, MSE represents the mean square error, MAE represents the mean absolute error, and RMSE represents the root mean square error.

[0109] Table 1. Experimental results of TIED-GAN and other baselines on the ETT data set

[0110]

[0111] Table 2. Experimental results of TIED-GAN and other baselines on the Electricity data set

[0112]

[0113] From the results of Tables 1 and 2, it can be seen that TIED-GAN performs better than other benchmark models on the power data set: compared with the best benchmark model, the proposed model reduces the root mean square error (RMSE) by about 50% and reduces the mean absolute error (MAE) by about 25%. This is mainly because the model disclosed in this embodiment makes full use of the noise information and the time interval information, which enables TIED-GAN to obtain a noise predictor that can produce better results. When dealing with power data information, CSDI and SSSD do not take the time interval information as an additional input, so the obtained experimental results are poor. This also indirectly shows that the time interval information is an indispensable part in dealing with the imputation of missing values in time series power data. It can be seen that the method of this embodiment is superior to the existing completion methods in the power data completion work.

[0114] Embodiment 2

[0115] In a typical implementation manner of the present invention, this embodiment discloses a power data missing value completion system based on time-aware diffusion, including:

[0116] A data acquisition module, configured to: acquire power data and perform preprocessing, and randomly mask a part of the power data to represent missing values;

[0117] A noise data prediction module, configured to: sample noise data, add the noise data to the power data in the masked part, and input it together with the original unmasked power data into a pre-constructed generative adversarial model based on time-aware diffusion for training to obtain predicted noise data;

[0118] A data completion module, configured to: remove the predicted noise data from the power data in the masked part with the added noise data to obtain the completion result of the missing values of the power data.

[0119] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for completing missing values ​​of power data based on time-aware diffusion, characterized in that: include: Obtain power data and preprocess it, and randomly mask a portion of the power data to represent missing values; Sampling noise data, adding the noise data to the masked power data, and inputting the noise data together with the unmasked original power data into a pre-built generative adversarial model based on time-aware diffusion for training to obtain predicted noise data; The predicted noise data is removed from the power data in the masked part with the noise data added, so as to obtain the completion result of the missing value of the power data; The generative adversarial model based on time-aware diffusion includes a noise predictor and a discriminator; the noise predictor is used to output predicted noise data based on the power data of the masked part of the noise data added and the unmasked original power data; the discriminator discriminates the real sampled noise and the generated noise data based on the noise data output by the noise predictor and the sampled noise data, and assists in generating a noise predictor; The noise predictor includes a time-aware transformer layer and a feature transformer layer; according to the missing characteristics of power data, the position encoding module of the transformer is rewritten in the time-aware transformer layer to capture information related to the missing time interval; the feature transformer layer takes the tensor of each time point as input to learn the time dependency.

2. A method for completing missing values ​​of power data based on time-aware diffusion according to claim 1, characterized in that: The discriminator includes a downsampling convolution block, a self-attention convolution block and an upsampling convolution block.

3. The method for completing missing values ​​of power data based on time-aware diffusion according to claim 1, characterized in that: During training, the loss function of the noise predictor is composed of the reconstruction error and the discriminator's judgment result.

4. A method for completing missing values ​​of power data based on time-aware diffusion as claimed in claim 3, characterized in that: The loss function of the noise predictor is: ; In the formula, is the reconstruction error, is the discriminator’s judgment result.

5. A method for completing missing values ​​of power data based on time-aware diffusion as claimed in claim 4, characterized in that: The reconstruction error It is the reconstruction error between the noise result output by the noise predictor and the actual sampled noise, which is calculated as follows: ; In the formula, is the original data, is the real sampled noise data, The predicted noise data generated by the noise predictor, t is the time step; For noise targets, Observe the data for a given condition.

6. A method for completing missing values ​​of power data based on time-aware diffusion as claimed in claim 4, characterized in that: The discriminator determines the result The calculation formula is: ; In the formula, is the real sampled noise data, Predicted noise data generated for the noise predictor.

7. The method for completing missing values ​​of power data based on time-aware diffusion according to claim 1, characterized in that: The power data includes multivariable power data of substations and daily power consumption data of power users; preprocessing includes data cleaning, data definition and storage; randomly masking a part of the information according to the missing rate to represent missing data; the sampled noise data is Gaussian noise.

8. A system for completing missing values ​​of power data based on time-aware diffusion, adopting a method for completing missing values ​​of power data based on time-aware diffusion as claimed in any one of claims 1 to 7, characterized in that: include: The data acquisition module is configured to: acquire power data and perform preprocessing, and randomly mask a portion of the power data to represent missing values; The noise data prediction module is configured to: sample noise data, add the noise data to the masked power data, and input the noise data together with the unmasked original power data into a pre-built time-aware diffusion-based generative adversarial model for training to obtain predicted noise data; The data completion module is configured to remove the predicted noise data from the power data in the masked part with the noise data added, so as to obtain the completion result of the missing value of the power data.

Citation Information

Patent Citations

  • Contrast-agent-free medical diagnostic imaging

    CA3104607A1

  • Method, device, and storage medium for deep learning based domain adaptation with data fusion for aerial image data analysis

    US20220092420A1