Energy big data processing method and device based on improved generative adversarial network, electronic equipment and storage medium
By improving the generation of adversarial network prediction and filling missing data in energy data, the problem of instability in the missing data filling accuracy in the prior art is solved, and high-accurate missing data filling processing is achieved.
Patent Information
- Application Number
- CN202510287830.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing similarity-based missing data filling methods have unstable accuracy and cannot achieve accurate missing data filling processing.
The improved generative adversarial network is adopted to obtain energy data and meteorological data, calculate the correlation, select target meteorological data, and input it into the preset generative adversarial network to predict and fill in the missing data in the energy data.
Without relying on subjective selection of similarity indicators, the generative adversarial network can automatically predict missing data, ensuring that the accuracy of the predicted data is objective and stable, and achieving accurate missing data filling processing.
Smart Images

Figure CN120217152A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to an energy big data processing method, device, electronic device and storage medium based on an improved generative adversarial network. Background Art
[0002] Data loss is a common problem in power system data management, which will affect the integrity of energy data such as load curves, photovoltaic power generation and wind power generation. This kind of loss is often caused by sensor failures, communication interruptions or equipment maintenance. Research shows that about 70% of the missing data segments last for no more than 4 hours. Effectively filling these short-term missing data is crucial for improving data quality and enhancing energy management decisions (such as load forecasting, power generation capacity assessment, etc.).
[0003] The existing methods for filling missing data are mainly divided into two categories: physical model-based methods and data-driven methods. Among them, physical model-based methods are limited by the incompleteness and dynamic changes of system models and are difficult to be widely applied. Therefore, currently, data-driven methods are more widely used. The technology widely applied in data-driven methods is similarity-based technology. Similarity-based methods restore the missing values in missing data samples by matching the missing data samples with other historical data samples. However, the accuracy of this method depends on the selection of similarity metrics. In each data processing process, similarity metrics need to be selected, and the selection of similarity metrics is subjective. The selection of different similarity metrics may result in different missing value prediction results, leading to unstable accuracy of this method and unable to achieve precise filling processing of missing data. Summary of the Invention
[0004] The present invention provides an energy big data processing method, device, electronic device and storage medium based on an improved generative adversarial network to solve the technical problem that the accuracy of similarity-based methods is unstable and precise filling processing of missing data cannot be achieved.
[0005] To solve the above technical problem, an embodiment of the present invention provides an energy big data processing method based on an improved generative adversarial network, including: Obtain energy data and corresponding meteorological data for a preset time period; wherein, the energy data includes: load operation data of the power system, photovoltaic power generation data, and wind power generation data; the meteorological data includes: temperature, humidity, atmospheric pressure, irradiance, wind speed, and rainfall; Calculate the correlation between the meteorological data and the energy data, and select the meteorological data of the data ratio as target meteorological data in the order from high to low according to the correlation according to a preset data ratio; Input the target meteorological data and the energy data into a preset generative adversarial network, so that the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the missing data values in the energy data to output complete energy data.
[0006] As a preferred solution, calculating the correlation between the meteorological data and the energy data, and selecting the meteorological data of the data ratio as the target meteorological data in descending order of correlation according to a preset data ratio includes: Calculate the Pearson correlation coefficient value between the meteorological data and the energy data, and obtain the correlation between the meteorological data and the energy data according to the Pearson correlation coefficient value; According to a preset data ratio, select the meteorological data of the data ratio as the target meteorological data in descending order of correlation, and perform data normalization processing on the target meteorological data and the energy data; Among them, the Pearson correlation coefficient value between the meteorological data and the energy data is calculated according to the following formula: Among them, u is the energy data, v is the meteorological data, and r uv is the Pearson correlation coefficient value between the energy data u and the meteorological data v, u i is the i-th data sample in the energy data u, v i is the i-th data sample in the meteorological data v, is the sample average value of the energy data u, is the sample average value of the meteorological data v, and n is the total number of samples.
[0007] As a preferred solution, the generative adversarial network includes: a generator and a discriminator; The generation of the generative adversarial network includes: Obtain a number of training samples; among them, the training samples include: real historical energy data, historical energy data samples after adding masks, and corresponding historical target meteorological data; According to each of the training samples, alternately iteratively train the generator and the discriminator in the generative adversarial network until the total loss function of the generative adversarial network converges; Among them, during each alternating training, the model parameters of the generator are fixed, and the current training samples are input into the generator, so that the generator predicts the missing data in the masked part of the historical energy data samples, generates predicted complete historical energy data, and inputs the predicted complete historical energy data into the discriminator, so that the discriminator discriminates the probability that the predicted complete historical energy data is real data. Then, according to the probability and the predicted complete historical energy data, the loss function value of the discriminator is calculated, and the model parameters of the discriminator are updated according to the loss function value of the discriminator; Fix the updated model parameters of the discriminator, calculate the loss function value of the generator according to the predicted complete historical energy data, the real historical energy data, and the probability, and then update the model parameters of the generator according to the loss function value of the generator.
[0008] As a preferred solution, the generation of the training samples includes: Obtain real historical energy data and corresponding historical meteorological data; Calculate the Pearson correlation coefficient value between the historical meteorological data and the historical energy data, and then obtain the correlation between the historical meteorological data and the historical energy data according to the Pearson correlation coefficient value; According to the data ratio, select the historical meteorological data of the data ratio as the historical target meteorological data in the order of decreasing correlation from high to low, and perform data normalization processing on the historical target meteorological data and the historical energy data; Generate a corresponding historical energy data curve according to the normalized historical energy data; Slide a moving window with a preset window length on the historical energy data curve in a preset sliding direction. Whenever the moving window slides a preset distance, the historical energy data in the moving window is used as a corresponding historical energy data sample; Add a mask with a preset length to each historical energy data sample, and then generate corresponding training samples according to the historical energy data sample after adding the mask, the real historical energy data, and the normalized historical meteorological data.
[0009] As a preferred solution, predicting the missing data values in the energy data according to the target meteorological data and the energy data, and filling the energy data with the predicted missing data values to output complete energy data includes: Extract the data features of the target meteorological data and the energy data; Perform data normalization processing on the data features, and then combine the normalized data features to generate corresponding combined data features; Predict the missing data values in the energy data according to the combined data features, and fill in the energy data with the predicted missing data values to output complete energy data.
[0010] As a preferred solution, the loss function of the generator is: where L G is the generator loss function, H is the length of the mask added to the historical energy data, O is the historical energy data, M is the dimension of the discriminator output, G(z) is the sample generated by the generator, D(G(z)) is the discriminator's discrimination result for G(z), and α is the weight coefficient; The loss function of the discriminator is: where L D is the discriminator loss function, D(O) is the discriminator's discrimination result for O, is a random sample, is the discriminator gradient of the discriminator network in the random sample space, β is the weight coefficient, and λ is a random number.
[0011] As a preferred solution, after the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills in the energy data with the predicted missing data values to output complete energy data, it further includes: Calculate the corresponding standard root mean square error, energy error, and model bias of the generative adversarial network according to the complete energy data, and then evaluate the model performance of the generative adversarial network according to the standard root mean square error, energy error, and model bias; Among them, the calculation formula of the standard root mean square error is: The calculation formula of the energy error is: The calculation formula of the model bias is: where ε RMSE is the standard root mean square error, ε EE is the energy error, ε bias is the model bias, n is the total number of samples, T i de is the time length of the missing data in the energy data, is the predicted missing data segment, is the true value of the energy data, It is the predicted value of energy data.
[0012] Based on the above embodiments, another embodiment of the present invention provides an energy big data processing device based on an improved generative adversarial network, including: a data acquisition module, a target meteorological data selection module, and a data filling module; The data acquisition module is used to acquire energy data and corresponding meteorological data for a preset period; wherein, the energy data includes: load operation data of the power system, photovoltaic output data, and wind power output data; the meteorological data includes: temperature, humidity, atmospheric pressure, irradiance, wind speed, and rainfall; The target meteorological data selection module is used to calculate the correlation between the meteorological data and the energy data, and select the meteorological data of the data ratio as the target meteorological data according to the preset data ratio in the order from high to low correlation; The data filling module is used to input the target meteorological data and the energy data into a preset generative adversarial network, so that the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the energy data according to the predicted missing data value, and outputs complete energy data.
[0013] Based on the above embodiments, another embodiment of the present invention provides an electronic device, the device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the energy big data processing method based on the improved generative adversarial network described in the above embodiments of the present invention.
[0014] Based on the above embodiments, another embodiment of the present invention provides a storage medium, the storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the storage medium is located to execute the energy big data processing method based on the improved generative adversarial network described in the above embodiments of the present invention.
[0015] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention provides a method for processing energy big data based on an improved generative adversarial network. Compared with the similarity-based method in the prior art, the present invention does not need to rely on the subjective selection of similarity metrics. In each process of processing missing data, the selected target meteorological data and energy data are input into a pre-trained generative adversarial network. The generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the missing data values in the energy data according to the predicted missing data values, and outputs complete energy data. Since the generative adversarial network is pre-trained and its model parameters are fixed, during the process of processing missing data, it is not necessary to make a subjective parameter selection by humans, and the missing data values in the energy data can be automatically predicted, which can ensure that the accuracy of the predicted data is objectively stable during the prediction process, and thus achieve accurate processing of missing data filling. Description of the Drawings
[0016] Figure 1 is a schematic flowchart of a method for processing energy big data based on an improved generative adversarial network provided by an embodiment of the present invention; Figure 2 is a schematic structural diagram of a generator of a generative adversarial network; Figure 3 is a schematic structural diagram of a discriminator of a generative adversarial network; Figure 4 is a framework diagram for processing energy data based on a generative adversarial network; Figure 5 is a schematic structural diagram of a device for processing energy big data based on an improved generative adversarial network provided by an embodiment of the present invention. Detailed Embodiments
[0017] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion.
[0019] In the description of the embodiments of the present application, technical terms such as "first" and "second" are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "a plurality of" is more than two, unless otherwise clearly and specifically defined.
[0020] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0021] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0022] In the description of the embodiments of the present application, the term "a plurality of" refers to more than two (including two). Similarly, "a plurality of groups" refers to more than two groups (including two groups), and "a plurality of pieces" refers to more than two pieces (including two pieces).
[0023] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installation", "connection", "connection", "fixation", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can also be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.
[0024] Embodiment 1 Please refer to Figure 1 , which is a schematic flowchart of a method for processing energy big data based on an improved generative adversarial network provided by an embodiment of the present invention, and includes the following specific steps: S1. Obtain energy data and corresponding meteorological data for a preset period; wherein, the energy data includes: load operation data of the power system, photovoltaic output data, and wind power output data; the meteorological data includes: temperature, humidity, atmospheric pressure, irradiance, wind speed, and rainfall; Specifically, first, collect the energy data and corresponding meteorological data within a preset time period as the original data set. The energy data includes the load operation data of the power system, the photovoltaic output data, and the wind power output data, and the meteorological data includes temperature, humidity, atmospheric pressure, irradiance, wind speed, and rainfall; among them, there are missing data in the energy data.
[0025] S2. Calculate the correlation between the meteorological data and the energy data, and select the meteorological data of the data ratio as the target meteorological data in descending order of correlation according to the preset data ratio; Preferably, calculating the correlation between the meteorological data and the energy data, and selecting the meteorological data of the data ratio as the target meteorological data in descending order of correlation according to the preset data ratio includes: calculating the Pearson correlation coefficient value between the meteorological data and the energy data, and obtaining the correlation between the meteorological data and the energy data according to the Pearson correlation coefficient value; selecting the meteorological data of the data ratio as the target meteorological data in descending order of correlation according to the preset data ratio, and performing data normalization processing on the target meteorological data and the energy data; among them, the Pearson correlation coefficient value between the meteorological data and the energy data is calculated according to the following formula: where u is the energy data, v is the meteorological data, and r uv is the Pearson correlation coefficient value between the energy data u and the meteorological data v, u i is the i-th data sample in the energy data u, v i is the i-th data sample in the meteorological data v, is the sample average value of the energy data u, is the sample average value of the meteorological data v, and n is the total number of samples.
[0026] Specifically, use the Pearson correlation coefficient PCC to perform correlation analysis on the energy data and the meteorological data, screen out the meteorological data with strong correlation with the energy data as the target meteorological data, and perform data preprocessing on the energy data and the target meteorological data. The specific process includes: (1) Correlation analysis: Calculate the Pearson correlation coefficient value between the meteorological data and the energy data, and obtain the correlation between the meteorological data and the energy data according to the Pearson correlation coefficient value.
[0027] Among them, the value of the Pearson correlation coefficient ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no linear correlation. The closer the absolute value of the coefficient is to 1, the greater the correlation between the two variables. Calculate the PCC coefficients between various meteorological data (temperature, humidity, atmospheric pressure, irradiance, wind speed, rainfall, etc.) and energy data (load operation data, photovoltaic output, and wind power output data, etc.), sort them according to the magnitude of the PCC coefficients, set the selected data ratio, eliminate the features with the smallest correlation, and select the meteorological data with the said data ratio as the target meteorological data.
[0028] The calculation formula of the Pearson correlation coefficient PCC is: In the formula, u is the energy data variable, v is the meteorological factor variable, and r uv is the Pearson correlation coefficient of variables u and v, and u i is the i-th data sample of variable u, and v i is the i-th data sample of variable v. is the average value of the variable u sample, is the average value of the variable v sample, and n is the total number of samples (2) Data preprocessing: Perform data normalization on the target meteorological data and the energy data, and convert the attributes of all target meteorological data and energy data into the same scale and range. The calculation formula for data normalization is: Among them, is the i-th original data, and x min is the i-th data after normalization, and x max is the minimum value in the original data, and x max is the maximum value in the original data.
[0029] Finally, the normalized target meteorological data forms the feature variables: Among them, L is the length of the time series, and F is the number of feature variables.
[0030] The normalized energy data can be expressed as Y = [y 1 , y 2 , L, y L . If the event is defined as data loss, the energy data can be expressed as: Y = [Y pre , Y de , Y post , where y pre is the energy data segment before the event, and Y de is the energy data segment during the event, and Ypost It is the energy data segment after the event.
[0031] By performing a correlation analysis on the meteorological data and the energy data, the following beneficial effects can be achieved: (1) Improve computational efficiency: When the dataset contains a large number of features, the computational cost will increase significantly. Feature screening can eliminate irrelevant features, reduce the number of features, thereby reducing the computational cost and accelerating the model training process.
[0032] (2) Reduce overfitting: Too many features may cause the model to be too complex and prone to overfitting. Feature screening simplifies the model structure by removing redundant and irrelevant features, reduces the risk of overfitting, and improves the generalization ability of the model.
[0033] (3) Enhance model performance: Retaining features with a high correlation with the target variable can make the model more focused on learning the factors that truly affect the prediction results, thereby improving the prediction accuracy of the model. In addition, reducing the number of features can also speed up the model training speed.
[0034] S3. Input the target meteorological data and the energy data into a preset generative adversarial network, so that the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the energy data with the predicted missing data values to output complete energy data.
[0035] Preferably, predicting the missing data values in the energy data according to the target meteorological data and the energy data, and filling the energy data with the predicted missing data values to output complete energy data includes: extracting the data features of the target meteorological data and the energy data; performing data normalization processing on the data features, and then combining the normalized data features to generate corresponding combined data features; predicting the missing data values in the energy data according to the combined data features, and filling the energy data with the predicted missing data values to output complete energy data.
[0036] Specifically, after normalizing the energy data and the target meteorological data, the normalized energy data and target meteorological data are input into a pre-trained generative adversarial network, so that the generative adversarial network extracts the data features of the target meteorological data and the energy data according to the target meteorological data and the energy data; the data features are normalized, and then the normalized data features are combined to generate corresponding combined data features; according to the combined data features, the missing data values in the energy data are predicted, and the energy data is filled with the predicted missing data values to output complete energy data.
[0037] Therefore, the present invention also has the ability to deeply extract features from bidirectional time series data and multiple meteorological variables. This feature enables the present invention to comprehensively consider the influence of time factors and environmental factors on energy data, thereby more accurately restoring the true details of the missing data section, effectively filling the data gap, and improving the integrity and usability of the data.
[0038] Preferably, the generative adversarial network includes: a generator and a discriminator; the generation of the generative adversarial network includes: obtaining a number of training samples; wherein, the training samples include: real historical energy data, historical energy data samples after adding masks, and corresponding historical target meteorological data; according to each of the training samples, the generator and the discriminator in the generative adversarial network are alternately iteratively trained until the total loss function of the generative adversarial network converges; wherein, in each alternate training, the model parameters of the generator are fixed, and the current training sample is input into the generator, so that the generator predicts the missing data in the masked part of the historical energy data sample to generate a predicted complete historical energy data, and the predicted complete historical energy data is input into the discriminator, so that the discriminator discriminates the probability that the predicted complete historical energy data is real data, and then according to the probability and the predicted complete historical energy data, the loss function value of the discriminator is calculated, and the model parameters of the discriminator are updated according to the loss function value of the discriminator; the updated model parameters of the discriminator are fixed, and according to the predicted complete historical energy data, the real historical energy data and the probability, the loss function value of the generator is calculated, and then the model parameters of the generator are updated according to the loss function value of the generator.
[0039] Preferably, the generation of the training samples includes: obtaining real historical energy data and corresponding historical meteorological data; calculating the Pearson correlation coefficient value between the historical meteorological data and the historical energy data, and then obtaining the correlation between the historical meteorological data and the historical energy data according to the Pearson correlation coefficient value; according to the data ratio, selecting the historical meteorological data of the data ratio as the historical target meteorological data in the order from high to low according to the correlation, and performing data normalization processing on the historical target meteorological data and the historical energy data; generating a corresponding historical energy data curve according to the normalized historical energy data; sliding a moving window with a preset window length on the historical energy data curve in a preset sliding direction, and whenever the moving window slides a preset distance, taking the historical energy data in the moving window as a corresponding historical energy data sample; adding a mask with a preset length to each historical energy data sample, and then generating a corresponding training sample according to the historical energy data sample after adding the mask, the real historical energy data, and the normalized historical meteorological data.
[0040] Preferably, the loss function of the generator is: where L G is the generator loss function, H is the length of the mask added to the historical energy data, O is the historical energy data, M is the dimension of the discriminator output, G(z) is the sample generated by the generator, D(G(z)) is the discriminator's discrimination result for G(z), and α is the weight coefficient; The loss function of the discriminator is: where L D is the discriminator loss function, D(O) is the discriminator's discrimination result for O, is a random sample, is the discriminator gradient of the discriminator network in the random sample space, β is the weight coefficient, and λ is a random number.
[0041] Specifically, the training process of the generative adversarial network is as follows: (1) Generate training samples: Obtain real historical energy data and corresponding historical meteorological data; calculate the Pearson correlation coefficient value between the historical meteorological data and the historical energy data, and then obtain the correlation between the historical meteorological data and the historical energy data according to the Pearson correlation coefficient value; according to the data ratio, select the historical meteorological data of the data ratio as the historical target meteorological data in the order from high to low according to the correlation, and perform data normalization processing on the historical target meteorological data and the historical energy data.
[0042] Generate a corresponding historical energy data curve based on the normalized historical energy data; slide a moving window of a preset window length on the historical energy data curve in a preset sliding direction. Whenever the moving window slides a preset distance, use the historical energy data within the moving window as a corresponding historical energy data sample; add a mask of a preset length to each historical energy data sample, and then generate a corresponding training sample based on the historical energy data sample after adding the mask, the real historical energy data, and the normalized historical meteorological data.
[0043] In a specific embodiment, slide a 24-hour moving window on the historical energy data curve, set the time shift of the moving window to one hour, generate a sample every time it moves forward one hour, and add a mask to cover the middle part of each sample to obtain a historical energy data sample. Place the missing data segment at the center of the 24-hour window so that the data lengths before and after the event are equal. Among them, use a boolean mask to represent the missing data segment, with the event period being 1 and 0 at other times. Generally, the length of the data masked by the mask should be greater than the length of the largest missing data segment to achieve a smooth transition between non-event periods and event periods. According to the relevant statistical analysis of the length of each missing load data segment in the real-world smart meter dataset, it is found that most of the missing load data segments are less than 4 hours.
[0044] (2) Alternating iterative training of the generator and the discriminator: Please refer to Figure 2 , which is a schematic structural diagram of the generator of the generative adversarial network. The following are Figure 2 the Chinese interpretations of each structure in
[0045] The structure of the generator is: 4 consecutive DBR layers, a transposed convolutional layer, and a Tanh activation function, where the DBR layer consists of a transposed convolutional layer, a batch normalization layer, and a ReLU activation function. The generator loss function is: In the formula, L G is the loss function of the generator, H is the length of the masked segment, O is the real load curve, M is the dimension of the discriminator output, G(z) is the sample generated by the generator, D(G(z)) is the discriminant result of the discriminator for G(z), and α is the weight coefficient.
[0046] After the data is input, it passes through 4 consecutive DBR layers. Through transposed convolution operations, data features are extracted and these features are combined. Then, through the batch normalization layer, the data features are normalized. The activation function layer is used to introduce non - linear factors, enabling the network to fit more complex functional relationships. The data finally enters the output layer, which is a convolutional layer matching the shape and dimension of the generated data. The output layer combines the extracted features and generates the missing segment data that meets the given conditions.
[0047] Please refer to Figure 3 , which is the structural schematic diagram of the discriminator of the generative adversarial network. The following are Figure 3 the Chinese interpretations of each structure in
[0048] The structure of the discriminator of the improved generative adversarial network constructed in step S3 is: 4 consecutive CBL layers, a convolutional layer, and a Sigmoid activation function, where the CBL layer consists of a convolutional layer, a batch normalization layer, and a Leakey ReLU activation function. Discriminator loss function: where, L D is the discriminator loss function, D(O) is the discriminator's discrimination result for O, is a random sample, is the discriminator gradient of the discriminator network in the random sample space, β is the weight coefficient, and λ is a random number.
[0049] The inputs of the discriminator include: the historical target meteorological data after feature screening, the predicted complete energy data value generated by the generator, and the actual energy data value concatenated into a matrix as the input of the discriminator. After convolution through 4 consecutive CBL layers, low - dimensional features of high - dimensional data are extracted. Finally, through the convolutional layer and the Sigmoid activation function, discrimination information is obtained to determine whether the input data is a generated sample or a real sample. The final output result is a scalar value representing the probability that the input data is real data.
[0050] Please refer to Figure 4 , which is the framework diagram of energy data processing based on the generative adversarial network. According to each of the training samples, the generator and discriminator in the generative adversarial network are alternately iteratively trained until the total loss function of the generative adversarial network converges; Among them, during each alternating training, the model parameters of the generator are fixed, and the current training samples are input into the generator, so that the generator predicts the missing data in the masked part of the historical energy data samples, generates the predicted complete historical energy data, and inputs the predicted complete historical energy data into the discriminator, so that the discriminator discriminates the probability that the predicted complete historical energy data is real data. Then, according to the probability and the predicted complete historical energy data, the loss function value of the discriminator is calculated, and the model parameters of the discriminator are updated according to the loss function value of the discriminator; the updated model parameters of the discriminator are fixed, and the loss function value of the generator is calculated according to the predicted complete historical energy data, the real historical energy data, and the probability, and then the model parameters of the generator are updated according to the loss function value of the generator.
[0051] The complete process of one iteration: First, fix the network parameters of the generator, calculate the gradient according to the discriminator loss function, update the network parameters of the discriminator, and train the discriminator to better distinguish real data and generated data; then fix the network parameters of the discriminator, calculate the gradient according to the generator loss function, update the network parameters of the generator, and train the generator to generate more realistic data. This process will be iterated continuously until the preset number of training rounds is reached or the total loss function is lower than the set threshold.
[0052] During the training process of CGAN, the update of the discriminator network parameters is realized by the gradient descent algorithm. The specific steps are as follows: Forward propagation: First, input the real data, the generated data, and their corresponding conditional information into the discriminator respectively to obtain the discrimination results; Calculate the loss: Then, calculate the loss function of the discriminator according to the discrimination results and the real labels (the real data label is 1, and the generated data label is 0). The loss function is used to measure the difference between the discriminator output and the real labels; Backward propagation: Next, use the backward propagation algorithm to calculate the gradient of the loss function with respect to the discriminator network parameters. This step is realized by the chain rule, and the gradient is calculated layer by layer from the output layer backward; Parameter update: Finally, update the network parameters of the discriminator according to the calculated gradient. This is usually realized by the gradient descent algorithm, such as optimization algorithms like SGD, Adam, etc. These algorithms will adjust the network parameters according to the magnitude and direction of the gradient to minimize the loss function; The network parameters of the generator and the discriminator: mainly include the weights, biases, etc. in the network architectures of the generator and the discriminator. The specific number of network parameters depends on its architecture, such as the number of network layers, the number of neurons in each layer, etc.
[0053] Preferably, after the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the energy data with the predicted missing data values to output complete energy data, the method further includes: calculating the standard root mean square error, energy error, and model bias corresponding to the generative adversarial network according to the complete energy data, and then evaluating the model performance of the generative adversarial network according to the standard root mean square error, energy error, and model bias; Among them, the calculation formula of the standard root mean square error is: The calculation formula of the energy error is: The calculation formula of the model bias is: Among them, ε RMSE is the standard root mean square error, ε EE is the energy error, ε bias is the model bias, n is the total number of samples, T i de is the time length of the missing data in the energy data, is the predicted missing data segment, is the true value of the energy data, is the predicted value of the energy data.
[0054] Specifically, according to the complete energy data, calculate the standard root mean square error, energy error, and model bias corresponding to the generative adversarial network, and then evaluate the model performance of the generative adversarial network according to the standard root mean square error, energy error, and model bias. The smaller the error value, the better the model effect and the smaller the error of the generated filling data. Their calculation formulas are as follows: Standard root mean square error: Energy error: Bias: In the formula, ε RMSE is the standard root mean square error, ε EE is the energy error, ε bias is the bias, n is the total number of samples, T i de is the duration length of the event, that is, the data missing time length, represents the energy data segment in the predicted event, represents the true value, Represents the predicted value.
[0055] In summary, the present invention provides a method for processing energy big data based on an improved generative adversarial network. Since the generative adversarial network is pre-trained and its model parameters are fixed, there is no need for subjective parameter selection by humans during the processing of missing data, and the missing data values in the energy data can be automatically predicted, ensuring the objectivity and stability of the accuracy of the predicted data during the prediction process, and then achieving accurate filling processing of the missing data.
[0056] In addition, the processed complete energy big data can help optimize the power distribution of the power system, support real-time state monitoring of the power grid (such as voltage, current, load distribution), optimize the power dispatching strategy, avoid overload or resource waste, and improve the processing efficiency of the power grid. In terms of prediction analysis, the complete energy big data can also perform power demand prediction or renewable energy generation prediction. Among them, power demand prediction can be divided into short-term load prediction and long-term trend analysis. Short-term load prediction uses historical load data and information such as weather and holidays to predict the power demand in the next few hours to days and optimize the power generation plan; long-term trend analysis analyzes the relationship between economic, population growth and energy consumption to support the expansion of the power grid or the planning of new energy infrastructure. In terms of equipment maintenance, the complete energy big data can also be used to analyze equipment operation data to predict faults and perform preventive maintenance.
[0057] Embodiment 2 Please refer to Figure 5 , which is a schematic structural diagram of an energy big data processing device based on an improved generative adversarial network provided by an embodiment of the present invention. The device includes: a data acquisition module, a target meteorological data selection module, and a data filling module; The data acquisition module is used to acquire energy data and corresponding meteorological data for a preset time period; wherein, the energy data includes: load operation data of the power system, photovoltaic output data, and wind power output data; the meteorological data includes: temperature, humidity, atmospheric pressure, irradiance, wind speed, and rainfall; The target meteorological data selection module is used to calculate the correlation between the meteorological data and the energy data, and select the meteorological data of the data ratio as the target meteorological data in descending order of correlation according to the preset data ratio; The data filling module is used to input the target meteorological data and the energy data into a preset generative adversarial network, so that the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the energy data with the predicted missing data values to output complete energy data.
[0058] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement this without creative efforts.
[0059] Those skilled in the art can clearly understand that for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiment, which will not be elaborated here.
[0060] Embodiment III Correspondingly, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the energy big data processing method based on the improved generative adversarial network described in the above-mentioned invention embodiments.
[0061] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The device may include, but is not limited to, a processor and a memory.
[0062] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the device, connecting various parts of the entire device through various interfaces and lines.
[0063] Embodiment IV Correspondingly, an embodiment of the present invention provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute the energy big data processing method based on the improved generative adversarial network described in the above embodiments of the present invention.
[0064] The memory can be used to store the computer program. By running or executing the computer program stored in the memory, and calling the data stored in the memory, various functions of the device can be realized. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0065] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0066] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A method for processing energy big data based on improved generative adversarial networks, characterized in that: include: Obtaining energy data and corresponding meteorological data for a preset period of time; wherein the energy data includes: load operation data of the power system, photovoltaic output data, and wind power output data; the meteorological data includes: temperature, humidity, atmospheric pressure, irradiance, wind speed, and rainfall; Calculating the correlation between the meteorological data and the energy data, and selecting the meteorological data with the data ratio as the target meteorological data according to the order of correlation from high to low according to the preset data ratio; The target meteorological data and the energy data are input into a preset generative adversarial network, so that the generative adversarial network predicts the missing data in the energy data according to the target meteorological data and the energy data, and fills the energy data according to the predicted missing data values, and outputs complete energy data.
2. The energy big data processing method based on improved generative adversarial network according to claim 1 is characterized in that: The calculating the correlation between the meteorological data and the energy data, and selecting the meteorological data with the data ratio as the target meteorological data according to the order of high to low correlation according to the preset data ratio, includes: Calculating a Pearson correlation coefficient value between the meteorological data and the energy data, and obtaining a correlation between the meteorological data and the energy data according to the Pearson correlation coefficient value; According to a preset data ratio, in order of relevance from high to low, meteorological data of the data ratio is selected as target meteorological data, and data normalization processing is performed on the target meteorological data and the energy data; The Pearson correlation coefficient between the meteorological data and the energy data is calculated according to the following formula: Among them, u is energy data, v is meteorological data, r uv is the Pearson correlation coefficient between energy data u and meteorological data v, u i is the i-th data sample in energy data u, v i is the i-th data sample in the meteorological data v, is the sample average of energy data u, is the sample mean of the meteorological data v, and n is the total number of samples.
3. The energy big data processing method based on improved generative adversarial network according to claim 2 is characterized in that: The generative adversarial network includes: a generator and an adversary; The generation of the generative adversarial network includes: Acquire a number of training samples; wherein the training samples include: real historical energy data, historical energy data samples with masks added, and corresponding historical target meteorological data; According to each of the training samples, the generator and the discriminator in the generative adversarial network are alternately and iteratively trained until the total loss function of the generative adversarial network converges; Wherein, in each alternating training, the model parameters of the generator are fixed, and the current training sample is input into the generator, so that the generator predicts the missing data of the masked part in the historical energy data sample, generates predicted complete historical energy data, and inputs the predicted complete historical energy data into the discriminator, so that the discriminator determines the probability that the predicted complete historical energy data is true data, and then calculates the loss function value of the discriminator according to the probability and the predicted complete historical energy data, and updates the model parameters of the discriminator according to the loss function value of the discriminator; The updated model parameters of the discriminator are fixed, and the loss function value of the generator is calculated according to the predicted complete historical energy data, the real historical energy data and the probability, and then the model parameters of the generator are updated according to the loss function value of the generator.
4. The energy big data processing method based on improved generative adversarial network according to claim 3 is characterized in that: The generation of the training samples includes: Obtain real historical energy data and corresponding historical meteorological data; Calculating a Pearson correlation coefficient value between the historical meteorological data and the historical energy data, and then obtaining a correlation between the historical meteorological data and the historical energy data according to the Pearson correlation coefficient value; According to the data ratio, in order of relevance from high to low, selecting the historical meteorological data of the data ratio as the historical target meteorological data, and performing data normalization processing on the historical target meteorological data and the historical energy data; Generate a corresponding historical energy data curve based on the normalized historical energy data; According to a preset sliding direction, a moving window of a preset window length is slid on the historical energy data curve, and whenever the moving window slides a preset distance, the historical energy data in the moving window is used as a corresponding historical energy data sample; A mask of a preset length is added to each historical energy data sample, and then corresponding training samples are generated according to the historical energy data sample after adding the mask, the real historical energy data and the normalized historical meteorological data.
5. The energy big data processing method based on improved generative adversarial network according to claim 4 is characterized in that: The step of predicting missing data values in the energy data according to the target meteorological data and the energy data, and performing data filling on the energy data according to the predicted missing data values to output complete energy data includes: extracting data features of the target meteorological data and the energy data; Performing data normalization processing on the data features, and then combining the normalized data features to generate corresponding combined data features; According to the combined data features, missing data values in the energy data are predicted, and data filling is performed on the energy data according to the predicted missing data values to output complete energy data.
6. The energy big data processing method based on improved generative adversarial network according to claim 5 is characterized in that: The loss function of the generator is: Among them, L G is the generator loss function, H is the length of the mask added to the historical energy data, O is the historical energy data, M is the dimension of the discriminator output, G(z) is the sample generated by the generator, D(G(z)) is the discriminator's judgment result on G(z), and α is the weight coefficient; The loss function of the discriminator is: Among them, L D is the discriminator loss function, D(O) is the discriminator’s judgment result on O, is a random sample, is the discriminator gradient of the discriminator network in the random sample space, β is the weight coefficient, and λ is a random number.
7. The energy big data processing method based on improved generative adversarial network according to claim 1 is characterized in that: After the generative adversarial network predicts missing data in the energy data according to the target meteorological data and the energy data, and fills the energy data with data according to the predicted missing data values, and outputs complete energy data, the method further includes: According to the complete energy data, the standard root mean square error, energy error and model deviation corresponding to the generative adversarial network are calculated, and then the model performance of the generative adversarial network is evaluated according to the standard root mean square error, energy error and model deviation; The calculation formula of the standard root mean square error is: The calculation formula of the energy error is: The calculation formula of the model deviation is: Among them, ε RMSE is the standard root mean square error, ε EE is the energy error, ε bias is the model deviation, n is the total number of samples, is the length of time that the energy data is missing, To predict missing data segments, is the true value of energy data, is the predicted value of energy data.
8. An energy big data processing device based on an improved generative adversarial network, characterized in that: include: Data acquisition module, target meteorological data selection module and data filling module; The data acquisition module is used to acquire energy data and corresponding meteorological data for a preset period of time; wherein the energy data includes: load operation data of the power system, photovoltaic output data and wind power output data; the meteorological data includes: temperature, humidity, atmospheric pressure, irradiance, wind speed and rainfall; The target meteorological data selection module is used to calculate the correlation between the meteorological data and the energy data, and select the meteorological data with the data ratio as the target meteorological data according to the order of correlation from high to low according to the preset data ratio; The data filling module is used to input the target meteorological data and the energy data into a preset generative adversarial network, so that the generative adversarial network predicts the missing data in the energy data based on the target meteorological data and the energy data, and fills the energy data according to the predicted missing data values, and outputs complete energy data.
9. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the energy big data processing method based on the improved generative adversarial network as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein, when the computer program is running, the device where the storage medium is located is controlled to execute the energy big data processing method based on the improved generative adversarial network as described in any one of claims 1 to 7.