Industrial user missing data filling method based on SA-GAN

By using visualization and SA-GAN models to process missing data from industrial users, the problem of insufficient accuracy in traditional methods has been solved, achieving higher accuracy in data imputation and load forecasting, thus contributing to the development of the electricity market.

CN116861177BActive Publication Date: 2026-01-20NANJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310823815.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-06
Publication Date
2026-01-20
Estimated Expiration
2043-07-06

AI Technical Summary

Technical Problem

Traditional methods for filling in missing data for industrial users are inaccurate and cannot meet the high-precision requirements of the electricity market.

Method used

A method based on SA-GAN is used to visualize the daily load data of industrial users. An improved grey relational algorithm is used to filter user datasets and historical datasets with high correlation. The SA-GAN model is used for image filling and inverse image processing to restore the load data.

Benefits of technology

It has enabled more accurate filling of missing data for industrial users, improved the accuracy of user load forecasting, and promoted the maturation of the electricity market.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116861177B_ABST
    Figure CN116861177B_ABST
Patent Text Reader

Abstract

The application provides an industrial user missing data filling method based on SA-GAN, which comprises the following steps: 1) constructing the industry correlation degree of the industrial user based on the grey correlation degree algorithm considering positive and negative correlation degrees, using the daily load data of each industrial user at the same time sequence to form a data set A for multiple industrial users with a correlation degree higher than a threshold value, and using historical daily load data to form a data set B for a single industrial user with a correlation degree not higher than the threshold value; 2) determining the abnormal data in the data set A and the data set B by using the quartile method, and setting the value of the abnormal data as 0; 3) converting the daily load data in the data set A and the data set B into RGB color pixel values, so that the data set A and the data set B are imaged to obtain an image A and an image B; 4) inputting the image A and the image B into an SA-GAN model, filling the image A and the image B to obtain an SA-GAN filled image E and an SA-GAN filled image F; and 5) performing an inverse imaging process on the image E and the image F to convert the image E and the image F into daily load data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of filling in missing data of industrial users, and particularly to an industrial user missing data filling method based on SA-GAN. BACKGROUND

[0002] With the development of modern science and technology industry, more and more new users access the integrated energy system, so in the direction of the electricity market, accurate prediction of power user load data is extremely important, and high-quality historical data can effectively improve the prediction accuracy, but the traditional filling method has poor precision and cannot meet the latest data filling requirements, and accurate data is a key step for the maturity of the electricity market. SUMMARY

[0003] The present application provides an industrial user missing data filling method based on SA-GAN to further mine data features and rules after imaging the industrial user daily load data, and uses SA-GAN which has achieved good results in the field of image filling to realize higher-precision industrial user missing data filling, assist in accurate prediction of user load, and promote the maturity of the electricity market, aiming to solve the problem of poor precision of traditional filling methods.

[0004] To achieve the above purpose, the present application adopts the following technical scheme:

[0005] An industrial user missing data filling method based on SA-GAN comprises the following steps:

[0006] S1, an industry correlation degree of industrial users is constructed based on a grey correlation degree algorithm considering positive and negative correlation degrees, for multiple industrial users with a correlation degree higher than a threshold value, daily load data of each industrial user at the same time sequence is used to form a data set A, and for a single industrial user with a correlation degree not higher than the threshold value, historical daily load data is used to form a data set B;

[0007] S2, abnormal data in the data set A and the data set B are determined by using the quartile method, and the value of the abnormal data is 0;

[0008] S3, the daily load data in the data set A and the data set B with the value 0 is converted into RGB color pixel values, so that the data set A and the data set B are imaged to obtain images A and B;

[0009] S4, the images A and B are input into the SA-GAN model, and the images A and B are filled to obtain images E and F filled by the SA-GAN;

[0010] S5, the images E and F are subjected to an inverse imaging process to be converted into daily load data again.

[0011] To optimize the above technical solution, the specific measures also include:

[0012] Furthermore, S1 specifically refers to:

[0013] S1.1, Combine the user's typical daily load sequence E with M typical daily load sequences P1, P2, ..., P from different industries. M Construct the sequence matrix O:

[0014]

[0015] In the formula, e t (t=1,2,...,T) represents the load value at the t-th time node in the typical daily load sequence E of the user, p t m (m = 1, 2, ..., M) represents the daily load sequence P of the m-th industry. m The load value at the t-th time node;

[0016] S1.2 Construct a typical daily load sequence E for users and a typical daily load sequence P for various industries. m The difference matrix R between (m = 1, 2, ..., M) is:

[0017]

[0018] Element s in the matrix t m This represents the t-th element e in the typical daily load sequence E of a user. t With the m-th industry daily load sequence P m The corresponding element p t m The absolute value of the difference between them;

[0019] S1.3 Calculate the minimum difference U between the two levels. min The maximum difference between the two levels U max ;

[0020]

[0021] S1.4 Calculate the correlation coefficient between the typical daily load sequence E of users and the sequences of various industries. Then, the correlation value δ between users and various industries is calculated. m ;

[0022]

[0023] In the formula, β represents the correlation coefficient between the t-th element in the typical daily load sequence E of users and the m-th daily load sequence of industries, β is the resolution, and δ is the correlation coefficient between the t-th element and the m-th daily load sequence of industries. mLet E be the correlation value between the daily load sequence of the m-th industry and the typical daily load sequence of users.

[0024] S1.5. Using the association sign function sgn(·) to determine the positive or negative correlation degree, the gray correlation degree considering positive and negative correlation degrees is calculated as follows:

[0025]

[0026] In the formula, This represents the correlation coefficient between the t-th element in the typical daily load sequence of users and the corresponding element in the m-th daily load sequence of industries, after considering negative correlation. This represents the correlation between the typical daily load sequence of users and the daily load sequence of the m-th industry, after considering negative correlation; Δe t It is the increment of the load value at time node t+1 relative to time node t in the typical daily load sequence E of a user; It is the increment of the load value at time node t+1 relative to time node t in the daily load sequence of the m-th industry;

[0027] S1.6, Based on the relevance of each power user For multiple industrial users with a correlation higher than the threshold, the daily load data of each industrial user in the same time series is used to form dataset A.

[0028]

[0029] In the formula, r represents the daily load data value of the user, C represents the length of the time series, and N is the total number of industrial users in the same time series;

[0030] For a single industrial user with a correlation degree not exceeding the threshold, historical daily load data is used to construct dataset B;

[0031]

[0032] q represents the user's daily load data value, C represents the time series length, and N is the total number of historical daily load curves for a single industrial user.

[0033] Furthermore, S2 specifically refers to:

[0034] S2.1. The quartile method is used to determine the median in dataset A and dataset B, determine the upper and lower limit data, and set the criteria for judging outlier data as exceeding the sum of 1.5 times the median and the upper limit data, and falling below the difference of 1.5 times the median and the lower limit data.

[0035] S2.2 Draw box plots for dataset A and dataset B respectively;

[0036] S2.3 Identify data in the box plot that are below 1.5 times the difference between the median and the lower limit data, or data that exceed 1.5 times the sum of the median and the upper limit data, and determine them as abnormal data;

[0037] S2.4. Treat the outlier data filtered from datasets A and B as missing data, and set the value of the missing data position to 0.

[0038] Furthermore, S3 specifically refers to:

[0039] S3.1 Use the x-axis to represent the sampling time point of each day from Monday to Sunday, the y-axis to represent the week number from the first week to the last week, and the pixel color to represent the normalized load value. Store all the basic parameters of the load data in the first row of the image. The basic parameters include the start time, end time, sampling period, image dimensions and color type.

[0040] S3.2 Statistical calculations are performed to obtain basic parameters, and all parameters are converted to integers. Each parameter occupies two pixels and is filled in the first row of the image. The remaining pixels in the first row are set to white, and R, G, and B are all set to 255.

[0041] S3.3, The load value of the user at each time point r∈[0,2] 24 The RGB color pixels of [-1] are obtained according to the following formulas:

[0042] R = (r / 256) / 256

[0043] G = (r / 256)%256

[0044] B = r%256

[0045] In the formula, / is the division operation for finding the quotient; % is the modulo operation for finding the remainder; R is the red channel, G is the green channel, and B is the blue channel;

[0046] S3.4 Load the load values ​​pixel by pixel starting from the second row of the image in chronological order. After loading is complete, fill the remaining spaces in the last row with black to obtain power industry user load data image A and image B.

[0047] Furthermore, S4 specifically refers to:

[0048] S4.1 Input the power industry user load data images A and B into the SA-GAN model. Both images A and B are N-dimensional with C channels.

[0049] S4.2 Perform graph convolution operation on image A and image B to obtain the convolution feature map R of image A and the convolution feature map Q of image B;

[0050] S4.3. Convert the convolutional feature map R into two feature spaces f and g; perform a 1×1 vector convolution operation to obtain the weight matrix. The convolutional feature map R is then transformed into a feature space h, and a 1×1 vector convolution operation is performed to obtain the weight matrix. in, The number of channels in the weight matrix. Where k is an integer;

[0051] S4.4. Transform the convolutional feature map Q into two feature spaces a and b; perform a 1×1 vector convolution operation to obtain the weight matrix. The convolutional feature map Q is then transformed into a feature space v, and a 1×1 vector convolution operation is performed to obtain the weight matrix.

[0052] S4.5, the weight matrix W g Transpose the matrix and use it with the weight matrix W. f Perform matrix multiplication and normalized exponentiation to obtain the weight matrix W. g With weight matrix W f Attention map S;

[0053] S4.6, Weight matrix W a Transpose the matrix and use it with the weight matrix W. b Perform matrix multiplication and normalized exponentiation to obtain the weight matrix W. a With weight matrix W b Attention map T;

[0054] S4.7. Connect the attention map S with the weight matrix W h Perform matrix multiplication, then convert the result of the multiplication into a feature space u, and use 1×1 vector convolution to obtain the self-attention feature map D;

[0055] S4.8. Connect the attention map T with the weight matrix W v Perform matrix multiplication, then convert the result of the multiplication into a feature space u, and use 1×1 vector convolution to obtain the self-attention feature map U;

[0056] S4.9 Multiply the self-attention feature map D by the learnable scalar θ, and then sum it with the convolutional feature map R to obtain the output feature map Y; multiply the self-attention feature map U by the learnable scalar θ, and then sum it with the convolutional feature map Q to obtain the output feature map Z;

[0057] S4.10. Perform deconvolution on the output feature map Y and the output feature map Z to obtain the images E and F after SA-GAN padding.

[0058] Furthermore, S5 specifically refers to:

[0059] Obtain the basic parameters stored in the first row of images E and F. The basic parameters include the start time, end time, sampling period, image dimensions, and color type of the load data. Starting from the second row, obtain the RGB color pixel values ​​for each pixel and convert them into load values. Calculate the pixel time tag based on the start time, end time, sampling period, and pixel position in the image. Output the load data when the time tag equals the end time.

[0060] The beneficial effects of this invention are:

[0061] The data imputation method proposed in this invention achieves higher accuracy in imputing missing data for industrial users. The grey relational algorithm, which considers positive and negative correlation, can quantitatively judge the similarity, overlap, and closeness of the connections between different load curves generated by various power users, thereby calculating the correlation value between them; thus helping to accurately predict user load and promoting the maturity of the power market. Attached Figure Description

[0062] Figure 1 Flowchart of a method for imputing missing data for industrial users based on SA-GAN;

[0063] Figure 2 A schematic diagram illustrating the principle of SA-GAN processing of image A. Detailed Implementation

[0064] The invention will now be described in further detail with reference to the accompanying drawings.

[0065] In one embodiment, the present invention proposes a method for imputing missing industrial user data based on SA-GAN, the flowchart of which is shown below. Figure 1 As shown, the specific process includes the following:

[0066] Constructing the industry correlation of industrial power users: Calculating the industry power user correlation value based on the grey relational degree algorithm that considers positive and negative correlation;

[0067] The industry characteristics of electricity users, to a certain extent, determine their electricity consumption patterns, thus having a significant impact on their electricity consumption characteristic curves. Furthermore, the supply and demand relationships among companies within the same industry also affect the load characteristics of different users. Therefore, we employ an improved grey relational analysis algorithm, which theoretically can quantify factors such as the similarity, overlap, and closeness of relationships between different load curves generated by various electricity users, thereby calculating their correlation values.

[0068] The typical daily load sequence E of the user is compared with the typical daily load sequences P1, P2, ..., P of M different industries. M Construct the sequence matrix O:

[0069]

[0070] In the formula, e t (t=1,2,...,T) represents the load value at the t-th time node in the typical daily load sequence E of the user, p t m (m = 1, 2, ..., M) represents the daily load sequence P of the m-th industry. m The load value at the t-th time node.

[0071] Construct a typical daily load sequence E for users and typical daily load sequences P for various industries. m The difference matrix R between (m = 1, 2, ..., M) is:

[0072]

[0073] Element s in the matrix t m This represents the t-th element e in the typical daily load sequence E of a user. t With the m-th industry daily load sequence P m The corresponding element p t m The absolute value of the difference between them.

[0074] Calculate the minimum difference U between the two levels min The maximum difference between the two levels U max

[0075]

[0076] Correlation coefficients between typical daily load series E of users and various industry series Then, the correlation value δ between users and various industries is calculated. m .

[0077]

[0078] In this formula, The correlation coefficient between the t-th element and the m-th industry daily load sequence in the typical user daily load sequence E, and β, a resolution that effectively suppresses distortion caused by excessively large ranges, are used; therefore, β is set to 0.5. Furthermore, δ... m δ represents the correlation between the daily load sequence of the m-th industry and the typical daily load sequence E of users. m The larger the value, the stronger the correlation.

[0079] Traditional grey relational analysis algorithms, due to the absolute value of the difference matrix, always result in a relational degree calculation greater than 0, contradicting the actual existence of negative relational degrees in the load sequences. Therefore, by introducing the relational sign function sgn(·) to determine the positive or negative relational degree, the improved grey relational degree calculation considering positive and negative relational degrees is shown in the following formula:

[0080]

[0081] In the formula, This represents the correlation coefficient between the t-th element in the typical daily load sequence of users and the corresponding element in the m-th daily load sequence of industries, after considering negative correlation. This represents the correlation between the typical daily load sequence of users and the daily load sequence of the m-th industry, after considering negative correlation; Δe t It is the increment of the load value at time node t+1 relative to time node t in the typical daily load sequence E of a user; It is the increment of the load value at time node t+1 relative to time node t in the m-th industry daily load sequence.

[0082] Based on the relevance of each electricity user For multiple industrial users with a correlation higher than the threshold, the daily load data of each industrial user in the same time series is used to form dataset A.

[0083]

[0084] In the formula, r represents the daily load data value of the user, C represents the length of the time series, and N is the total number of industrial users in the same time series;

[0085] For a single industrial user with a correlation degree not exceeding the threshold, historical daily load data is used to construct dataset B;

[0086]

[0087] q represents the user's daily load data value, C represents the time series length, and N is the total number of historical daily load curves for a single industrial user.

[0088] Use the quartile method to identify outliers in datasets A and B, and define outliers as missing data.

[0089] Power industry users use daily load curves as the basic data. The quartile method is used to determine the median of the daily load curve data in datasets A and B, and to determine the upper and lower limit data. The criteria for judging abnormal data are set as exceeding the sum of 1.5 times the median and the upper limit data, and falling below the difference of 1.5 times the median and the lower limit data.

[0090] Draw box plots for dataset A and dataset B respectively.

[0091] In a box plot, identify data points that are below 1.5 times the difference between the median and the lower limit, or data points that exceed 1.5 times the sum of the median and the upper limit, and classify them as outliers.

[0092] The outlier data selected from datasets A and B are marked as missing data.

[0093] The resulting matrices A and B are the load datasets after outlier identification, where both outliers and missing values ​​have been labeled.

[0094] Set the value at the location of the missing data to 0.

[0095] During the forward conversion of the load data visualization, the x-axis represents the sampling time point for each day from Monday to Sunday, the y-axis represents the week number from the first week to the last week, and the pixel color represents the normalized load value. All basic parameters of the load data are stored in the first row of the image, used for subsequent reverse conversion to parse pixel colors and restore the load data. Basic parameters include the start time, end time, sampling period, image dimensions (width and height), and color type. The specific steps for converting load data to RGB pixels are as follows.

[0096] The datasets A and B, which contain 0 values, are directly visualized.

[0097] Basic parameters are obtained through statistical calculations and converted to integers. Each parameter occupies two pixels and is filled in the first row of the image. The remaining pixels in the first row are set to white (R, G, and B are all set to 255). The data types of the basic parameters are mainly date / time, integers, and real numbers. The date uses standard UNIX epoch time, converting the time data to seconds (s) and then to an integer.

[0098] For the user's load value at each time point r∈[0,2] 24 The RGB color pixels of [-1] can be obtained according to the formula.

[0099] R = (r / 256) / 256

[0100] G = (r / 256)%256

[0101] B = r%256

[0102] In the formula: / is the division operation for finding the quotient; % is the modulo operation for finding the remainder; R is the red channel, G is the green channel, and B is the blue channel.

[0103] Load information is loaded pixel by pixel, starting from the second row of the image, in chronological order. After loading is complete, the remaining spaces in the last row are filled with black, thus completing the visualization of the load data.

[0104] Input the power industry user load data images A and B into the SA-GAN model. Both images A and B are N-dimensional with C channels.

[0105] Perform graph convolution operations on images A and B to obtain the convolution feature map R of image A and the convolution feature map Q of image B.

[0106] Feature extraction learning is performed on convolutional feature maps A and B respectively. First, the convolutional feature map R is transformed into two feature spaces f and g; then, a 1×1 vector convolution operation is performed to obtain the weight matrix. The convolutional feature map R is then transformed into a feature space h, and a 1×1 vector convolution operation is performed to obtain the weight matrix. in, The number of channels in the weight matrix. Where k is an integer, k = 1, 2, 4, 8.

[0107] The convolutional feature map Q is transformed into two feature spaces a and b; a 1×1 vector convolution operation is performed to obtain the weight matrix. The convolutional feature map Q is then transformed into a feature space v, and a 1×1 vector convolution operation is performed to obtain the weight matrix.

[0108] The weight matrix W g Transpose the matrix and use it with the weight matrix W. f Perform matrix multiplication and normalized exponentiation to obtain the weight matrix W. g With weight matrix W f Attention map S.

[0109] The weight matrix W a Transpose the matrix and use it with the weight matrix W. b Perform matrix multiplication and normalized exponentiation to obtain the weight matrix W. a With weight matrix W b Attention map T.

[0110] Connect the attention map S with the weight matrix W h Perform matrix multiplication, then convert the result into a feature space u, and use 1×1 vector convolution to obtain the self-attention feature map D.

[0111] Connect the attention map T with the weight matrix W v Perform matrix multiplication, then convert the result into a feature space u, and use 1×1 vector convolution to obtain the self-attention feature map U.

[0112] The self-attention feature map D is multiplied by a learnable scalar θ and then summed with the convolutional feature map R to obtain the output feature map Y; the self-attention feature map U is multiplied by a learnable scalar θ and then summed with the convolutional feature map Q to obtain the output feature map Z, where θ is a learnable scalar initialized to 0. Introducing the learnable θ allows the SA-GAN network to initially rely on cues from local neighborhoods, and then gradually learn to assign more weights to non-local cues.

[0113] Deconvolve the output feature map Y and the output feature map Z to obtain the images E and F after SA-GAN padding.

[0114] The imbalanced learning rate of the generator and discriminator updates that occur during image imputation is addressed by introducing TTUR.

[0115] To address the problem of unstable training in image inpainting, spectral normalization of the generator and discriminator is introduced.

[0116] Deconvolution is performed on the output feature map Y corresponding to image A and the output feature map Z corresponding to image B to obtain the SA-GAN repaired images E and F.

[0117] By reversing the process of visualizing industrial user load data, the E and F images are reconstructed into load data, thus completing the missing data filling for datasets A and B.

[0118] Retrieve the basic parameters stored in the first row of the RGB format images E and F of the load data. These basic parameters include the start and end times of the load data, the sampling period, the image's dimensions (width and height), and the color type.

[0119] Starting from the second row, retrieve the parameters from the image one by one.

[0120] The time label is calculated based on the start time, end time, sampling period, and pixel position in the image. Load data is output when the time label equals the end time.

[0121] The completed power industry user datasets A and B are obtained.

[0122] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for imputing missing industrial user data based on SA-GAN, characterized in that, Includes the following steps: S1. Construct the industry correlation degree of industrial users based on the grey relational degree algorithm that considers positive and negative correlation. For multiple industrial users with correlation degree higher than the threshold, use the daily load data of each industrial user in the same time series to form dataset A. For a single industrial user with correlation degree not higher than the threshold, use historical daily load data to form dataset B. S2. Use the quartile method to identify outliers in datasets A and B, and set the value of outliers to 0. S3. Convert the daily load data in datasets A and B, which contain 0 values, into RGB color pixel values ​​to visualize datasets A and B, resulting in image A and image B. S4. Input images A and B into the SA-GAN model to fill in images A and B, obtaining SA-GAN-filled images E and F; S4 specifically involves: S4.1 Input the power industry user load data images A and B into the SA-GAN model. Both images A and B are... N dimension C aisle; S4.2 Perform graph convolution operation on images A and B to obtain the convolution feature map of image A. R Convolutional feature map of image B Q ; S4.3, Convolution feature map R Transformed into two feature spaces f and g Perform a 1×1 vector convolution operation to obtain the weight matrix. , Then convolution feature maps R Transform into feature space h Perform a 1×1 vector convolution operation to obtain the weight matrix. ,in, The number of channels in the weight matrix. ,in k It is an integer; S4.4, Convolution feature map Q Transformed into two feature spaces a and b Perform a 1×1 vector convolution operation to obtain the weight matrix. , Then convolution feature maps Q Transform into feature space v Perform a 1×1 vector convolution operation to obtain the weight matrix. ; S4.5, Weight matrix Transpose and then use the weight matrix. Perform matrix multiplication and normalized exponentiation to obtain the weight matrix. With weight matrix attention map S ; S4.6, Weight matrix Transpose and then use the weight matrix. Perform matrix multiplication and normalized exponentiation to obtain the weight matrix. With weight matrix attention map T ; S4.7, Connect the attention map S with the weight matrix Perform matrix multiplication, and then transform the result of the multiplication into a feature space. u The self-attention feature map is obtained by using 1×1 vector convolution. D ; S4.8, Attention Map T With weight matrix Perform matrix multiplication, and then transform the result of the multiplication into a feature space. u The self-attention feature map is obtained by using 1×1 vector convolution. U ; S4.9 Self-Attention Feature Map D Multiply by the learnable scalar Then convolutional feature map R Summing yields the output feature map Y; self-attention feature map U Multiply by the learnable scalar Then convolutional feature map Q Summing yields the output feature map Z; S4.

10. Perform deconvolution on the output feature map Y and the output feature map Z to obtain the image E and the image F after SA-GAN padding; S5. Perform the inverse image conversion process on images E and F to reconstruct daily load data.

2. The method for imputing missing industrial user data based on SA-GAN as described in claim 1, characterized in that, S1 specifically refers to: S1.1, The typical daily load sequence of users E and M Typical daily load sequences of different industries P 1, P 2, ..., P M Construct a sequence matrix O : In the formula, e t (t=1,2,..., T (This refers to a typical daily load sequence for users) E The Middle t Load values ​​at each time point, p t m ( m =1,2,..., M ) is the first m Daily load sequence of individual industries P m The Middle t Load values ​​at each time point; S1.2 Constructing a typical daily load sequence for users E Typical daily load sequences of various industries P m (m= 1 , 2 ,...,M) The difference matrix between R : Elements in the matrix s t m Represents a typical daily load sequence for users E The Middle t element e t With the m Daily load sequence of individual industries P m Corresponding element p t m The absolute value of the difference between them; S1.3 Calculate the minimum difference U between the two levels. min The maximum difference between the two levels U max ; S1.4 Calculate the typical daily load sequence for users. E Correlation coefficients between various industry series This allows for the calculation of the correlation scores between users and various industries. ; In the formula, Representative user's typical daily load sequence E The Middle t The element and the first m Correlation coefficients between daily load series of individual industries It's about resolution. For the first m Daily load sequences for various industries and typical daily load sequences for users E The correlation value between them; S1.

5. Using the association sign function sgn(·) to determine the positive or negative correlation degree, the gray correlation degree considering positive and negative correlation degrees is calculated as follows: In the formula, This represents the first day of the typical daily load sequence of users after considering negative correlation. t The element and the first m The correlation coefficients of elements corresponding to daily load sequences of each industry; This indicates the relationship between the typical daily load sequence of users after considering negative correlation and the first... m The correlation value of the daily load series of each industry; This is a typical daily load sequence for users. E The Middle t +1 time node relative to the first t The increment of load value at each time node; It is the first m The first of the daily load sequences of the industry t+ 1. Time node relative to the first t The increment of load value at each time node; S1.6, Based on the relevance of each power user For multiple industrial users with a correlation higher than the threshold, the daily load data of each industrial user in the same time series is used to form dataset A. In the formula, r This represents the user's daily load data value. C Indicates the length of the time series. N This represents the total number of industrial users within the same time series. For a single industrial user with a correlation degree not exceeding the threshold, historical daily load data is used to construct dataset B; q This represents the user's daily load data value. C Indicates the length of the time series. N This represents the total number of historical daily load curves for a single industrial user.

3. The method for imputing missing industrial user data based on SA-GAN as described in claim 1, characterized in that, S2 specifically refers to: S2.

1. The quartile method is used to determine the median in dataset A and dataset B, determine the upper and lower limit data, and set the criteria for judging outlier data as exceeding the sum of 1.5 times the median and the upper limit data, and falling below the difference of 1.5 times the median and the lower limit data. S2.2 Draw box plots for dataset A and dataset B respectively; S2.3 Identify data in the box plot that are below 1.5 times the difference between the median and the lower limit data, or data that exceed 1.5 times the sum of the median and the upper limit data, and determine them as abnormal data; S2.

4. Treat the outlier data filtered from datasets A and B as missing data, and set the value of the missing data position to 0.

4. The method for imputing missing industrial user data based on SA-GAN as described in claim 1, characterized in that, S3 specifically refers to: S3.1, Use x The axis represents the sampling time point for each day from Monday to Sunday. y The axis represents the week number from the first week to the last week, and the pixel color represents the normalized load value. All the basic parameters of the load data are stored in the first row of the image. The basic parameters include the start time, end time, sampling period, image dimensions and color type. S3.2 Statistical calculations are performed to obtain basic parameters, and all parameters are converted to integers. Each parameter occupies two pixels and is filled in the first row of the image. The remaining pixels in the first row are set to white, and R, G, and B are all set to 255. S3.3, User load value at each time point The RGB color pixels are obtained according to the following formulas: In the formula, / represents the division operation to find the quotient; % is the modulo operation for finding the remainder; R is the red channel, G is the green channel, and B is the blue channel; S3.4 Load the load values ​​pixel by pixel starting from the second row of the image in chronological order. After loading is complete, fill the remaining spaces in the last row with black to obtain power industry user load data image A and image B.

5. The method for imputing missing industrial user data based on SA-GAN as described in claim 1, characterized in that, S5 specifically refers to: Obtain the basic parameters stored in the first row of images E and F. The basic parameters include the start time, end time, sampling period, image dimensions, and color type of the load data. Starting from the second row, obtain the RGB color pixel values ​​for each pixel and convert them into load values. Calculate the pixel time tag based on the start time, end time, sampling period, and pixel position in the image. Output the load data when the time tag equals the end time.

Citation Information

Patent Citations

  • Data filling method and device based on landmarks

    CN111177135A

  • Face image eye completion method based on self-attention mechanism model generative adversarial network

    CN111738940A