A power consumption prediction method based on user clustering to expand data

The electricity consumption time series is processed through EEMD, PCA and K-Means clustering methods, the IMF sequence of similar users is expanded, and the electricity consumption prediction is predicted using the CNN-LSTM model, which solves the problems of data correlation and time difference in the prior art, and improves prediction accuracy and efficiency.

CN114819392BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210548819.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-27
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

In the short-term prediction of electricity loads in the prior art, data correlation and time difference have a great impact, resulting in insufficient model training capabilities and high prediction randomness, and lack of characteristics of relationships between users using electricity.

Method used

The electricity consumption time series is decomposed by empirical modal decomposition method (EEMD), and combined with principal component analysis (PCA) and K-Means clustering methods, dimensionality reduction and clustering of user data are performed. Then, the IMF sequence of similar users is expanded according to the user clustering results, and the power consumption prediction is performed as input to the convolutional neural network (CNN) fused long-term short-term memory artificial neural network (LSTM) model.

Benefits of technology

It improves the accuracy and efficiency of the electricity use prediction model, makes full use of the behavioral similarity between user electricity use data, enhances the integrity of the data and the prediction accuracy of the model, and saves program running time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114819392B_ABST
    Figure CN114819392B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting electricity consumption based on user clustering extended data, which belongs to the field of power data analysis and processing; the method comprises decomposing the user's electricity consumption time series characteristics by using EEMD to obtain an IMF sequence; reducing the dimension of the decomposition result by using the PCA K-Means clustering method to obtain the user IMF clustering result; using the PCA K-Means clustering method again on the IMF clustering result to obtain the user clustering result; according to the user clustering result, expanding the IMF sequence belonging to the same type of users; inputting the extended data into a convolutional neural network fused with a long short-term memory artificial neural network model to train the network model; inputting the user's IMF sequence to be tested and the electricity consumption time series of the same type of users into the trained network model to obtain the user's electricity consumption prediction result. The present invention uses the relationship between users to expand the data, which improves the accuracy and stability of the prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power data analysis and processing, and relates to a power consumption prediction method based on user clustering to expand data. Background Art

[0002] The accuracy of short-term power consumption prediction determines both the financial performance of wholesale power and capacity market participants and the reliability of power system operation. With the development of current computer technology, more artificial neural networks (ANNs) are introduced to improve the accuracy of charge prediction. The academic community generally believes that using neural networks to solve prediction problems is effective, which naturally includes the problem of predicting power load time series.

[0003] Since deep learning models can capture hidden and non-linear features in data, many types of neural networks have been used for power load prediction. For example, feedforward neural networks, radial basis function networks, spiking neural networks, and recurrent neural networks, etc. Among them, the most popular learning method is the backpropagation algorithm. However, compared with classical methods, the application research of such methods in short-term load prediction problems is relatively less, and most of the work only gives a general description of the model framework and its results, which makes their implementation quite difficult.

[0004] In addition, in order to improve the accuracy of the charge prediction model, some studies consider the approximation and decomposition of time series. The decomposition of time series is the extraction of trends, periodicity, and random fluctuations. Clustering methods are usually required in this process to identify the time differences of time series. The K-means algorithm is a very commonly used clustering algorithm, while DTW (Dynamic time warping) is a very commonly used method for measuring the similarity of time series. The fusion of DTW and K-means is a classic time series clustering method. Although this fusion method can successfully identify the time differences between time series and bring good classification results, its time overhead is large and the complexity is high. The method of fusing PCA and K-means is fast in calculation speed, but the clustering effect is not as good as that of the method of fusing DTW and K-means. If the efficiency can be improved while avoiding the influence of time differences on sample data based on this method of fusing PCA and K-means, it will be more beneficial to the clustering effect.

[0005] In addition, the considered time series components do not necessarily exist in every time series. There may be no components with trends or periodic fluctuations in these time series, or both. Some literature considers the time series itself, but does not consider other data characteristics. In the field of electricity consumption prediction, characteristics such as weather and traffic are generally augmented, and the training process is completed with the augmented data characteristics. However, this method lacks the characteristics of the relationships between electricity users, resulting in insufficient model training ability. Moreover, due to the lack of multiple feature dependencies, the randomness of model prediction is relatively high. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide an electricity consumption prediction method based on user clustering extended data of electricity consumption time series, which includes technologies such as EEMD (Empirical Mode Decomposition), PCA (Principal Component Analysis), CNN (Convolutional Neural Network), and LSTM (Long Short-Term Memory Artificial Neural Network), and can solve the problems of insufficient electrolysis data for neural networks, the correlation between data, separating random and non-random sequences, and the prediction of electricity consumption in communities.

[0007] In traditional technologies, prediction is generally carried out based on the obtained data set, and usually only the EEMD decomposition result is directly input into the prediction model without considering other characteristics. However, the present invention believes that users of the same type have a certain similarity in behavior and can complement each other's missing data. Therefore, the present invention considers using data of users of the same type as the input of the prediction model. The present invention mainly uses K-Means clustering twice on the EEMD decomposition result. The first time is used to obtain the clustering result of the IMF sequence of users, and the second time is used to obtain the clustering result of the users themselves. According to the second user clustering result, the IMF sequence after the first clustering is extended, and the extended result is used as the input of the CNN-LSTM model to train the model, thereby improving the prediction effect of the model; at the same time, the present invention also creates its own data frame for each type of user, so that user data can be extended in the format of the data frame. Whether in the training stage or the testing stage, only one EEMD decomposition process is required for the data, which can save the corresponding program running time.

[0008] To achieve the above purpose, the present invention provides the following technical solutions:

[0009] An electricity consumption prediction method based on user clustering extended data, the method includes the following steps:

[0010] S1: Obtain a user data set, where the user data set includes user electricity consumption data for each period; and preprocess the electricity consumption data of the user to obtain a user electricity consumption time series;

[0011] S2: Decomposing the user's electricity consumption time series using the empirical mode decomposition method to obtain the user's IMF series;

[0012] S3: Use the K-Means clustering method based on principal component analysis to reduce the dimension of the user's IMF sequence, obtain the clustering result of the user's IMF sequence, and use it as the updated user data set;

[0013] S4: using a K-Means clustering method based on principal component analysis to cluster the user IMF sequence to obtain a user clustering result;

[0014] S5: according to the user clustering result, the IMF sequence of users of the same type is expanded, and the electricity consumption time series of users of the same type are added to the IMF sequence of the users;

[0015] S6: input the expanded IMF sequence as the updated user data set into the convolutional neural network fused with long short-term memory artificial neural network model to train the network model;

[0016] S7: Input the IMF sequence of the user in the current period and the electricity consumption time series of similar users in the current period into the trained network model to obtain the electricity consumption prediction result of the user in the next period.

[0017] The beneficial effects of the present invention are:

[0018] 1) The present invention makes full use of user electricity consumption data, obtains the predicted value of the current user's electricity consumption data to be tested based on the historical electricity consumption data of similar users, and provides an auxiliary basis for the electricity consumption prediction of community users;

[0019] 2) According to the characteristics of similarity of electricity consumption behavior of electricity users, the present invention uses the electricity consumption data of other users in the same category to supplement the electricity consumption data of the current user, which is not only conducive to the training of the network model, but also conducive to enhancing the integrity of the electricity consumption data to be tested, thereby improving the prediction accuracy of the network model;

[0020] 3) The present invention makes full use of the EEMD method, provides optimization for the PCA-K-means clustering method, and also uses the decomposition results in the prediction model to obtain better time series clustering results. In addition, the one-time EEMD preprocessing also saves program running time.

[0021] The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0023] Figure 1 It is a flowchart of an electricity consumption prediction method based on user clustering to expand data in an embodiment of the present invention;

[0024] Figure 2 It is a schematic diagram of the update process of the user dataset in an embodiment of the present invention;

[0025] Figure 3 It is a schematic diagram of the data frame structure of similar users in an embodiment of the present invention;

[0026] Figure 4 It is a schematic diagram of the electricity consumption prediction process in an embodiment of the present invention;

[0027] Figure 5 It is a network model structure diagram in an embodiment of the present invention. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0029] Figure 1 It is a flowchart of an electricity consumption prediction method based on user clustering to expand data in an embodiment of the present invention, as Figure 1 shown, the prediction method includes:

[0030] S1: Obtain a user dataset, where the user dataset includes user electricity consumption data for each period; and preprocess the electricity consumption data of the user to obtain a user electricity consumption time series;

[0031] In an embodiment of the present invention, first, a user dataset needs to be obtained. The user dataset includes the historical power consumption data of the user and the current user data. Among them, the historical power consumption data is the power consumption data of the user in a historical time period. Here, the historical time period refers to a time period that has passed, and part or all of the corresponding power consumption data is known. For example, assume there are N users in total. In the historical time period from t - m to t + m, each user should have a load value at time stamps t - m, t - m + 1... t + m. Of course, this load value may be missing. If it is missing, it can be filled in by corresponding methods. The current user data is the power consumption data of the user in the current time period. Here, the current time period refers to the current real-time time period. The present invention needs to use the IMF sequence of the current user A in the current time period T combined with the power consumption time series of the same type of users {N(A)} in the current time period T to predict the power consumption data of the current user A in the next time period T + 1.

[0032] In an embodiment of the present invention, after preprocessing the obtained user dataset, the corresponding user power consumption time series can be obtained, which can be expressed as:

[0033]

[0034]

[0035] Among them, is the power consumption data of the Nth community user in the t time period, is the power consumption data of the Nth community user at the time stamp t + m.

[0036] Among them, in an embodiment of the present invention, the power consumption data of N users in the current area is read. If it is found that the load value at the current moment is missing, that is, assume that in the user data missing means then let If there are multiple consecutive missing values, they will be directly filled with the average value of the user. Among them, represents the load value of the nth user at the t moment, t - 1 represents the moment before the t moment, t + 1 represents the moment after the t moment, where t ∈ (0, L), and L is the length of the user power consumption time series.

[0037] In a preferred embodiment of the present invention, considering that if only the historical power consumption data of the user is used to predict the current power consumption data, then this data lacks interpretability, that is, the current influencing factors are not considered, and the natural law of the user's power consumption data at the objective level cannot be reflected. This prediction result only has mathematical significance and no substantial significance. Therefore, when collecting the user's historical data, the present invention also takes into account other factors that affect the user's power consumption changes, including temperature factors and holiday factors. Therefore, when collecting the user's power consumption data, in addition to collecting the user's basic information and load, it is also necessary to collect the temperature data and holiday data corresponding to the user's power consumption data. For example, the power consumption load of user A at time t is a, the temperature is b, and the time is c. The temperature data and time data can reflect the changes in the user's power consumption from the side, so as to improve the accuracy of the prediction result of the user's power consumption data.

[0038] S2: Decompose the user power consumption time series by using the empirical mode decomposition method to obtain the user's IMF sequence;

[0039] In step S2, it may specifically include:

[0040] S21: Set the number of IMF components according to the accuracy required by the network model and the effectiveness of the decomposition by the empirical mode decomposition method to determine the dimension of the user's IMF data set;

[0041] S22: Use the empirical mode decomposition method to repeatedly add white noise to each user power consumption time series, and take the average value of the calculated IMF components as the final result.

[0042] In the embodiment of the present invention, considering that for the preliminary user data set, it is necessary to set the number of IMF components according to the accuracy required by the model and the effectiveness of the EEMD decomposition method to determine the dimension of the user data set. Before preparing to perform the EEMD decomposition method, it is also necessary to standardize each time series:

[0043]

[0044] where y i represents the i-th element after standardization, x is the input time series, and x i is the i-th element corresponding to the input time series.

[0045] EEMD extracts a finite number of EM (empirical mode) IMFs i (t) from the original data x(t) and obtains the r I (t) remainder:

[0046]

[0047] The EEMD method is based on repeatedly adding white noise with an infinitesimal amplitude to the signal and calculating the average of the obtained intrinsic mode functions as the final result:

[0048]

[0049]

[0050] where, IMF ji (t) is the i-th intrinsic mode function in the j-th decomposition period, and r jI (t) is the EM and the remainder obtained through different decompositions; j = 1, 2,..., J is the number of decomposition periods (added to the white noise signal).

[0051] S3: Dimension reduction is performed on the user's IMF sequence using the K-Means clustering method based on principal component analysis to obtain the clustering result of the user's IMF sequence, and it is used as the updated user dataset;

[0052] In step S3, it may specifically include:

[0053] S31: Perform principal component analysis (PCA) dimension reduction on each historical IMF sequence in step S22 to obtain a two-dimensional IMF dataset; assuming there are M IMF components, then the two-dimensional IMF dataset can be expressed as:

[0054]

[0055]

[0056] where, represents the first-dimensional data of the M-th IMF component, represents the second-dimensional data of the M-th IMF component.

[0057] In step S31, it may specifically include:

[0058] S311: Define the dimension of the principal component analysis output as two, and obtain the average value of each IMF component obtained in step S22;

[0059] S312: Move the origin to the center of the user dataset, calculate the sample covariance matrix, and find the eigenvectors and eigenvalues of the covariance matrix;

[0060] S313: Select the eigenvector P corresponding to the larger eigenvalue, and perform sample mapping using matrix multiplication, that is, Y = PX, to obtain the two-dimensional dataset after principal component analysis dimension reduction.

[0061] S32: Apply the K-means clustering method to the two-dimensional IMF dataset obtained in step S31 to obtain the clustering results of the IMF sequences, that is, the clustering results of a certain level of IMF components corresponding to each user;

[0062] Among them, applying the K-means clustering method to the two-dimensional IMF dataset, that is, the Y dataset, to obtain the clustering results of a certain level of IMF components corresponding to each user can be expressed as That is the clustering result of the M-th level IMF of the N-th user.

[0063] S33: Use the clustering results of the IMF sequences to replace the user power consumption time series to form an updated user dataset.

[0064] According to the above process, the updated user dataset can be changed to:

[0065]

[0066] Among them, represents the dataset of the N-th user. Here, the user dataset is no longer reflected by the time series, but by the clustering results of the IMF sequences Reflected as:

[0067]

[0068] Among them, is the M-th IMF class number of the users in the N-th cell, manifested as the clustering results of the IMF components.

[0069] S4: Apply the K-Means clustering method based on principal component analysis to the clustering results of the user IMF sequences to obtain the user clustering results;

[0070] In step S4, it may specifically include:

[0071] S41: Perform principal component analysis dimensionality reduction on the clustering results of the IMF sequences again;

[0072] S42: Apply the K-means clustering method to the clustering results of the dimensionality-reduced IMF sequences to obtain the user clustering results based on the time series.

[0073] Similarly, step S41 can adopt similar steps as steps S311 - S313 to obtain the user clustering results based on the time series.

[0074] S5: According to the user clustering results, expand the IMF sequences of users belonging to the same class, and add the power consumption time series of users in the same class to the user IMF sequences;

[0075] In step S5, it specifically includes:

[0076] S51: Divide the user data according to the user clustering results, that is, create its own data frame for each class of users. The column label of the data frame is the user number, and the data column is the corresponding user electricity consumption data;

[0077] S52: Expand the user electricity consumption data according to the user class. That is, if user n is in class m, then the data set of user n in class m includes: the data of user n in class m and the data of other users belonging to class m;

[0078] S53: Use the expanded data frame obtained in step S52 as the expanded feature vector, and use the IMF data set in step S22 to replace the user's own electricity consumption time series, that is, expand the user's electricity consumption data with the data frame similar to each IMF itself.

[0079] Among them, each expanded IMF sequence contains the user's IMF sequence and the electricity consumption time series of its same-class users. For example, assume that in an electricity consumption network scenario, in one classification, it includes user A, user B, and user C; user A has 4 levels of IMF components. Then the expanded IMF sequence of user A includes: the IMF1 component of user A, the electricity consumption time series of user B, and the electricity consumption time series of user C; the IMF2 component of user A, the electricity consumption time series of user B, and the electricity consumption time series of user C; the IMF3 component of user A, the electricity consumption time series of user B, and the electricity consumption time series of user C; the IMF4 component of user A, the electricity consumption time series of user B, and the electricity consumption time series of user C; This expansion method can effectively enhance the user data set and improve the accuracy of the model.

[0080] S6: Input the expanded IMF sequence as the user data set after being updated again into the convolutional neural network fused with the long short-term memory artificial neural network model to train the network model;

[0081] In the embodiment of the present invention, it is necessary to update the user data set again according to the IMF clustering results and construct a training set based on this:

[0082]

[0083] Among them, is the load value of the Nth user in the same class C as the target user at the t-M time stamp.

[0084] In step S6, it specifically includes: S61: Construct a convolutional neural network fused with a long short-term memory artificial neural network model, where the long short-term memory artificial neural network includes 3 consecutive LSTM layers;

[0085] S62: Input the user data set after being updated again in step S5 into the network model to obtain the corresponding prediction result;

[0086] S63: Compare the corresponding prediction result with the actual power consumption result, and train the network model until the network model meets the accuracy requirement.

[0087] During the training process, the IMF components, corresponding temperature and time of the user at the previous moment, the temperature and time of the user at the next moment, and the user time series and corresponding temperature and time of the same type of users at the previous moment can be used to predict the user time series of the user at the next moment. Compare the predicted user time series with the actual user time series to train the network model so that the network model can meet the prediction requirements. In this process, on the one hand, the network model is needed to learn the relationship between the current user and the power consumption data of other users, and on the other hand, the network model is also needed to learn the relationship between the temperature and time during power consumption and the user's power consumption. Through the combination of these two aspects, the prediction model not only fully considers the characteristics of the relationship between power consumption users, but also considers the temperature factor and time factor that affect load changes, making the result predicted by the model not only interpretable, but also improving the prediction accuracy of the network model.

[0088] S7: Input the user's to-be-tested IMF sequence and the historical IMF sequences of the same type of users into the trained network model to obtain the power consumption prediction result of the user.

[0089] In the embodiment of the present invention, when predicting and processing the to-be-tested power consumption data of the user, only the to-be-tested power consumption data of the user needs to be decomposed by EEMD, and it is not necessary to decompose the power consumption data of the same type of users by EEMD. Input the to-be-tested IMF sequence of the user and the historical IMF sequences of the same type of users into the trained network model, and the network model can output the power consumption prediction result of the to-be-tested power consumption data of the user.

[0090]

[0091] Among them, represents the power consumption prediction result, i.e., the load value, of the target user A at the moment t + 1. is the clustering result of the IMF component of the target user A at the moment t. is the load value of the Nth user of the same type C as the target user A at the moment t.

[0092] Because the to-be-tested IMF sequence contains multiple classifications, in order to implement the prediction of the parallel CNN-LSTM model, it is necessary to synthesize the prediction results, that is, the synthesis formula is:

[0093]

[0094] Among them, yout is the comprehensive prediction result, is the prediction result of the i-th IMF, where M is the number of IMFs; T 1 represents the user's current temperature, T 2 represents the current time; f represents the relationship function between the user, temperature factor, and time factor, which is obtained through the training and iteration process of the network model; by synthesizing different IMF sequences and combining the user's current temperature factor and time factor; performing inverse variance normalization on the data to obtain the final prediction result.

[0095] It can be understood that in the embodiment of the present invention, the prediction method not only considers the trend that can reflect the prediction sequence but also considers the trends of other time series that directly or indirectly affect it. The relationship between time series can be "bidirectional", that is, one time series can be affected by several other time series simultaneously. Based on this relationship, on the one hand, the present invention uses the IMF sequences of similar users to complement the IMF sequences of its own users. When training the network model, an extended user dataset is used to ensure the possibility of effective network prediction. On the other hand, the present invention also uses the IMF sequences of similar users to collaboratively predict the electricity consumption data to be measured, improving the accuracy of the prediction result.

[0096] Figure 2 is a schematic diagram of the update process of the user dataset in the embodiment of the present invention, as Figure 2 shown. First, the user dataset to be input is required, where the input user dataset contains N users, and each user has a load value at each timestamp within the time period t - m to t + M. Secondly, perform EEMD analysis on the input user dataset, and this process can extract IMF components and form IMF sequences. Thirdly, perform PCA dimensionality reduction on each level of IMF sequence. Then, perform K-means clustering on the dimensionality-reduced IMF sequences to output the updated user dataset, where the output user dataset still contains N users, and each user's feature belongs to the class number of different levels of IMF sequences, converting the time series into the clustering result of IMF sequences. For example, a certain user's time series is decomposed to obtain M IMFs, that is, M-level IMFs, and other users also have M IMFs. Perform clustering on the M-level IMF to obtain the class number, and the new feature of the user is called IMF M, which contains the M-level IMF class number of this user.

[0097] Figure 3 is a schematic diagram of the data frame structure of similar users in the embodiment of the present invention, as Figure 3As shown, by combining the data of similar users, the power consumption time series of the target user is moved to the first piece of data, and then the extended data frames of each IMF sequence of this user are used as the extended data to train the network model. Obviously, using clustering analysis techniques, such as the K-means algorithm, the time series of different users with similarities can be combined into subgroups. In this way, we can expand the target user dataset, that is, improve the prediction accuracy, and make power consumption predictions according to the subgroup population. At this time, the data frames to which the users belong are divided according to the L-class numbers, and the data of the target user is transferred to the first place in the data frame of this user. For further model optimization, it is necessary to create extended data frames for each IMF of the user, that is, one original data frame becomes the number of data frames equal to the number of IMFs.

[0098] It can be understood that when new data appears for a user, it is necessary to extract the IMF sequence from the new data of the current user, and there is no need to extract the IMF sequence from the new data of similar users. Only by combining the new data of similar users with the IMF components of the new data of the current user and inputting them into the trained model can the power consumption prediction result of the current user be obtained. This method simplifies the data processing flow and can also improve the prediction accuracy.

[0099] Since the general EEMD prediction model does not consider the problem of multiple features; and the applied neural network only includes the LSTM layer and does not have other ability layers to assist in improving the model prediction effect. To solve these problems, the CNN-LSTM model of the present invention attempts to expand the original dataset with multiple features to obtain better prediction results. Figure 4 is a schematic diagram of the power consumption prediction process in an embodiment of the present invention, as Figure 4 shown, taking Figure 3 the data frame as the input of the CNN-LSTM model. Since there are N data frames, each data frame needs to be processed separately and its own prediction result is obtained separately; finally, each prediction result is combined to obtain the final prediction result.

[0100] Push the data frame into the CNN-LSTM model for prediction. Based on Keras in Python, as Figure 5 shown, the network model in the embodiment of the present invention consists of Conv1D and MaxPooling1D as parts of the CNN, which can effectively extract the potential relationship between continuous data and discontinuous data in the feature map to form a feature vector; the input of the Conv1D structure is (N, C in , L) and the output is (N, C out , L out ) can be described as:

[0101]

[0102] Among them, ★ is an effective cross - correlation operator, and N i is the batch size, C represents the number of features, that is, the number of users in the same category, L is the length of the input time series, kerner_size is the convolution kernel size, and C in is the number of input channels; is the bias of the output result; weight represents the weight; T 1 represents the user's temperature at the current time, and T 2 represents the current time.

[0103] Regarding the input (N, C, L) and output (N, C, L out ) of Maxpolling1D, it can be described as:

[0104]

[0105] Among them, stride is the step size. Then, the feature vectors are constructed in a time - series manner and used as input data, and 3 LSTM network layers are adopted for short - term load forecasting. Each LSTM layer calculates:

[0106] z t = g(W z x t + R z y t-1 + b z ) input

[0107] i t = σ(W ix x t + R i y t-1 + p i ⊙ c t-1 + b i ) input gate

[0108] f t = σ(W fx x t + R f y t-1 + p f ⊙ c t-1 + b f ) forget gate

[0109] c t = f t ⊙ c t-1 + i t ⊙ z t cell state

[0110] o t = σ(W ox x t + Ro y t-1 + p i ⊙ c t + b o ) Output gate

[0111] y t = o t ⊙ h(c t ) Output

[0112] where x t is the input vector of time t, W is the input weight matrix, R is the squared recurrent weight matrix, p is the peephole connection weight vector, b is the bias vector. σ is the sigmoid function, and ⊙ is the Hadamard product.

[0113] Because the parallel CNN-LSTM model prediction is implemented, the prediction results need to be integrated. That is, the integration formula is:[[]]

[0114]

[0115] where, y out is the integrated prediction result, is the prediction result of the i-th IMF, M is the number of IMFs; f represents the relationship function between the user and the temperature factor and the time factor; the inverse variation of data standardization is implemented to obtain the final prediction result.

[0116] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "coaxial", "bottom", "one end", "top", "middle", "the other end", "upper", "one side", "top", "inner", "outer", "front", "center", "both ends", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0117] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "set", "connected", "fixed", "rotated", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. It can be the communication inside two elements or the interaction relationship between two elements. Unless otherwise clearly limited, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0118] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A power consumption prediction method based on user clustering to expand data, characterized in that, the method comprises the following steps: S1: Obtain a user data set, where the user data set includes user power consumption data for each period; and preprocess the power consumption data of the users to obtain a user power consumption time series; S2: Decompose the user power consumption time series by using the empirical mode decomposition method to obtain the IMF sequence of the user; S3: Reduce the dimension of the user's IMF sequence by using the K-Means clustering method based on principal component analysis, obtain the clustering result of the user IMF sequence, and use it as the updated user data set; S4: Use the K-Means clustering method based on principal component analysis for the clustering result of the user IMF sequence to obtain the user clustering result; S5: According to the user clustering result, expand the IMF sequences of users belonging to the same category, and add the power consumption time series of the same-category users to the user's IMF sequence; in step S5, it specifically includes: S51: Divide the user data according to the user clustering result, that is, create its own data frame for each category of users, the column label of the data frame is the user number, and the data column is the corresponding user power consumption data; S52: Expand the user power consumption data according to the user category, that is, if user n is in category m, then the data set of user n in category m includes: the data of user n in category m and the data of other users belonging to category m; S53: Use the expanded data frame obtained in step S52 as the expanded feature vector, and use the IMF data set in step S2 to replace the user's own power consumption time series, that is, expand the user's power consumption data with the data frame similar to each IMF itself; S6: Input the expanded IMF sequence as the user data set updated again into the convolutional neural network fusion long short-term memory artificial neural network model to train the network model; S7: Input the IMF sequence of the user in the current period and the power consumption time series of the same-category users in the current period into the trained network model to obtain the power consumption prediction result of the user in the next period; the specific steps of step S7 include: During the training process of the network model, use the IMF component of the user at the previous moment and the corresponding temperature and time, the temperature and time of the user at the next moment, combine the user time series of the same-category users at the previous moment and the corresponding temperature and time, predict the user time series of the user at the next moment, and compare the predicted user time series with the real user time series to train the network model so that the network model can meet the prediction requirements; The power consumption prediction result of the user in the next period is expressed as: Among them, y out is the comprehensive prediction result, is the prediction result of the i-th IMF, and M is the number of IMFs; T 1 represents the user's current temperature, T 2 represents the current time; f represents the relationship function between the user and the temperature factor and time factor, and this function is obtained through the training and iteration process of the network model.

2. A power consumption prediction method based on user clustering to expand data according to claim 1, characterized in that, in step S1, the preprocessing method includes: filling null values with the average value and converting the time series data to hourly units.

3. A power consumption prediction method based on user clustering to expand data according to claim 1, characterized in that, in step S2, it specifically includes: S21: Set the number of IMF components according to the required accuracy of the network model and the effectiveness of decomposition by the empirical mode decomposition method to determine the dimension of the user's IMF dataset. S22: Use the empirical mode decomposition method to repeatedly add white noise to each user's power consumption time series, and take the average value of the calculated IMF components as the final result, including the user's historical IMF sequence and the user's to-be-tested IMF sequence.

4. A power consumption prediction method based on user clustering extended data according to claim 3, characterized in that, in step S3, it specifically includes: S31: Perform principal component analysis dimensionality reduction on each historical IMF sequence in step S22 to obtain a two-dimensional IMF dataset. S32: For the two-dimensional IMF dataset obtained in step S31 use the K-means clustering method to obtain the clustering result of the IMF sequence, that is, the clustering result of a certain level of IMF component corresponding to each user. S33: Use the clustering result of the IMF sequence to replace the user's power consumption time series to form an updated user dataset.

5. A power consumption prediction method based on user clustering extended data according to claim 4, characterized in that, in step S4, it specifically includes: S41: Perform principal component analysis dimensionality reduction on the clustering result of the IMF sequence obtained in step S33 again. S42: Use the K-means clustering method for the clustering result of the dimensionality-reduced IMF sequence to obtain the user's clustering result based on the time series.

6. A power consumption prediction method based on user clustering extended data according to claim 1, characterized in that, in step S6, it specifically includes: S61: Construct a convolutional neural network fusion long short-term memory artificial neural network model, where the long short-term memory artificial neural network includes 3 consecutive LSTM layers. S62: Input the user dataset updated again in step S5 into the network model to obtain the corresponding prediction result. S63: Compare the corresponding prediction result with the actual power consumption result, and train the network model until the network model meets the accuracy requirements.

Citation Information

Patent Citations

  • Power consumer characteristic analysis method and system based on BEMD and kmeans

    CN111898857A

  • Power load probability prediction method based on user power consumption behavior clustering

    CN113361776A