VMD-GRU Short-Term Load Forecasting Method Based on Whale Optimization Algorithm and Type Recognition
Through the whale optimization algorithm and type recognition VMD-GRU short-term load prediction method, the parameters are optimized by random forest classification and whale optimization algorithm, combined with the GRU-Attention neural network for weighted prediction, the problem of insufficient non-stationarity and robustness of load data in the existing technology is solved, and higher prediction accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202311279093.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-09-28
AI Technical Summary
The existing short-term load prediction methods are not effective when processing non-stationary and nonlinear load data. The parameter selection depends on experience, fail to fully consider the differences in electricity loads used in industries, residents, and public places, are not robust enough, and fail to effectively respond to volatility and seasonal impacts caused by climate and policy.
The VMD-GRU short-term load prediction method using whale optimization algorithm and type recognition is used to determine the load-day category through the random forest classification model, and the whale optimization algorithm is used to search for the optimal parameters of VMD decomposition, and weighted prediction is used to improve the robustness and pertinence of the model.
It improves the accuracy and robustness of short-term load prediction, can better adapt to load changes in different scenarios, and enhances pattern recognition capabilities.
Smart Images

Figure CN117195947B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power grids, and particularly relates to a short-term load forecasting method based on the whale optimization algorithm and type recognition for VMD-GRU. Background Art
[0002] Nowadays, the proportion of new energy in the power grid is increasing. Environmental factors such as weather, atmospheric humidity, and wind speed will affect the power generation of new energy. Therefore, new energy power has the disadvantage of instability, which requires a relatively accurate load forecasting method to assist in the decision-making of power dispatching and reduce the unstable impact of new energy power on the power grid.
[0003] In the related art, when performing short-term load forecasting, first, the variational mode decomposition is used to decompose the load curve for a period of time into several smooth and more predictable components, then the GRU (Gate Recurrent Unit) neural network is used for forecasting, and finally, the forecasting results based on each component are re-accumulated and integrated to obtain the final load forecasting result.
[0004] However, there are still many problems in the existing short-term load forecasting methods: First, short-term load data has the characteristics of non-stationarity and non-linearity, and the GRU neural network model is more suitable for processing stationary and linear data; Second, in the neural network model algorithm, parameters such as the learning rate, the number of neurons in the hidden layer, and the number of iterations can only be selected through experience and cannot meet the learning and training of different load data curves; Third, the electricity load characteristics of industries, residents, and public places are very different, and the power load usually has large fluctuations, seasonality, and periodicity due to factors such as climate and policies. The existing power load forecasting methods do not consider these factors; Fourth, the robustness of the currently commonly used neural network methods needs to be strengthened. Summary of the Invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a short-term load forecasting method based on the whale optimization algorithm and type recognition for VMD-GRU. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] The present invention provides a short-term load forecasting method based on the whale optimization algorithm and type recognition for VMD-GRU, including:
[0007] Obtain the characteristic quantities of n1 load days before the prediction day t and the load day categories of n2 load days before the prediction day t, and form the feature X label ;
[0008] For the feature X labelInput it into a pre-trained random forest classification model to obtain the load day category label of the to-be-predicted day t t = τ, τ ∈ [1, …, k], 1, …, k represent k types of load day categories;
[0009] Input the load amounts and feature amounts of the n3 load days before the to-be-predicted day t into each pre-trained component model under each load day category, and determine the current prediction result corresponding to each load day category according to the outputs of each pre-trained component model under each load day category;
[0010] Obtain the accuracy rate obtained by using the validation set for testing during the training process of the random forest classification model;
[0011] According to the accuracy rate and the load day category label t = τ to determine the weights, and weight the current prediction results corresponding to all load day categories to obtain the load prediction result of the to-be-predicted day t.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0013] The present invention provides a VMD-GRU short-term load prediction method based on the whale optimization algorithm and type recognition. Before prediction, first use the random forest model to determine the load day category label t = τ of the to-be-predicted day, and then use k pre-trained GRU-Attention neural network models corresponding to k types of load day categories respectively to perform load prediction. Finally, weight the current prediction results output by each GRU-Attention neural network model to obtain the load prediction result of the to-be-predicted day. Since in the training process of the present invention, the clustering algorithm is used to distinguish the categories of load days, and the whale optimization algorithm is used to search for the optimal parameters of VMD decomposition and the training parameters of each GRU-Attention model, the prediction difficulty of the GRU-Attention model is reduced, and the obtained result is the weighted output of multiple different category neural network models, which improves the robustness of the neural network for load prediction. Compared with the prior art, it has stronger pattern recognition ability, stronger pertinence to the prediction scenario, improves the robustness of the neural network in actual use, and improves the accuracy rate of load prediction in specific scenarios.
[0014] The following will further describe the present invention in detail with reference to the drawings and embodiments. Description of the Drawings
[0015] Figure 1 is a flowchart of a VMD-GRU short-term load prediction method based on the whale optimization algorithm and type recognition provided by an embodiment of the present invention;
[0016] Figure 2 It is a flowchart of the whale optimization algorithm provided by an embodiment of the present invention;
[0017] Figure 3 It is a schematic structural diagram of the GRU network model provided by an embodiment of the present invention;
[0018] Figure 4 It is a diagram of the calculation results of the silhouette coefficient provided by an embodiment of the present invention;
[0019] Figure 5a It is a diagram of the clustering results of the first type of load day category provided by an embodiment of the present invention;
[0020] Figure 5b It is a diagram of the clustering results of the second type of load day category provided by an embodiment of the present invention;
[0021] Figure 6a It is a schematic diagram of the change of the envelope entropy value during the search for the optimal parameters of VMD decomposition provided by an embodiment of the present invention;
[0022] Figure 6b It is a schematic diagram of the change of the penalty factor during the search for the optimal parameters of VMD decomposition provided by an embodiment of the present invention;
[0023] Figure 6c It is a schematic diagram of the change of the number of decompositions during the search for the optimal parameters of VMD decomposition provided by an embodiment of the present invention;
[0024] Figure 7 It is a schematic diagram of the VMD decomposition result provided by an embodiment of the present invention;
[0025] Figure 8a It is a schematic diagram of the change of the fitness function during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention;
[0026] Figure 8b It is a schematic diagram of the change of the batch size during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention;
[0027] Figure 8c It is a schematic diagram of the change of the number of nodes in the fully connected layer during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention;
[0028] Figure 8d It is a schematic diagram of the change of the learning rate during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention;
[0029] Figure 8eIt is a schematic diagram of the change in the number of training times during the search for the optimal training parameters of the GRU-Attention network model provided by the embodiments of the present invention;
[0030] Figure 8f It is a schematic diagram of the change in the number of hidden nodes during the search for the optimal training parameters of the GRU-Attention network model provided by the embodiments of the present invention. Detailed implementation manners
[0031] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0032] Figure 1 It is a flowchart of a VMD-GRU short-term load forecasting method based on the whale optimization algorithm and type recognition provided by the embodiments of the present invention. As Figure 1 shown, the embodiments of the present invention provide a VMD-GRU short-term load forecasting method based on the whale optimization algorithm and type recognition, including:
[0033] S1. Obtain the feature quantities of n1 load days before the prediction day t and the load day categories of n2 load days before the prediction day t, and form the feature X label ;
[0034] S2. Input the feature X label into a pre-trained random forest classification model to obtain the load day category label of the prediction day t t =τ, τ∈[1,…,k], 1,…,k represent k load day categories;
[0035] S3. Input the load quantities and feature quantities of n3 load days before the prediction day t into each pre-trained component model under each load day category, and determine the current prediction result corresponding to each load day category according to the outputs of each pre-trained component model under each load day category;
[0036] S4. Obtain the accuracy rate obtained by using the validation set for testing during the training process of the random forest classification model;
[0037] S5. Determine the weight according to the accuracy rate and the load day category label t =τ, and weight the current prediction results corresponding to all load day categories to obtain the load prediction result of the prediction day t.
[0038] It should be noted that in step S3, the current prediction result corresponding to each load day category is obtained by weighted summation of the outputs of each pre-trained component model under that load day category.
[0039] Optionally, in step S1, the steps of obtaining the feature quantities of the n1 load days before the prediction day t and the load day categories of the n2 load days before the prediction day t include:
[0040] S301. Obtain the feature quantities of the 3 load days before the prediction day t respectively the feature quantities of the 2 load days before and the feature quantities of the 1 working day before
[0041] S302. Obtain the load day category labels of the 5 load days before the prediction day t respectively t-5 the load day category labels of the 4 load days before t-4 the load day category labels of the 3 load days before t-3 the load day category labels of the 2 load days before t-2 and the load day category labels of the 1 working day before t-1 ; where the feature quantities include the light intensity G, environmental temperature T, relative humidity H, and wind speed V of the load day.
[0042] S303. Generate features based on the feature quantities of the 3 load days before the prediction day t and the load day categories of the 5 load days before the prediction day t
[0043]
[0044] Optionally, in step S3, the steps of inputting the load quantities and feature quantities of the n3 load days before the prediction day t into each pre-trained component model under each load day category and determining the current prediction result corresponding to each load day category according to the outputs of each pre-trained component model under each load day category include:
[0045] S301. Obtain the load quantities and feature quantities of the 20 load days before the prediction day t and form a matrix S M ;
[0046] S302. Obtain each pre-trained component model under the k load day categories, and the component model is a GRU-Attention model;
[0047] S303. Convert the matrix S M into a tensor form S tensor , and input it into each pre-trained component model under each load day category;
[0048] S304. Accumulate the outputs of each pre-trained component model under each load day category to obtain the current prediction result corresponding to each load day category.
[0049] Optionally, in step S5, according to the accuracy rate and the load day category label t = τ to determine the weight, and weight the current prediction results corresponding to all load day categories to obtain the load prediction result of the day t to be predicted, the steps include:
[0050] S501. Obtain the current prediction result output by the τ-th pre-trained GRU-Attention neural network model corresponding to the load day category τ
[0051] S502. Obtain the current prediction result output by the e-th pre-trained GRU-Attention neural network model corresponding to other load day categories e where e ≠ τ;
[0052] S503. Weight the current prediction results corresponding to all load day categories according to the following formula:
[0053]
[0054] In the formula, Acc represents the accuracy rate obtained by using the validation set for testing during the training of the random forest classification model, and P f represents the load prediction result of the day t to be predicted.
[0055] In this embodiment, the above component model is trained according to the following steps:
[0056] Collect N sample data including sample feature quantities and load quantities at a time resolution of 15 minutes to construct the first data set U = {(X, Y)} = {(x1, y1), (x2, y2),..., (x N , y N )}, where the feature quantity data set X = {x1, x2,..., x N}, x1, x2,..., x N represent 1, 2,..., N sample feature quantities, and the load quantity data set Y = {P1, P2,..., P N}, P1, P2,..., P N represent the load quantities corresponding to x1, x2,..., x N respectively;
[0057] Convert the first data set U into a second data set in units of days where X day , Y day represent the feature quantity set in units of days and the load quantity data set in units of days respectively, represent 1, 2,..., N / 96 sample feature quantities in units of days respectively, respectively represent the load amounts in days corresponding to ;
[0058] After clustering the second data set U day to obtain k types of load day categories, based on the second data set U day generate the load amount sequence and the feature amount sequence for each load day category, and perform variational mode decomposition on each load amount sequence to obtain multiple IMF components;
[0059] Construct a component data set according to the feature amount sequence and the IMF components, and use the component data set to train and obtain each component model for each load day category respectively.
[0060] Optionally, after clustering the second data set U day to obtain k types of load day categories, based on the second data set U day generate the load amount sequence for each load day category, and the steps of performing variational mode decomposition on each load amount sequence to obtain multiple IMF components include:
[0061] Use the K-means algorithm to cluster the load amount data set Y in days day to obtain k types of load day categories;
[0062] Based on the second data set U day generate the load amount sequence for each load day category;
[0063] Use the whale optimization algorithm to search for the optimal parameters corresponding to the variational mode decomposition of the load amount sequence for each load day category. The optimal parameters include: the number of modal components M and the penalty factor α;
[0064] Based on the optimal parameters, perform variational mode decomposition on the load amount sequence for each load day category to obtain multiple IMF components.
[0065] Specifically, the weather conditions of a certain area can be collected from a weather website to obtain the sample feature amounts for training, that is, the four features of light irradiance G, ambient temperature T, relative humidity H, and wind speed V. In this embodiment, the time resolution of the load prediction output is 15 minutes. Therefore, the four sample feature amounts and the load amount need to be collected every 15 minutes to form the first data set U. The first data set U can be expressed as U = {(X, Y)} = {(x1, y1), (x2, y2),..., (x N , y N )}, the feature amount data set X = {x1, x2,..., x i , …, x N} contains N sample feature amounts, where x i = (G i,T i ,H i ,V i ) represents the i-th sample feature quantity in the 4D sample space, G i ,T i ,H i ,V i respectively represent the light irradiance, ambient temperature, relative humidity, and wind speed in the i-th sample feature quantity; the load quantity dataset Y = {P1, P2,..., P i ,..., P n} contains the load quantities corresponding to the 1st, 2nd,..., N-th sample feature quantities. That is to say, each sample feature in X can find the corresponding load quantity in Y.
[0066] Furthermore, based on the above setting of time resolution, the sample feature quantities and load quantities at 96 time points in a day should be collected. The data of these 96 points are arranged in order in the dataset and altogether contain days of data. Therefore, the second dataset in units of days is represented as where X day , Y day respectively represent the set of feature quantities in units of days and the load quantity dataset in units of days, respectively represent the 1st, 2nd,..., N / 96 sample feature quantities in units of days, respectively represent the load quantities in units of days corresponding to .
[0067] Next, use the K-means clustering algorithm to cluster the load days based on the second dataset U day to distinguish different load day categories and label each load day. It should be noted that when constructing the second dataset U day in units of days, an accumulation operation can be performed on the light irradiance G within a day, and an average operation can be performed on the ambient temperature T, relative humidity H, and wind speed V. The specific calculation method is:
[0068]
[0069]
[0070]
[0071]
[0072] where represents rounding down, that is, ignoring the decimal part and only taking the integer part. In the second dataset U day among
[0073] In addition, it is also necessary to normalize Y j as follows:
[0074]
[0075] where d = 1, 2,..., 96, represents the maximum load among the 96 points within the j-th day, represents the minimum load among the 96 points within the j-th day.
[0076] After normalization, the load on the j-th day Then the load dataset in days can be expressed as
[0077] The purpose of using the K-means clustering algorithm for clustering is to classify into k different load day categories, and the clustering can be specifically carried out according to the following steps:
[0078] Step 1: Initialization. Let t = 0, randomly determine a clustering number k, and randomly select k sample data as the initial clustering centers
[0079] Step 2: Start clustering the samples. For the fixed class centers calculate the distance from each sample to the class center:
[0080]
[0081] where, represents the center of the l-th load day category G l .
[0082] After calculating the distances, assign each sample data to the class with the nearest center to obtain the clustering result C (t) .
[0083] Step 3: For the clustering result C (t) , calculate the mean of the sample data in each current class as the new class center If the iteration converges or meets the iteration stop condition, output C (t) as the clustering of the sample set. If the above two conditions are not met, let t = t + 1, return to Step 2 and continue to iterate until the conditions are met, and finally obtain the clustering result
[0084] Of course, several clustering numbers k can be set, and the silhouette coefficient is calculated for each clustering result to measure the quality of the clustering result. The silhouette coefficient is:
[0085]
[0086] Among them, a(i) is the within-cluster dissimilarity, which is the average distance from sample i to other samples in the same cluster, and b(i) is the between-cluster dissimilarity, which is the average distance from sample data i to all samples in other clusters. The closer the silhouette coefficient is to 1, the more reasonable the sample clustering is. When selecting k, follow this standard and select the k value with the silhouette coefficient closest to 1 as the final number of clusters.
[0087] In summary, finally, k different load day categories can be obtained through clustering. The load day types corresponding to the load days are denoted as label = l, l = 1,..., k, and the obtained categories are used as labels to generate the load quantity sequences under each load day category.
[0088] Furthermore, variational mode decomposition is performed on the load quantity sequences under different load day categories. In this process, the decomposition effect of the load quantity sequences will be affected by relevant parameters, and the determination of parameters generally has strong subjectivity, which will affect the decomposition effect of the load quantity sequences and the final load prediction accuracy. Therefore, before performing variational mode decomposition, this embodiment first uses the WOA (Whale Optimization Algorithm) to search for the optimal parameters corresponding to the load quantity sequences under each load day category, including: the number of modal components M and the penalty factor α.
[0089] Figure 2 is a flowchart of a whale optimization algorithm provided by an embodiment of the present invention. Please refer to Figure 2 . During the search process, the minimum value of the envelope entropy representing the sparse characteristics of the original signal is selected as the fitness function. When there is more noise and less characteristic information in the modal components, the value of the envelope entropy is larger; on the contrary, the value of the envelope entropy is relatively small. Therefore, this embodiment uses minimizing the envelope entropy as the fitness function.
[0090] Based on the k types of load day categories obtained through clustering, the component data set can be expressed as:
[0091] U′ = {U (1) , U (2) ,..., U (l) , …, U (k)}
[0092] Among them, U (1) , U (2) ,..., U (l) , …, U (k) respectively include the characteristic quantity sequences and load quantity sequences under the 1st, 2nd,..., kth load day categories.
[0093] It should be noted that in this embodiment, only the load quantity sequence is subjected to variational mode decomposition, and the individual characteristic quantity sequences are not decomposed.
[0094] Before decomposition, it is still necessary to normalize the data in the load quantity sequence. This normalization is different from the normalization of the load data for each day in the above text on a daily basis. Instead, it is the normalization of all the load data included in the entire component dataset U':
[0095]
[0096] At this time, the load quantity sequence under the l-th load day category can be expressed as: where c is the product of the number of days corresponding to this load day type and 96, indicating that there are c load value points in this type of load data. Subsequent variational mode decomposition is based on the sequence for operation.
[0097] Next, it is necessary to calculate the corresponding envelope entropy to represent the fitness function to be used in the whale optimization algorithm. The calculation formula of the envelope entropy is as follows:
[0098]
[0099] where p x is the normalized form of c(x), and c(x) is the envelope signal after Hilbert demodulation of the M modal components generated after performing VMD (Variational Mode Decomposition) on the sequence .
[0100] Calculate the envelope entropy values of all the IMF components obtained by VMD decomposition when the whale population is in a position (corresponding to a parameter combination M and α), and take the smallest one among all as the local minimum entropy value, denoted as min L E p IMF . The component corresponding to this value is the component that best contains the local characteristic information of the load sequence. Set the fitness function as the local minimum entropy value that appears in the optimization process, and minimize this value as the best optimization goal. Use the whale optimization algorithm to optimize the number of modal components M and the penalty parameter α.
[0101] The whale optimization algorithm realizes the purpose of optimal search by simulating the predation behaviors of whales such as searching for prey, surrounding prey, and attacking prey with a bubble net. Since the position of the prey in the solution space is unknown to the whale population before solving the problem, at the beginning, the position of the optimal whale individual in the current population is set as the position of the prey, and other whales in the population defend towards the position of the optimal whale. Its mathematical model is:
[0102] D = |C·F * (t) - F(t)|
[0103] F(t + 1) = F * (t) - A•D
[0104] where t is the current iteration number, F * represents the position of the optimal whale in the current whale population, F represents the current whale position, and A and C are vectors, that is:
[0105] A = 2ar - a
[0106] C = 2r
[0107] where a is the convergence factor, and as the predation iteration of the whale population progresses, the value of a linearly decreases from 2 to 0; r represents a random number vector with values in the interval [0, 1].
[0108] There are two mechanisms in total for describing the bubble net predation behavior of the whale population. The first is the shrinking encircling mechanism, which is achieved by reducing the value of a in A = 2ar - a. The fluctuation range of |A| will also decrease due to the reduction of a. That is to say, |A| is a random value within the interval [-a, a], where a linearly decreases from 2 to 0 during the iteration. When defining the random value of |A| within the interval [-1, 1], the new position of the whale individual can be defined at a certain position between the original position of the whale and the current optimal whale position. The second is the spiral update position mechanism. This method first calculates the distance between the whale at the position [X, Y] and the prey at the position [X * , Y * . A spiral equation is used between the whale and the prey to imitate the spiral movement of the whale, that is:
[0109] F(t + 1) = D′·e bh ·cos(2πh) + F * (t)
[0110] where D′ = |F * (t) - F(t)| represents the distance between the optimal whale individual and the current whale individual in the t-th iteration, b is a constant, and h is a random number within the interval [-1, 1]. Assume that there is a 50% probability of randomly selecting between the shrinking encircling mechanism and the spiral update position mechanism to update the position of the whale individual during the optimization process. Its mathematical model is:
[0111]
[0112] where q is a random number within the interval [0, 1].
[0113] During the process of whales searching for prey, they can utilize the change of matrix A for global exploration as the iteration progresses. In fact, whales will randomly explore the solution space based on each other's positions. Therefore, when |A| > 1 or |A| < -1, the algorithm performs local exploitation operations to update the positions of whales, moving away from the current individual. In contrast to local exploitation, during the global exploration stage, the positions of whale individuals are updated based on randomly selected whale individuals, rather than the currently found optimal whale individual. This mechanism focuses on exploration, so when |A| > 1, the whale optimization algorithm performs global exploration operations. During the prey search stage, since the prey position is unknown to the whale population, whales need to obtain the prey position through collective cooperation. Whales use the random individual positions in the population as navigation targets to search for food, and its mathematical model is described as:
[0114] D = |C • F rand - F|
[0115] F(t + 1) = F rand - A · D
[0116] where F rand represents the position of a randomly selected whale individual in the current whale population.
[0117] The whale optimization algorithm starts execution from the positions of the given initial whale population. At each iteration, whale individuals update their own positions using the position information of randomly selected positions or the position information of the whale individual with the maximum fitness value obtained so far in the iteration. As the parameter a linearly decreases from 2 to 0, the transition between the global exploration stage and the local exploitation stage of the algorithm is realized. When |A| > 1, a whale is randomly selected from the population. When |A| < 1, the whale with the currently maximum fitness value is selected to update the position of the current whale individual. Given the value of q, the whale optimization algorithm has the ability to switch between the shrinking surrounding mechanism and the spiral position update mechanism. Finally, the whale optimization algorithm terminates when a termination condition is met.
[0118] Before starting the optimization, the position of the whale population vector, that is, the optimization target [M, α], needs to be initialized, and then the local minimum entropy min L E p IMFCalculate the fitness of each whale as the fitness function. When performing VMD transformation, process the signal according to the position of each whale, and then calculate the envelope entropy value corresponding to each whale individual and record the position of the optimal individual. At this time, it is necessary to judge whether the position of the current optimal individual meets the set stop condition. If it meets, exit the algorithm and output the optimal target value. If it does not meet, the parameters a, A, C, and h in the algorithm need to be updated, and then operations such as surrounding the prey, bubble net attack, and searching for the prey are performed according to the values of |A| and q in the algorithm. After that, return to the operation of calculating and updating the position of the optimal individual until the stop condition is met and output the data set of the l-th category. The optimal parameters [M (l)* , α (l)* corresponding to the VMD decomposition. Then, process the data sets of other load day categories according to the above process to obtain the VMD optimal parameters [M (1)* , α (1)* of the load amount sequences under all load day categories, [M (2)* , α (2)* ,..., [M (k)* , α (k)* .
[0119] After searching for the optimal parameters of the VMD decomposition, perform variational mode decomposition according to the optimal parameters. For the sequence , the constraint expression for signal decomposition using VMD is:
[0120]
[0121]
[0122] Among them, {u m}, {ω b} are the expressions of the m-th modal component and the center frequency after decomposition. The parameter M (l)* represents the optimal number of modal decompositions of the load amount sequence of the l-th load day category obtained by the whale optimization algorithm. δ(t) represents the Dirac function, and * is the convolution operator. All the modal components obtained by the decomposition are consistent with the sequence . The solutions of the above two equations can be obtained by introducing Lagrange multipliers to transform the constrained problem into an unconstrained problem:
[0123]
[0124] Among them, λ is the Lagrange operator, and α (l)* is the optimal penalty factor for the VMD decomposition of the load amount sequence of the l-th load day category.
[0125] Then, after using ADMM for optimization iteration, the modal component u b, the respective modal frequencies ω can be obtained b and the Lagrange operator λ, as follows:
[0126]
[0127]
[0128]
[0129] where γ represents the noise tolerance, which can meet the requirements of VMD signal decomposition and fidelity. u i (ω), and after Fourier transform of λ(ω), we can obtain u i (t), and λ(t). The process of decomposing the load sequence using VMD can be described as follows: First, initialize the relevant parameters such as λ 1 etc., set the maximum number of iterations T
[0130] and the expected decomposition accuracy ε, and then use the above equations (1), (2), and (3) to update the parameters u m , ω m , λ. Then calculate the decomposition accuracy. If it does not meet and n < N, continue to calculate and update the parameters. If it meets, end the algorithm and output the results of the modal components and the center frequencies.
[0131] Furthermore, the steps of constructing a component dataset according to the feature quantity sequence and IMF components and training each component model under each load day category using the component dataset include:
[0132] According to the feature quantity sequence and IMF components, respectively construct the component datasets for training the corresponding GRU-Attention neural network models where, represents the r-th IMF component obtained by variational mode decomposition of the load quantity sequence under the l-th load day category, and X (l) represents the feature quantity sequence under the l-th load day category, l = 1, 2,..., k, r = 1, 2,..., M;
[0133] Using the component dataset for the r-th IMF component obtained by variational mode decomposition of the load quantity sequence under the l-th load day category Train the corresponding GRU-Attention neural network model to obtain M component models under the l-th load day category.
[0134] Specifically, after performing variational mode decomposition on the load quantity sequences under all load day categories, multiple IMF components are obtained. Based on the IMF components and the feature quantity sequences, a component data set is constructed to train the GRU-Attention neural network model corresponding to the r-th IMF component obtained after variational mode decomposition of the load quantity sequences under each load day category. Exemplarily, take the r-th IMF component obtained after variational mode decomposition of the load quantity sequence under the l-th load day category r = 1, 2,..., M (l)* , combined with the feature quantity sequence X in (l) to form a complete component data set Then divide the component data set, where 75% is used as the training set, denoted as Another 25% is used as the test set, denoted as
[0135] Optionally, before the step of training the GRU-Attention neural network model corresponding to the r-th IMF component obtained after variational mode decomposition of the load quantity sequence under the l-th load day category using the component data set the r-th IMF component obtained after variational mode decomposition of the load quantity sequence under the l-th load day category it also includes:
[0136] Use the whale optimization algorithm to search for the optimal training parameters of the corresponding GRU-Attention neural network model.
[0137] In this embodiment, the training parameters include: learning rate, number of iterations, number of selected samples, number of hidden layer neuron nodes, and number of nodes in the fully connected layer.
[0138] Specifically, this embodiment uses the whale optimization algorithm to find the optimal parameters in the training process of each GRU-Attention model. This step is similar to the above-mentioned use of the whale optimization algorithm to search for the optimal parameters of VMD. First, select the fitness function of the whale optimization algorithm, and use the root mean square error between the predicted value and the true value as the fitness function, that is, RMSE is the fitness function. In the data set , is the true value in the data set, let be the predicted value predicted by the GRU-Attention model, then RMSE is:
[0139]
[0140] where Q is the length of the component .
[0141] After determining the fitness function, it is also necessary to determine what the optimization objective of the whale optimization is. For this neural network, the parameters to be optimized are: the learning rate lr during model training, the number of iterations Nep of model training, the number of samples batchsize selected in one training, the number of hidden layer neurons hidnum in the GRU network, and the number of nodes fcnum in the fully connected layer. Integrating these several parameters into a vector [lr, Nep, batchsize, hidnum, fcnum] is the optimization objective of the whale optimization algorithm here. The remaining operations are the same as those in step 3 for finding the optimal parameters in VMD transformation. After determining the optimization objective, set the optimization range for each parameter, and then start the optimization iteration. When training, process the signal according to the position of each whale, then calculate the RMSE corresponding to each whale individual and record the position of the optimal individual. At this time, it is necessary to determine whether the current position of the optimal individual meets the set stop condition. If it meets, exit the algorithm and output the optimal objective value. If it does not meet, it is necessary to update the four parameters a, A, C, and h in the algorithm, and then perform operations such as surrounding the prey, bubble net attack, and searching for the prey according to the magnitudes of the values of |A| and q in the algorithm. After that, return to the operation of calculating and updating the position of the optimal individual until the stop condition is met and output the r-th component after VMD decomposition of the load sequence under the l-th load day category The optimal value set for GRU - Attention model training .
[0142] During the training process of the GRU - Attention model, it is necessary to convert the initial array types in the training dataset and the test dataset into tensor types that can be calculated by the pytorch framework. In this embodiment, the method of time - sliding window is adopted. Integrate the load and feature quantities of the past 20 time points of the time point to be predicted and convert them into tensor types. Finally, a training dataset with a dimension of 3 and a size of [r train , 20, 5] and a test dataset with a size of [r test , 20, 5] are formed, where r train is equal to the length of the training dataset minus 20, and r test is equal to the length of the test dataset minus 20; [r train , 20, 5] and [r testThe meaning of 20 in [[0, 20, 5]] is that the time step used in constructing the dataset with a time-sliding window is 20, and the meaning of 5 is that the dataset contains a total of 5 features that can be used in training, including: irradiance, relative humidity, ambient temperature, wind speed, and the load of the past 20 time points at the current time point.
[0143] Next, the training dataset is put into the GRU-Attention network model for training. The training parameters of the GRU-Attention network model include the learning rate, the number of training epochs, the batch size, the number of hidden layers, and the size of the fully connected layer. GRU is one of the variants of the recurrent neural network RNN. Since the power load has a periodic pattern, it can well capture the long-term and short-term dependencies in the time series.
[0144] Figure 3 is the schematic structural diagram of the GRU network model provided by the embodiment of the present invention. As Figure 3 shown, the key difference between the GRU model and the ordinary RNN is that GRU supports gating of the hidden state, and the calculation formula of its internal logical structure is as follows:
[0145] R t = σ(X t W xr + H t-1 W hr + b r );
[0146] Z t = σ(X t W xz + H t-1 W hz + b z );
[0147]
[0148]
[0149] R t 、Z t 、 and H t are the reset gate, update gate, candidate hidden state, and hidden state of the GRU model respectively. W xr 、W hr 、W xz 、W hz 、W xh 、W hh 、b r 、b z 、b hThey are the reset gate weight parameter, update gate weight parameter, candidate hidden state weight parameter, reset gate bias parameter, update gate bias parameter, and candidate hidden state bias parameter respectively. ⊙ represents the element-wise product operator, and σ is the sigmoid function, that is:
[0150]
[0151] tanh is the hyperbolic tangent function, which is:
[0152]
[0153] The input of the GRU model is X t , and the output of the final model is H t . Next, put the output of the GRU into the Attention mechanism.
[0154] The Attention mechanism assigns different weights to feature vectors, pays sufficient attention to important features, and ignores irrelevant information, thereby highlighting relatively key features. To implement the Attention mechanism, first calculate the attention probability distribution value e i , and its calculation method is:
[0155] e i = U · tanh(W eh H i + b e );
[0156] Among them, H t represents the output of the GRU network, U and W eh are weight matrices, b e is the bias term, and tanh is the hyperbolic tangent activation function. Then calculate the matching score a t for each time step. Use softmax to normalize the matching degree between the outputs of all time steps and the attention probability distribution value e i at the current time step to obtain the matching scores of each time step for all time steps:
[0157] a t = σ(e i );
[0158] Finally, the weighted sum of the outputs H t of each time step and the matching score a t gives the output S t of the Attention layer:
[0159]
[0160] Then the output of Attention is placed in the fully connected layer, and the predicted result can be output after calculation in the fully connected layer.
[0161] The training process of the GRU-Attention model is actually the problem of solving the parameters in the forward propagation process mentioned above. After the forward propagation has an output, the loss function can be used to calculate the loss between the predicted value and the accurate value. The loss function is the mean square error (MSE), and its calculation formula is:
[0162]
[0163] After calculating the loss value, backpropagation is required to modify the weights and biases of the neural network using the Adam optimizer to minimize the loss. The number of training rounds is calculated according to the optimal parameters obtained in step 6 above. After training for so many rounds, you can stop training, save the model, and perform variational mode decomposition on the data set of the lth load day category to obtain the rth component. The trained model is recorded as This is the rth component after the variational modal decomposition of the data set of the lth load category Similarly, in the process of training the model, all components of all categories need to be trained in step 7, and all trained models need to be saved.
[0164] Taking the production load power of a cement plant in Northwest China in 2022 as an example, we first filter out bad points in the load data to avoid affecting the data normalization operation. After removing the bad points, we get the data set U = {X, Y}. We perform clustering on the data set and calculate the silhouette coefficient after dividing it into multiple load day categories. Figure 4 See Figure 4 ,When the original load data is divided into two load day categories, the ,silhouette coefficient obtained by clustering is closest to 1, and the result k = 2.
[0165] Figure 5a This is a clustering result diagram of the first load day category provided by an embodiment of the present invention. Figure 5b This is a clustering result diagram of the second load day category provided by the embodiment of the present invention. Figure 5a-5b , of the two load day categories obtained by clustering, the data of the first load day category is relatively stable, while the data of the second load day category is more volatile. According to the clustering results, the original data set U = {X, Y} is reconstructed to obtain the data set of load day categories The original data set U = {X, Y} is classified to obtain the data set U (1) and U (2) .
[0166] Then, for the dataset U (1) and U (2) search for the optimal parameters of VMD decomposition respectively. Taking the dataset U (1) as an example, Figure 6a is a schematic diagram of the change of the envelope entropy value during the search for the optimal parameters of VMD decomposition provided by an embodiment of the present invention, Figure 6b is a schematic diagram of the change of the penalty factor during the search for the optimal parameters of VMD decomposition provided by an embodiment of the present invention, Figure 6c is a schematic diagram of the change of the number of decompositions during the search for the optimal parameters of VMD decomposition provided by an embodiment of the present invention. The obtained result is that the optimal parameters for the dataset of the first load day category are [M (1)* , α (1)* = [3, 2023], and the optimal parameters for the dataset of the second load day category are [M (2)* , α (2)* = [3, 2067]. According to these parameters, perform VMD decomposition on the first and second load day categories respectively.
[0167] Figure 7 is a schematic diagram of the VMD decomposition result provided by an embodiment of the present invention. As Figure 7 shown, the components obtained by decomposing the load quantity sequence under the first load day category are The components obtained by decomposing the load quantity sequence under the second load day category are
[0168] Figure 8a is a schematic diagram of the change of the fitness function during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention, Figure 8b is a schematic diagram of the change of the batch size during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention, Figure 8c is a schematic diagram of the change of the number of nodes in the fully connected layer during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention, Figure 8d is a schematic diagram of the change of the learning rate during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention, Figure 8e is a schematic diagram of the change of the number of training times during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention, Figure 8f is a schematic diagram of the change of the number of hidden nodes during the search for the optimal training parameters of the GRU-Attention network model provided by an embodiment of the present invention. Please refer to Figure 8a-8f , and the obtained optimal parameter set for this component data is
[0169] Optimization operations are performed on other categories and other components, and the results are as follows: The optimal parameter set for the second component of the first load day category is [0.00789, 43, 68, 8, 3], the optimal parameter set for the third component of the first load day category is [0.00856, 43, 73, 6, 3], the optimal parameter set for the first component of the second load day category is [0.00621, 48, 64, 8, 5], the optimal parameter set for the second component of the second load day category is [0.00921, 45, 74, 12, 7], and the optimal parameter set for the third component of the second load day category is [0.00941, 43, 69, 6, 3].
[0170] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A short-term load forecasting method based on whale optimization algorithm and type recognition for VMD-GRU, characterized in that, Including: Obtain the day to be predicted before characteristic quantities of load days and the day to be predicted before load day categories of load days to form features ; Input the feature into a pre-trained random forest classification model to obtain the load day category of the day to be predicted , , indicating types of load day categories; Input the load amounts and characteristic quantities of the load days before the to-be-predicted day into each pre-trained component model under each load day category, and determine the current prediction result corresponding to each load day category according to the outputs of each pre-trained component model under each load day category; Obtaining the accuracy rate obtained by testing the random forest classification model using a validation set during training; According to the accuracy rate and the load day category Determine the weights, and weight the current prediction results corresponding to all load day categories to obtain the load prediction result for the day to be predicted ; The component model is trained according to the following steps: Collect at a time resolution of 15 minutes N sample data containing sample feature quantities and load quantities to construct a first data set , where the feature quantity data set , represents containing sample feature quantities, and the load quantity data set , represents the load quantities corresponding to respectively; Convert the first data set into a second data set in days , where and respectively represent a set of feature quantities in days and a data set of load quantities in days, respectively represent sample feature quantities in days, respectively represent the load quantities in days corresponding to ; For the second data set perform clustering to obtain types of load days. After that, based on the second data set generate a load quantity sequence and a feature quantity sequence for each type of load day, and perform variational mode decomposition on each load quantity sequence to obtain multiple IMF components; According to the characteristic quantity sequence and the IMF components, a component data set is constructed, and each component model under each load day category is respectively trained by using the component data set; For the second data set perform clustering to obtain types of load days. After that, based on the second data set generate a load quantity sequence for each type of load day, and perform variational mode decomposition on each load quantity sequence to obtain multiple IMF components, including: Use the K-means algorithm to cluster the load data set in units of days to obtain types of load day categories; Based on the second data set Generate a load quantity sequence for each load day category; Using the whale optimization algorithm, search for the optimal parameters when performing variational mode decomposition on the load quantity sequence under each load day category. The optimal parameters include: the number of modal components M and the penalty factor ; Based on the optimal parameters, perform variational mode decomposition on the load quantity sequence under each load day category to obtain multiple IMF components.
2. The short-term load forecasting method based on whale optimization algorithm and type recognition of VMD-GRU according to claim 1, characterized in that Obtain the day to be predicted before characteristic quantities of load days and the day to be predicted before steps of load day categories of load days, including: Obtain the to-be-predicted date separately The characteristic quantities of the first 3 load days , the characteristic quantities of the first 2 load days and the characteristic quantities of the first 1 working day ; Obtain the days to be predicted respectively The load day categories of the first 5 load days , the load day categories of the first 4 load days , the load day categories of the first 3 load days , the load day categories of the first 2 load days and the load day category of the first 1 working day ; wherein, the characteristic quantity includes the illumination amplitude of the load day G , ambient temperature T , relative humidity H , wind speed V ; Based on the day to be predicted The feature quantities of the previous 3 load days and the day to be predicted The load day categories of the previous 5 load days are used to generate features .
3. The short-term load forecasting method based on whale optimization algorithm and type recognition according to claim 2, characterized in that Input the load amounts and characteristic quantities of the load days before the prediction date into each pre-trained component model under each load day category, and determining the current prediction result corresponding to each load day category according to the outputs of each pre-trained component model under each load day category, includes: Obtain the day to be predicted The load amounts and characteristic amounts of the previous 20 load days form a matrix ; Obtain each pre-trained component model under a load day category, where the component model is a GRU-Attention model; Convert the said matrix into a tensor form , and input it into each pre-trained component model under each load day category; Accumulating the outputs of each pre-trained component model under each load day category to obtain the current prediction result corresponding to each load day category.
4. The short-term load forecasting method based on whale optimization algorithm and type recognition for VMD-GRU according to claim 3, wherein According to the accuracy rate and the load day category determine weights, and weight the current prediction results corresponding to all load day categories to obtain the load prediction result of the day to be predicted The steps include: Obtain the load day category The corresponding current prediction result ; Obtain other load day categories The corresponding current prediction result , where ; Weighting the current prediction results corresponding to all load day categories according to the following formula: ; In the formula, represents the accuracy obtained by testing the random forest classification model using a validation set during the training process, represents the day to be predicted and the load prediction result.
5. The short-term load forecasting method based on whale optimization algorithm and type recognition of VMD-GRU according to claim 1, characterized in that According to the characteristic quantity sequence and the IMF step of constructing a component data set based on components and respectively training each component model under each load day category by using the component data set, includes: Construct component datasets for training the corresponding GRU-Attention neural network model respectively according to the feature quantity sequence and the IMF components wherein, represents the th component obtained by variational mode decomposition of the load quantity sequence under the th IMF load day category, represents the feature quantity sequence under the th load day category, ; Using the component dataset , for the th load day category, after variational mode decomposition of the load volume sequence, the th IMF component corresponding GRU-Attention neural network model is trained to obtain the th load day category component models.
6. The short-term load forecasting method based on whale optimization algorithm and type recognition of VMD-GRU according to claim 5, wherein Using the component data set , for the th load day category, after performing variational mode decomposition on the load quantity sequence, before the steps of training the th IMF component corresponding GRU-Attention neural network model, it further includes: Search using the whale optimization algorithm The optimal training parameters of the corresponding GRU-Attention neural network model.
7. The short-term load forecasting method based on whale optimization algorithm and type recognition of VMD-GRU according to claim 6, characterized in that, The training parameters include: learning rate, number of iterations, number of selected samples, number of hidden layer neuron nodes, and number of nodes in the fully connected layer.
8. The short-term load forecasting method based on whale optimization algorithm and type recognition of VMD-GRU according to claim 1, wherein The current prediction result corresponding to each load day category is obtained by weighted summation of the outputs of each pre-trained component model under this load day category.
Citation Information
Patent Citations
Short-term load prediction method for optimizing SVM based on MWOA algorithm
CN110516831A
Short-term power load prediction method based on whale optimization algorithm and DRESN
CN116681159A