Maximum load prediction method based on EMD-KPCA-improved BP neural network
By combining EMD, KPCA and improved BP neural networks in power load prediction, a nested prediction model is formed, which solves the problem that traditional load prediction methods are difficult to deal with nonlinear changes and emergencies, and improves prediction accuracy and adaptability.
Patent Information
- Application Number
- CN202510131931.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Traditional load prediction methods rely on historical data and linear models, and are difficult to effectively deal with nonlinear changes and emergencies, and cannot adapt to the rapid changes in power demand and the uncertainty of new energy access.
The maximum load prediction method based on EMD-KPCA-improved BP neural network is adopted, and key features are extracted through EMD decomposition and KPCA analysis, and combined with LSTM network and BP neural network for prediction, forming a nested prediction model.
The accuracy of maximum power load data prediction is improved, and more accurate data preprocessing methods and improved BP neural network prediction model are provided, which can effectively handle nonlinear changes and emergencies and adapt to the rapid changes in power demand.
Smart Images

Figure CN120124786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, and in particular to a maximum load prediction method based on EMD (Empirical Mode Decomposition)-KPCA (Kernel Principal Component Analysis)-improved BP (Back Propagation) neural network. Background Art
[0002] Power load prediction is an important part of power system planning and operation, and plays an irreplaceable role in ensuring the safe, stable and economic operation of the power system. Accurate load prediction can help grid operators reasonably arrange the start-stop plans of generating units, ensure that there is enough reserve capacity to cope with emergencies at different time periods, and thus maintain the continuity and stability of power supply. Load prediction is the basis for dispatching decisions on the power generation side and is conducive to promoting the orderly development of the power market.
[0003] Due to the characteristics of strong intermittency and large fluctuations of renewable energy such as wind energy and solar energy, traditional load prediction methods rely on historical data and linear models, and their performance is limited when dealing with non-linear changes and emergencies, making it difficult to adapt to the rapid changes in current power demand and the uncertain load prediction requirements of new energy access. Traditional load prediction can no longer meet the actual needs. Summary of the Invention
[0004] In view of the above analysis, an embodiment of the present invention aims to provide a maximum load prediction method based on EMD-KPCA-improved BP neural network to solve the technical problems that traditional load prediction methods rely on historical data and linear models, have limited performance when dealing with non-linear changes and emergencies, and are difficult to adapt to the rapid changes in power demand and the uncertainty of new energy access.
[0005] The object of the present invention is mainly achieved through the following technical solutions:
[0006] The present invention provides a maximum load prediction method based on EMD-KPCA-improved BP neural network, including the following steps:
[0007] Obtain the basic data affecting the power system load and the corresponding maximum load data in multiple historical periods in the region, and perform data preprocessing including EMD decomposition and linear kernel KPCA analysis in sequence, and generate a first training sample set based on the preprocessed basic data and maximum load data;
[0008] Train the LSTM network using the first training sample set, and use the trained LSTM network to obtain the predicted values of each variable in the basic data and the corresponding maximum load data in the future time period. After performing reverse linear kernel KPCA analysis and EMD decomposition processing on the predicted values in the future time period in sequence, obtain the predicted basic data and maximum load data;
[0009] The basic data and the corresponding maximum load data that affect the power system load within the historical period, combined with the predicted basic data and maximum load data, generate a second training sample set; use the second training sample set to train the BP neural network to obtain a trained power load prediction model; wherein, the power load prediction model is a BP neural network nested with an LSTM network;
[0010] Obtain the basic data that affects the power system load in the area in real time and input it into the trained power load prediction model to obtain the predicted value of the maximum load data of the regional power system.
[0011] Furthermore, perform EMD data preprocessing on each variable of the basic data and the maximum load data in each historical period in sequence, including:
[0012] The first step: Define the m historical data of each variable obtained as the original time series x(t), and the residual component r 0 (t)=x(t), and initialize the iteration number i = 1;
[0013] The second step: In each iteration, after extracting an IMF from the current residual component, update the current residual component as the residual component after the previous iteration minus the currently extracted IMF;
[0014] The third step: Find all the local maximum points and local minimum points in the current residual component r i-1 (t);
[0015] The fourth step: Use cubic spline interpolation to fit all the local maximum points and local minimum points of the current residual component respectively to obtain the upper envelope e max (t) and the lower envelope e min (t), and calculate the average value of the upper and lower envelopes
[0016] The fifth step: Subtract the average value of the upper and lower envelopes from the current residual component to obtain the component h i (t)=r i-1 (t)-m i (t); if the component h i (t) meets the preset conditions, then use it as the i-th IMF component c i(t), otherwise, it will be regarded as a new residual component and the second to fifth steps will be repeated until the i-th IMF component is obtained;
[0017] Step 6: Based on the obtained i-th IMF component, update the residual component r i (t) = r i-1 (t) - c i (t); If r i (t) is not a monotonic function or a constant, then i = i + 1, and return to the second step to continue the process. Otherwise, stop the EMD decomposition process, and the m data of each variable in each historical period are expressed as the sum of p IMF components and the residual signal of the last iteration As the IMF components of each variable in each historical period;
[0018] For the n variables of the basic data and the corresponding maximum load data in each historical period, each containing m data, after EMD decomposition, p×n×m IMF components are obtained.
[0019] Furthermore, perform linear kernel KPCA analysis on the p×n×m IMF components, including:
[0020] For the p×n×m IMF variables as sample points C = [c 1 , c 2 ,... c pnm , construct the kernel matrix K of the sample points, and its matrix element K ij is expressed as K ij = C T *C = k(c i , c j );
[0021] Construct the centered kernel matrix K' = HKH, where, I is the identity matrix, and B is the all-ones vector;
[0022] Perform eigenvalue decomposition on the centered kernel matrix K', and obtain the eigenvalues λ 1 , λ 2 ,... λ N and their corresponding eigenvectors α 1 , α 2 ,..., α n ;
[0023] Select the first k largest eigenvalues and the corresponding eigenvectors as the principal components, and obtain the principal component values after projection in the principal component space where, α j is the eigenvector corresponding to the j-th principal component, λ j is the corresponding eigenvalue, and c new is a sample point in C.
[0024] Further, the principal component values corresponding to the basic data and the principal component values corresponding to the maximum load data are used as labels to form sample data, and all the sample data forms a first training sample set.
[0025] Further, training the LSTM network using the first training sample set includes:
[0026] Initializing the weights and biases of the LSTM network;
[0027] Inputting the principal component values corresponding to each historical data in the first training sample set into the LSTM network to obtain predicted values, and comparing them with the corresponding true values in the first training sample set to construct a loss function;
[0028] Performing training using the gradient descent method and updating the weights and biases;
[0029] When the loss function is less than the first preset value, stop the iteration, determine the weight matrices and bias vector values of the input gate, forget gate, and output gate, and obtain the trained LSTM network.
[0030] Further, during the training process of the BP neural network using the second training sample set, an adaptive weighted composite loss function is used, as follows:
[0031]
[0032]
[0033] L total = μ 1 L 1 + μ 2 L 2 ;
[0034] where x m,i is the i-th true value, z 2,i is the i-th predicted value, f is the number of samples in the second training sample set, and μ 1 , μ 2 are the adaptive weights of the two loss functions L 1 and L 2 respectively.
[0035] Further, after each iteration of training the BP neural network, update μ 1 , μ 2 , as follows:
[0036]
[0037] is L at the t-th iteration.1 The adaptive weight of the function is the L at the t-th iteration 1 value of the loss function is the L at the t-th iteration 2 value of the loss function; is the L at the (t + 1)-th iteration 1 value of the adaptive weight of the function is the L at the (t + 1)-th iteration 2 value of the adaptive weight of the function;
[0038] After each iteration update, calculate the total loss value L in the current iteration process total .
[0039] Furthermore, in each iteration process of training the BP neural network, weight and bias updates are performed as follows:
[0040]
[0041]
[0042] where η is the learning rate, and W 1 , W' 1 are the weight matrices before and after the update from the input layer to the hidden layer respectively, and W 2 , W' 2 are the weight matrices before and after the update from the hidden layer to the output layer respectively, b 1 is the bias vector of the hidden layer, and b 2 is the bias vector of the output layer.
[0043] Furthermore, the loss function of the LSTM network is as follows:
[0044]
[0045] where c is the sample data volume, y zs,i is the i-th true value, h t,i is the corresponding i-th predicted value, and δ is the threshold parameter;
[0046] The threshold parameter δ is determined by the quantile method.
[0047] Furthermore, the preset conditions include: within m data of each variable of the basic data and the maximum load data, the sum of the number of local maximum points and the number of local minimum points is equal to or at most differs by one from the number of zero-crossing points, and the average value of the upper envelope formed by the local maximum points and the lower envelope formed by the local minimum points is zero at any time.
[0048] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0049] 1. By combining EMD, KPCA, and an improved BP neural network, the present invention can effectively handle non-linear changes and emergencies, and improve the accuracy of maximum power load data prediction.
[0050] 2. A more precise data preprocessing method is provided. Through EMD decomposition for the first data preprocessing and KPCA dimensionality reduction for the second data preprocessing, the idea of raising and lowering the dimension of the original data is used to screen out the heterogeneous features of the data and the relevant historical data that have the most significant impact on the dependent variable to the greatest extent. It can effectively extract the key features in multi-variable data, remove redundant information and noise, reduce the dimension of the data, and retain the main information, providing a more refined data basis for subsequent prediction steps.
[0051] 3. An improved BP neural network prediction model is provided. By nesting an LSTM network in the BP neural network for prediction, various factors affecting power load are comprehensively considered, providing a richer fitting data set and improving the prediction accuracy of the power load prediction model for the maximum load data of the power system.
[0052] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained from the content specifically pointed out in the specification and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings are only for the purpose of showing specific embodiments and are not considered as limiting the present invention. Throughout the drawings, the same reference signs represent the same components.
[0054] Figure 1 It is a flowchart of a maximum load prediction method based on EMD-KPCA-improved BP neural network in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0056] Compared with the existing methods, the present invention first increases the preprocessing of the original data to determine the data in the historical data of various influencing factors that significantly affect the power load change. Secondly, an improved BP neural network is used, considering the influence effects of various influencing factors on the power load, thereby improving the prediction accuracy of the power load.
[0057] A specific embodiment of the present invention discloses a maximum load prediction method based on EMD-KPCA-improved BP neural network, which includes the following steps:
[0058] Step S1: Obtain the basic data affecting the power system load and the corresponding maximum load data in multiple historical periods in the region, and perform data preprocessing including EMD decomposition and linear kernel KPCA analysis in sequence. Generate the first training sample set based on the preprocessed basic data and maximum load data;
[0059] Step S2: Use the first training sample set to train the LSTM network. Use the trained LSTM network to obtain the predicted values of each variable in the basic data and the corresponding maximum load data in the future time period. After performing reverse linear kernel KPCA analysis and EMD decomposition processing on the predicted values in the future time period in sequence, obtain the predicted basic data and maximum load data;
[0060] Step S3: Combine the basic data affecting the power system load and the corresponding maximum load data in the historical period with the predicted basic data and maximum load data to generate the second training sample set; use the second training sample set to train the BP neural network to obtain the trained power load prediction model; wherein, the power load prediction model is a BP neural network nested with an LSTM network;
[0061] Step S4: Real-time obtain the basic data affecting the power system load in the region and input it into the trained power load prediction model to obtain the predicted value of the maximum load data of the regional power system.
[0062] Step S1 includes steps S11-S14.
[0063] Step S11: Obtain the basic data affecting the power system load and the corresponding maximum load data in multiple historical periods in the region.
[0064] Determine the basic data and the corresponding maximum load data affecting the power system load situation in multiple historical periods in the regional power system. The basic data includes short-term climate change characteristics and long-term change characteristics within the same specific time period t;
[0065] The short-term climate change characteristics are the climate characteristics affecting the power system load, specifically including the highest temperature, lowest temperature, average temperature, highest humidity, lowest humidity, and average humidity in the same time period as the maximum load data in the historical period in this region;
[0066] The long-term change characteristics are the characteristics affecting the load in the same time period as the maximum load data in the historical period. Exemplarily, such as the total industrial output value and the installed capacity of power plants.
[0067] Obtain the basic data for multiple time periods in a certain area, as well as the corresponding maximum load data.
[0068] The above data is sourced from the data published by the National Bureau of Statistics, local statistical bureaus, the National Energy Statistical Yearbook, local energy statistical yearbooks, the National Meteorological Bureau or local meteorological bureaus.
[0069] The maximum load data within the historical period is obtained through the National Bureau of Statistics, the National Energy Statistical Yearbook, and the statistical bureau or energy statistical yearbook of this region.
[0070] The highest temperature, lowest temperature, average temperature, highest humidity, lowest humidity, and average humidity in the same time period of this region are obtained through the National Meteorological Bureau or the meteorological bureau of this region.
[0071] The total industrial output value and installed power generation capacity in the same time period of this region are obtained through the National Bureau of Statistics, the National Energy Statistical Yearbook, and the statistical bureau or energy statistical yearbook of this region.
[0072] Assume that a total of n variables of basic data and corresponding maximum load data are finally obtained in this step, and each variable contains m pieces of data.
[0073] In the present invention, n - 1 pieces of basic data obtained are, by way of example, as shown in Table 1, and the basic data is 8.
[0074] The above-obtained basic data can be expanded or reduced according to the needs of the research.
[0075] The maximum load data corresponding to the basic data also includes m pieces of corresponding data within a historical period.
[0076] Table 1: Exemplary description of obtaining n variables of basic data and corresponding maximum load data
[0077]
[0078] The function of step S11 is to collect and determine the basic data and corresponding maximum load data within the historical period that affect the power system load, providing a data basis for subsequent data preprocessing and training of the power load prediction model.
[0079] Next, perform data preprocessing on the n variables of the obtained basic data and corresponding maximum load data, including the first data preprocessing (EMD data preprocessing in step S12), and the second data preprocessing (principal component analysis data preprocessing in step S13).
[0080] Step S12: For the n (n≥9) variables of the basic data affecting the power system load and the corresponding maximum load data determined in Step S11, perform EMD decomposition on the data set containing m (m≥1) historical data for each variable, and conduct the first data preprocessing.
[0081] The purpose of the first data preprocessing is to decompose the m historical data of each variable of the above basic data and the corresponding maximum load data, so that the historical data of each variable can be decomposed into p IMF (Intrinsic Mode Function) components, facilitating the second data preprocessing of principal component analysis in Step S13.
[0082] Exemplarily, p≥5.
[0083] Perform EMD data preprocessing on each variable of the basic data and the maximum load data in each historical cycle in sequence, including:
[0084] First step: Define the m historical data of each variable obtained as the original time series x(t), and the residual component r 0 (t)=x(t), and initialize the iteration number i = 1;
[0085] Second step: In each iteration, after extracting an IMF from the current residual component, update the current residual component as the residual component after the previous iteration minus the currently extracted IMF;
[0086] Third step: Find all the local maximum points and local minimum points in the current residual component r i-1 (t);
[0087] Fourth step: Use cubic spline interpolation to fit all the local maximum points and local minimum points of the current residual component respectively, to obtain the upper envelope e max (t) and the lower envelope e min (t), and calculate the average value of the upper and lower envelopes
[0088] Fifth step: Subtract the average value of the upper and lower envelopes from the current residual component to obtain the component h i (t)=r i-1 (t)-m i (t); If the component h i (t) meets the preset conditions, then regard it as the i-th IMF component c i (t), otherwise regard it as the new residual component and repeat the second to fifth steps until the i-th IMF component is obtained;
[0089] Sixth step: Based on the obtained i-th IMF component, update the residual component r ir(t) = r i-1 (t) - c i (t); If r i (t) is not a monotonic function or a constant, then i = i + 1, and return to the second step to continue processing. Otherwise, stop the EMD decomposition process, and obtain m data in each historical period of each variable, which is expressed as the sum of p IMF components and the residual signal of the last iteration As the IMF components of each variable in each historical period;
[0090] For the n variables of the basic data and the corresponding maximum load data in each historical period, each containing m data, after EMD decomposition, p × n × m IMF components are obtained.
[0091] First step, specifically.
[0092] Perform EMD decomposition on the data set of each variable containing m (m ≥ 1) historical data. Taking the m data of the first variable (for example: average temperature) as an example, let the m data sequences be x(t), and let the initial residual component be r 0 (t) = x(t), and initialize the number of iterations;
[0093] At the beginning of the EMD decomposition, the residual component is equal to the original data sequence, and x(t) is the time series data of the original data.
[0094] During the EMD decomposition process, the residual component represents the residual component after extracting an intrinsic mode function, contains the original data that has not been decomposed, and combines all parts of the original data that have not been extracted as IMF.
[0095] Second step, specifically.
[0096] In each iteration, after extracting an IMF from the current residual component, update the residual component to the current residual component minus the newly extracted IMF, as shown in formula (1):
[0097] r i-1 (t) = r i-2 (t) - IMF i-1 Formula (1);
[0098] Among them, r i-1 (t) is the current residual component, r i-2 (t) is the residual component of the (i - 2)-th iteration, and IMF i-1 is the IMF extracted in the current iteration.
[0099] The significance of the residue component in EMD decomposition lies in allowing the decomposition of complex non-linear and non-stationary time series data into a series of simple IMF components with different frequencies. These IMF components can be analyzed independently or used for further data processing and analysis. The residue component is an intermediate product in the EMD decomposition process and is updated with each extraction of an IMF until the decomposition process ends.
[0100] Step 3, specifically.
[0101] For the current residue component r i-1 (t), find all local maximum points and local minimum points.
[0102] Exemplarily, it is implemented through the findpeaks function built into matlab. The code is as follows:
[0103] % Assume the signal data is x(t) and the time axis is t
[0104] t(1:500); % Time axis
[0105] x(t)=fault_data(:,1); %
[0106] x(t)=x(t)(1:500,:)'; % Ensure x(t) is a row vector
[0107] This code first defines the time axis t and the signal data x(t), and then calls the findpeaks function respectively to obtain the indices of the local maximum points and local minimum points.
[0108] Among them, for the local minimum points, the problem of finding local maximum points is converted by taking the negative of the signal.
[0109] % Solve for local maximum points and local minimum points
[0110] [~,indmax]=findpeaks(x(t)); % Index of local maximum points
[0111] [~,indmin]=findpeaks(-x(t)); % Index of local minimum points
[0112] Step 4, specifically.
[0113] Use cubic spline interpolation to fit all local maximum points and local minimum points of the current residue signal respectively to obtain the upper envelope e max (t) and the lower envelope e min (t).
[0114] Exemplarily, it is implemented through the spline function in matlab. The code is as follows:
[0115] % Cubic spline interpolation fitting (including endpoints)
[0116] s_max = spline([1, indmax, length(y)], [y(1), y(indmax), y(end)], 1:length(y)); % Code for the upper envelope
[0117] s_min = spline([1, indmin, length(y)], [y(1), y(indmin), y(end)], 1:length(y)); % Code for the lower envelope
[0118] The first argument of the spline function is a vector containing the indices of all the points participating in the interpolation; the second argument is the corresponding signal values; the third argument is the time axis specifying where to calculate the interpolation result.
[0119] Finally, calculate the average value of the upper and lower envelopes.
[0120] Step 5, specifically.
[0121] Judge whether the component h i (t) obtained by subtracting the average value of the upper and lower envelopes from the current residual component satisfies the preset conditions, and determine whether to repeat steps 2 to 5 to further decompose the IMF components.
[0122] The preset conditions include: within each variable of the basic data and m data of the maximum load data, the sum of the number of local maximum points and the number of local minimum points is equal to or at most one different from the number of zero-crossing points, and the average value of the upper envelope formed by the local maximum points and the lower envelope formed by the local minimum points is zero at any time.
[0123] The average value of the upper envelope formed by the local maximum values and the lower envelope formed by the local minimum values is zero at any time, that is, the upper and lower envelopes are locally symmetric about the time axis.
[0124] Step 6, specifically.
[0125] When the remaining residual component r n (t) is a monotonic function or a constant, the EMD decomposition ends.
[0126] At this time, the m data of the first variable are expressed as the sum of p IMF components and the final residual component, as follows:
[0127]
[0128] Repeat the process from the first step to the sixth step for n variables, obtaining m pieces of data in each of the n variables. The sum of the p IMF components decomposed from the m pieces of data plus the final residual signal is the result of EMD decomposition for each variable.
[0129] After EMD decomposition, corresponding to the n variables, p×n×m components are obtained.
[0130] The function of step S12 is that the historical data of each variable of the basic data and the corresponding maximum load data are decomposed into multiple simple IMF components. Each variable is expressed as the sum of multiple IMF components and the final residual component, which is used for subsequent kernel principal component analysis (KPCA). This decomposition helps to explain the internal patterns and trends in the data and provides a more refined data basis for maximum load prediction.
[0131] Step S13: For the data after EMD decomposition, use the linear kernel KPCA (kernel principal component analysis) method to perform the second data preprocessing.
[0132] Map the data after EMD decomposition to a high-dimensional space, and then perform linear kernel KPCA analysis in the high-dimensional space to determine the component values that significantly affect the historical daily maximum load data among the p×n×m components obtained after EMD decomposition.
[0133] Perform linear kernel KPCA analysis on the p×n×m IMF components, including:
[0134] For the p×n×m IMF variables as sample points C = [c 1 , c 2 ,... c pnm , construct the kernel matrix K of the sample points, and its matrix element K ij is expressed as K ij = C T * C = k(c i , c j );
[0135] Construct the centered kernel matrix K' = HKH, where I is the identity matrix and B is the all-ones vector;
[0136] Perform eigenvalue decomposition on the centered kernel matrix K' to obtain the eigenvalues λ 1 , λ 2 ,... λ N and their corresponding eigenvectors α 1 , α 2 ,...,, α n ;
[0137] Select the top k largest eigenvalues and their corresponding eigenvectors as the principal components, and obtain the principal component values after projection in the principal component space. Among them, α j is the eigenvector corresponding to the j-th principal component, and λ j is the corresponding eigenvalue, and c new is a sample point in C.
[0138] The element K of the kernel matrix K ij can be expressed as K ij = C T * C = k(c i , c j ), that is, the result of multiplying the sample matrix C by its transpose C T . Apply K' = HKH to centralize the kernel matrix.
[0139] Perform eigenvalue decomposition on the centralized kernel matrix K'. Find all the eigenvalues of K' and their corresponding eigenvectors. These eigenvectors will be used in the subsequent dimensionality reduction process. Obtain the eigenvalues λ 1 , λ 2 ,... λ N and their corresponding eigenvectors α 1 , α 2 ,..., α n .
[0140] Specifically, solve the following system of equations: K'α = λα. Solving this homogeneous linear system of equations can be done by using ready-made library functions, such as numpy.linalg.eigh or scipy.linalg.eigh in Python. Among them, α is the eigenvector and λ is the corresponding eigenvalue.
[0141] Sort according to the magnitude of the eigenvalues, and select the eigenvectors corresponding to the top k largest eigenvalues as the principal components. In this model, it is set to select all the eigenvalues that can explain 95% of the data variation as the principal components.
[0142] In linear kernel KPCA principal component analysis, the value of k depends on the cumulative explained variance of the data. Exemplarily, the general steps to determine the value of k are as follows:
[0143] (1) Calculate the eigenvalues: Perform eigenvalue decomposition on the kernel matrix to obtain a series of eigenvalues.
[0144] (2) Cumulative contribution rate: Calculate the cumulative contribution rate of the variance of each eigenvalue. This is usually achieved by dividing each eigenvalue by the sum of all eigenvalues.
[0145] (3) Accumulated contribution rate: Starting from the largest eigenvalue, accumulate the variance contribution rate of each eigenvalue until the cumulative contribution rate reaches or exceeds 95%.
[0146] (4) Determine the value of k: The serial number corresponding to the last eigenvalue used in the accumulation process is the value of k.
[0147] Exemplarily, if the cumulative variance contribution rate of the first 3 largest eigenvalues is 90%, and the cumulative variance contribution rate of the first 4 eigenvalues is 96%, then the value of k should be 4, because this is the smallest integer that makes the cumulative variance contribution rate reach 95%.
[0148] In practical applications, the value of k may vary from 2 to 10 or more, depending on the characteristics of the dataset and the degree of information retention required. Usually, it is necessary to calculate and compare the cumulative variance contribution rates under different k values to determine the most appropriate k value.
[0149] Project the total of p×n×m IMF components obtained after EMD decomposition of the n variables of the basic data and the corresponding maximum load data, with each variable containing m pieces of data, into the kernel function. For one sample point c new , its value after projection in the principal component space is where α j is the eigenvector corresponding to the jth principal component, and λ j is the corresponding eigenvalue.
[0150] Step S14: Generate the first training sample set from the preprocessed basic data and maximum load data.
[0151] Use the principal component values corresponding to the basic data and the principal component values corresponding to the maximum load data as labels to form sample data, and all the sample data forms the first training sample set.
[0152] The first training sample set is used to train the LSTM network. 70% of this sample set is used as the training set, and 30% is used as the test set.
[0153] The function of step S1 is to collect and preprocess the historical data affecting the power system load, including climate characteristics, economic characteristics, and maximum load data, and extract key features through the EMD-KPCA method to prepare the first training sample set for constructing the power load prediction model.
[0154] Step S2: Specifically.
[0155] First, construct an LSTM network. The LSTM network includes three gates: the forget gate, the input gate, and the output gate, as well as a cell state.
[0156] The forget gate determines which information should be discarded, as follows:
[0157] f t = sigmoid(W 1*[h t-1 ,x t +b 1 ) Formula (3);
[0158] Among them, f t is the output of the forget gate, W 1 is the weight matrix of the output data of the forget gate, b 1 is the bias vector of the output data of the forget gate, h t-1 is the hidden state at the previous moment, x t is the current input.
[0159] The input gate determines which new information should be added to the cell state. The formula consists of two parts: the activation value of the input gate, as follows:
[0160] i t = sigmoid(W 2 *[h t-1 ,x t +b 2 ) Formula (4);
[0161] Among them, i t is the activation value, W 2 is the weight matrix of the activation value, b 2 is the bias vector of the activation value.
[0162] New candidate value, as follows:
[0163] C′ t = tanh(W 3 *[h t-1 ,x t +b 3 ) Formula (5);
[0164] Among them, C′ t is the new candidate value, W 3 is the weight matrix of the new candidate value, b 3 is the bias vector of the new candidate value.
[0165] The updated value of the cell state, as follows:
[0166] C t = f t *C t-1 + i t *C t Formula (6);
[0167] Among them, C t is the updated value of the cell state.
[0168] The output gate determines which part of the cell state is output, as follows:
[0169] o t = sigmoid(W 4 * [h t-1 , x t + b 4 ) Formula (7);
[0170] Among them, o t is the cell state output by the output gate, W 4 is the weight matrix of the cell output state, b 4 is the bias vector of the cell output state.
[0171] The final output value is as follows:
[0172] h t = o t * tanh(C t ) Formula (8);
[0173] Among them, h t is the output value of the LSTM model.
[0174] Training the LSTM network using the first training sample set includes:
[0175] Initializing the weights and biases of the LSTM network;
[0176] Inputting the principal component values corresponding to each historical data in the first training sample set into the LSTM network to obtain predicted values, and comparing them with the corresponding true values in the first training sample set to construct a loss function;
[0177] Performing training using the gradient descent method and updating the weights and biases;
[0178] When the loss function is less than the first preset value, stop the iteration, determine the weight matrix and bias vector values of the input gate, forget gate, and output gate, and obtain the trained LSTM network.
[0179] The loss function is used to measure the difference between the predicted output and the true label. Training is performed using the gradient descent method until the loss function L is less than the first set value, then stop the iteration and determine the weight matrix and bias vector values of the input gate, forget gate, and output gate at this time.
[0180] Set the initial weights and biases. The weights W i (i = 1, 2, 3, 4) are the weights of each link involved in this step. The weights are drawn from the standard normal distribution. Initialize the weights, set the weight initialization to 0.01, and the bias initialization to 0.
[0181] At this time, input the time series x corresponding to a historical data in the first training sample set tInput, a predicted value h of the output will be obtained t For each time series x t The predicted value h of the LSTM training model below t,i Compare with the corresponding true value to construct a loss value function.
[0182] The loss function of the LSTM network is as follows:
[0183]
[0184] Where c is the sample data volume, y zs,i is the i-th true value, h t,i is the corresponding i-th predicted value, and δ is the threshold parameter;
[0185] The threshold parameter δ is determined by the quantile method.
[0186] Then perform gradient calculation as shown below:
[0187] For the forget gate:
[0188]
[0189] For the input gate:
[0190]
[0191] Where i t is the activation value of the input gate.
[0192] For the candidate memory:
[0193]
[0194] For the output gate:
[0195]
[0196]
[0197] For the cell state:
[0198]
[0199] Where C t is the value of the cell state after update, C t+1 is the value of the cell state updated by the combined action of the forget gate and the input gate, combining the cell state of the previous moment and the new information of the current moment.
[0200] Then perform weight and bias update, and the general formula is as follows:
[0201]
[0202] Wherein, W' is the updated weight; b' is the updated bias, and η is the learning rate.
[0203] Exemplarily, the learning rate η is set to 0.001.
[0204] Iteration stops until the loss function Loss is less than the first set value.
[0205] Exemplarily, the first preset value is set to 0.01; when the iteration stops, the weight matrices and bias vector values of the input gate, forget gate, and output gate at this time are determined to obtain a trained LSTM network.
[0206] Use the trained LSTM model to predict the predicted values of n variables in the future T time periods for the basic data and the corresponding maximum load data. Taking the first variable as an example for prediction, assume that the time series corresponding to the future T time periods is x 1 to x T , that is, the input x t where t = 1 to t = T, input x 1 , and at this time, the first predicted value h 1 is output.
[0207] Add h 1 to the original data sequence X 1 and remove the earliest input value to keep the length of the input sequence unchanged, which can be achieved through the sequence formula:
[0208] X t-1 = X t-1 [1:]+[h t (at this time t = 2), Formula (21);
[0209] Then repeat this step until the prediction for the future T time periods is completed, and the predicted values of the first variable in the future T time periods are obtained. Then, by analogy, the predicted values of the n - 1 variables of the other basic data and the corresponding maximum load data are obtained, where each variable has predicted values in the future T time periods.
[0210] Use the trained LSTM network to obtain the predicted values of each variable in the basic data and the corresponding maximum load data in the future time period. The predicted values obtained by using the trained LSTM network are the data after EMD decomposition and then linear kernel KPCA analysis for data preprocessing;
[0211] Perform reverse linear kernel KPCA analysis on the predicted values in the future time period, and then perform EMD decomposition processing to obtain the predicted basic data and maximum load data.
[0212] The function of step S2 is to train the LSTM network using the first training sample set that has undergone EMD decomposition and then linear kernel KPCA analysis for data preprocessing. By iteratively optimizing the network parameters until the loss function Loss of the LSTM network is less than the first preset value, a trained LSTM network is obtained. The trained LSTM network can accurately predict the basic data and maximum load data of the power system load in the future time period based on the basic data and maximum load data that affect the power system load within the historical period.
[0213] Step S3 includes steps S31 - S32.
[0214] Step S31: Generate the second training sample set.
[0215] The basic data and corresponding maximum load data that affect the power system load within the historical period, combined with the predicted basic data and maximum load data, are used to generate the second training sample set;
[0216] Step S32: Use the second training sample set to train the BP neural network to obtain a trained power load prediction model.
[0217] First, initialize the BP neural network. The BP neural network includes an input layer, a hidden layer, and an output layer.
[0218] The activation function of the hidden layer is tan - sigmoid, as follows:
[0219]
[0220] where x is the input sample of the BP neural network.
[0221] The output range of this function is (-1, 1), and it has good gradient characteristics near the origin, which helps to accelerate the training process. The activation function of the output layer is purelin linear mapping, as follows:
[0222] f(x) = x formula (23);
[0223] Then, set the initial weights and biases, as follows:
[0224] w ij ~N(0, σ 2 ) formula (24);
[0225] The weights are drawn from the standard normal distribution, set to 0.01 in this model, and the biases are initialized to 0. First, use the first 70% of the x uv and x m historical data as the training set, and the last 30% of the data as the test set.
[0226] Next, set the formulas for the input layer, hidden layer, and output layer. During the forward propagation process, the input data is passed layer by layer through the network until it reaches the output layer. The specific steps are as follows:
[0227] From the input layer to the hidden layer, as follows:
[0228] z 1 = W 1 x + b 1 Formula (25);
[0229] a 1 = tanh(z 1 ) Formula (26);
[0230] Among them, W 1 is the weight matrix from the input layer to the hidden layer, b 1 is the bias vector of the hidden layer, z 1 is the net input of the hidden layer, a 1 is the activation output of the hidden layer.
[0231] From the hidden layer to the output layer, as follows:
[0232] z 2 = W 2 * a 1 + b 2 Formula (27);
[0233] Among them, W 2 is the weight matrix from the hidden layer to the output layer, b 2 is the bias vector of the output layer, z 2 is the net input of the output layer and also the final output.
[0234] Train the BP neural network and predict the change trend of the power system load based on the BP neural network.
[0235] Construct the loss function of the BP neural network. This loss function is used to measure the difference between the predicted output and the true label. Use the gradient descent method for training until the loss function is less than the second set value, then stop the iteration and determine the weight matrix and bias vector from the input layer to the hidden layer, as well as the weight matrix and bias vector from the hidden layer to the output layer at this time.
[0236] During the training process of the BP neural network using the second training sample set, use the adaptive weighted composite loss function, as follows:
[0237]
[0238] L total = μ 1 L 1 + μ 2 L2 Formula (30);
[0239] where x m,i is the i-th true value, z 2,i is the i-th predicted value, f is the number of samples in the second training sample set, μ 1 , μ 2 are the adaptive weights of the two loss functions L 1 and L 2 respectively.
[0240] The adaptive weights μ 1 and μ 2 of the two loss functions L 1 , μ 2 are initially set to 0.5. μ 1 , μ 2 reflect the importance of the two loss terms in the current batch of data.
[0241] After each iteration of training the BP neural function network, update μ 1 , μ 2 as follows:
[0242]
[0243] is the adaptive weight of the L 1 function at the t-th iteration, is the value of the L 1 loss function at the t-th iteration, is the value of the L 2 loss function at the t-th iteration; is the value of the adaptive weight of the L 1 function at the (t + 1)-th iteration, is the value of the adaptive weight of the L 2 function at the (t + 1)-th iteration.
[0244] After each iteration update, calculate the total loss value L total in the current iteration process.
[0245] During each iteration of training the BP neural network, update the weights and biases as follows:
[0246]
[0247]
[0248] where η is the learning rate, W 1 , W′ 1 are the weight matrices before and after the update from the input layer to the hidden layer respectively, W 2 , W′2 are the weight matrices before and after the update from the hidden layer to the output layer, and b 1 is the bias vector of the hidden layer, and b 2 is the bias vector of the output layer.
[0249] Exemplarily, the learning rate η is set to 0.001.
[0250] When the total loss value is less than the second preset value, the iteration ends.
[0251] Exemplarily, the second preset value is set to 0.01. When the iteration ends, the weight matrix and bias vector from the input layer to the hidden layer, and the weight matrix and bias vector from the hidden layer to the output layer at this time are determined.
[0252] The function of step S3 is to use the second training sample set generated by combining historical data and LSTM network prediction data to train the BP neural network, optimize the BP neural network parameters through the adaptive weighted composite loss function and the gradient descent method, and finally obtain a BP neural network model of the nested LSTM network that can accurately predict the power load, as the trained power load prediction model.
[0253] Step S4, specifically.
[0254] Obtain the basic data affecting the load of the power system in the area in real time, input it into the trained power load prediction model, and obtain the predicted value of the maximum load data of the regional power system.
[0255] The power load prediction model is a BP neural network of a nested LSTM network, which is an improved BP neural network.
[0256] The original basic data in a specific historical period in the area obtained in real time is first input into the LSTM network therein to obtain the predicted basic data, and the predicted basic data is input into the BP neural network therein to obtain the predicted value of the maximum load of the power system in the future time period.
[0257] In summary, a method for predicting the maximum load based on EMD-KPCA-improved BP neural network according to an embodiment of the present invention has the following beneficial effects:
[0258] 1. By combining EMD, KPCA and the improved BP neural network, the present invention can effectively handle non-linear changes and emergencies, and improve the accuracy of predicting the maximum power load data;
[0259] 2. A more precise data preprocessing method is provided. Through EMD decomposition for the first data preprocessing and KPCA dimensionality reduction for the second data preprocessing, the idea of raising and lowering the dimension of the original data is used to screen out the heterogeneous features of the data and the relevant historical data that have the most significant impact on the dependent variable to the greatest extent. It can effectively extract the key features in multivariate data, remove redundant information and noise, reduce the dimension of the data, and retain the main information, providing a more refined data basis for the subsequent prediction steps.
[0260] 3. An improved BP neural network prediction model is provided. By nesting an LSTM network in the BP neural network for prediction, various influencing factors of the power load are comprehensively considered, providing a richer fitting data set and improving the prediction accuracy of the power load prediction model for the maximum load data of the power system.
[0261] As described above, only the preferred specific embodiments of the present invention are given, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A maximum load prediction method based on EMD-KPCA-improved BP neural network, characterized in that: The steps include: Obtain basic data affecting the load of the power system and corresponding maximum load data in multiple historical periods in the region, perform data preprocessing including EMD decomposition and linear kernel KPCA analysis in sequence, and generate a first training sample set based on the preprocessed basic data and maximum load data; Using the first training sample set to train the LSTM network, using the trained LSTM network to obtain the predicted value of each variable in the basic data and the corresponding maximum load data in the future time period, and performing inverse linear kernel KPCA analysis and EMD decomposition processing on the predicted values of the future time period in turn to obtain the predicted basic data and maximum load data; The basic data affecting the load of the power system in the historical period and the corresponding maximum load data are combined with the predicted basic data and maximum load data to generate a second training sample set; The BP neural network is trained using the second training sample set to obtain a trained power load prediction model; wherein the power load prediction model is a BP neural network with a nested LSTM network; The basic data affecting the load of the power system in the region is obtained in real time and input into the trained power load prediction model to obtain the predicted value of the maximum load data of the regional power system.
2. The method according to claim 1, characterized in that: The EMD data preprocessing is performed on each variable of the basic data and maximum load data in each historical period, including: The first step is to define the m historical data of each variable as the original time series x(t), the residual component r0(t) = x(t), and initialize the number of iterations i = 1; Step 2: In each iteration, after extracting an IMF from the current residual component, the current residual component is updated to be the residual component after the previous iteration minus the currently extracted IMF; Step 3: Find the current residual component r i-1 All local maximum and local minimum points in (t); Step 4: Use the cubic spline interpolation method to fit all local maximum and local minimum points of the current residual component to obtain the upper envelope e max (t) and the lower envelope e min (t), calculate the average value of the upper and lower envelopes Step 5: Subtract the average value of the upper and lower envelopes from the current residual component to obtain the component h i (t) = r i-1 (t)-m i (t); if the component h i (t) satisfies the preset conditions, then it is taken as the i-th IMF component c i (t), otherwise it will be regarded as a new residual component and steps 2 to 5 will be repeated until the i-th IMF component is obtained; Step 6: Update the residual component r based on the obtained i-th IMF component i (t) = r i-1 (t)-c i (t); if r i If (t) is not a monotonic function or a constant, then i = i + 1, and return to the second step to continue processing, otherwise stop the EMD decomposition process, and obtain m data in each historical period of each variable, which is represented by the sum of p IMF components and the residual signal of the last iteration. as the IMF component for each variable in each historical period; For the basic data and the corresponding maximum load data in each historical period, the n variables, each containing m data, are decomposed by EMD to obtain p×n×m IMF components.
3. The method according to claim 2, characterized in that: Perform linear kernel KPCA analysis on the p×n×m IMF components, including: For p×n×m IMF variables as sample points C=[c1,c2,...c pnm ], construct the kernel matrix K of the sample point, whose matrix elements K ij Represented as K ij =C T *C=k(c i , c j ); Construct a centralized kernel matrix K′=HKH, where I is the identity matrix, B is the all-1 vector; Perform eigenvalue decomposition on the centralized kernel matrix K′ to obtain eigenvalues λ1, λ2, ...λ N and its corresponding eigenvectors α1, α2, …, α n ; Select the first k largest eigenvalues and corresponding eigenvectors as principal components, and obtain the principal component values after projection into the principal component space. Among them, α j is the eigenvector corresponding to the jth principal component, λ j is the corresponding eigenvalue, c new is a sample point in C.
4. The method according to claim 1, characterized in that: The principal component values corresponding to the basic data and the principal component values corresponding to the maximum load data are used as labels to form sample data, and all sample data form a first training sample set.
5. The method according to claim 4, characterized in that: Training the LSTM network using the first training sample set includes: Initialize the weights and biases of the LSTM network; Input the principal component value corresponding to each historical data in the first training sample set into the LSTM network to obtain a predicted value, and compare it with the corresponding true value in the first training sample set to construct a loss function; Perform gradient descent training and update weights and biases; When the loss function is less than the first preset value, the iteration is stopped, and the weight matrix and bias vector values of the input gate, the forget gate, and the output gate are determined to obtain a trained LSTM network.
6. The method according to claim 1, characterized in that: In the process of training the BP neural network using the second training sample set, an adaptive weighted composite loss function is used, as follows: L total =μ1L1+μ2L2; Among them, x m,i is the i-th true value, z 2,i is the i-th prediction value, f is the number of samples in the second training sample set, μ1 and μ2 are the adaptive weights of the two loss functions L1 and L2 respectively.
7. The method according to claim 6, characterized in that: After each iteration of training the BP neural network, μ1 and μ2 are updated as follows: is the adaptive weight of the L1 function at the tth iteration, is the L1 loss function value at the tth iteration, is the L2 loss function value at the tth iteration; is the adaptive weight value of the L1 function at the t+1th iteration, is the adaptive weight value of the L2 function at the t+1 iteration; After each iteration update, calculate the total loss value L in the current iteration process total .
8. The method according to claim 7, characterized in that: During each iteration of training the BP neural network, weights and biases are updated as follows: Where η is the learning rate, W1 and W′1 are the weight matrices before and after the update from the input layer to the hidden layer, W2 and W′2 are the weight matrices before and after the update from the hidden layer to the output layer, b1 is the bias vector of the hidden layer, and b2 is the bias vector of the output layer.
9. The method according to claim 5, characterized in that: The loss function of the LSTM network is as follows: Among them, c is the sample data size, y zs,i is the ith true value, h t,i is the corresponding i-th predicted value, δ is the threshold parameter; The threshold parameter δ is determined using the quantile method.
10. The method according to claim 2, characterized in that: The preset conditions include: within the m data of each variable of the basic data and the maximum load data, the number of the local maximum points and the number of the minimum points and the number of zero-crossing points are equal to or differ by at most one, and the average values of the upper envelope formed by the local maximum points and the lower envelope formed by the local minimum points are zero at any time.
Citation Information
Patent Citations
Power load prediction method based on empirical mode decomposition and long-short-term memory network
CN112488415A
Short-term power load prediction method based on CNN-LSTM-ARIMA combination
CN113705915A
Power transmission line icing thickness prediction method based on EMD-KPCA-LSTM
CN118568451A
Secondary network thermal load prediction method based on CEEMDAN-KPCA-LSTM algorithm
CN119129825A
Lithium-battery temperature field spatio-temporal modeling method based on EMD-lstm
WO2024227376A1