Short-term load prediction method based on VMD-NRBO optimization neural network fusion method

The load data noise is removed by VMD, combined with the Transformer-BiLSTM model and the NRBO optimization algorithm, the problems of volatility and uncertainty in the microgrid load prediction are solved, and high-precision and robust load prediction are achieved.

CN119994868APending Publication Date: 2025-05-13STATE GRID CHONGQING ELECTRIC POWER COMPANY +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510048279.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the load prediction in the microgrid is predicted in the prior art, it is difficult to effectively deal with load fluctuations and uncertainties, resulting in low prediction accuracy and difficulty in optimizing hyperparameters, affecting the generalization ability of the model.

Method used

Variable modal decomposition (VMD) is used to remove high-frequency noise and redundant information in load data, combine the Transformer encoder and the bidirectional long and short-term memory (BiLSTM) decoder to build a neural network model, and optimize hyperparameters through the Newton-Lavson optimization algorithm (NRBO).

Benefits of technology

By removing noise and redundant information, the stability and accuracy of load prediction are improved. The NRBO optimization algorithm effectively reduces the difficulty of hyperparameter selection and improves the robustness and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994868A_ABST
    Figure CN119994868A_ABST
Patent Text Reader

Abstract

The invention relates to a short-term load prediction method based on a VMD-NRBO optimization neural network fusion method, belongs to the field of power systems, and aims to solve the problem of insufficient load prediction accuracy of the power systems. The method comprises the following steps: firstly, decomposing historical load data by using variational mode decomposition (VMD), effectively removing high-frequency noise and redundant information, and retaining key signal components; and then, inputting the decomposed modal components into a Transform encoder-bidirectional long short-term memory network BiLSTM decoder, fusing the modal components with a neural network model for training, and optimizing hyper-parameters of the neural network model through a Newton-Raphson optimization algorithm NRBO (Newton-Raphson Optimization). And finally, using the trained neural network model to predict future load data, and carrying out reverse normalization processing on a prediction result to obtain a final prediction value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power systems and relates to a Transformer-bidirectional long short-term memory (BiLSTM) fusion neural network load forecasting method based on Newton-Raphson optimization algorithm (NRBO) and variational mode decomposition (VMD). Background Art

[0002] With the widespread access of renewable energy, the load fluctuation problem faced by microgrids is becoming increasingly prominent due to the changes in renewable energy output, which brings great challenges to the dispatching and operation of power systems. The task of the present invention is to study the short-term load forecasting method based on NRBO and VMD fusion neural network for the problem of power system load uncertainty, and propose a high-precision and high-Robust load forecasting model to solve the problem of insufficient accuracy of power system load forecasting. The technical process includes multiple key steps. First, VMD is used to decompose the historical load data, and the load data is decomposed into multiple modal components according to different frequencies by VMD, effectively removing high-frequency noise and redundant information, and retaining key signal components; secondly, the Transformer encoder is combined with the BiLSTM decoder to construct a Transformer-BiLSTM fusion model, and the feature information of the time series data is extracted by the Transformer encoder, and the long-term dependency in the data is captured by the two-way learning ability of BiLSTM; finally, the NRBO optimization algorithm is applied to optimize the hyperparameters, and the Newton-Raphson optimization algorithm is used to intelligently optimize the hyperparameters of the model, reducing the difficulty of manual selection and ensuring that the hyperparameter configuration can adapt to different load forecasting scenarios. This technical method can effectively remove noise and redundant components in load data, and improve the stability and accuracy of prediction results.

[0003] Electric load has strong volatility and uncertainty, especially in microgrid environment. Due to the volatility of renewable energy, load forecasting faces great challenges. Existing load forecasting methods often fail to fully deal with these fluctuations, resulting in low prediction accuracy. The autoregressive integrated moving average (ARIMA) model is a classic time series analysis method, which is suitable for linear and steady-state time series data. However, the ARIMA method has great limitations in dealing with nonlinear, seasonal and trend changes in load data. In particular, when there are emergencies or significant periodic fluctuations in load data, the prediction ability of the ARIMA model will be greatly reduced. Wavelet transform can perform multi-resolution analysis when processing composite data and has good performance in signal decomposition. However, when processing high-dimensional data, the selection of wavelet basis functions is more complicated, and there is a boundary effect in the denoising process, which often leads to signal distortion at the boundary. Support vector machine (SVM) can model load data through nonlinear kernel functions, but it has certain limitations in dealing with long-term dependencies in time series. In particular, in high-dimensional time series data, SVM may not be able to capture complex time series characteristics. Traditional regression models, such as linear regression and polynomial regression models, are usually only applicable to situations where the data relationship is relatively simple. For complex nonlinear load changes, regression models often cannot be effectively fitted, resulting in the inability to generalize the model to new data. Although existing load forecasting methods can solve some problems to a certain extent, they still have problems such as insufficient handling of load volatility and uncertainty, incomplete noise removal, insufficient capture of long-term dependencies, difficulty in hyperparameter optimization, poor generalization ability, and low computational efficiency. These deficiencies make it difficult for traditional methods to cope with complex and dynamic load forecasting tasks in microgrids. Therefore, new methods are urgently needed to overcome these challenges and improve the accuracy, robustness, and real-time performance of load forecasting. Summary of the invention

[0004] In view of this, the object of the present invention is to provide a short-term load forecasting method based on VMD-NRBO optimized neural network fusion method.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method includes the following steps:

[0007] Step 1: Perform variational mode decomposition (VMD) on the historical load data to decompose the load data into multiple modal components according to different frequencies, effectively remove high-frequency noise and redundant information, and retain key signal components;

[0008] Step 2: Input the decomposed modal components into the Transformer encoder-bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, BiLSTM) decoder fusion neural network model for training, and optimize the hyperparameters of the neural network model through the Newton-Raphson Optimization algorithm (Newton-Raphson Optimization, NRBO);

[0009] Step 3: Use the neural network model trained in step 2 to predict future load data, and perform denormalization on the prediction results to obtain the final prediction value.

[0010] Further, the VMD decomposition method in step 1 is:

[0011] The number of K decomposed intrinsic mode function (IMF) components is preset;

[0012] With the goal of minimizing bandwidth, the limited bandwidth and center frequency of each IMF are continuously optimized to achieve optimal decomposition of the signal.

[0013] Furthermore, the Transformer encoder-BiLSTM decoder fusion neural network model in step 2 includes:

[0014] Transformer encoder, used to extract feature information of time series data;

[0015] BiLSTM decoder, used to capture long-term dependencies in the data;

[0016] The fully connected layer is used to output the final prediction value.

[0017] Further, the Transformer encoder includes:

[0018] A multi-head attention mechanism that determines the importance of each position relative to other positions and thus weights and aggregates parts of the input sequence;

[0019] A positional encoder that adds position information to each position in the input sequence.

[0020] Further, the BiLSTM decoder comprises:

[0021] The forward LSTM network is used to process the forward part of the input data;

[0022] The backward LSTM network is used to process the reverse part of the input data.

[0023] Furthermore, the NRBO optimization algorithm in step 2 is used to optimize the following hyperparameters:

[0024] The number of hidden layer units;

[0025] Maximum training cycle Epoch;

[0026] Initial learning rate.

[0027] Further, the denormalization processing method in step three is:

[0028] Multiply the predicted result by the difference between the maximum and minimum values ​​of the original data and add the maximum value of the original data.

[0029] The short-term load forecasting system based on VMD-NRBO optimization neural network fusion method includes:

[0030] Data processing module, used to perform VMD decomposition on historical load data;

[0031] The neural network model training module is used to input the decomposed modal components into the Transformer encoder-BiLSTM decoder fusion neural network model for training, and optimize the hyperparameters of the neural network model through the NRBO optimization algorithm;

[0032] The load forecasting module is used to use the trained neural network model to predict future load data and perform denormalization on the prediction results to obtain the final prediction value.

[0033] The beneficial effects of the present invention are:

[0034] (1) The VMD decomposition method solves the difficulty of extracting data features caused by the volatility of load data. It removes high-frequency segments from the decomposed modal components, reconstructs and superimposes the data waveform to be smoother, retains key information, and removes the interference of noise signals on model prediction.

[0035] (2) Based on the Transformer encoder-decoder structure, BiLSTM is used to replace the attention layer in the original Transformer decoder, and residual connection is combined to process the input sequence data, retaining the encoder information, improving information capture and processing capabilities, and solving the long-term dependency problem in sequence data.

[0036] (3) The NRBO optimization algorithm solves the difficulty of manually selecting network model hyperparameters and the incompatibility of the empirical selection method for specific prediction scenarios, which effectively improves the prediction accuracy of the model. The MAE, MAPE, MSE and RMSE are reduced by 42.38%, 53.40%, 73.94% and 48.94% respectively, and the R2 is increased by 15.18%.

[0037] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0039] Figure 1 It is the Transformer model structure;

[0040] Figure 2 It is a BiLSTM structure;

[0041] Figure 3 It is the Transformer-BiLSTM model structure;

[0042] Figure 4 It is the overall framework of VMD-NRBO-Transformer-BiLSTM;

[0043] Figure 5 This is the overall framework flow chart of the VMD-NRBO-Transformer-BiLSTM short-term load forecasting model. DETAILED DESCRIPTION

[0044] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0045] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0046] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0047] 1. VMD principle

[0048] VMD variational mode decomposition is Yanovsky et al. [Variational destriping in remote sensing imagery:total variation with L1 fidelity] A data processing technology proposed in 2014. It is highly robust when facing dynamically unstable, nonlinear and complex signals. It can decompose the original signal according to certain constraints and achieve signal denoising.

[0049] 1.1 Variational Constraint Model Construction

[0050] The number of K decomposed Intrinsic mode function (IMF) components is pre-set, and the limited bandwidth and center frequency of each IMF are continuously optimized to achieve optimal decomposition of the signal with the goal of minimizing the bandwidth. The model expression is as follows:

[0051]

[0052] In the formula, st means subject to constraints; u k is the kth IMF component; ω k is the center frequency of the corresponding IMF component; δ(t) is the Dirac function; To find the partial derivative with respect to time t; is the convolution operation symbol; f(t) is the original signal; j is the imaginary unit.

[0053] 1.2 Constraint model solution

[0054] In order to eliminate the above constraints, the constrained variational problem is unconstrained by introducing the Lagrangian multiplier λ(t) and the second-order penalty factor α, and a new augmented Lagrangian optimization model is obtained, which is expressed as follows:

[0055]

[0056] The alternating direction multiplier method is used to solve the unconstrained problem, combined with Fourier isometric transformation, to obtain the k and ω kPerform alternating optimization, the expression is as follows:

[0057]

[0058] Where: τ is the noise tolerance; and for u k The Fourier transform of (t) and λ(t).

[0059] 2. BiLSTM network model based on Transformer encoder

[0060] Based on the Transformer encoder-decoder structure, the present invention designs a network model structure embedded with BiLSTM. In the encoding link, the Transformer encoder is used to extract features of the input data, and the decoding part is improved by using the BiLSTM layer and the fully connected layer to obtain a new Transformer-Encoder-BiLSTM-Decoder structure network model.

[0061] 2.1 Transformer encoder feature extraction

[0062] The structure of the Transformer model is as follows Figure 1 As shown in the figure, it includes two parts, the encoder and the decoder. The encoder part contains N = 6 identical layers, each layer consists of a head attention mechanism and a position-by-position feedforward neural network sublayer. At the same time, residual connections and layer normalization operations are used to process the output of each sublayer. Such a structure is conducive to capturing the dependencies of all positions in the input sequence. The decoder is also composed of N = 6 identical layers, each layer contains three sublayers: a masked self-attention layer, an encoding-decoding attention layer, and a position-by-position feedforward neural network. Each sublayer is also followed by residual connections and layer normalization. This structure ensures that the decoder takes into account previous outputs when generating sequences and avoids the influence of future information.

[0063] The structure of the Transformer encoder used in the present invention is as follows.

[0064] 1) Positional Encoding

[0065] Traditional RNN networks can automatically input position information, but Transformer lacks built-in sequence position information. It is necessary to add position information to each position in the input sequence through a position encoder to identify different position markers of the input sequence data and ensure the complete characteristics of the input information.

[0066]

[0067] Where: pos represents the position of the sequence; i represents the dimension; d model represents the embedding space dimension.

[0068] 2) Multi-head attention mechanism (Self-Attention)

[0069] The self-attention mechanism determines the importance of each position relative to other positions by calculating the attention score between the query, key, and value, and then weights and aggregates the parts of the input sequence.

[0070]

[0071] Where: Q, K, V represent Query, Key, and Value matrices respectively; d K is the dimension of Key. The sequence information after position encoding is divided into Q, K, and V as the input of the encoder.

[0072] In order to capture richer sequence features, Transformer uses multiple self-attention mechanisms for parallel calculations. Each head captures information in different subspaces by learning different weights, and the results of the h attention heads are spliced, as shown in formula (8). Since the spliced ​​results are not organically integrated, a linear transformation is required.

[0073] MultiHead(Q,K,V)=Concat(head 1 ,head 2 ,...,head h )W O (8)

[0074] head i =Attention(Q i ,K i ,V i ) (9)

[0075] Q i ,K i ,V i =QW i Q ,KW i K ,VW i V (10)

[0076] Where: head i is the output of attention head i; Q i , K i 、V i As the input of attention head i; Wi Q , W i K , W i V and W O is the parameter matrix of linear transformation, where W O Used for linear transformation after concatenating the output results of multiple attention heads.

[0077] 2.2 BiLSTM Network

[0078] As a special RNN network that can better describe the correlation of sequence data, LSTM network is widely used in the processing of time series data. By introducing input gate, forget gate and output gate mechanism to control the flow of feature information, it solves the problems faced by traditional RNN networks such as gradient explosion and disappearance. BiLSTM network contains two types of LSTM networks, one type of forward LSTM is used to process the forward part of input data, and the other type is used to process the reverse part, which enhances the ability to capture information.

[0079] The BiLTSM network simultaneously obtains past and future information and captures the bidirectional dependency of input data. The structure is as follows: Figure 2 The input information is passed to the forward BiLSTM and the backward BiLSTM at the same time, and the output results of each at the same time are concatenated to obtain the final output y t .

[0080]

[0081] Where: and is the state phasor of the LSTM unit in the forward propagation layer at the corresponding moment; and is the LSTM unit state phasor of the backward propagation at the corresponding moment; and are the input weight matrices for the forward and backward propagation layers respectively; and are the forget weight matrices for the forward and backward propagation layers respectively; and are the biases of the forward and backward propagation layers, respectively.

[0082] 2.3 Transformer-BiLSTM Model Architecture

[0083] In the traditional Transformer model, when machine translating, the masked multi-head attention mechanism can observe the previous text content when the text sequence is generated, ensuring that the generated text content is coherent and reasonable. However, when predicting time series data, the mapping relationship between the time series sliding window data is different from machine translation. The model input is known historical data, and future time data is predicted. Therefore, BiLSTM is used to replace the attention layer in the original Transformer decoder, and the residual connection is combined to process the input sequence data. While retaining the encoder information, the long-term dependency problem in the sequence data is solved. In order to prevent overfitting of the model training, a Dropout layer is set in the BiLSTM layer to randomly discard the output of some neurons. Finally, the final prediction value is obtained through a fully connected feedforward neural network. The structure of the fused Transformer-BiLSTM model is as follows: Figure 3 shown.

[0084]

[0085] Where: x is the output of the Transformer encoder.

[0086] 3. NRBO Optimization Algorithm

[0087] NRBO defines the search path by discovering the search area using the Newton–Raphson Search Rule (NRSR) and the Trap Avoidance Operator (TAO).

[0088] 3.1 Population Initialization

[0089] NRBO searches for the optimal solution by generating an initial random population within the candidate solution boundary. Based on M populations, each with N decision variables, the initial random population is generated as follows.

[0090]

[0091] Where: is the position of the nth dimension in the mth population; rand is a random number between 0 and 1; lb and ub are the lower and upper boundaries respectively. The generated population matrix is ​​shown below.

[0092]

[0093] 3.2 Newton-Raphson Search Rule (NRSR)

[0094] The Newton-Raphson method (NRM) uses the calculation of Taylor series to update the position of the solution, and repeats it to find the optimal solution, which can enhance the exploration trend and accelerate the convergence speed. Starting from a random initial solution, it updates the position of the next solution along a certain direction. The position update process is as follows.

[0095]

[0096] Where: x it is the position of the current solution; x it+1 is the position of the solution after update; f(x) is the fitness function.

[0097] According to Taylor series expansion, the parameter ρ is introduced to guide the correct evolution direction of the population, thereby improving the efficiency of NRBO. And according to the NRM proposed in the literature [A variant of Newton's method with accelerated third-order convergence. Appl. Math. Lett][Newton's method. A Contemporary Study of Iterative Methods], it is improved:

[0098]

[0099] ΔX=rand(1,dim)·|X bet -X it | (19)

[0100]

[0101] Y wor =r 3 ·(Mean(Z it+1 +X it )+r 3 ΔX) (21)

[0102] Y bet =r 3 ·(Mean(Z it+1 +X it )-r 3 ΔX) (22)

[0103]

[0104] Where: a and b are random numbers between 0 and 1; is the population n after the it-th iteration; X bet and X bet They are The better and worse populations nearby; r 1 and r 2 is a random integer between 0 and M; Y wor and Y bet It is Z it+1 and X itThe two positions generated; r 3 is a random number between 0 and 1. So formula (17) is updated as:

[0105]

[0106] Use optimal position X bet replace Get the new position update formula:

[0107]

[0108] Formula (24) is biased towards local search and has limitations in global search; Formula (25) focuses on global search and has limitations in local search. NRBO combines the two to achieve the best search effect, so the new position of the next iteration is:

[0109]

[0110] Where: K it is the number of iterations; is the maximum number of iterations; r 4 A random number between 0 and 1.

[0111] 3.3 Trap Avoidance Operator (TAO)

[0112] TAO is an improved enhancement operator [Gradient-based optimizer: a new metaheuristic optimization algorithm] that can improve the efficiency of NRBO in dealing with practical problems. bet and Get a better solution.

[0113]

[0114] μ 1 =3β·rand+(1-β) (30)

[0115] μ 2 =β·rand+(1-β) (31)

[0116] Where: rand represents a uniform random number between (0,1); θ 1 and θ 2 are uniform random numbers between (1,1) and (0.5,0.5); μ 1 and μ 2 is a random number; β is a binary number. The randomness of the parameters can diversify the population and avoid local optimality.

[0117] 4. VMD-NRBO-Transformer-BiLSTM combined prediction model

[0118] 4.1 Model parameter optimization

[0119] There are many hyperparameters in the Transformer-BiLSTM network model, such as training cycle, number of hidden layer units, learning rate, regularization parameter, etc. The selection of hyperparameters affects the training effect of the model. The artificial selection of hyperparameters is complicated and inefficient. The general selection of hyperparameters based on experience cannot adapt to different prediction scenarios. Therefore, the present invention uses the NRBO optimization algorithm to optimize the model hyperparameters, and optimizes the three hyperparameters of the number of hidden layer units, maximum training cycle Epoch and initial learning rate of the Transformer-BiLSTM network model. In addition, the Adam optimizer is also used to adaptively adjust the learning rate of each parameter to speed up the convergence speed and improve the model performance.

[0120] 4.2 Overall Framework

[0121] The overall framework of the VMD-NRBO-Transformer-BiLSTM combined network model is as follows Figure 4 As shown in the figure. The model is divided into three modules. First, the original data is decomposed into multiple sub-sequence data through VMD, and then divided into training set and test set. After data normalization and data flattening, it is input into the Transformer-BiLSTM network for model training and testing. The output prediction data of each sub-sequence is denormalized and superimposed to obtain the prediction result. At the same time, the NRBO optimization algorithm is used to optimize the initial hyperparameters of the network model to improve the prediction accuracy of the model.

[0122] Figure 5 It is the overall framework flow chart of the present invention.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A short-term load forecasting method based on VMD-NRBO optimized neural network fusion method, characterized by: The following steps are involved: Step 1: Perform variational mode decomposition (VMD) on the historical load data to decompose the load data into multiple modal components according to different frequencies, effectively remove high-frequency noise and redundant information, and retain key signal components; Step 2: Input the decomposed modal components into the Transformer encoder-bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, BiLSTM) decoder fusion neural network model for training, and optimize the hyperparameters of the neural network model through the Newton-Raphson Optimization algorithm (Newton-Raphson Optimization, NRBO); Step 3: Use the neural network model trained in step 2 to predict future load data, and perform denormalization on the prediction results to obtain the final prediction value.

2. The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method according to claim 1 is characterized in that: The VMD decomposition method in step 1 is: The number of K decomposed intrinsic mode function (IMF) components is preset; With the goal of minimizing bandwidth, the limited bandwidth and center frequency of each IMF are continuously optimized to achieve optimal decomposition of the signal.

3. The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method according to claim 1 is characterized in that: The Transformer encoder-BiLSTM decoder fusion neural network model in step 2 includes: Transformer encoder, used to extract feature information of time series data; BiLSTM decoder, used to capture long-term dependencies in the data; The fully connected layer is used to output the final prediction value.

4. The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method according to claim 3 is characterized by: The Transformer encoder includes: A multi-head attention mechanism that determines the importance of each position relative to other positions and thus weights and aggregates parts of the input sequence; A positional encoder that adds position information to each position in the input sequence.

5. The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method according to claim 3 is characterized by: The BiLSTM decoder includes: The forward LSTM network is used to process the forward part of the input data; The backward LSTM network is used to process the reverse part of the input data.

6. The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method according to claim 1 is characterized by: The NRBO optimization algorithm in step 2 is used to optimize the following hyperparameters: The number of hidden layer units; Maximum training cycle Epoch; Initial learning rate.

7. The short-term load forecasting method based on VMD-NRBO optimized neural network fusion method according to claim 1 is characterized by: The denormalization processing method in step 3 is: Multiply the predicted result by the difference between the maximum and minimum values ​​of the original data and add the maximum value of the original data.

8. A short-term load forecasting system based on VMD-NRBO optimized neural network fusion method is characterized by: include: Data processing module, used to perform VMD decomposition on historical load data; The neural network model training module is used to input the decomposed modal components into the Transformer encoder-BiLSTM decoder fusion neural network model for training, and optimize the hyperparameters of the neural network model through the NRBO optimization algorithm; The load forecasting module is used to use the trained neural network model to predict future load data and perform denormalization on the prediction results to obtain the final prediction value.

Citation Information

Cited By

  • Sea surface weak target detection method based on improved bidirectional long and short time memory network

    CN120178200A

  • Faint target detection method on sea surface based on improved bidirectional long short-term memory network

    CN120178200B

  • LSTM daily runoff prediction method based on MFF and NRBO

    CN120975337A