Power load prediction method based on NIWPSO + CNN + LSTM + Attention
By introducing the NIWPSO algorithm to optimize hyperparameters in the power load prediction model, and combining CNN, LSTM and Attention mechanisms, the problems of low prediction accuracy and difficult parameter optimization in the existing technology are solved, and more efficient power load prediction is achieved.
Patent Information
- Application Number
- CN202510398562.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-05-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the existing power load prediction methods deal with load data volatility and instability, the prediction accuracy is not high, and the model parameters are difficult to optimize, and rely on manual tuning and inefficient efficiency.
The CNN+LSTM+Attention model optimized based on the NIWPSO algorithm is adopted, and the hyperparameters are automatically searched through NIWPSO, combined with CNN extraction of spatial features, LSTM processing time series information and Attention mechanism weighting important features to improve prediction accuracy.
Automatically find the optimal hyperparameter combination in multi-dimensional hyperparameter space, improve model adaptability and prediction accuracy, reduce the dependence of traditional manual tuning, and improve the model's performance under complex data.
Smart Images

Figure CN119917842A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a method for predicting power load based on NIWPSO+CNN+LSTM+Attention. Background Art
[0002] Accurate prediction of power load is crucial to ensure the safe, reliable and economical and efficient operation of the power system. Power load forecasting is to use statistics, machine learning and other methods to mine the key factors affecting the load from historical data such as meteorology and date, and establish a load forecasting model. Power load forecasting methods have undergone a transformation from traditional statistical methods to modern deep learning methods. As a typical deep learning method, convolutional neural network (CNN) can fully solve the nonlinear problems in large-scale load data, so it has been widely used in power load forecasting.
[0003] However, a single CNN model may not work well when dealing with large volatility and instability in load data. To overcome this problem, many studies have proposed methods of combining other models. Studies have shown that combining CNN and LSTM can simultaneously process spatial features and time series features, thereby improving prediction accuracy. The Attention mechanism has achieved remarkable success in various deep learning tasks in recent years, especially when processing sequence data. Attention can automatically focus on the most important part of the input data for the current prediction task, and add a mechanism for dynamic weight adjustment to the model. In power load forecasting, the combination of CNN+LSTM+Attention is a very effective model architecture. CNN is responsible for extracting spatial features, LSTM processes time series information, and the Attention mechanism helps the model focus on the most critical features. This combination can effectively improve the accuracy of power load forecasting, especially when facing multi-dimensional and complex load data, it can better capture the patterns and trends therein, thereby providing more accurate prediction results for power dispatching and management.
[0004] In power load forecasting, too many model parameters will significantly increase the difficulty of optimization, which is manifested in the fact that the parameter tuning process is complex and highly dependent on experience. For example, key parameters such as the number and size of kernels in the convolution layer, the number of units in the LSTM layer, the number of layers, and the Dropout rate in the CNN-LSTM model require repeated experiments. Hyperparameter adjustment mostly relies on manual methods, which is not only inefficient, but also difficult to find the optimal hyperparameter combination in large-scale complex data, and often cannot be effectively balanced. Traditional parameter adjustment methods are inefficient in high-dimensional parameter space, prone to local optimality, and have a high risk of overfitting, resulting in poor prediction accuracy and stability. Summary of the invention
[0005] The purpose of the present invention is to provide a power load forecasting method based on NIWPSO+CNN+LSTM+Attention to solve the problems raised in the above background technology.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: A power load forecasting method based on NIWPSO+CNN+LSTM+Attention, the method comprising: S100, acquiring power data, and performing preprocessing and standardization operations on the power data; S200, taking the preprocessed data as input, constructing a CNN-LSTM-Attention model; S300, use NIWPSO algorithm to optimize the hyperparameters of CNN-LSTM-Attention model; S400, set up residual monitoring and dynamic update mechanism, monitor residual in real time during model prediction, and automatically trigger retraining process once it exceeds the preset threshold.
[0007] Preferably, S100 includes: Get the original input data , where x includes relevant information such as power load and temperature, data standardization is adopted according to the formula to achieve normalization of values: ; Get standardized data , and use the standardized data as the input of the model; in, Represents the original data The mean of Represents the original data The standard deviation of .
[0008] Preferably, the CNN-LSTM-Attention model in S200 includes: a CNN layer, a LSTM layer and an Attention layer; S201, the CNN layer focuses on extracting data information features: According to the formula: , extract local features of input data; in, represents the output of the convolutional layer as the feature of the input data, Represents the input data, which is the standardized data. Represents the weight of the convolution kernel. The convolution kernel slides on the input data and extracts features by performing convolution operations with local areas of the input data. Represents the bias term of the convolution layer. After each convolution operation, a bias is added to improve the flexibility of the model; According to the formula: , the local features Perform pooling operations; in Represents the output of the pooling layer. Pooling operations (such as maximum pooling or average pooling) will reduce the size of the feature map by downsampling the local area, thereby reducing the amount of computation and avoiding overfitting; S202, the LSTM layer is used to learn the rules of CNN layer output: Based on the input gate, output gate, memory cell unit, and forget gate, LSTM completes the learning of the hidden value at the previous moment and updates the hidden layer state information; S203, the Attention layer weights the hidden states of different time steps according to the input LSTM hidden state, thereby generating a context vector focusing on important time steps: Get the LSTM layer Hidden state at the moment , then according to the formula: ; Confirm the Attention weight; among them, Indicates The Attention weight at a certain moment indicates the degree of influence of that moment on the final output. According to the hidden state The calculated score is usually computed by a simple linear transformation or inner product, Indicates the length of the sequence; According to the formula: ; Normalize the Attention weights, where represents the context vector obtained by weighted summation, which is the sum of all hidden states According to the Attention weight The weighted results, through Attention weighting, can selectively focus on information at certain time steps, thereby better capturing important moments in the sequence; S204, the output layer converts the context vector Input to the fully connected layer for the final prediction output: According to the formula: ; Confirm the final output prediction ;in, represents the context vector, which is the output of the Attention layer. and Represent the weight matrix and bias term of the output layer respectively, Represents the fully connected layer operation, usually a linear transformation, and the final output Represents the model’s final judgment on a task.
[0009] Preferably, S300 includes: S301, obtaining the hyperparameters of the CNN-LSTM-Attention model as an optimization object, wherein the hyperparameters include: the number and size of convolution kernels in the CNN layer, the number of units, the number of layers, and the Dropout rate in the LSTM layer, and the Dropout rate in the Attention layer; By randomly generating the initial values of the hyperparameters as the initialization coordinates of the particles in the particle swarm optimization algorithm; In the initial stage, when the random number x>0.95, according to the formula , mutate one dimension of the particle with a probability of 5%; in, Indicates that the group Particle Dimension mutation operation; S302, obtaining the prediction result output by the model in S204, taking the mean absolute percentage error of the prediction result as the fitness value, and calculating the fitness value according to the following formula: ; in, represents the true value of the validation sample, represents the predicted value of the validation sample, represents the number of validation samples, Represents the fitness value, which is obtained by calculating the MAPE value of the model corresponding to each particle. The fitness value is inversely proportional to the prediction effect of the model. The lower the fitness, the better the prediction effect of the model. S303, the individual optimal solution of each particle Set to the current position of the particle, and calculate the fitness value of each particle, where the individual optimal solution of the particle with the smallest fitness value is the current population The optimal solution of S304, compare the fitness value of each particle with Compare and keep the better one Similarly, the fitness value of each particle is compared with Compare and keep the better one ; According to the mutation formula and fitness calculation formula in S301 and S302, the position and velocity of the particles are updated. NIWPSO optimizes the global and local search balance of the traditional PSO algorithm by improving the nonlinear adjustment strategy of the inertia weight ω and introducing mutation operations:
[0010]
[0011]
[0012] in, and are the maximum and minimum values of the inertia weight, t is the current iteration number, is the maximum number of iterations, the control factor Used to adjust the smoothness of the curve, the coefficient before the variable iteration is 1.5 to ensure Between 0.3 and 0.7; At this time, when the termination condition is met, the optimal particles are used to construct the CNN-LSTM-Attention model; S305: Input the test set into the constructed CNN-LSTM-Attention model for prediction, and output the power load prediction value.
[0013] Preferably, S400 includes: S401, residual calculation: real-time calculation of predicted values With the true value The absolute error , and calculate the root mean square error within the sliding window: ; S402. Threshold setting: Based on the 95% quantile of historical error distribution or dynamic adjustment formula: , determine the threshold ;in, represents the smoothing coefficient; S403, trigger mechanism: when continuous Residuals at each time point Or the single point residual exceeds the maximum threshold When , the model update is triggered; S404, incremental training and parameter optimization: The improved NWPSO algorithm is used to dynamically search the model hyperparameters. The formula is: ; S405. Fine-tune model weights based on new data: .
[0014] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the above-mentioned power load forecasting method based on NIWPSO+CNN+LSTM+Attention.
[0015] A computer device comprises a memory, a processor and a computer program stored in the memory and running on the processor. When the processor executes the program, the steps in the above-mentioned power load forecasting method based on NIWPSO+CNN+LSTM+Attention are implemented.
[0016] Compared with the prior art, the beneficial effects achieved by the present invention are: The present invention innovatively adopts the NIWPSO optimization algorithm to apply to the CNN+LSTM+Attention model for power load forecasting, which can automatically search and find the best hyperparameter combination in the multi-dimensional hyperparameter space, avoiding the drawbacks of traditional manual adjustment of hyperparameters, and making the model more adaptable when facing large-scale complex data. In terms of feature extraction, CNN is used to automatically learn the spatial features in the input data, and LSTM effectively models the temporal features of the time series. At the same time, the Attention mechanism is introduced to further optimize the weighted processing of the model at different times and features, so that each prediction step can focus on the most important information.
[0017] The PSO algorithm is optimized innovatively, and a nonlinear inertia weight strategy is adopted. Dynamic adjustments are made based on the algorithm iteration process and particle search status to prevent the PSO algorithm from falling into a local optimal solution too early. At the same time, the "mutation" idea of the genetic algorithm is used to perform a one-dimensional mutation operation on the particles to expand the search range, helping the algorithm to efficiently obtain the global optimal solution in the complex solution space and ensure the optimization of model parameters.
[0018] The organic combination of CNN, LSTM and Attention mechanism enables the model to simultaneously process spatial features, time series features and weighted important features, thereby significantly improving the accuracy of power load forecasting. In particular, when power load data is highly volatile and unstable, the model can better adapt to complex change patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 It is a flow chart of a power load forecasting method based on NIWPSO+CNN+LSTM+Attention of the present invention; Figure 2It is the optimization flow chart of NIWPSO algorithm of the present invention; Figure 3 It is a flow chart of power load prediction of the present invention. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] See also Figure 1-Figure 3 , the present invention provides a technical solution:
[0022] Embodiment 1: The present invention aims to address the limitations of existing power load forecasting technologies in terms of accuracy, stability and hyperparameter adjustment, especially the low efficiency of traditional manual hyperparameter adjustment methods, the difficulty of the model in effectively processing spatial and temporal features at the same time, and the lack of adaptability to complex fluctuating data. A power load forecasting model combining hyperparameter optimization and multi-layer feature extraction is proposed. The implementation method of each step is described in detail below: A power load forecasting method based on NIWPSO+CNN+LSTM+Attention, the method comprising: S100, acquiring power data, and performing preprocessing and standardization operations on the power data; Preferably, S100 includes: In order to improve training efficiency and enhance evaluation accuracy, data standardization is used to normalize values and obtain original input data. , where x includes relevant information such as power load and temperature, data standardization is adopted according to the formula to achieve normalization of values: ; Get standardized data , and use the standardized data as the input of the model; where, Represents the original data The mean of Represents the original data The standard deviation of .
[0023] S200, taking the preprocessed data as input, constructing a CNN-LSTM-Attention model; Preferably, the CNN-LSTM-Attention model in S200 includes: a CNN layer, a LSTM layer and an Attention layer; S201, the CNN layer focuses on extracting data information features: According to the formula: , extract local features of input data; in, represents the output of the convolutional layer as the feature of the input data, Represents the input data, which is the standardized data. Represents the weight of the convolution kernel. The convolution kernel slides on the input data and extracts features by performing convolution operations with local areas of the input data. Represents the bias term of the convolution layer. After each convolution operation, a bias is added to improve the flexibility of the model; According to the formula: , the local features Perform pooling operations; in Represents the output of the pooling layer. Pooling operations (such as maximum pooling or average pooling) will reduce the size of the feature map by downsampling the local area, thereby reducing the amount of computation and avoiding overfitting; S202, construct the LSTM layer to learn the rules of CNN layer output: In terms of the LSTM layer, the long short-term memory neural network can capture the long-term and solve the gradient vanishing or gradient exploding problem of RNN during gradient descent. Based on the input gate, output gate, memory cell unit, and forget gate, LSTM can complete the learning of the hidden value of the previous moment and update the hidden layer state information.
[0024] formula:
[0025]
[0026]
[0027]
[0028]
[0029] , , : The weight matrix and bias term corresponding to the input gate.
[0030] : Hyperbolic tangent activation function, used to ensure that the candidate memory unit value range is between [-1,1].
[0031] : The state of the memory unit at the current moment, which is a weighted combination of the state of the memory unit at the previous moment and the candidate memory unit at the current moment.
[0032] : The state of the memory unit at the previous moment.
[0033] : The output of the forget gate determines how much information is forgotten.
[0034] : The output of the input gate, which determines how much new information is stored.
[0035] Memory Unit Combined with the control of the forget gate and the input gate, it is the core of LSTM and is responsible for preserving and updating long-term memory. : Output gate output, controlling the final hidden state .
[0036] , , : The weight matrix and bias term of the output gate.
[0037] : The output hidden state of LSTM, which is composed of the output gate and memory unit The value of is jointly determined.
[0038] : The memory unit is activated by the hyperbolic tangent function.
[0039] S203, the Attention layer weights the hidden states of different time steps according to the input LSTM hidden state, thereby generating a context vector focusing on important time steps: Get the LSTM layer Hidden state at the moment , then according to the formula: ; Confirm the Attention weight; among them, Indicates The Attention weight at a certain moment indicates the degree of influence of that moment on the final output. According to the hidden state The calculated score is usually computed by a simple linear transformation or inner product, Indicates the length of the sequence; According to the formula: ; Normalize the Attention weights, where represents the context vector obtained by weighted summation, which is the sum of all hidden states According to the Attention weight The weighted results, through Attention weighting, can selectively focus on information at certain time steps, thereby better capturing important moments in the sequence; S204, the output layer converts the context vector Input to the fully connected layer for the final prediction output: According to the formula: ; Confirm the final output prediction ;in, represents the context vector, which is the output of the Attention layer. and Represent the weight matrix and bias term of the output layer respectively, Represents the fully connected layer operation, usually a linear transformation, and the final output Represents the model’s final judgment on a task.
[0040] S300, use NIWPSO algorithm to optimize the hyperparameters of CNN-LSTM-Attention model; Preferably, S300 includes: S301, obtaining the hyperparameters of the CNN-LSTM-Attention model as an optimization object, wherein the hyperparameters include: the number and size of convolution kernels in the CNN layer, the number of units, the number of layers, and the Dropout rate in the LSTM layer, and the Dropout rate in the Attention layer; By randomly generating the initial values of the hyperparameters as the initialization coordinates of the particles in the particle swarm optimization algorithm; In the initial stage, when the random number x>0.95, according to the formula , mutate one dimension of the particle with a probability of 5%; in, Indicates that the group Particle Dimension mutation operation; S302, obtaining the prediction result output by the model in S204, taking the mean absolute percentage error of the prediction result as the fitness value, and calculating the fitness value according to the following formula: ; in, represents the true value of the validation sample, represents the predicted value of the validation sample, represents the number of validation samples, Represents the fitness value, which is obtained by calculating the MAPE value of the model corresponding to each particle. The fitness value is inversely proportional to the prediction effect of the model. The lower the fitness, the better the prediction effect of the model. S303, the individual optimal solution of each particle Set to the current position of the particle, and calculate the fitness value of each particle, where the individual optimal solution of the particle with the smallest fitness value is the current population The optimal solution of S304, compare the fitness value of each particle with Compare and keep the better one Similarly, the fitness value of each particle is compared with Compare and keep the better one ; According to the mutation formula and fitness calculation formula in S301 and S302, the position and velocity of the particles are updated. NIWPSO optimizes the global and local search balance of the traditional PSO algorithm by improving the nonlinear adjustment strategy of the inertia weight ω and introducing mutation operations:
[0041]
[0042]
[0043] in, and are the maximum and minimum values of the inertia weight, t is the current iteration number, is the maximum number of iterations, the control factor Used to adjust the smoothness of the curve, the coefficient before the variable iteration is 1.5 to ensure Between 0.3 and 0.7; At this time, when the termination condition is met, the optimal particles are used to construct the CNN-LSTM-Attention model; S305: Input the test set into the constructed CNN-LSTM-Attention model for prediction, and output the power load prediction value.
[0044] S400, set up residual monitoring and dynamic update mechanism, monitor residual in real time during model prediction, and automatically trigger retraining process once it exceeds the preset threshold; Preferably, S400 includes: S401, residual calculation: real-time calculation of predicted values With the true value The absolute error , and calculate the root mean square error within the sliding window: ; S402. Threshold setting: Based on the 95% quantile of historical error distribution or dynamic adjustment formula: , determine the threshold ;in, represents the smoothing coefficient; S403, trigger mechanism: when continuous Residuals at each time point Or the single point residual exceeds the maximum threshold When , the model update is triggered; S404, incremental training and parameter optimization: The improved NWPSO algorithm is used to dynamically search the model hyperparameters. The formula is: ; S405. Fine-tune model weights based on new data: .
[0045] Technical Point 1: A power load forecasting method based on NIWPSO+CNN+LSTM+Attention. The present invention innovatively proposes to apply the improved NIWPSO algorithm to the power load forecasting model of CNN (convolutional neural network) + LSTM (long short-term memory network) + Attention (attention mechanism). The global and local search capabilities of NIWPSO are used to optimize the parameters of LSTM. CNN is responsible for extracting spatial features from power load data, LSTM is used to process long-term dependencies in time series data, and the Attention mechanism focuses on key information in the data. The three are combined to form a powerful power load forecasting model framework, and NIWPSO optimizes model parameters and improves forecasting performance.
[0046] Technical Point 2: This invention innovatively proposes to improve the particle swarm algorithm (PSO) by introducing nonlinear inertia weight and adaptive mutation operation. The algorithm achieves a balance between global search and local development by dynamically adjusting the inertia weight, and combines the adaptive mutation mechanism to enhance the ability to jump out of the local optimum. Nonlinear inertia weight adjustment: Abandon the linear decrease of inertia weight with the number of iterations in the traditional PSO algorithm, and adopt a nonlinear change of inertia weight. Through the formula
[0047] Adjust, where ωmax and ωmin are the maximum and minimum values of the inertia weight respectively, t is the current number of iterations, is the maximum number of iterations, and k is the control factor (valued at 0.6). In the initial search stage, the inertia weight coefficient decreases nonlinearly, giving the algorithm a stronger global search capability; in the later stage, the search is slow, enhancing the local search capability.
[0048] Adaptive mutation operation: Refer to the "mutation" operation of the genetic algorithm to mutate one dimension of the particle.
[0049]
[0050] in, Indicates that the group Particle dimensional mutation operation, when Random number between It changes when the value is set, and rand is a random number that changes in the range of [0,1].
[0051] Embodiment 2: The computer-readable storage medium of this embodiment stores a computer program thereon, which, when executed by a processor, implements the steps in a power load forecasting method based on NIWPSO+CNN+LSTM+Attention in Embodiment 1.
[0052] The computer-readable storage medium of this embodiment may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal; the computer-readable storage medium of this embodiment may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card, a secure digital card, a flash memory card, etc. equipped on the terminal; further, the computer-readable storage medium may also include both an internal storage unit of the terminal and an external storage device.
[0053] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0054] Embodiment 3: The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of a power load forecasting method based on NIWPSO+CNN+LSTM+Attention in Embodiment 1 are implemented.
[0055] In this embodiment, the processor may be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, readily available programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0056] Those skilled in the art will appreciate that the contents disclosed in the embodiments may be provided as methods, systems, or computer program products. Therefore, the present solution may take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware. Moreover, the present solution may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.
[0057] The present solution is described with reference to the method according to the embodiment of the present solution and the flowchart and / or block diagram of the computer program product. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions; these computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 one or more processes and / or methods Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0058] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 one or more processes and / or methods Figure 1 A function specified in one or more boxes.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 one or more processes and / or methods Figure 1 The steps for the functions specified in one or more boxes.
[0060] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0061] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A power load forecasting method based on NIWPSO+CNN+LSTM+Attention, characterized by: The method comprises: S100, acquiring power data, and performing preprocessing and standardization operations on the power data; S200, taking the preprocessed data as input, constructing a CNN-LSTM-Attention model; S300, using the NIWPSO algorithm to optimize the hyperparameters of the CNN-LSTM-Attention model; S400, set up residual monitoring and dynamic update mechanism, monitor residual in real time during model prediction, and automatically trigger retraining process once it exceeds the preset threshold.
2. The power load forecasting method based on NIWPSO+CNN+LSTM+Attention according to claim 1, characterized in that: The S100 includes: Get the original input data , then the data standardization method is adopted according to the formula to achieve the normalization of the values: ; Get standardized data , and use the standardized data as the input of the model; in, Represents the original data The mean of Represents the original data The standard deviation of .
3. The power load forecasting method based on NIWPSO+CNN+LSTM+Attention according to claim 1, characterized in that: The CNN-LSTM-Attention model in S200 includes: a CNN layer, a LSTM layer and an Attention layer; S201, the CNN layer focuses on extracting data information features: According to the formula: , extract local features of input data; in, represents the output of the convolutional layer as the feature of the input data, Represents the input data, which is the standardized data. Represents the weight of the convolution kernel. The convolution kernel slides on the input data and extracts features by performing convolution operations with local areas of the input data. Represents the bias term of the convolutional layer; According to the formula: , the local features Perform pooling operations; in Represents the output of the pooling layer; S202, the LSTM layer is used to learn the rules of CNN layer output: Based on the input gate, output gate, memory cell unit, and forget gate, LSTM completes the learning of the hidden value at the previous moment and updates the hidden layer state information; S203, the Attention layer weights the hidden states of different time steps according to the input LSTM hidden state, thereby generating a context vector focusing on important time steps: Get the LSTM layer Hidden state at the moment , then according to the formula: ; Confirm the Attention weight; among them, Indicates The Attention weight at a certain moment indicates the degree of influence of that moment on the final output. According to the hidden state The calculated score, Indicates the length of the sequence; According to the formula: ; Normalize the Attention weights, where represents the context vector obtained by weighted summation, which is the sum of all hidden states According to the Attention weight Weighted results; S204, the output layer converts the context vector Input to the fully connected layer for the final prediction output: According to the formula: ; Confirm the final output prediction ;in, represents the context vector, which is the output of the Attention layer. and Represent the weight matrix and bias term of the output layer respectively, Represents the fully connected layer operation, the final output Represents the model’s final judgment on a task.
4. The power load forecasting method based on NIWPSO+CNN+LSTM+Attention according to claim 1, characterized in that: The S300 includes: S301, obtaining the hyperparameters of the CNN-LSTM-Attention model as an optimization object; Among them, the hyperparameters include: the number and size of convolution kernels in the CNN layer, the number of units, number of layers and Dropout rate in the LSTM layer, and the Dropout rate in the Attention layer; By randomly generating the initial values of the hyperparameters as the initialization coordinates of the particles in the particle swarm optimization algorithm; In the initial stage, when the random number x>0.95, according to the formula , mutate one dimension of the particle with a probability of 5%; in, Indicates that the group Particle Dimension mutation operation; S302, obtaining the prediction result output by the model in S204, taking the mean absolute percentage error of the prediction result as the fitness value, and calculating the fitness value according to the following formula: ; in, represents the true value of the validation sample, represents the predicted value of the validation sample, represents the number of validation samples, Represents the fitness value, where the fitness value is inversely proportional to the prediction effect of the model; S303, the individual optimal solution of each particle Set to the current position of the particle, and calculate the fitness value of each particle, where the individual optimal solution of the particle with the smallest fitness value is the current population The optimal solution of S304, compare the fitness value of each particle with Compare and keep the better one Similarly, the fitness value of each particle is compared with Compare and keep the better one ; According to the mutation formula and fitness calculation formula in S301 and S302, the position and speed of the particles are updated. NIWPSO improves the inertia weight The nonlinear adjustment strategy and the introduction of mutation operation are used to optimize the global and local search balance of the traditional PSO algorithm: ; ; ; in, and are the maximum and minimum values of the inertia weight, t is the current iteration number, is the maximum number of iterations, the control factor Used to adjust the smoothness of the curve, the coefficient before the variable iteration is 1.5 to ensure Between 0.3 and 0.7; At this time, when the termination condition is met, the optimal particles are used to construct the CNN-LSTM-Attention model; S305: Input the test set into the constructed CNN-LSTM-Attention model for prediction, and output the power load prediction value.
5. The power load forecasting method based on NIWPSO+CNN+LSTM+Attention according to claim 1, characterized in that: The S400 includes: S401, residual calculation: real-time calculation of predicted values With the true value The absolute error , and calculate the root mean square error within the sliding window: ; S402. Threshold setting: Based on the 95% quantile of historical error distribution or dynamic adjustment formula: , determine the threshold ;in, represents the smoothing coefficient; S403, trigger mechanism: when continuous Residuals at each time point Or the single point residual exceeds the maximum threshold When , the model update is triggered; S404, incremental training and parameter optimization: The improved NWPSO algorithm is used to dynamically search the model hyperparameters. The formula is: ; S405. Fine-tune model weights based on new data: 。 6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the power load forecasting method based on NIWPSO+CNN+LSTM+Attention as described in any one of claims 1 to 5 are implemented.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps in the power load forecasting method based on NIWPSO+CNN+LSTM+Attention are implemented as described in any one of claims 1-5.
Citation Information
Patent Citations
Short-term power load prediction method of improved CNN-LSTM algorithm
CN118504613A
Power load prediction method based on SL-SSA optimization combination model hyper-parameter
CN118630730A
Ultra-short-term power load prediction method and system based on attention mechanism and long-short-term memory network, storage medium and electronic equipment
CN119337121A
Mountainous area slope displacement prediction method based on mi-GRA and improved PSO-lstm
WO2024001942A1
Cited By
Power consumption prediction method based on intelligent algorithm
CN120657765A
Power load prediction method of electric energy meter data system and medium
CN120806279A
Power load prediction method for electric energy meter data system and medium
CN120806279B