Power clearing price prediction method based on improved whale migration algorithm and attention mechanism
Patent Information
- Application Number
- CN202611000472.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]第一,在模型优化层面,传统元启发式算法,如粒子群算法PSO、灰狼算法GWO、鲸鱼优化算法WOA等,在优化神经网络超参数时普遍存在种群多样性低、易陷入局部最优和收敛速度慢等问题
[0063](1)改进鲸鱼迁徙算法IWMA通过Chebyshev混沌初始化、透镜成像学习、非均匀变异和自适应拉普拉斯交叉四项改进策略,提高了超参数优化的全局搜索能力和收敛精度。在所采用的CEC2005标准测试函数实验条件下,IWMA在代表性函数上的平均适应度和稳定性指标优于GWO、WOA、ALO、IVY和原始WMA;其中在所采用实验条件下,F1和F2函数上获得的平均适应度和标准差均为0.00E+00,表明该算法在相应实验条件下能够收敛至已知全局最优附近。
Smart Images

Figure CN122820264A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electricity market forecasting technology, specifically relating to an electricity clearing price forecasting method based on an improved whale migration algorithm and attention mechanism. Background Technology
[0002] The rapid development of the electricity spot market has made clearing price forecasting an important basis for market participants to optimize their decisions. Clearing price series exhibit complex characteristics such as strong non-stationarity, high volatility, and frequent extreme peaks. Existing forecasting models mainly suffer from the following shortcomings.
[0003] First, at the model optimization level, traditional metaheuristic algorithms, such as Particle Swarm Optimization (PSO), Grey Wolf Algorithm (GWO), and Whale Optimization Algorithm (WOA), generally suffer from low population diversity, susceptibility to local optima, and slow convergence when optimizing neural network hyperparameters. While the standard Whale Migration Algorithm (WMA) achieves a balance between exploration and exploitation by simulating the seasonal migration behavior of humpback whales, its initialization relies on a pseudo-random number generator, which can easily lead to uneven population distribution; the leader update strategy lacks directionality and adaptability, and population diversity rapidly declines in the later stages, making premature convergence likely.
[0004] Second, at the signal decomposition level, although a single CEEMDAN decomposition can suppress mode aliasing, its analytical depth is limited for high-frequency non-stationary components with insufficient frequency resolution. It is difficult to fully explore the complex time-frequency structure contained in the clearing price at extreme fluctuations, resulting in a large prediction error of the model at extreme price moments.
[0005] Third, at the network structure level, traditional fully connected networks lack adaptive feature selection capabilities when facing non-stationary subsequences with multi-frequency aliasing, and cannot dynamically adjust the information flow to adapt to the differentiated modeling needs of components at different frequency scales.
[0006] Fourth, at the feature fusion level, existing models mostly use equal-weighted summation to integrate the prediction results of each final component, ignoring the dynamic differences in the information contribution of different components under different prediction scenarios; their ability to perceive historical time steps is limited, making it difficult to accurately focus on key time nodes during periods of drastic price fluctuations.
[0007] Therefore, there is an urgent need for a power clearing price prediction method that can simultaneously improve hyperparameter optimization capabilities, enhance the analytical depth of complex time-frequency structures, strengthen the adaptive modeling capabilities of non-stationary sequences, and dynamically integrate multi-scale component information. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention provides a method for predicting electricity clearing prices based on an improved whale migration algorithm and attention mechanism.
[0009] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0010] Firstly, this invention provides a method for predicting electricity clearing prices based on an improved whale migration algorithm and a two-layer attention mechanism. This method constructs a collaborative prediction framework of "adaptive decomposition—intelligent optimization—gated modeling—attention fusion," specifically including the following steps.
[0011] Step 1: Obtain historical clearing price data and related characteristic data of the electricity spot market, and preprocess the data;
[0012] Furthermore, the relevant characteristic data includes at least one of the following: thermal power unit operating capacity, power load, power transmission plan, total thermal power plan, and total new energy plan;
[0013] Furthermore, the preprocessing includes missing value processing, outlier processing, normalization processing, and time order partitioning;
[0014] Furthermore, the aforementioned temporal division includes dividing the samples into a training set, a validation set, and a test set according to chronological order;
[0015] Furthermore, the validation set is used for hyperparameter optimization;
[0016] Furthermore, the test set is used for the final predictive performance evaluation;
[0017] Step 2: First CEEMDAN Decomposition: The clearing price series is decomposed using the adaptive noisy complete set empirical mode decomposition (CEEMDAN). Gaussian white noise is adaptively added to obtain... One intrinsic mode function (IMF) component and one residual component are used to suppress mode aliasing and noise residual problems in traditional EMD and EEMD.
[0018] Step 3: Construct the basic prediction model of the improved whale migration algorithm-optimized deep neural network IWMA-DNN. The IWMA introduces Chebyshev chaotic mapping initialization, lens imaging learning strategy, non-uniform mutation strategy and adaptive Laplacian cross operation on the basis of the standard whale migration algorithm WMA. The hyperparameters of the DNN are globally optimized using the improved whale migration algorithm IWMA, and IWMA-DNN sub-prediction models are established for each IMF component and residual component. The basic prediction value is obtained by fusion through an adaptive weight strategy.
[0019] Specifically, it includes the following sub-steps:
[0020] Step 3.1: A 4th-order Chebyshev chaotic map is used for population initialization, replacing the uniform random distribution in the standard WMA. This generates an initial whale population with ergodicity and randomness, preventing the population from clustering in local areas during the initialization phase. The mathematical expression for the Chebyshev chaotic map is: (1), where, For the first Chaotic variables generated in the next iteration. ∈ When used for population initialization, a fourth-order Chebyshev mapping is employed to generate chaotic sequences; if used for population initialization, the generated sequences are linearly mapped to the search space or... Interval.
[0021] Step 3.2: In the early stages of algorithm iteration, the position of the leader whale is updated with a high probability using a lens imaging learning strategy. The position update formula is as follows: (2), of which For the updated position of the leader whale, For the first Individual position, and These are the lower and upper bounds of the search space, respectively. The parameters are dynamically adjusted for lens imaging. In a preferred embodiment, lens imaging mapping is performed on the current optimal solution according to the dynamic parameter k to generate candidate solutions. Use a smaller size in the early stages of the iteration. Alternatively, a larger mapping perturbation amplitude can be used to expand the search range, gradually increasing in the later stages of iteration. Or reduce the mapping perturbation amplitude to enhance local fine-grained search capabilities; when When the search space boundary is exceeded, Perform boundary truncation or mapping back to the preset search space.
[0022] Step 3.3: In the later stages of the algorithm iteration, a non-uniform mutation strategy is used with a higher probability to update the position of the leader whale. The position update formula is as follows: (3), among which, It is a non-uniform variation perturbation factor. for Random numbers between This represents the current iteration number. The maximum number of iterations, Parameters used to control the rate of decay of variation intensity. and These are the lower and upper bounds of the search space, respectively.
[0023] Step 3.4: Perform an adaptive Laplace crossover operation on the current optimal and suboptimal solutions with preset probabilities. The crossover formula is as follows: (4), (5), among which, and These are the current optimal solution and the second-best solution, respectively. It follows a Laplace distribution; by utilizing the peak-heavy tail characteristic of the Laplace distribution, it can simultaneously achieve large-step global jumps and small-step local fine-tuning. If the updated individual exceeds the boundary of the search space, it will be truncated or mapped back to the preset boundary range.
[0024] Step 3.5: The switching probability between the lens imaging learning strategy and the non-uniform mutation strategy is dynamically adjusted with the number of iterations. The following probability is preferred: (6), The probability of employing a global exploration strategy. This represents the maximum number of iterations. In the early stages of iteration, This term is close to 1; in the later stages of the iteration, This value approaches 0, thus enabling a smooth transition from global exploration to local development.
[0025] Step 3.6: Optimize the four hyperparameters of the DNN—learning rate, batch size, number of network layers, and number of neurons per layer—using the IWMA algorithm. Reconstruct the hyperparameter optimization problem into a global optimization problem in a high-dimensional non-convex space. Determine the optimal combination of hyperparameters based on the training set loss and the validation set prediction error as fitness evaluation criteria.
[0026] Step 3.7: Establish IWMA-DNN sub-prediction models for each IMF component and residual component, and assign adaptive weights to each component using a prediction value optimization strategy. Calculate the weights based on the mean absolute error (MAE) of each component on the training set, preferably using normalized weights that are inversely proportional to the MAE. (7), among which, For the first The weight of each component, For the first Each component corresponds to the mean absolute error of the sub-prediction model on the training set. To prevent the stability constant from having a denominator of zero, components with smaller prediction errors are given higher weights. Finally, the prediction results of each component are weighted and fused to obtain the base prediction value.
[0027] Step 4: Calculate the prediction error contribution of each IMF component and residual component, and identify the high-frequency IMF components from among them. Arrange the high-frequency IMF components in descending order of prediction error contribution and accumulate them sequentially. Select the components whose cumulative contribution first reaches or exceeds a preset threshold. The high-frequency IMF component is used as the object of secondary CEEMDAN decomposition. The object of secondary CEEMDAN decomposition is decomposed into sub-IMF components. The sub-IMF components, the IMF components that have not undergone secondary decomposition, and the residual components constitute the final set of components that participate in prediction fusion.
[0028] Based on the basic prediction model trained in step 3, the root mean square error (RMSE) of each IMF component and its contribution to the total RMSE are statistically analyzed. The IMF components are then sorted in descending order of contribution, and the top components with a cumulative contribution exceeding 50% are selected. The high-frequency IMF components are subjected to a second CEEMDAN decomposition to obtain a refined set of sub-IMF components; the low-frequency IMF components and residual components retain the results of the first decomposition.
[0029] The contribution of each component error can be expressed as: (8), among which, For the first The error contribution of each component is considered. Preferably, the components with a cumulative contribution exceeding 50% are selected. Each component is used as a secondary decomposition object.
[0030] The termination condition for the quadratic decomposition is: calculating the correlation coefficient between two consecutive sub-IMFs. ,when Higher than the preset threshold When the correlation between the continuous sub-IMFs is high and it becomes difficult to further separate the independent oscillation modes, the iterative decomposition is terminated; otherwise, the decomposition continues. Among these, The value ranges from 0.6 to 0.8, with 0.7 being preferred.
[0031] In the denominator When representing the contribution of statistical error, consider There are one IMF component and one residual component; however, the candidates for the second-order CEEMDAN decomposition are limited to the high-frequency IMF component, and the residual component is not considered as a second-order decomposition object.
[0032] Step 5: Construct a gated dynamic network (GDNN) prediction sub-model for each final component in the final component set. The GDNN is a gated unit containing update gates and reset gates embedded after the intermediate hidden layer of the DNN, and uses residual form gated updates to modulate the inter-layer information flow.
[0033] The forward propagation process of the gating unit is as follows: (9), (10) (11), (12);
[0034] Expressions in residual form: (13), among which, Let i be the hidden state of the i-th final component in the l-th layer. To update the door, To reset the door, This is a candidate hidden state. It is the Sigmoid activation function. For element-wise multiplication, , , and , , These are the learnable parameters in the gating unit;
[0035] GDNN employs a grouped parameter sharing strategy: the final components in the final component set are divided into high-frequency, mid-frequency, and low-frequency groups according to their frequency characteristics. Gating parameters are shared within each group, while the gating parameters between groups are independent. The gating unit uses a central placement scheme, inserted after the intermediate hidden layer of the DNN; this intermediate hidden layer is the nth hidden layer, where n is the lowest frequency. The integer obtained by rounding up. This represents the total number of hidden layers in the DNN. It enables the gating unit to operate on the higher-level semantic hidden states that have undergone partial nonlinear transformation.
[0036] Step 6: In the hidden state sequence output by GDNN A time attention mechanism is applied, and attention weights at different time steps are calculated through a query-key-value mechanism. A causal masking mechanism is also used to prevent information leakage at future moments.
[0037] Step 6.1: Linear projection, using three sets of learnable weight matrices to project the hidden state sequence. Mapped to queries respectively ,key Sum : (14) (15) (16); among which, , , For learnable parameter matrix, For projection dimensions;
[0038] Step 6.2: Calculate the attention score using the scaled dot product method for the query. AND key Correlation score matrix between them: (17) Among them, This is a scaling factor used to prevent the dot product result from becoming too large, causing the softmax to enter the gradient saturation region;
[0039] Step 6.3: Causal Masking. A causal masking mask is applied to the score matrix, setting the scores corresponding to future time steps to negative infinity. This ensures that each time step can only focus on itself and its preceding historical information, conforming to the causal constraints of time series prediction. The score matrix after causal masking is represented as follows: (18);
[0040] Step 6.4: Attention weight normalization. Apply the softmax function to each row of the masked score matrix to obtain the normalized attention weight matrix: (19);
[0041] Step 6.5: Weighted aggregation, using attention weights to aggregate the value vector Perform a weighted summation to generate a time series context vector: (20) (twenty one);
[0042] Step 6.6: Residual Connections and Layer Normalization. To alleviate gradient vanishing in deep networks and accelerate training convergence, a residual connection and layer normalization structure is adopted: (twenty two);
[0043] make Consistent with the hidden state dimension; or through linear projection Mapped to After establishing the same dimensions, perform residual joins, where... To perform standardization along the feature dimension, the calculation formula is as follows: (23), among which, and For learnable scaling and translation parameters in layer normalization, is the numerical stability constant.
[0044] Step 7: Apply component attention mechanism to the prediction features of each final component, introduce energy perception enhancement, calculate the attention fusion weight of each final component through a feedforward neural network, realize adaptive weighted fusion of multi-scale component prediction results, and obtain the final clearing price prediction value.
[0045] Energy-sensing enhancement is introduced into the prediction features of each final component, and the energy of each component is calculated: (24), among which, For the length of the training set, The value of the i-th component participating in the final fusion at time t is given. This final component includes the sub-IMF component obtained from the second decomposition, the IMF component not subjected to the second decomposition, and the residual term. After logarithmic smoothing of the energy, we obtain: (25);
[0046] The enhanced feature is obtained by concatenating the energy feature with the predicted feature: (26), among which, Let i be the dynamic prediction feature of the i-th component at time t. To dynamically predict feature dimensions, the feature concatenation operation is represented; attention scores are calculated through a two-layer feedforward neural network, and the fusion weights are obtained after softmax normalization. (27) (28), among which, , , and For the learnable parameters in the component attention network, the final predicted value is obtained by weighted summation of the prediction results of each component: (29), among which, The total number of components that ultimately participate in the fusion after selective secondary decomposition includes the sub-IMF components obtained from the secondary decomposition, the IMF components that have not undergone secondary decomposition, and the residual terms. Let be the fusion weight of the i-th component at time t. Let be the predicted value of the i-th component at time t. This is the projected final clearing price.
[0047] Step 8: Based on the composite loss function between the predicted and true values, the DNN parameters, gating unit parameters, temporal attention parameters, and component attention parameters are jointly optimized using the backpropagation algorithm until convergence.
[0048] The composite loss function is: ,in, ∈(0,1), preferred =0.7; This represents the mean squared error, which measures the squared deviation between the predicted and actual values. The mean absolute percentage error is used to measure the proportion of the prediction error relative to the actual value. Through this composite loss function, the model can simultaneously take into account the overall fitting accuracy and relative error control, thereby improving the stability and accuracy of the clearing price prediction results.
[0049] Furthermore, the optimization process of DNN hyperparameters by IWMA in step 3.6 is as follows: the learning rate, batch size, number of network layers, and number of neurons per layer of the DNN are taken as variables to be optimized, and the IWMA algorithm performs a global search for optimization within the preset range of hyperparameter values, with the goal of minimizing the prediction error on the validation set, to determine the optimal combination of hyperparameters.
[0050] Furthermore, the selective quadratic decomposition strategy described in step 4 is as follows: Calculate the RMSE contribution of each IMF component on the training set, sort the components in descending order of contribution, and select the components with a cumulative contribution exceeding 50%. Each component is used as a secondary decomposition object.
[0051] Furthermore, the gating unit described in step 5 adopts a centrally located scheme, inserting the DNN's first... After the hidden layer, among which This represents the total number of hidden layers in the DNN. This indicates rounding up. It causes the gating unit to act on the higher-level semantic hidden state that has undergone partial nonlinear transformation.
[0052] Furthermore, the energy in the energy perception enhancement described in step 7 is defined as the average power of each final component on the complete training set. After logarithmic smoothing, it is used as a static statistical feature and concatenated with the dynamic prediction feature, so that the component attention network can simultaneously perceive the dynamic prediction quality of the component and its intrinsic energy characteristics.
[0053] Secondly, this invention provides an electricity clearing price prediction system based on an improved whale migration algorithm and attention mechanism, comprising:
[0054] The data acquisition and preprocessing module is used to acquire historical clearing price data and related feature data of the electricity spot market, and to perform preprocessing; the preprocessing includes missing value imputation, outlier handling and normalization;
[0055] A single CEEMDAN decomposition module is used to perform a single decomposition on the clearing price series to obtain... One IMF component and one residual component;
[0056] The IWMA-DNN basic prediction module is used to construct a deep neural network basic prediction model for an improved whale migration algorithm. The IWMA introduces a Chebyshev chaotic mapping initialization strategy, a lens imaging learning strategy, a non-uniform mutation strategy, and an adaptive Laplacian cross operation strategy on the basis of the standard whale migration algorithm WMA. The IWMA-DNN basic prediction module establishes IWMA-DNN sub-prediction models for each IMF component and residual component, and obtains the basic prediction value by fusing them through an adaptive weight strategy.
[0057] The selective quadratic decomposition module is used to calculate the prediction error contribution of each IMF component and the residual component, and select some high-frequency IMF components from each IMF component for secondary CEEMDAN decomposition based on the prediction error contribution.
[0058] The gated dynamic network module is used to build a GDNN prediction sub-model for each final component in the final component set;
[0059] The temporal attention module is used to calculate the attention weights at each time step on the hidden state sequence output by the GDNN through a query-key-value mechanism.
[0060] The component attention module is used to calculate the attention fusion weights for the prediction features of each final component and generate the final prediction value.
[0061] The end-to-end training module is used to jointly optimize system parameters based on a composite loss function.
[0062] Compared with the prior art, the present invention has the following advantages:
[0063] (1) The improved Whale Migration Algorithm (IWMA) improves the global search capability and convergence accuracy of hyperparameter optimization through four improvement strategies: Chebyshev chaotic initialization, lens imaging learning, non-uniform mutation, and adaptive Laplacian crossover. Under the experimental conditions of the CEC2005 standard test function, IWMA outperforms GWO, WOA, ALO, IVY, and the original WMA in terms of average fitness and stability on representative functions. Among them, under the experimental conditions, the average fitness and standard deviation obtained on the F1 and F2 functions are both 0.00E+00, indicating that the algorithm can converge to the vicinity of the known global optimum under the corresponding experimental conditions.
[0064] (2) The selective quadratic decomposition strategy refines the high-frequency IMF components with high error contribution after the first CEEMDAN decomposition, effectively deepening the analysis level of extreme price fluctuation structure, while keeping the first decomposition results unchanged for low-frequency components and residual components, thus avoiding unnecessary computational complexity growth.
[0065] (3) Gated Dynamic Network (GDNN) replaces the traditional fully connected structure with an adaptive gating mechanism, which can adaptively weight the input features of non-stationary subsequences and dynamically adjust the information flow; the residual form of gating update is conducive to maintaining the original hidden state information and improving gradient propagation.
[0066] (4) The dual-layer attention mechanism consists of time attention and component attention connected in series, which realize the differentiated weighting of key information from the time dimension and frequency dimension, respectively. Time attention can dynamically perceive the information value of different historical moments, and causal masking ensures the causal rationality of the prediction process; component attention introduces an energy perception enhancement mechanism, which can adaptively adjust the fusion weight of each frequency component according to the characteristics of electricity price fluctuations.
[0067] (5) IWMA optimization, prediction value optimization strategy, selective quadratic decomposition, GDNN and two-layer attention mechanism form a progressive synergistic effect. Experimental results show that the CEEMDAN-IWMA-DNN basic model has lower MAE and RMSE than benchmark methods such as BiGRU and Transformer in normal scenarios; on this basis, after adding selective quadratic decomposition, GDNN and two-layer attention mechanism, the complete model shows high prediction accuracy and robustness in multiple typical time periods or typical scenarios, and is suitable for the task of predicting the clearing price of the electricity spot market. Attached Figure Description
[0068] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0069] Figure 1 This is an overall flowchart of the electricity clearing price prediction method of the present invention;
[0070] Figure 2 This is the flowchart of the improved whale migration algorithm IWMA;
[0071] Figure 3 This is the overall architecture of the CEEMDAN-IWMA-DNN model;
[0072] Figure 4 This is a flowchart of selective quadratic CEEMDAN decomposition;
[0073] Figure 5 This is a schematic diagram of the gated dynamic network (GDNN) structure;
[0074] Figure 6 This is a flowchart of the two-layer attention mechanism fusion process;
[0075] Figure 7 It is GEEMDAN 2 - Complete architecture diagram of GDNN-TLAttention model. Detailed Implementation
[0076] To gain a deeper understanding of this invention, we will provide a comprehensive and detailed description. However, this invention has various implementations and is not limited to the specific examples listed herein. These examples are presented to enhance a full understanding of the disclosure of this invention.
[0077] Example 1: Improved Whale Migration Algorithm IWMA and its Performance Verification
[0078] like Figure 2 As shown, this embodiment provides an improved whale migration algorithm, IWMA. Based on the standard whale migration algorithm WMA, this algorithm introduces four improvement strategies: Chebyshev chaotic initialization, lens imaging learning, non-uniform mutation, and adaptive Laplace crossover, to enhance the algorithm's population diversity, global search capability, local exploitation capability, and convergence stability.
[0079] First, a fourth-order Chebyshev chaotic map is used to generate the initial population, replacing the uniform random initialization method in the standard WMA. The Chebyshev chaotic sequence has strong ergodicity and randomness, which can make the initial whale individuals more evenly distributed in the search space and avoid the population from being overly concentrated in local areas during the initialization stage. Its mathematical expression is shown in Equation (1).
[0080] Secondly, in the early stages of algorithm iteration, the position of the leader whale is updated with a high probability using a lens imaging learning strategy. The lens imaging learning strategy expands the search range and improves the algorithm's ability to escape local optima by constructing candidate positions corresponding to the current leader whale position. Its position update formula is shown in equation (2). For the updated position of the leader whale, Let i be the position of the i-th individual. and Let be the lower and upper bounds of the search space, respectively, and k be the dynamic adjustment parameter for lens imaging. According to equation (2), the smaller k is, the greater the perturbation amplitude of the lens imaging; the larger k is, the closer the generated position is to the center of the search space. Therefore, a smaller k can be used in the early stage of iteration to enhance the global exploration capability, and k can be gradually increased in the later stage of iteration to improve the local fine search capability.
[0081] Furthermore, in the later stages of the algorithm iteration, a non-uniform mutation strategy is used with a high probability to perturb the position of the leader whale. The amplitude of the non-uniform mutation decreases non-linearly with the number of iterations, allowing the algorithm to maintain a large search step size in the early stages and gradually shift to local development in the later stages. Its position update formula is shown in Equation (3).
[0082] Furthermore, an adaptive Laplace crossover operation is performed on the current optimal and suboptimal solutions with preset probabilities, and the crossover formulas are shown in equations (4) and (5). The Laplace distribution has a peaked and heavy-tailed characteristic, which can simultaneously generate large-step global jumps and small-step local perturbations, thereby enhancing the algorithm's ability to search across regions and fine-tune locally. The switching probability between the lens imaging learning strategy and the non-uniform mutation strategy is shown in equation (6).
[0083] To verify the overall performance of the IWMA algorithm, this embodiment selects several representative standard test functions from the CEC2005 benchmark test function set for comparative experiments. The test functions cover single-peak, multi-peak, and composite terrain types. The comparison algorithms include GWO, WOA, ALO, IVY, and the original WMA. Each algorithm is run independently multiple times under the same parameter settings, and the mean (Ave), standard deviation (Std), and optimum (Best) are statistically analyzed. Experimental results show that, under the experimental conditions adopted, IWMA outperforms the comparison algorithms in terms of average fitness, stability, and optimum on the representative test functions. This indicates that the combination of Chebyshev chaotic initialization, lens imaging learning, non-uniform mutation, and adaptive Laplacian crossover can effectively improve the convergence accuracy, search stability, and ability to escape local optima of the WMA algorithm.
[0084] Example 2: Construction of the CEEMDAN-IWMA-DNN Basic Prediction Model
[0085] like Figure 3 As shown, this embodiment constructs a CEEMDAN-IWMA-DNN basic prediction model to achieve basic prediction of the clearing price in the electricity spot market. The model includes a data preprocessing layer, a primary CEEMDAN decomposition layer, an IWMA-DNN sub-prediction layer, and a weighted fusion layer of predicted values.
[0086] The experimental data came from a power generation company in Shanxi Province, selecting electricity spot market clearing price data from March 1, 2023 to February 28, 2024, with a sampling period of 15 minutes. Relevant features included the operating capacity of thermal power units, electricity load, power transmission plans, total thermal power plans, and total renewable energy plans. After missing value imputation, outlier handling, and normalization, the raw data was divided into training, validation, and test sets in chronological order.
[0087] In the decomposition layer, CEEMDAN is used to perform a first decomposition of the clearing price series, resulting in... There are 7 IMF components and 1 residual component. In this embodiment, one CEEMDAN decomposition yields 7 IMF components and 1 residual component, i.e. =7. Through a single CEEMDAN decomposition, the non-stationarity of the original clearing price series can be reduced, and the feature representation ability at different frequency scales can be enhanced.
[0088] In the optimization layer, the IWMA algorithm described in Example 1 is used to optimize the key hyperparameters of the DNN model. The hyperparameters to be optimized include the learning rate, batch size, number of network layers, and number of neurons per layer. IWMA encodes these hyperparameters as candidate solution positions and uses the training set loss and validation set prediction error as the fitness evaluation criteria to perform global optimization within a preset search space to obtain the optimal combination of hyperparameters.
[0089] In the prediction layer, IWMA-DNN sub-prediction models are established for each IMF component and residual component. Each sub-prediction model takes the historical features extracted by the sliding window as input and the predicted value of the corresponding component at the next time step as output. After the prediction of each component is completed, an adaptive weighted fusion is performed using a prediction value optimization strategy. The weight calculation formula is shown in Equation (7). The component with the smaller prediction error receives a higher weight, thus obtaining the basic prediction value of CEEMDAN-IWMA-DNN.
[0090] Experimental results show that compared with benchmark prediction models such as BiGRU and Transformer, the CEEMDAN-IWMA-DNN basic prediction model has reduced MAE and RMSE, indicating that the fusion strategy of CEEMDAN decomposition, IWMA hyperparameter optimization and prediction value optimization can improve the accuracy of clearing price prediction in conventional scenarios.
[0091] Example 3: Construction of the complete CEEMDAN²-GDNN-TLAttention model
[0092] Building upon Example 2, and addressing the issue of insufficient prediction accuracy during extreme fluctuations, this example further introduces selective quadratic CEEMDAN decomposition, gated dynamic network GDNN, temporal attention mechanism, and component attention mechanism to construct a complete prediction model: CEEMDAN²-GDNN-TLAttention. The overall framework of the complete model is as follows: Figure 7 As shown.
[0093] First, such as Figure 4 As shown, based on the CEEMDAN-IWMA-DNN basic prediction model trained in Example 2, the prediction error contribution of each IMF component is statistically analyzed. The error contribution of each component is calculated according to Equation (8). The IMF components are sorted from high to low according to their error contribution, and the components with a cumulative contribution exceeding a preset threshold of 50% are selected. A secondary CEEMDAN decomposition is performed on the high-frequency IMF components. In this embodiment, IMF1, IMF2, and IMF3 are preferably selected as the secondary decomposition objects, i.e. =3; IMF4 to IMF7 and the residual components remain unchanged from the first decomposition result. During the second decomposition, the correlation coefficient between two consecutive sub-IMFs is calculated. ,when > When the correlation between adjacent sub-IMFs is high, further decomposition makes it difficult to obtain independent oscillation modes, and the secondary decomposition is terminated at this point; the preferred threshold is... =0.7. After selective quadratic decomposition, the total number of components participating in the final prediction fusion is denoted as . .
[0094] Secondly, such as Figure 5 As shown, a GDNN prediction sub-model is established for each component after selective quadratic decomposition. GDNN embeds gating units between the hidden layers of the DNN. The gating units include update gates and reset gates, and their forward propagation process is shown in equations (9) to (13). Through the gating mechanism, GDNN can adaptively adjust the information flow of different frequency components; through residual gating updates, GDNN can enhance gradient propagation ability and alleviate the gradient vanishing problem in the training process of deep networks. In order to reduce the number of model parameters and maintain the differentiated expressive ability of different frequency components, the components are divided into high-frequency group, mid-frequency group and low-frequency group according to frequency characteristics. The gating parameters are shared within the group and the parameters are independent between the groups.
[0095] Again, such as Figure 6 As shown, a temporal attention mechanism is applied to the hidden state sequence output by the GDNN. Let the input time window length be... The hidden state sequence output by GDNN is The hidden state sequence is mapped to the query matrix Q and the key matrix using a learnable weight matrix. The sum matrix V is used, and the attention score is calculated using a scaled dot product. To ensure the causal rationality of the temporal prediction, causal masking is applied to the attention score matrix, ensuring that each time step can only focus on current and historical information, and cannot use future information. Then, the attention weight matrix is obtained through softmax normalization, and the value vectors are weighted and aggregated to obtain the temporal context vector. Finally, residual connections and layer normalization structures are used to improve the model's training stability and convergence speed.
[0096] Finally, as Figure 6 As shown, a component attention mechanism is applied to the predicted features of each final component. First, the energy of each component on the training set is calculated. Energy characteristics were obtained through logarithmic smoothing. Then, the energy features are concatenated with the dynamic prediction features to obtain the enhanced features. Then, the attention score of each component at time t is calculated using a two-layer feedforward neural network, and the fusion weights are obtained by softmax normalization. Final predicted value The result is obtained by weighted summation of the prediction results of each component, as shown in Equation (29). This component attention mechanism can adaptively adjust the fusion weights according to the differences in the contributions of different components at different prediction times, thereby improving the prediction accuracy in extreme fluctuation scenarios.
[0097] During the end-to-end training phase, the model outputs the predicted clearing price. With true clearing price The results are compared and optimized using the composite loss function shown in Equation (30). The composite loss function consists of a weighted average of MSE and MAPE, where λ∈(0,1), preferably λ=0.7. MSE measures the squared deviation between the predicted and true values, while MAPE measures the proportional deviation of the prediction error relative to the true value. Through this composite loss function, the model can simultaneously consider both overall fitting accuracy and relative error control.
[0098] In summary, this invention achieves high-precision prediction of clearing prices in the electricity spot market through the collaborative design of improved whale migration algorithm, selective quadratic CEEMDAN decomposition, gated dynamic network, and two-layer attention mechanism. It is particularly suitable for clearing price sequences with strong non-stationarity, high volatility, and extreme peak characteristics.
[0099] Contents not described in detail in this specification are prior art known to those skilled in the art. Although illustrative specific embodiments of the invention have been described above to facilitate understanding by those skilled in the art, it should be understood that the invention is not limited to the scope of the specific embodiments. Various modifications are readily apparent to those skilled in the art as long as they fall within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of this invention are protected.
Claims
1. A method for predicting electricity clearing prices based on an improved whale migration algorithm and a two-layer attention mechanism, characterized in that, Includes the following steps: Step 1: Obtain historical clearing price data and related characteristic data of the electricity spot market, and preprocess the data; Step 2: The clearing price series is decomposed using CEEMDAN (Adaptive Noise Complete Set Empirical Mode Decomposition). Gaussian white noise is adaptively added to obtain... One intrinsic mode function (IMF) component and one residual component; Step 3: Construct the basic prediction model of the improved whale migration algorithm-optimized deep neural network IWMA-DNN. The IWMA introduces Chebyshev chaotic mapping initialization, lens imaging learning strategy, non-uniform mutation strategy and adaptive Laplacian cross operation on the basis of the standard whale migration algorithm WMA. The hyperparameters of the DNN are globally optimized using the improved whale migration algorithm IWMA, and IWMA-DNN sub-prediction models are established for each IMF component and residual component. The basic prediction value is obtained by fusion through an adaptive weight strategy. Step 4: Calculate the prediction error contribution of each IMF component and residual component, and identify the high-frequency IMF components from among them. Arrange the high-frequency IMF components in descending order of prediction error contribution and accumulate them sequentially. Select the components whose cumulative contribution first reaches or exceeds a preset threshold. The high-frequency IMF component is used as the object of secondary CEEMDAN decomposition. The object of secondary CEEMDAN decomposition is decomposed into sub-IMF components. The sub-IMF components, the IMF components that have not undergone secondary decomposition, and the residual components constitute the final set of components that participate in prediction fusion. Step 5: Construct a gated dynamic network (GDNN) prediction sub-model for each final component in the final component set. The GDNN embeds a gated unit containing update gates and reset gates after the intermediate hidden layer of the DNN, and uses residual-form gated updates to modulate the inter-layer information flow. Step 6: Apply a temporal attention mechanism to the hidden state sequence output by the GDNN, calculate the attention weights at different time steps through a query-key-value mechanism, and use a causal masking mechanism to prevent information leakage at future time steps; Step 7: Apply component attention mechanism to the prediction features of each final component, introduce energy perception enhancement, calculate the attention fusion weight of each final component through a feedforward neural network, realize adaptive weighted fusion of multi-scale component prediction results, and obtain the final clearing price prediction value.
2. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, In step 1, the relevant feature data includes at least one of the following: thermal power unit operating capacity, power load, power transmission plan, total thermal power plan, and total new energy plan; The preprocessing includes missing value handling, outlier handling, normalization, and time order partitioning; The time-order partitioning includes dividing the samples into a training set, a validation set, and a test set according to time sequence. The validation set is used for hyperparameter optimization, and the test set is used for final prediction performance evaluation.
3. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, Step 3 specifically includes the following sub-steps: Step 3.1: Initialize the population using a 4th-order Chebyshev chaotic map, and then map each hyperparameter to be optimized to its corresponding search space based on its upper and lower bounds to generate an initial whale population; the mathematical expression for the Chebyshev chaotic map is: (1), where, For the first Chaotic variables generated in the next iteration; when ∈ When used for population initialization, a fourth-order Chebyshev mapping is employed to generate chaotic sequences; if used for population initialization, the generated sequences are linearly mapped to the search space or... interval; Step 3.2: In the early stages of algorithm iteration, a lens imaging learning strategy is used to update the position of the leader whale. The position update formula is as follows: (2), where, For the updated position of the leader whale, For the first Individual position, and These are the lower and upper bounds of the search space, respectively. The parameters are dynamically adjusted for lens imaging; lens imaging mapping is performed on the current optimal solution according to the dynamic parameter k to generate candidate solutions. Use small dynamic parameters in the early stages of iteration. To broaden the search scope, the search area is gradually increased in the later stages of iteration. Or reduce the mapping perturbation amplitude to enhance local fine-grained search capabilities; when When the search space boundary is exceeded, Perform boundary truncation or mapping back to the preset search space; Step 3.3: In the later stages of the algorithm iteration, a non-uniform mutation strategy is used to update the position of the leader whale. The position update formula is as follows: (3), among which, It is a non-uniform variation perturbation factor. for Random numbers between This represents the current iteration number. The maximum number of iterations, Parameters used to control the rate of decay of variation intensity; and These are the lower and upper bounds of the search space, respectively. Step 3.4: Apply a preset probability to the current optimal solution. and suboptimal solutions Perform an adaptive Laplace crossover operation, with the crossover formula as follows: (4), (5), among which, For random variables that follow a Laplace distribution; by utilizing the peaked and heavy-tailed characteristics of the Laplace distribution, both large-step global jumps and small-step local fine-tuning are achieved. If the updated individual exceeds the search space boundary, it is truncated or mapped back to the preset boundary range. Step 3.5: The switching probability between the lens imaging learning strategy and the non-uniform mutation strategy is dynamically adjusted with the number of iterations. (6), among which, The probability of employing a global exploration strategy. This represents the maximum number of iterations; in the early stages of iteration, This term is close to 1; in the later stages of the iteration, This value approaches 0, thus enabling a smooth transition from global exploration to local development; Step 3.6: Optimize the four hyperparameters of the DNN—learning rate, batch size, number of network layers, and number of neurons per layer—using the IWMA algorithm. Reconstruct the hyperparameter optimization problem into a global optimization problem in a high-dimensional non-convex space. Determine the optimal combination of hyperparameters based on the training set loss and the validation set prediction error as fitness evaluation criteria. Step 3.7: Establish IWMA-DNN sub-prediction models for each IMF component and residual component, and assign adaptive weights to each component using a prediction value optimization strategy; calculate the weights based on the mean absolute error (MAE) of each component on the training set, preferably using normalized weights that are inversely proportional to the MAE. (7), among which, For the first The weight of each component, For the first Each component corresponds to the mean absolute error of the sub-prediction model on the training set. To prevent the stability constant from having a denominator of zero, the component with the smaller prediction error is given a higher weight. Finally, the prediction results of each component are weighted and fused to obtain the basic prediction value.
4. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, The execution probabilities of the lens imaging learning strategy and the non-uniform mutation strategy are dynamically adjusted with the number of iterations. In the early stage of the iteration, the execution probability of the lens imaging learning strategy is higher than that of the non-uniform mutation strategy, while in the later stage of the iteration, the execution probability of the non-uniform mutation strategy is higher than that of the lens imaging learning strategy, so as to achieve a dynamic balance between global exploration and local development.
5. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, The hyperparameters used in step 3 to globally optimize the hyperparameters of the DNN using the Improved Whale Migration Algorithm (IWMA) include: learning rate, batch size, number of network layers, and number of neurons per layer; the fitness function of the hyperparameters is determined based on the prediction error on the validation set. The adaptive weighting strategy includes: based on the average absolute error of the IWMA-DNN sub-prediction model corresponding to each IMF component and residual component on the training or validation set. Calculate fusion weights , ;in, To prevent the stability constant from being zero, and to give higher weights to components with smaller prediction errors, the prediction results of each component are finally weighted and fused to obtain the basic prediction value.
6. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, In step 4, the contribution of prediction error Represented as: (8), among which, Indicates the first The root mean square error of each component; the preset threshold The threshold is 50%; the termination condition for the second-order CEEMDAN decomposition includes: calculating the correlation coefficient between two consecutive sub-IMF components. ,when Greater than the preset relevant threshold When the time is reached, the iterative decomposition stops; where, The value range is from 0.6 to 0.
8.
7. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, Step 5 includes gating units for updating and resetting gates, and uses residual gating updates to modulate the inter-layer information flow. Specifically: The forward propagation process of the gating unit is as follows: (9), (10) (11), (12); Expressions in residual form: (13), among which, Let i be the hidden state of the i-th final component in the l-th layer. To update the door, To reset the door, This is a candidate hidden state. It is the Sigmoid activation function. For element-wise multiplication, , , and , , These are the learnable parameters in the gating unit; GDNN employs a grouped parameter sharing strategy: the final components in the final component set are divided into high-frequency, mid-frequency, and low-frequency groups according to their frequency characteristics. Gating parameters are shared within each group, while the gating parameters between groups are independent. The gating unit uses a central placement scheme, inserted after the intermediate hidden layer of the DNN; this intermediate hidden layer is the nth hidden layer, where n is the lowest frequency. The integer obtained by rounding up. This represents the total number of hidden layers in the DNN.
8. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, Step 6 specifically includes the following steps: Step 6.1: Linear projection, using three sets of learnable weight matrices to project the hidden state sequence. Mapped to queries respectively ,key Sum : (14) (15) (16); among which, , , For learnable parameter matrix, For projection dimensions; Step 6.2: Calculate the attention score using the scaled dot product method for the query. AND key Correlation score matrix between them: (17) Among them, This is a scaling factor used to prevent the dot product result from becoming too large, causing the softmax to enter the gradient saturation region; Step 6.3: Causal Masking. A causal masking mask is applied to the score matrix, setting the scores corresponding to future time steps to negative infinity. This ensures that each time step can only focus on itself and its preceding historical information, conforming to the causal constraints of time series prediction. The score matrix after causal masking is represented as follows: (18); Step 6.4: Attention weight normalization. Apply the softmax function to each row of the masked score matrix to obtain the normalized attention weight matrix: (19); Step 6.5: Weighted aggregation, using attention weights to aggregate the value vector Perform a weighted summation to generate a time series context vector: (20) (twenty one); Step 6.6: Residual Connections and Layer Normalization. To alleviate gradient vanishing in deep networks and accelerate training convergence, a residual connection and layer normalization structure is adopted: (twenty two); make Consistent with the hidden state dimension; or through linear projection Mapped to After establishing the same dimensions, perform residual joins, where... To perform standardization along the feature dimension, the calculation formula is as follows: (23), among which, and For learnable scaling and translation parameters in layer normalization, is the numerical stability constant.
9. The electricity clearing price prediction method based on an improved whale migration algorithm and a two-layer attention mechanism according to claim 1, characterized in that, Step 7 includes Energy-sensing enhancement is introduced into the prediction features of each final component, and the energy of each component is calculated: (24), among which, For the length of the training set, The value of the i-th component participating in the final fusion at time t is given. This final component includes the sub-IMF component obtained from the second decomposition, the IMF component not subjected to the second decomposition, and the residual term. After logarithmic smoothing of the energy, we obtain: (25); The enhanced feature is obtained by concatenating the energy feature with the predicted feature: (26), among which, Let i be the dynamic prediction feature of the i-th component at time t. To dynamically predict feature dimensions, the feature concatenation operation is represented; attention scores are calculated through a two-layer feedforward neural network, and the fusion weights are obtained after softmax normalization. (27) (28), among which, , , and For the learnable parameters in the component attention network, the final predicted value is obtained by weighted summation of the prediction results of each component: (29), among which, The total number of components that ultimately participate in the fusion after selective secondary decomposition includes the sub-IMF components obtained from the secondary decomposition, the IMF components that have not undergone secondary decomposition, and the residual terms. Let be the fusion weight of the i-th component at time t. Let be the predicted value of the i-th component at time t. This is the projected final clearing price.
10. A power clearing price prediction system based on an improved whale migration algorithm and attention mechanism, characterized in that, include: The data acquisition and preprocessing module is used to acquire historical clearing price data and related characteristic data of the electricity spot market and perform preprocessing. A single CEEMDAN decomposition module is used to perform a single decomposition on the clearing price series to obtain... One IMF component and one residual component; The IWMA-DNN basic prediction module is used to construct a deep neural network basic prediction model for an improved whale migration algorithm. The IWMA introduces a Chebyshev chaotic mapping initialization strategy, a lens imaging learning strategy, a non-uniform mutation strategy, and an adaptive Laplacian cross operation strategy on the basis of the standard whale migration algorithm WMA. The IWMA-DNN basic prediction module establishes IWMA-DNN sub-prediction models for each IMF component and residual component, and obtains the basic prediction value by fusing them through an adaptive weight strategy. The selective quadratic decomposition module is used to calculate the prediction error contribution of each IMF component and the residual component, and select some high-frequency IMF components from each IMF component for secondary CEEMDAN decomposition based on the prediction error contribution. The gated dynamic network module is used to build a GDNN prediction sub-model for each final component in the final component set; The temporal attention module is used to calculate the attention weights at each time step on the hidden state sequence output by the GDNN through a query-key-value mechanism. The component attention module is used to calculate the attention fusion weights for the prediction features of each final component and generate the final prediction value. The end-to-end training module is used to jointly optimize system parameters based on a composite loss function.