A short-term power load forecasting method and system based on collaborative optimization
The Antlion algorithm, which optimizes the hyperparameters of the power load forecasting model through an adaptive information sharing mechanism, solves the problems of low efficiency and poor forecasting effect in existing technologies, achieves more efficient and accurate load forecasting, and improves the operational stability of the power system.
Patent Information
- Application Number
- CN202511241248.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing AI-based power load forecasting algorithms suffer from low hyperparameter optimization efficiency and are prone to getting trapped in local optima when faced with high-dimensional and complex search spaces. Furthermore, the lack of information sharing mechanisms among individual algorithms leads to poor performance of the forecasting models when dealing with complex time-series information.
The antlion algorithm based on an adaptive information sharing mechanism is used for hyperparameter optimization. By mapping the hyperparameters of the load prediction model to the multidimensional location vector of the antlion population, an objective function is established, and the optimal hyperparameter combination is optimized by iteratively solving the problem using the antlion algorithm based on the adaptive information sharing mechanism.
It significantly improves the efficiency and accuracy of hyperparameter optimization, enhances the accuracy and robustness of load forecasting, and can better capture the nonlinear characteristics and complex time-series relationships in load data, thereby improving the operating efficiency and stability of the power system.
Smart Images

Figure CN120744465B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power load forecasting technology, and in particular to a short-term power load forecasting method and system based on collaborative optimization. Background Technology
[0002] Electricity load forecasting is primarily used to formulate electricity production plans and arrange short-term power system operation modes, making it an essential and crucial part of power system operation. With the deepening of research on load forecasting, the research trend has gradually shifted from classical forecasting algorithms based on statistical thinking to artificial intelligence forecasting algorithms based on machine learning.
[0003] While existing AI prediction algorithms can utilize data features for model building, the strong volatility and nonlinearity of load data lead to limitations in data processing capabilities, insufficient feature capture, and difficulty in modeling complex temporal information. Deep learning models, though capable of handling complex temporal relationships, are highly dependent on hyperparameter selection. Traditional hyperparameter optimization methods, such as grid search and random search, are inefficient and prone to getting trapped in local optima. Intelligent optimization algorithms, such as particle swarm optimization and genetic algorithms, while improving hyperparameter search efficiency to some extent, suffer from slow convergence and less than ideal optimization results when facing high-dimensional and complex search spaces. Furthermore, most optimization algorithms lack information-sharing mechanisms during individual search processes, resulting in weak collaboration among individuals and difficulty in fully utilizing global information to guide the search in complex search spaces. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a short-term power load forecasting method and system based on collaborative optimization, which can effectively cope with fluctuations in power load and improve the accuracy and robustness of load forecasting.
[0005] In a first aspect, the present invention provides a short-term power load forecasting method based on collaborative optimization, the method comprising:
[0006] The load-related data of the power system is acquired, and the load-related data is preprocessed to obtain a data vector. The load-related data includes load data, meteorological data, and date data.
[0007] The data vector is input into a preset load prediction model to obtain the power load prediction result. The load prediction model is constructed using a deep neural network model based on a self-attention mechanism and the hyperparameters are optimized using an antlion algorithm with an adaptive information sharing mechanism.
[0008] The steps of hyperparameter optimization using the antlion algorithm with an adaptive information sharing mechanism include:
[0009] The hyperparameters of the load prediction model are mapped to the multidimensional location vectors of the antlion population, and an objective function is established based on the model performance and the rationality of the hyperparameters.
[0010] The Antlion algorithm, based on an adaptive information sharing mechanism, iteratively solves the objective function to obtain the optimal combination of hyperparameters.
[0011] Further, the step of preprocessing the load-related data to obtain a data vector:
[0012] Wavelet decomposition and reconstruction are performed on the load data to obtain the reconstructed load sequence;
[0013] Perform sine and cosine position encoding on the date data to obtain the time encoding vector;
[0014] The meteorological data, the reconstructed load sequence, and the time-coded vector are normalized to obtain a data vector;
[0015] The reconstructed load sequence is represented by the following formula:
[0016]
[0017] In the formula, Indicates the reconfigured load sequence. Describes the wavelet transform function. Represents load data, Represents a linear correction unit. This represents the coefficients of the third-level wavelet decomposition. This represents the coefficients of the fourth-level wavelet decomposition. This represents the standard deviation function.
[0018] Furthermore, the step of mapping the hyperparameters of the load prediction model to a multidimensional location vector of the antlion population, and establishing the objective function based on model performance and the rationality of the hyperparameters, includes:
[0019] The hyperparameters of the load prediction model are mapped to the multidimensional location vector of the antlion population, and the antlion population is initialized. The hyperparameters include the number of attention heads, the dimension of the hidden layer, and the learning rate.
[0020] The ratio of the number of attention heads to the dimension of the hidden layer is used as the redundancy of the model parameters, and the difference between the learning rate and the learning rate baseline is used as the learning rate deviation.
[0021] The objective function is to minimize the cross-entropy loss, model parameter redundancy, and learning rate deviation of the load prediction model on the validation set.
[0022] The objective function is represented by the following formula:
[0023]
[0024] In the formula, CE represents the cross-entropy loss. The cross-entropy loss weights are represented by HD, which indicates the redundancy of the model parameters. XD represents the model parameter redundancy weights, and XD represents the learning rate deviation. Indicates the weight of the learning rate deviation. This represents the minimum value function.
[0025] Furthermore, the step of iteratively solving the objective function using the antlion algorithm based on the adaptive information sharing mechanism to obtain the optimal hyperparameter combination includes:
[0026] Based on the random walk of ants, calculate the antlion position matrix and the corresponding antlion fitness matrix, and select non-elite antlions with fitness below the fitness threshold from the antlion position matrix.
[0027] The positions of the selected non-elite antlions are updated based on the positions of the elite antlions, and the antlion position matrix is updated based on the updated positions of the non-elite antlions.
[0028] Based on the updated antlion position matrix, the antlion positions and the optimal antlion positions are obtained;
[0029] The antlion position is optimized based on the preset exploration weight, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position. The exploration weight includes local exploration weight and global exploration weight.
[0030] The optimized antlion position is iteratively calculated according to the above steps until the iteration stops, and the optimal hyperparameter combination is obtained.
[0031] Furthermore, the step of optimizing the antlion position based on a preset exploration weight, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position includes:
[0032] Calculate the first difference between the antlion's historical best position and the antlion's current position, and calculate the second difference between the antlion's best position and the antlion's historical best position;
[0033] The antlion position is optimized based on the product of the local exploration weight and the first difference, and the product of the global exploration weight and the second difference, to obtain the optimized antlion position.
[0034] The optimized antlion position is represented by the following formula:
[0035]
[0036] In the formula, This represents the position of the antlion after the t-th iteration of optimization. This represents the position of the antlion before the t-th iteration of optimization. Indicates the local exploration weight, Indicates the global exploration weight. This represents the historical best position of the antlion in the t-th iteration. This represents the optimal position of the antlion in the t-th iteration.
[0037] Furthermore, the step of optimizing the antlion position based on preset exploration weights, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position also includes:
[0038] The information sharing intensity of each antlion is calculated based on the individual fitness and the average fitness of the group.
[0039] Based on the comparison between the information sharing intensity and the intensity threshold, the preset exploration weight is adjusted, and the antlion position is optimized based on the adjusted exploration weight.
[0040] The intensity of information sharing is represented by the following formula:
[0041]
[0042] In the formula, f represents the information sharing strength of the j-th antlion in the t-th iteration, α represents the adjustment factor, β represents the sensitivity factor, and f m,t f represents the average fitness of the antlions in the current population during the t-th iteration. t,j Let represent the individual fitness of the j-th antlion in the t-th iteration.
[0043] Furthermore, the step of adjusting the preset exploration weights based on the comparison relationship between the information sharing intensity and the intensity threshold includes:
[0044] Determine whether the information sharing intensity is greater than a first intensity threshold. If it is greater than the first intensity threshold, increase the preset local exploration weight by a preset ratio to obtain the adjusted local exploration weight. Otherwise, determine whether the information sharing intensity is less than a second intensity threshold.
[0045] If it is less than the second intensity threshold, the preset global exploration weight is increased by a preset ratio to obtain the adjusted global exploration weight.
[0046] Furthermore, the step of iteratively solving the objective function using the antlion algorithm based on the adaptive information sharing mechanism to obtain the optimal hyperparameter combination also includes:
[0047] Based on the fitness improvement rate of the current hyperparameter combination compared to the previous iteration, the current number of attention heads in the current hyperparameter combination is updated to obtain the updated number of attention heads.
[0048] The current hyperparameter combination is updated based on the updated number of attention heads;
[0049] The updated number of attention heads is represented by the following formula:
[0050]
[0051] In the formula, This indicates the updated number of attention heads, where h represents the current number of attention heads. Indicates the fitness improvement rate. This represents the search intensity coefficient.
[0052] Furthermore, the iteration stopping condition is expressed by the following formula:
[0053]
[0054] In the formula, Let represent the mean absolute error of the validation set in the t-th iteration. This represents the mean absolute error of the validation set for the first t-5 iterations.
[0055] Secondly, the present invention provides a short-term power load forecasting system based on collaborative optimization, the system comprising:
[0056] The data preprocessing module is used to acquire load-related data of the power system and perform data preprocessing on the load-related data to obtain a data vector. The load-related data includes load data, meteorological data, and date data.
[0057] The load forecasting module is used to input the data vector into a preset load forecasting model to obtain the power load forecasting result. The load forecasting model is constructed using a deep neural network model based on a self-attention mechanism and uses an antlion algorithm with an adaptive information sharing mechanism for hyperparameter optimization.
[0058] The steps of hyperparameter optimization using the antlion algorithm with an adaptive information sharing mechanism include:
[0059] The hyperparameters of the load prediction model are mapped to the multidimensional location vectors of the antlion population, and an objective function is established based on the model performance and the rationality of the hyperparameters.
[0060] The Antlion algorithm, based on an adaptive information sharing mechanism, iteratively solves the objective function to obtain the optimal combination of hyperparameters.
[0061] This invention provides a short-term power load forecasting method and system based on collaborative optimization. The technical advantages of this invention include:
[0062] ① Optimize hyperparameter efficiency: The antlion algorithm with an adaptive information sharing mechanism provides efficient hyperparameter search capabilities, which can significantly shorten the hyperparameter optimization time compared with traditional grid search and random search methods, while ensuring the global optimality of the optimization results.
[0063] ② Improve prediction accuracy: This invention optimizes the hyperparameters of the load prediction model through the Antlion algorithm with an adaptive information sharing mechanism, enabling the model to more accurately capture the nonlinear characteristics and complex time series relationships in the load data, thereby significantly improving prediction accuracy.
[0064] ③ Improve system operating efficiency: The load prediction model optimized by the Antlion algorithm with an adaptive information sharing mechanism in this invention can converge quickly and achieve high-precision prediction, thereby adapting to changes in different load characteristics, which helps the real-time load scheduling and management of the power system and improves the overall operating efficiency and stability. Attached Figure Description
[0065] Figure 1 This is a flowchart illustrating the short-term power load forecasting method based on collaborative optimization in an embodiment of the present invention.
[0066] Figure 2 This is a schematic diagram of the structure of the short-term power load forecasting system based on collaborative optimization in an embodiment of the present invention;
[0067] Figure label:
[0068] 10. Data preprocessing module; 20. Load forecasting module. Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0070] Please see Figure 1 The first embodiment of the present invention proposes a short-term power load forecasting method based on collaborative optimization, which includes steps S10 to S20:
[0071] Step S10: Obtain load-related data of the power system and perform data preprocessing on the load-related data to obtain a data vector. The load-related data includes load data, meteorological data, and date data.
[0072] Step S20: Input the data vector into a preset load prediction model to obtain the power load prediction result. The load prediction model is constructed using a deep neural network model based on a self-attention mechanism and uses the Antlion algorithm with an adaptive information sharing mechanism for hyperparameter optimization.
[0073] In this invention, a deep learning model is used to predict power load. The power load-related data used include load data, meteorological data, and date data. Load data refers to the load value per hour or per minute; meteorological data refers to environmental parameters related to load fluctuations, such as temperature, humidity, and wind speed; and date data refers to time attributes related to load fluctuations, such as holidays, seasons, and weekdays / weekends.
[0074] Before inputting this data into the load forecasting model, data preprocessing is required. Specific preprocessing steps include:
[0075] Wavelet decomposition and reconstruction are performed on the load data to obtain the reconstructed load sequence;
[0076] Perform sine and cosine position encoding on the date data to obtain the time encoding vector;
[0077] The meteorological data, the reconstructed load sequence, and the time-coded vector are normalized to obtain a data vector.
[0078] In this embodiment, the load data is first decomposed and reconstructed using wavelet decomposition:
[0079]
[0080] In the formula, Indicates the reconfigured load sequence. Describes the wavelet transform function. Represents load data, Represents a linear correction unit. This represents the coefficients of the third-level wavelet decomposition. This represents the coefficients of the fourth-level wavelet decomposition. This represents the standard deviation function.
[0081] In this system, wavelet decomposition coefficients reflect the short-term fluctuation characteristics of the load data, standard deviation is used to dynamically adjust the enhancement amplitude of high-frequency components, and linear correction units are used to filter out negative detail noise while retaining effective high-frequency information. This is achieved by superimposing the low-frequency trend of the original load. With enhanced key high-frequency details Generate reconstructed load data This allows the reconstructed load sequence to retain the overall trend while highlighting important fluctuation characteristics.
[0082] For date data, sine and cosine positional encoding is used to convert the linear time s into a periodic two-dimensional vector, enabling subsequent load forecasting models to perceive the daily periodic variation of the load, such as morning and evening peak hours. The time-series feature encoding formula is as follows:
[0083]
[0084] Among them, 1440 represents the total number of minutes in a day, through... Map time s to The radian range allows timestamps to be transformed into periodic function inputs. Sine and cosine positional encoding is a low-dimensional continuous representation that can more naturally express the cyclical nature of time.
[0085] Finally, the meteorological data, reconstructed load series, and time-coded vectors are normalized to obtain a data vector, which is then used as input data for the load forecasting model.
[0086] In this embodiment, the load forecasting model is constructed using the Transformer deep neural network model based on the self-attention mechanism. Since the performance of the model depends on accurate hyperparameter tuning, in order to solve the problem that existing optimization methods are inefficient in complex hyperparameter spaces and are prone to getting trapped in local optima, thus directly affecting the performance and application effect of the forecasting model, this invention proposes an Antlion algorithm based on an adaptive information sharing mechanism to tune the hyperparameters of the load forecasting model, thereby improving the adaptability of the forecasting model to short-term load data and the forecasting accuracy.
[0087] In a preferred embodiment, the steps of hyperparameter optimization using the antlion algorithm with an adaptive information sharing mechanism include:
[0088] The hyperparameters of the load prediction model are mapped to the multidimensional location vectors of the antlion population, and an objective function is established based on the model performance and the rationality of the hyperparameters.
[0089] The Antlion algorithm, based on an adaptive information sharing mechanism, iteratively solves the objective function to obtain the optimal combination of hyperparameters.
[0090] In this embodiment, initialization is performed first, which involves defining the search space and dynamic constraint rules for the model's hyperparameters and initializing the antlion population. Specifically, the hyperparameters of the Transformer model are mapped to multidimensional position vectors of the antlion population. The model's hyperparameters include the number of attention heads, the hidden layer dimension, and the learning rate. The dynamic constraint rule for the number of attention heads is to initialize it to uniform sampling {2, 4, 6, 8}, and reduce the number of heads if memory usage is >80%. The dynamic constraint rule for the hidden layer dimension is that the hidden layer dimension is an integer multiple of the number of attention heads, and the hidden layer dimension is reduced if memory usage is >80%. The dynamic constraint rule for the learning rate is logarithmically uniform sampling [1e-5, 1e-3], and the search range is narrowed to ±10% of the current value when the gradient oscillation amplitude is too large. Then, an initial population is randomly generated, with each individual representing a set of hyperparameter combinations.
[0091] In this embodiment, an objective function for iterative optimization of the parameters is established. This objective function is constructed based on model performance and the rationality of hyperparameters. Specifically, the ratio of the number of attention heads to the dimension of the hidden layer is used as the model parameter redundancy (HD).
[0092]
[0093] In the formula, Heads represents the number of attention heads, and dmodel represents the dimension of the hidden layer;
[0094] The difference between the learning rate and the baseline learning rate is used as the learning rate deviation XD:
[0095]
[0096] In the formula, θ represents the learning rate, θ base This represents the baseline learning rate.
[0097] Then, the objective function F is to minimize the cross-entropy loss, model parameter redundancy, and learning rate deviation of the load prediction model on the validation set:
[0098]
[0099] In the formula, CE represents the cross-entropy loss. The cross-entropy loss weights are represented by HD, which indicates the redundancy of the model parameters. XD represents the model parameter redundancy weights, and XD represents the learning rate deviation. Indicates the weight of the learning rate deviation. This represents the minimum value function.
[0100] In this embodiment, the performance of the prediction model is measured by cross-entropy loss, the redundancy of model parameters is used to avoid excessive head count leading to wasted computing resources, and the learning rate deviation is used to prevent the learning rate from deviating too much from the base value, thereby enhancing the stability of training.
[0101] After the above initialization steps are completed, this embodiment uses the Antlion algorithm based on an adaptive information sharing mechanism to iteratively solve the objective function, thereby obtaining the optimal hyperparameter combination. The specific steps include:
[0102] Based on the random walk of ants, calculate the antlion position matrix and the corresponding antlion fitness matrix, and select non-elite antlions with fitness below the fitness threshold from the antlion position matrix.
[0103] The positions of the selected non-elite antlions are updated based on the positions of the elite antlions, and the antlion position matrix is updated based on the updated positions of the non-elite antlions.
[0104] Based on the updated antlion position matrix, the antlion positions and the optimal antlion positions are obtained;
[0105] The antlion position is optimized based on the preset exploration weight, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position. The exploration weight includes local exploration weight and global exploration weight.
[0106] The optimized antlion position is iteratively calculated according to the above steps until the iteration stops, and the optimal hyperparameter combination is obtained.
[0107] In this embodiment, the AntLion Optimizer (ALO) mainly consists of five steps: ants randomly walking, antlions setting traps, antlions using traps to catch ants, capturing prey, and resetting traps. In this algorithm, ants choose to randomly walk to simulate movement, realizing the antlion's global search. The antlion's position is equivalent to the solution to the problem. The antlion preys on ants with high fitness and replaces their position to continuously update the optimal solution. The optimization steps of the antlion algorithm are briefly explained below.
[0108] In each optimization step, ants update their positions through random walks, and each random walk remains within the search space. The ants' random walks are influenced by antlion traps, which can be understood as ants randomly walking within a hypersphere surrounding a selected antlion. To simulate the hunting ability of antlions, assuming each ant is trapped in a selected antlion, the ALO algorithm uses a roulette wheel selection method to choose antlions based on the ants' fitness during the optimization process. This mechanism provides a high chance for antlions with higher fitness to capture ants. Antlions can build traps proportional to their fitness. Once an antlion realizes there is an ant in the trap, it will shoot sand outwards from the center of the trap. This behavior will deflect trapped ants attempting to escape. Data modeling represents this behavior as the radius of the hypersphere where the ants' random walks adaptively decreases.
[0109] The final stage of the hunt is when an ant reaches the bottom of the trap and is caught by the antlion. To simulate this process, it is assumed that prey-catching behavior occurs when an ant becomes healthier than its corresponding antlion. The antlion is then required to update its position to the latest position of the captured ant to increase its chances of catching new prey. That is:
[0110]
[0111] In the formula, This represents the position of the j-th antlion in the t-th iteration. This represents the position of the i-th ant in the t-th iteration.
[0112] Elite selection is a key feature of the ALO algorithm. The best antlion obtained in each iteration is saved and considered an elite. Since the elite is the most fit antlion, it can influence the movement of all ants during the iteration process. Therefore, suppose each ant, along with the elite, randomly moves around a selected ant using a roulette wheel selection method, as shown below:
[0113]
[0114] In the formula, where The antlion is randomly selected by the roulette wheel during the t-th iteration. It is the random walk around the elite in the t-th iteration. This represents the position of the i-th ant in the t-th iteration.
[0115] Based on the above iterative optimization steps and according to the above dynamic constraint rules, the first generation of ants and antlions are initialized, the fitness of ants and antlions is calculated, the best antlion is found and assumed to be elite, and when the termination condition is not met, for each ant, an antlion is selected using the roulette wheel selection method, and the position of the ant is updated according to the above steps; when the termination condition is met, the fitness of all ants is calculated, if the fitness of the ant corresponding to the antlion becomes higher, the antlion is replaced with the corresponding ant, and if the antlion becomes more fit than the elite, the elite is updated.
[0116] During the optimization process, the position of each ant is stored in an ant position matrix:
[0117]
[0118] In the formula, M Ant It is the ant position matrix, A a,b Let n represent the variable in the b-th dimension (b=1,…,d) of the a-th ant (a=1,…,n), where n represents the number of ants, d represents the number of ant dimensions, and the ant position refers to the parameter of a specific solution. The ant position matrix stores the positions of all ants during the optimization process, which is also the variable of all solutions.
[0119] To evaluate each ant, a fitness function was used during the optimization process, and the fitness of all ants is stored in the fitness matrix:
[0120]
[0121] In the formula, It is the ant fitness matrix. The fitness function is represented by the mean absolute error (MAE). Preferably, the fitness function is the mean absolute error (MAE).
[0122] Assume that the antlions are also hidden somewhere in the search space, and use an antlion position matrix and an antlion fitness matrix to store the position and fitness of each antlion:
[0123]
[0124]
[0125] In the formula, This represents the antlion position matrix. This represents the antlion fitness matrix. AL represents the fitness function. c,bLet n represent the value of the b-th dimension of the c-th antlion, n represent the number of antlions, and d represent the number of dimensions of the antlion. It should be noted that the antlion position matrix and the ant position matrix mentioned above both use n and d to represent the number of elements and the number of dimensions in the matrix. The specific values of n and d can be set according to the actual situation in different position matrices, and their specific values are not limited to be the same.
[0126] Based on the iterative optimization steps of the conventional ALO algorithm, it is known that the location quality of the antlion algorithm affects the optimization trend and accuracy of the algorithm. If the ants cannot find a better solution near the antlion, the slow update of the antlion's position will cause vicious development in the local area to varying degrees, leading the algorithm to get stuck in a local optimum. Therefore, this embodiment introduces a sharing mechanism of optimal information of elite antlions on the basis of the antlion algorithm. When the algorithm is about to get stuck in a local optimum, the antlions with poor fitness are reassigned to better positions to continue attracting ants, thereby improving the ability to escape the local optimum. Specifically, after the antlion position matrix is updated and sorted according to fitness values, the s-dimensional variable contained in the k-th non-elite antlion is selected and reassigned to the corresponding variable of the elite antlion after a small-scale mutation. The formula can be expressed as:
[0127]
[0128] In the formula, This represents the position of the k-th non-elite antlion after updating in the s-th dimension. Let represent the position of the elite antlion in the s-th dimension. R represents a random number following a standard normal distribution, thus ensuring that the population has a certain degree of mutation while carefully searching for the optimal solution. Then, based on the updated position variables of the non-elite antlions, the antlion position matrix is updated to obtain the updated antlion position matrix, thereby determining the antlion position of each antlion in the current iteration.
[0129] After obtaining the antlion positions through the above steps, the elite antlion positions are selected from the antlion position matrix as the optimal antlion positions. Simultaneously, the historical optimal antlion positions are determined based on the positions obtained in previous iterations. Finally, the antlion positions obtained in the current iteration are optimized based on preset exploration weights, the optimal antlion positions, and the historical optimal antlion positions. Specific steps include:
[0130] Calculate the first difference between the antlion's historical best position and the antlion's current position, and calculate the second difference between the antlion's best position and the antlion's historical best position;
[0131] The antlion position is optimized by multiplying the local exploration weight by the first difference and the global exploration weight by the second difference, resulting in an optimized antlion position.
[0132] In this embodiment, the optimized antlion position can be represented as:
[0133]
[0134] In the formula, This represents the position of the antlion after the t-th iteration of optimization. This represents the position of the antlion before the t-th iteration of optimization. This represents the historical best position of the antlion in the t-th iteration. This represents the optimal position of the antlion in the t-th iteration. Indicates the local exploration weight, Represents the global exploration weight, the preferred one. , .
[0135] In a preferred embodiment, the local exploration weight and the global exploration weight are determined based on the information sharing intensity of the antlion, and the specific steps include:
[0136] The information sharing intensity of each antlion is calculated based on the individual fitness and the average fitness of the group.
[0137] Based on the comparison between the information sharing intensity and the intensity threshold, the preset exploration weight is adjusted, and the antlion position is optimized based on the adjusted exploration weight.
[0138] In this embodiment, although the information sharing mechanism can alleviate the local optimum problem to some extent, the collaborative ability between individuals is limited in high-dimensional and complex search spaces, resulting in slow convergence. The current information sharing mechanism does not dynamically adjust for each individual; some antlions may overly rely on the optimal individual, leading to a narrow search range. To enable individuals with poor fitness to obtain better search directions by sharing information with elite antlions, this embodiment adjusts the individual search strategy of antlions based on the information sharing intensity among them. Specifically, it is assumed that in each iteration, the information sharing intensity of the antlions is calculated based on their fitness:
[0139]
[0140] In the formula, f represents the information sharing strength of the j-th antlion in the t-th iteration, α represents the adjustment factor, β represents the sensitivity factor, and f m,t f represents the average fitness of the antlions in the current population during the t-th iteration. t,j Let represent the individual fitness of the j-th antlion in the t-th iteration.
[0141] Preferably, during the calculation process It's usually set to 1; to quickly capture periodic patterns, you can increase it. To accelerate model convergence. The range of values for β is... When the fitness of the populations approaches the same level, it is necessary to increase... To enhance sensitivity to subtle differences.
[0142] The core role of information sharing intensity in location search is to dynamically adjust individual search strategies. Based on the fitness differences among individual antlions, the intensity of their learning from elite antlions is dynamically adjusted. By comparing individual fitness with the group's average fitness, the magnitude of information sharing intensity can be determined. A higher information sharing intensity indicates that the current individual is better than the group's average level, meaning better fitness. In this case, local search can be enhanced to preserve advantageous features. Conversely, a lower information sharing intensity indicates that the current individual is worse than the group's average level. In this case, global search can be enhanced to explore new areas. Therefore, in this embodiment, the exploration weights can be adjusted based on the information sharing intensity. Specific steps include:
[0143] Determine whether the information sharing intensity is greater than a first intensity threshold. If it is greater than the first intensity threshold, increase the preset local exploration weight by a preset ratio to obtain the adjusted local exploration weight. Otherwise, determine whether the information sharing intensity is less than a second intensity threshold.
[0144] If it is less than the second intensity threshold, the preset global exploration weight is increased by a preset ratio to obtain the adjusted global exploration weight.
[0145] In this embodiment, a threshold range for information sharing intensity is preset, namely [second intensity threshold, first intensity threshold]. Based on the calculated information sharing intensity of each antlion, it is determined whether it falls within the threshold range. If it is within the preset range, that is, the information sharing intensity is greater than the second intensity threshold and less than the first intensity threshold, it indicates that the current antlion's fitness is appropriate. In this case, the antlion's position is optimized using a preset exploration weight. For cases where the information sharing intensity is not within the threshold range, it can be divided into two cases: information sharing intensity greater than the first intensity threshold and information sharing intensity less than the second intensity threshold.
[0146] If the information sharing intensity is greater than or equal to the first intensity threshold, it indicates that the current individual is better than the group average. In this case, the local search can be enhanced to preserve advantageous features. Therefore, the preset local exploration weight can be increased by a preset proportion, such as by 20%, while the global exploration weight can remain unchanged. Alternatively, the preset global exploration weight can be decreased proportionally, such as by 20%, to further enhance local exploration. If the information sharing intensity is less than the second intensity threshold, it indicates that the individual is below the group average, and global exploration needs to be enhanced. In this case, the preset global exploration weight can be increased by a preset proportion, such as by 20%, while the local exploration weight can remain unchanged. Alternatively, the preset local exploration weight can be decreased proportionally, such as by 20%, to further enhance global exploration.
[0147] This embodiment introduces a dynamic adaptive information sharing mechanism, which automatically adjusts the intensity of elite information transmission by sensing individual fitness differences in real time. This effectively solves the problem of search stagnation in traditional algorithms when the ant fitness is lower than that of the antlion, achieves a dynamic balance between global exploration and local development, and significantly improves the algorithm's ability to escape local extremes and the efficiency of hyperparameter optimization.
[0148] In a preferred embodiment, after the antlion algorithm using an adaptive information sharing mechanism iterates through the hyperparameters once, the method further includes:
[0149] Based on the fitness improvement rate of the current hyperparameter combination compared to the previous iteration, the current number of attention heads in the current hyperparameter combination is updated to obtain the updated number of attention heads.
[0150] The current hyperparameter combination is updated based on the updated number of attention heads.
[0151] In this embodiment, after one iteration of the Antlion algorithm based on the adaptive information sharing mechanism, a set of hyperparameter combinations is obtained. This set of hyperparameter combinations includes the number of attention heads, the hidden layer dimension, and the learning rate. Specifically, the number of attention heads is updated based on the fitness improvement rate of the current hyperparameter combination compared to the previous iteration, according to the ALO algorithm.
[0152]
[0153] In the formula, This indicates the updated number of attention heads, where h represents the current number of attention heads. This represents the fitness improvement rate of the current hyperparameter combination compared to the previous iteration's hyperparameter combination. Its value is the difference between the fitness of the current hyperparameter combination and the fitness of the previous iteration's hyperparameter combination. This represents the search intensity coefficient. Wherein, The default value is 0.1, which is used to control the adjustment range.
[0154] According to the update formula above, when the fitness of the latest iteration's hyperparameter combination improves (i.e., ΔF > 0), the number of heads is increased by a preset ratio to enhance model capacity; conversely, the number of heads is reduced to decrease redundancy. Furthermore, since the hidden layer dimension is an integer multiple of the number of attention heads, after updating the number of attention heads, the hidden layer dimension also needs to be adaptively adjusted to obtain an updated hyperparameter combination. Based on this updated hyperparameter combination, the next iteration is performed until the iteration stops.
[0155] In this embodiment, the iteration stopping condition is expressed by the following formula:
[0156]
[0157] In the formula, Let represent the mean absolute error of the validation set in the t-th iteration. This represents the mean absolute error of the validation set for the first t-5 iterations.
[0158] By comparing the current MAE with the MAE five iterations ago, if the relative rate of change between the two is less than 1%, the model is considered to have converged, and hyperparameter optimization is stopped.
[0159] After obtaining the optimal hyperparameter combination of the load forecasting model through the above steps, the model can be trained and tested using a dataset built based on historical data. During training, elastic batch training is used, and the batch size is automatically adjusted according to memory usage to prevent memory overflow. During testing, the mean absolute error (MAE) is used as the validation metric. Finally, the trained load forecasting model is obtained.
[0160] Finally, the data preprocessing results in a data vector, which is then input into the trained load forecasting model to obtain short-term power load forecasting results. It should be noted that the model training steps can refer to conventional model training procedures, and will not be elaborated upon here.
[0161] This embodiment provides a short-term power load forecasting method based on collaborative optimization. The present invention optimizes the hyperparameters of the load forecasting model through the Antlion algorithm with an adaptive information sharing mechanism, which improves the efficiency and global search capability of the optimization algorithm in complex search spaces, significantly improves the convergence speed and accuracy of hyperparameter optimization, and ensures the global optimality of the optimization results. This enables the model to more accurately capture the nonlinear characteristics and complex time series relationships in the load data, thereby significantly improving the prediction accuracy and further ensuring the operating efficiency and stability of the power system.
[0162] Please see Figure 2Based on the same inventive concept, the second embodiment of this invention proposes a short-term power load forecasting system based on collaborative optimization, comprising:
[0163] The data preprocessing module 10 is used to acquire load-related data of the power system and perform data preprocessing on the load-related data to obtain a data vector. The load-related data includes load data, meteorological data, and date data.
[0164] The load forecasting module 20 is used to input the data vector into a preset load forecasting model to obtain the power load forecasting result. The load forecasting model is constructed using a deep neural network model based on a self-attention mechanism and uses an antlion algorithm with an adaptive information sharing mechanism for hyperparameter optimization.
[0165] The steps of hyperparameter optimization using the antlion algorithm with an adaptive information sharing mechanism include:
[0166] The hyperparameters of the load prediction model are mapped to the multidimensional location vectors of the antlion population, and an objective function is established based on the model performance and the rationality of the hyperparameters.
[0167] The Antlion algorithm, based on an adaptive information sharing mechanism, iteratively solves the objective function to obtain the optimal combination of hyperparameters.
[0168] The technical features and effects of the short-term power load forecasting system based on collaborative optimization proposed in this invention are the same as those of the method proposed in this invention, and will not be repeated here. Each module in the above-mentioned short-term power load forecasting system based on collaborative optimization can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or it can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0169] In summary, the present invention proposes a short-term power load forecasting method and system based on collaborative optimization. The method acquires load-related data of the power system and preprocesses this data to obtain a data vector. The load-related data includes load data, meteorological data, and date data. The data vector is then input into a preset load forecasting model to obtain a power load forecasting result. The load forecasting model is constructed using a deep neural network model based on a self-attention mechanism and employs an antlion algorithm with an adaptive information sharing mechanism for hyperparameter optimization. The step of using the antlion algorithm with an adaptive information sharing mechanism for hyperparameter optimization includes: mapping the hyperparameters of the load forecasting model to a multi-dimensional position vector of the antlion population, and establishing an objective function based on model performance and hyperparameter rationality; iteratively solving the objective function using the antlion algorithm with an adaptive information sharing mechanism to obtain the optimal hyperparameter combination. This invention introduces a global and local information sharing mechanism for dynamic adaptive change in the hyperparameter optimization algorithm, which improves the efficiency and global search capability of the optimization algorithm in complex search spaces, enhances the convergence speed and accuracy of hyperparameter optimization, and enables the load forecasting model to better capture the nonlinear and temporal characteristics in load data, significantly improving the model's performance in short-term load forecasting.
[0170] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0171] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A short-term power load forecasting method based on collaborative optimization, characterized in that, include: The load-related data of the power system is acquired, and the load-related data is preprocessed to obtain a data vector. The load-related data includes load data, meteorological data, and date data. The data vector is input into a preset load prediction model to obtain the power load prediction result. The load prediction model is constructed using a deep neural network model based on a self-attention mechanism and the hyperparameters are optimized using an antlion algorithm with an adaptive information sharing mechanism. The steps of hyperparameter optimization using the antlion algorithm with an adaptive information sharing mechanism include: The hyperparameters of the load prediction model are mapped to the multidimensional location vectors of the antlion population, and an objective function is established based on the model performance and the rationality of the hyperparameters. The Antlion algorithm, based on an adaptive information sharing mechanism, iteratively solves the objective function to obtain the optimal combination of hyperparameters; The steps of mapping the hyperparameters of the load prediction model to a multidimensional location vector of the antlion population, and establishing the objective function based on model performance and the rationality of the hyperparameters, include: The hyperparameters of the load prediction model are mapped to the multidimensional location vector of the antlion population, and the antlion population is initialized. The hyperparameters include the number of attention heads, the dimension of the hidden layer, and the learning rate. The ratio of the number of attention heads to the dimension of the hidden layer is used as the redundancy of the model parameters, and the difference between the learning rate and the learning rate baseline is used as the learning rate deviation. The objective function is to minimize the cross-entropy loss, model parameter redundancy, and learning rate deviation of the load prediction model on the validation set. The objective function is represented by the following formula: In the formula, CE represents the cross-entropy loss. The cross-entropy loss weights are represented by HD, which indicates the redundancy of the model parameters. XD represents the model parameter redundancy weights, and XD represents the learning rate deviation. Indicates the weight of the learning rate deviation. This represents the minimum value function.
2. The short-term power load forecasting method based on collaborative optimization according to claim 1, characterized in that, The step of preprocessing the load-related data to obtain a data vector: Wavelet decomposition and reconstruction are performed on the load data to obtain the reconstructed load sequence; Perform sine and cosine position encoding on the date data to obtain the time encoding vector; The meteorological data, the reconstructed load sequence, and the time-coded vector are normalized to obtain a data vector; The reconstructed load sequence is represented by the following formula: In the formula, Indicates the reconfigured load sequence. Describes the wavelet transform function. Represents load data, Represents a linear correction unit. This represents the coefficients of the third-level wavelet decomposition. This represents the coefficients of the fourth-level wavelet decomposition. This represents the standard deviation function.
3. The short-term power load forecasting method based on collaborative optimization according to claim 1, characterized in that, The steps of the Antlion algorithm based on the adaptive information sharing mechanism to iteratively solve the objective function and obtain the optimal hyperparameter combination include: Based on the random walk of ants, calculate the antlion position matrix and the corresponding antlion fitness matrix, and select non-elite antlions with fitness below the fitness threshold from the antlion position matrix. The positions of the selected non-elite antlions are updated based on the positions of the elite antlions, and the antlion position matrix is updated based on the updated positions of the non-elite antlions. Based on the updated antlion position matrix, the antlion positions and the optimal antlion positions are obtained; The antlion position is optimized based on the preset exploration weight, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position. The exploration weight includes local exploration weight and global exploration weight. The optimized antlion position is iteratively calculated according to the above steps until the iteration stops, and the optimal hyperparameter combination is obtained.
4. The short-term power load forecasting method based on collaborative optimization according to claim 3, characterized in that, The step of optimizing the antlion position based on a preset exploration weight, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position includes: Calculate the first difference between the antlion's historical best position and the antlion's current position, and calculate the second difference between the antlion's best position and the antlion's historical best position; The antlion position is optimized based on the product of the local exploration weight and the first difference, and the product of the global exploration weight and the second difference, to obtain the optimized antlion position. The optimized antlion position is represented by the following formula: In the formula, This represents the position of the antlion after the t-th iteration of optimization. This represents the position of the antlion before the t-th iteration of optimization. Indicates the local exploration weight, Indicates the global exploration weight. This represents the historical best position of the antlion in the t-th iteration. This represents the optimal position of the antlion in the t-th iteration.
5. The short-term power load forecasting method based on collaborative optimization according to claim 3, characterized in that, The step of optimizing the antlion position based on the preset exploration weight, the optimal antlion position, and the historical optimal antlion position to obtain the optimized antlion position further includes: The information sharing intensity of each antlion is calculated based on the individual fitness and the average fitness of the group. Based on the comparison between the information sharing intensity and the intensity threshold, the preset exploration weight is adjusted, and the antlion position is optimized based on the adjusted exploration weight. The intensity of information sharing is represented by the following formula: In the formula, f represents the information sharing strength of the j-th antlion in the t-th iteration, α represents the adjustment factor, β represents the sensitivity factor, and f m,t f represents the average fitness of the antlions in the current population during the t-th iteration. t,j Let represent the individual fitness of the j-th antlion in the t-th iteration.
6. The short-term power load forecasting method based on collaborative optimization according to claim 5, characterized in that, The step of adjusting the preset exploration weights based on the comparison relationship between the information sharing intensity and the intensity threshold includes: Determine whether the information sharing intensity is greater than a first intensity threshold. If it is greater than the first intensity threshold, increase the preset local exploration weight by a preset ratio to obtain the adjusted local exploration weight. Otherwise, determine whether the information sharing intensity is less than a second intensity threshold. If it is less than the second intensity threshold, the preset global exploration weight is increased by a preset ratio to obtain the adjusted global exploration weight.
7. The short-term power load forecasting method based on collaborative optimization according to claim 6, characterized in that, The step of iteratively solving the objective function using the antlion algorithm based on the adaptive information sharing mechanism to obtain the optimal hyperparameter combination further includes: Based on the fitness improvement rate of the current hyperparameter combination compared to the previous iteration, the current number of attention heads in the current hyperparameter combination is updated to obtain the updated number of attention heads. The current hyperparameter combination is updated based on the updated number of attention heads; The updated number of attention heads is represented by the following formula: In the formula, This indicates the updated number of attention heads, where h represents the current number of attention heads. Indicates the fitness improvement rate. This represents the search intensity coefficient.
8. The short-term power load forecasting method based on collaborative optimization according to claim 3, characterized in that, The iteration stopping condition is expressed by the following formula: In the formula, Let represent the mean absolute error of the validation set in the t-th iteration. This represents the mean absolute error of the validation set for the first t-5 iterations.
9. A short-term power load forecasting system based on collaborative optimization, characterized in that, The system is applied to the method as described in any one of claims 1 to 8, comprising: The data preprocessing module is used to acquire load-related data of the power system and perform data preprocessing on the load-related data to obtain a data vector. The load-related data includes load data, meteorological data, and date data. The load forecasting module is used to input the data vector into a preset load forecasting model to obtain the power load forecasting result. The load forecasting model is constructed using a deep neural network model based on a self-attention mechanism and uses an antlion algorithm with an adaptive information sharing mechanism for hyperparameter optimization. The steps of hyperparameter optimization using the antlion algorithm with an adaptive information sharing mechanism include: The hyperparameters of the load prediction model are mapped to the multidimensional location vectors of the antlion population, and an objective function is established based on the model performance and the rationality of the hyperparameters. The Antlion algorithm, based on an adaptive information sharing mechanism, iteratively solves the objective function to obtain the optimal combination of hyperparameters.
Citation Information
Patent Citations
Short-term power load prediction method
CN117709393A