A multi-objective microgrid power load prediction method
By filling in data gaps using MSG-GAN, combining time-varying filtering and fuzzy entropy aggregation, and utilizing multi-objective attention mechanisms and MOMPA parameter optimization methods, the accuracy and stability of microgrid power load forecasting are improved. This solves the problem of imperfect multi-objective optimization in existing technologies and achieves better microgrid scheduling support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAIYIN INSTITUTE OF TECHNOLOGY
- Filing Date
- 2026-02-27
- Publication Date
- 2026-05-29
Smart Images

Figure CN122118679A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of load forecasting for microgrid power systems, and specifically to a multi-objective microgrid power load forecasting method. Background Technology
[0002] The global energy system is rapidly shifting towards a cleaner and lower-carbon model, and microgrids, as a key platform for absorbing distributed energy and improving energy efficiency, play a crucial role in the development of smart grids. Microgrid load forecasting is fundamental to achieving optimized microgrid scheduling, energy storage configuration, and reliable power supply; its accuracy directly impacts the economic benefits, stability, and environmental performance of microgrid operation. Therefore, improving the accuracy of microgrid load forecasting is of great value for promoting the large-scale application of distributed energy, reducing energy waste, and maintaining the safe and efficient operation of the energy system. However, microgrid loads are influenced by multiple factors, including fluctuations in distributed power generation output, the randomness of user electricity consumption, changes in meteorological conditions, and interaction with the power grid, exhibiting significant nonlinear, time-varying, and multi-scale fluctuation characteristics. Furthermore, different operating scenarios have different requirements for forecast accuracy, peak value fitting capability, and model robustness, making it challenging to achieve high-precision load forecasting under multiple objectives.
[0003] Traditional microgrid power load forecasting methods mainly include statistical models and classical data-driven methods. Statistical models use mathematical tools such as time series analysis (e.g., ARIMA models) and regression analysis for forecasting. While simple in principle and computationally low, they have limited ability to describe the complex nonlinear relationships of multi-factor intertwined microgrid loads and are difficult to adapt to dynamically changing operating environments. Classical data-driven methods rely on historical load and related factor data, using machine learning models such as support vector machines and shallow neural networks for fitting, which improves the capture of nonlinear features. However, these traditional methods still have significant shortcomings: First, they usually only focus on minimizing prediction errors, failing to consider the important requirements of peak load accuracy and fluctuation trend consistency in microgrid scheduling; second, they are sensitive to data quality, and their generalization performance is limited in real-world scenarios with scarce data and high noise levels; furthermore, they do not sufficiently explore the multi-scale characteristics of loads and their spatiotemporal correlations, and the prediction accuracy and stability cannot yet meet the actual requirements of complex microgrid operation.
[0004] Current data-driven microgrid load forecasting methods still face several bottlenecks in practical applications. On the one hand, feature selection often relies on human experience, failing to effectively quantify the differences in the impact of various factors such as weather, user behavior, and distributed power generation on load, resulting in low effectiveness of input features and limiting further improvement in forecast accuracy. On the other hand, multi-objective optimization mechanisms are still underdeveloped, making it difficult to achieve a balance between mutually constraining objectives such as forecast error, peak fitting accuracy, and fluctuation robustness, thus failing to adapt to the scheduling needs of diverse microgrid operating scenarios. Furthermore, model structure and hyperparameter settings typically rely on manual adjustments, lacking intelligent optimization mechanisms, making it difficult to quickly respond to dynamic changes in microgrid load, resulting in unreliable forecast stability under different operating conditions. Therefore, developing a microgrid power load forecasting method that can accurately extract load features, effectively coordinate multi-objective requirements, and achieve intelligent model optimization has become an urgent task for improving microgrid operation and management and facilitating the transition to clean and low-carbon energy. Summary of the Invention
[0005] Purpose of the Invention: The technical problem to be solved by the present invention is to address the shortcomings of the existing technology by providing a multi-objective microgrid power load forecasting method. This method addresses the problems of imperfect multi-objective optimization mechanisms, insufficient feature effectiveness, and reliance on manual model optimization in existing methods, which lead to poor prediction accuracy and stability, and difficulty in adapting to the scheduling needs of microgrids in various scenarios. In this way, the invention improves the multi-objective adaptability, prediction accuracy, and operational stability of microgrid load forecasting, and provides reliable support for microgrid optimized scheduling, energy storage configuration, and safe and efficient operation decisions.
[0006] The method includes the following steps:
[0007] Step 1: Collect power load data of the microgrid, and use multi-scale gradient transfer generative adversarial network MSG-GAN to fill in missing values in the acquired charging load data by real-time focusing to reduce computational costs.
[0008] Step 2: Decompose the power load data using a method combining time-varying filter eigenmode decomposition (TVF-FMD) and fuzzy entropy aggregation.
[0009] Step 3: Input the decomposed power load data into the Scaleformer model improved by introducing a multi-objective attention mechanism for prediction, and obtain the final charging load prediction result.
[0010] Step 4: Optimize the parameters in the TVF-FMD method using the Multi-Objective Marine Predator Algorithm (MOMPA) with the objective function of minimizing local fuzzy entropy and decomposed error;
[0011] Step 5: Optimize the parameters in the improved Scaleformer model using the spatiotemporal correlation weighted fluctuation higher-order moment error and the multi-scale peak weighted nonlinear error as objective functions.
[0012] Step 1 includes:
[0013] Step 1-1: First, analyze and process the obtained basic data. Input the initial data into the generator to obtain the data, and then input it into the discriminator to verify the accuracy of the obtained data and whether it meets the requirements.
[0014] ,
[0015] The generator outputs noise x to obtain load data G(Z). The obtained data is put into the discriminator D to determine whether the input data comes from the true distribution or the generated distribution. The generator G takes the latent variable z as input and outputs synthetic data G(z). The discriminator D receives the true data x or the generated data G(z) and outputs the probability estimate of the true data. To generate the objective function in an adversarial network that minimizes the generator G and maximizes the discriminator D, Let x be the expected value of the random variable.
[0016] Steps 1-2: The discriminator simultaneously receives intermediate outputs from different levels of the generator. It then discriminates the obtained data. For the data G(Z) provided by the generator, the discriminator optimizes its own capabilities while the generator minimizes its loss, continuously improving the quality of the generated data. The loss function is expanded to a weighted sum of losses at each scale.
[0017] ,
[0018] Where S is the scale index. The weights are represented by the discriminator, which simultaneously receives intermediate outputs from different levels of the generator to form gradient information for a multi-scale discrimination mechanism, which is then backpropagated from the discriminator's levels to the generator. The objective function (value function) of the multi-scale generative adversarial network (MSG-GAN) is used to measure the game objectives of the generator G and the discriminator D. Represents a random variable that follows the true distribution of data. Take the expected value; The discriminator believes It is a probability estimate of real data; This represents the expected value of a random variable z that follows a noise distribution; This represents the probability estimate that the discriminator considers the generated data to be real data.
[0019] Steps 1-3, combining signal decomposition methods, involve the following data processing steps:
[0020] Original load sequence Decomposed into K intrinsic modal components:
[0021] ,
[0022] in It is a fragment; This represents the k-th intrinsic mode component; The residual term; the MSG-GAN generator output is:
[0023] ,
[0024] in, This represents the k-th intrinsic mode component output by the generator; This represents the IMF composite data output by the k-th generator; the composite complete payload sequence is then obtained.
[0025] ,
[0026] in, This represents the complete loading sequence synthesized; The sequence of intrinsic modal components generated by the generator is represented; for missing values in the load data, a mask matrix M' is introduced, and the reconstruction loss is defined. :
[0027] ,
[0028] Where x represents the actual load data; This represents the synthesized load data output by the generator; where This is element-wise multiplication;
[0029] Introducing cycle consistency loss :
[0030] ,
[0031] in, This indicates a filter generator used to repair or filter data; This indicates a destruction generator used to simulate situations where data is missing or corrupted.
[0032] Step 2 includes:
[0033] Step 2-1: Obtain the signal at time t ;
[0034] Step 2-2: Calculate the instantaneous amplitude and instantaneous frequency of the input signal using the Hilbert transform:
[0035] ,
[0036] in, yes Hilbert transform, Indicates Cauchy's principal value; Representing the integral variable Signal strength at the corresponding time point; Representing the integral variable The differential symbol;
[0037] instantaneous amplitude Represented as:
[0038] ,
[0039] instantaneous frequency The calculation formula is:
[0040] ;
[0041] in, The sign for the differential of the time variable t;
[0042] Step 2-3: Determine the initial decomposition interval of the characteristic mode function: based on the signal... The time-domain characteristics determine the initial decomposition interval. ,in and These are the start and end times of the signal, respectively.
[0043] Steps 2-4: Apply time-varying filtering: Within the initial decomposition interval, process the signal... Applying time-varying filtering, the filtering function for:
[0044] ,
[0045] in Let exp be the standard deviation of the time-varying filter, and let exp represent the natural exponential function.
[0046] Step 2-5: Perform Eigenmode Decomposition: Perform eigenmode decomposition on the filtered signal to obtain a series of eigenmode functions IMFi(t), with the following expression:
[0047] ,
[0048] in For the first The number of modes of a characteristic mode function. ϕi represents the modal coefficients. For modal basis functions, , Modal time position;
[0049] Steps 2-6: Calculate the fuzzy entropy of each characteristic mode function: For each characteristic mode function... Calculate fuzzy entropy :
[0050] ,
[0051] Where M and N are the number of rows and columns in the fuzzy partition, respectively. For fuzzy membership degree;
[0052] Step 2-7: Introduce fuzzy entropy aggregation for feature fusion: Aggregate the fuzzy entropies of each feature mode function to obtain the fused feature vector F(t):
[0053] ,
[0054] Among them, w i The weights are for the i-th feature mode function.
[0055] Step 3 includes:
[0056] Step 3-1, TimeNorm (Time Series Normalization Layer):
[0057] ,
[0058] in, and These represent the learnable scaling factor and offset factor, respectively. It is a constant;
[0059] Step 3-2, Multi-scale attention mechanism;
[0060] Trend branches:
[0061] ,
[0062] Among them, Q T Represents the query matrix. This indicates a one-dimensional convolutional layer with 168 output channels, used to extract trend features; K represents the residual components after decomposition. T V represents the bond matrix. T Represents a value matrix; This indicates the attention output for the trend branch. Indicates the dimensions of the query matrix and the key matrix;
[0063] Seasonal branches:
[0064] ,
[0065] ,
[0066] in, The attention mask matrix represents the seasonal branch and is used to mask information during non-peak periods; This represents the time point of the i-th time step; Indicates the preset peak period; This represents the query matrix after merging multiple branches; This indicates the attention output for trend branches; This represents the attention output for seasonal branches; W represents the attention output of the residual branch. T W S W R It is a learnable weight matrix, and b is the bias term;
[0067] Step 3-3: Construct a power load forecasting model based on fusion features.
[0068] Step 3-3 includes: constructing an electricity load prediction model using the fused feature vector F(t), and the output of the prediction model is the predicted electricity load. The calculation formula is:
[0069] ,
[0070] in, This represents the mapping function of the prediction model.
[0071] Step 4 includes:
[0072] Step 4-1: Define the dual optimization objectives of TVF-FMD, setting the minimum local fuzzy entropy and the minimum decomposition error as the dual objective functions;
[0073] The formula for the local fuzzy entropy of objective function 1 is as follows:
[0074] ,
[0075] in, is the objective function value of the local fuzzy entropy; t is the center time of the local time window, and the length of the local window L is preset according to the time domain characteristics of the signal. Let be the fuzzy membership degree of the signal feature within the local window at time t in the fuzzy grid at row i and column j. The natural logarithm of fuzzy membership degree;
[0076] Step 4-2, the decomposition error formula for objective function 2 is as follows:
[0077] ,
[0078] in, To decompose the objective function value of the error; The original power load sequence at time t; Let be the value of the i-th characteristic mode function at time t; It is an L2 norm;
[0079] Step 4-3: Determine the parameters and range to be optimized for TVF-FMD, and clarify the core adjustable parameters and preset search range of TVF-FMD, as shown in the following formula:
[0080] ,
[0081] in, The parameter vector to be optimized for TVF-FMD; The bandwidth is the time-varying filter bandwidth; The modal decomposition layer number control coefficient; The threshold for fuzzy entropy aggregation;
[0082] Step 4-4: Initialize the MOMPA algorithm, construct the population structure, and set the core operating parameters of MOMPA. The individual encoding formula for the population is as follows:
[0083] ,
[0084] in , For the i-th individual in the population; Let be the time-varying filter bandwidth parameter value for the i-th individual; Let be the modal decomposition layer control coefficient value for the i-th individual; is the fuzzy entropy aggregation threshold parameter value for the i-th individual; N is the population size;
[0085] The formulas for setting the core parameters are as follows:
[0086] ,
[0087] Where T is the maximum number of iterations; P is the predator's step size; and d is the dimension of the parameter to be optimized.
[0088] Steps 4-5: MOMPA two-stage search optimization, guided by Pareto leading edge leaders, performs a global parameter search. The global exploration formula for the leader stage is as follows:
[0089] ,
[0090] in, For the updated combination of parameters for the i-th individual in the population; This refers to the i-th individual in the population before the update. It is a 3-dimensional random vector uniformly distributed in the interval [0,1]. This indicates that the corresponding parameter components are multiplied one by one; To be a leading individual at the forefront of Pareto;
[0091] Steps 4-6, the follower phase, involve fine-tuning local parameters based on the optimal individual in the population. The adaptive control parameter formulas are as follows:
[0092] ,
[0093] Wherein, CF is the adaptive control parameter, with a value range of [0,1]; This represents the current iteration number; This represents the maximum number of iterations.
[0094] The formula for updating local parameters is as follows:
[0095] ,
[0096] in, It is the best individual in the current population;
[0097] Steps 4-6: Pareto non-dominated sorting and optimal parameter selection. The non-dominated solution set is selected to determine the optimal parameters for TVF-FMD. The selection criteria formula is as follows:
[0098] and ,
[0099] in To combine the objective function values;
[0100] The formula for determining the optimal parameters is as follows:
[0101] ,
[0102] in, The optimal parameter combination for TVF-FMD; the Pareto front is the optimal solution set obtained after non-dominated sorting; argmin represents the parameter vector that minimizes the objective function value;
[0103] Steps 4-6: TVF-FMD parameters are fixed, and the optimal parameters are determined. The final TVF-FMD parameter formula is as follows:
[0104] ,
[0105] in, This represents the optimal value for the optimized time-varying filter bandwidth. The optimal value of the control coefficient for the number of modal decomposition layers is obtained after optimization. The optimal value for the fuzzy entropy aggregation threshold is the result of optimization.
[0106] Step 5 includes:
[0107] Step 5-1: Define the IScaleformer dual optimization objectives;
[0108] Objective function 1 is the higher-order moment error of spatiotemporal correlation weighted fluctuations, as shown in the following formula:
[0109] ,
[0110] The formulas for higher-order moments are as follows:
[0111] ,
[0112] The spatiotemporal weight formula is as follows:
[0113] ,
[0114] in, Here, S represents the ST-FHME value, and T represents the time duration. For spatiotemporal correlation weights, To predict the k-th order central moment of the load, The actual load is represented by the k-th order central moment, where k is the order of higher-order moments, and W is the length of the calculation window. This represents the average actual load within the window. The spatiotemporal correlation coefficient;
[0115] Step 5-2: Objective function 2 is the multi-scale peak-weighted nonlinear error, as shown in the following formula:
[0116] ,
[0117] The peak weight formula is as follows:
[0118] ,
[0119] The formula for the nonlinear penalty term is as follows:
[0120] ,
[0121] in, For MS-PLWE-NL value, For multi-scale layers, Let m be the length of the sequence at scale m. For peak weight, To predict load, This is the actual load. The peak penalty coefficient, Peak threshold This represents the maximum load value at the m-th scale. This is the average load.
[0122] Step 5-3: Determine the parameters to be optimized for IScaleformer:
[0123] ,
[0124] in, For spatial feature dimensions, For multi-scale convolution kernel number, For the number of spatiotemporal attention heads, The spatiotemporal fusion coefficient; This represents the set of parameters that need to be optimized in the IScaleformer model;
[0125] Step 5-4, MOMPA parameter optimization:
[0126] The initialization formula is as follows:
[0127] ,
[0128] in, For the i-th parameter combination;
[0129] Parameter update, global exploration formula is as follows:
[0130] ,
[0131] The formula for partial development is as follows:
[0132] ,
[0133] in, It is a 4-dimensional random vector. Represents element-wise product. As a leading figure in Pareto;
[0134] Step 5-5: Optimal parameter selection and solidification.
[0135] In step 5-5, the optimal parameters are selected and solidified using the following formula:
[0136] ,
[0137] ,
[0138] in, For the optimal parameter combination, They are respectively , , , The optimal value.
[0139] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method.
[0140] The present invention also provides a storage medium storing a computer program or instructions that, when the computer program or instructions are run on a computer, execute the steps of the method described.
[0141] Compared with the prior art, the present invention has the following beneficial effects: 1. By using a multi-scale gradient transfer generative adversarial network to perform mask-guided generative imputation of missing load records, the generator follows a smooth Wasserstein metric in the latent space, and the discriminator synchronously receives multi-resolution feature maps, so that the imputation curve maintains spectral consistency with the real distribution in both low-frequency trends and high-frequency fluctuations, thereby eliminating the conditional distribution drift caused by data gaps and providing unbiased samples for subsequent modeling.
[0142] 2. Through a collaborative framework of time-varying filtering eigenmode decomposition and fuzzy entropy aggregation, the original charging load sequence is projected segment by segment onto the instantaneous frequency axis. The filtering window is adaptively determined based on the local bandwidth. After obtaining several intrinsic modes, the complexity rate of each mode is measured by fuzzy entropy. Then, entropy-weighted fusion is performed to compress the high-dimensional mode space into three interpretable components: trend, period, and residual. This not only preserves the main energy but also suppresses non-stationarity and impulse noise, significantly reducing the redundancy of the downstream model input.
[0143] 3. Embedding a time normalization layer and a multi-scale local-global attention mechanism in the improved Scaleformer: Time normalization calculates the sliding statistic along the time axis, avoiding the destruction of sequence autocorrelation by traditional batch normalization; local-global attention uses a wide convolutional kernel to capture long-period behavior for large-scale trends, introduces a learnable mask for peak periods, enhances the sensitivity to early and late peaks, achieves efficient modeling of long sequence dependencies, alleviates gradient vanishing, and improves the ability to characterize complex load patterns.
[0144] 4. The multi-objective marine predator algorithm is used to jointly optimize the hyperparameters of the decomposition and prediction modules. The TVF-FMD parameter update is driven by minimizing both modal complexity and decomposition error. The IScaleformer structure adjustment is driven by simultaneously optimizing the spatiotemporal correlation weighted high-order moment error (ST-FHME) and the multi-scale peak weighted nonlinear error (MS-PLWE-NL). Attached Figure Description
[0145] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0146] Figure 1 This is a schematic diagram of the overall structure of the present invention.
[0147] Figure 2 This is an exploded view of the present invention.
[0148] Figure 3 This is a schematic diagram illustrating the multi-objective optimization of the present invention.
[0149] Figure 4 This is a graph showing the error index of the predicted values. Detailed Implementation
[0150] like Figure 1 As shown, this embodiment provides a microgrid power load forecasting method based on multiple objectives, including the following steps:
[0151] S1. Collect power load data of the microgrid, and use multi-scale gradient transfer generative adversarial network (MSG-GAN) to fill in missing values in the acquired charging load data by real-time focusing to reduce computational costs and improve the accuracy of data filling.
[0152] The specific steps are as follows:
[0153] Step 1-1: First, analyze and process the obtained basic data. Input the initial data into the generator to obtain the data, and then input it into the discriminator to verify the accuracy of the obtained data and whether it meets the requirements.
[0154] ,
[0155] The generator outputs noise x to obtain load data G(Z). The obtained data is put into the discriminator D to determine whether the input data comes from the true distribution or the generated distribution. The generator G takes the latent variable z as input and outputs synthetic data G(z). The discriminator D receives the true data x or the generated data G(z) and outputs the probability estimate of the true data. To generate the objective function in an adversarial network that minimizes the generator G and maximizes the discriminator D, Let x be the expected value of the random variable.
[0156] Steps 1-2: The discriminator simultaneously receives intermediate outputs from different levels of the generator. It then discriminates the obtained data. For the data G(Z) provided by the generator, the discriminator optimizes its own capabilities while the generator minimizes its loss, continuously improving the quality of the generated data. The loss function is expanded to a weighted sum of losses at each scale.
[0157] ,
[0158] Where S is the scale index. The weights are represented by the discriminator, which simultaneously receives intermediate outputs from different levels of the generator to form gradient information for a multi-scale discrimination mechanism, which is then backpropagated from the discriminator's levels to the generator. The objective function (value function) of the multi-scale generative adversarial network (MSG-GAN) is used to measure the game objectives of the generator G and the discriminator D. Represents a random variable that follows the true distribution of data. Take the expected value; The discriminator believes It is a probability estimate of real data; This represents the expected value of a random variable z that follows a noise distribution; This represents the probability estimate that the discriminator considers the generated data to be real data.
[0159] Steps 1-3, combining signal decomposition methods, involve the following data processing steps:
[0160] Original load sequence Decomposed into K intrinsic modal components:
[0161] ,
[0162] in It is a fragment; This represents the k-th intrinsic mode component; The residual term; the MSG-GAN generator output is:
[0163] ,
[0164] in, This represents the k-th intrinsic mode component output by the generator; This represents the IMF composite data output by the k-th generator; the composite complete payload sequence is then obtained.
[0165] ,
[0166] in, This represents the complete loading sequence synthesized; The sequence of intrinsic modal components generated by the generator is represented; for missing values in the load data, a mask matrix M' is introduced, and the reconstruction loss is defined. :
[0167] ,
[0168] Where x represents the actual load data; This represents the synthesized load data output by the generator; where This is element-wise multiplication;
[0169] Introducing cycle consistency loss :
[0170] ,
[0171] in, This indicates a filter generator used to repair or filter data; This indicates a destruction generator used to simulate situations where data is missing or corrupted.
[0172] S2, such as Figure 2 As shown, the power load data is decomposed using a method combining time-varying filter feature mode decomposition (TVF-FMD) and fuzzy entropy aggregation.
[0173] The specific steps are as follows:
[0174] Step 2-1: Obtain the signal at time t ;
[0175] Step 2-2: Calculate the instantaneous amplitude and instantaneous frequency of the input signal using the Hilbert transform:
[0176] ,
[0177] in, yes Hilbert transform, Indicates Cauchy's principal value; Representing the integral variable Signal strength at the corresponding time point; Representing the integral variable The differential symbol;
[0178] instantaneous amplitude Represented as:
[0179] ,
[0180] instantaneous frequency The calculation formula is:
[0181] ;
[0182] in, The sign for the differential of the time variable t;
[0183] Step 2-3: Determine the initial decomposition interval of the characteristic mode function: based on the signal... The time-domain characteristics determine the initial decomposition interval. ,in and These are the start and end times of the signal, respectively.
[0184] Steps 2-4: Apply time-varying filtering: Within the initial decomposition interval, process the signal... Applying time-varying filtering, the filtering function for:
[0185] ,
[0186] in Let exp be the standard deviation of the time-varying filter, and let exp represent the natural exponential function.
[0187] Step 2-5: Perform Eigenmode Decomposition: Perform eigenmode decomposition on the filtered signal to obtain a series of eigenmode functions IMFi(t), with the following expression:
[0188] ,
[0189] in For the first The number of modes of a characteristic mode function. ϕi represents the modal coefficients. For modal basis functions, , Modal time position;
[0190] Steps 2-6: Calculate the fuzzy entropy of each characteristic mode function: For each characteristic mode function... Calculate fuzzy entropy :
[0191] ,
[0192] Where M and N are the number of rows and columns in the fuzzy partition, respectively. For fuzzy membership degree;
[0193] Step 2-7: Introduce fuzzy entropy aggregation for feature fusion: Aggregate the fuzzy entropies of each feature mode function to obtain the fused feature vector F(t):
[0194] ,
[0195] Among them, w i The weights are for the i-th feature mode function.
[0196] S3. The decomposed power load data is fed into the Scaleformer model (IScaleformer) improved by introducing a multi-objective attention mechanism for prediction, and the final charging load prediction result is obtained.
[0197] The specific steps are as follows:
[0198] Step 3-1, TimeNorm (Time Series Normalization Layer):
[0199] ,
[0200] in, and These represent the learnable scaling factor and offset factor, respectively. It is a constant;
[0201] Step 3-2, Multi-scale attention mechanism;
[0202] Trend branches:
[0203] ,
[0204] Among them, Q T Represents the query matrix. This indicates a one-dimensional convolutional layer with 168 output channels, used to extract trend features; K represents the residual components after decomposition. T V represents the bond matrix. T Represents a value matrix; This indicates the attention output for the trend branch. Indicates the dimension of the query / key matrix;
[0205] Seasonal branches:
[0206] ,
[0207] ,
[0208] in, The attention mask matrix represents the seasonal branch and is used to mask information during non-peak periods; This represents the time point of the i-th time step; Indicates the preset peak period; This represents the query matrix after merging multiple branches; This indicates the attention output for trend branches; This represents the attention output for seasonal branches; W represents the attention output of the residual branch. T W S W R It is a learnable weight matrix, and b is the bias term;
[0209] Step 3-3: Construct a power load prediction model based on fused features: A power load prediction model is constructed using the fused feature vector F(t). The output of the prediction model is the predicted power load. The calculation formula is:
[0210] ,
[0211] in, This represents the mapping function of the prediction model.
[0212] S4, such as Figure 3 As shown, the parameters in the TVF-FMD method are optimized by using the Multi-Objective Marine Predator Algorithm (MOMPA) with the objective function of minimizing the local fuzzy entropy and the decomposed error.
[0213] The specific steps are as follows:
[0214] Step 4-1: Define the dual optimization objectives of TVF-FMD, setting the minimum local fuzzy entropy and the minimum decomposition error as the dual objective functions;
[0215] The formula for the local fuzzy entropy of objective function 1 is as follows:
[0216] ,
[0217] in, t is the objective function value of the local fuzzy entropy; M is the number of rows in the fuzzy partition; N is the number of columns in the fuzzy partition; t is the center time of the local time window, and the length of the local window L is preset according to the time domain characteristics of the signal; Let be the fuzzy membership degree of the signal feature within the local window at time t in the fuzzy grid at row i and column j. The natural logarithm of fuzzy membership degree;
[0218] Step 4-2, the decomposition error formula for objective function 2 is as follows:
[0219] ,
[0220] in, To decompose the objective function value of the error; The original power load sequence at time t; Let be the value of the i-th characteristic mode function at time t; It is an L2 norm;
[0221] Step 4-3: Determine the parameters and range to be optimized for TVF-FMD, and clarify the core adjustable parameters and preset search range of TVF-FMD, as shown in the following formula:
[0222] ,
[0223] in, The parameter vector to be optimized for TVF-FMD; The bandwidth is the time-varying filter bandwidth; The modal decomposition layer number control coefficient; The threshold for fuzzy entropy aggregation;
[0224] Step 4-4: Initialize the MOMPA algorithm, construct the population structure, and set the core operating parameters of MOMPA. The individual encoding formula for the population is as follows:
[0225] ,
[0226] in , For the i-th individual in the population; Let be the time-varying filter bandwidth parameter value for the i-th individual; Let be the modal decomposition layer control coefficient value for the i-th individual; is the fuzzy entropy aggregation threshold parameter value for the i-th individual; N is the population size;
[0227] The formulas for setting the core parameters are as follows:
[0228] ,
[0229] Where T is the maximum number of iterations; P is the predator's step size; and d is the dimension of the parameter to be optimized.
[0230] Steps 4-5: MOMPA two-stage search optimization, guided by Pareto leading edge leaders, performs a global parameter search. The global exploration formula for the leader stage is as follows:
[0231] ,
[0232] in, For the updated combination of parameters for the i-th individual in the population; This refers to the i-th individual in the population before the update. It is a 3-dimensional random vector uniformly distributed in the interval [0,1]. This indicates that the corresponding parameter components are multiplied one by one; To be a leading individual at the forefront of Pareto;
[0233] Steps 4-6, the follower phase, involve fine-tuning local parameters based on the optimal individual in the population. The adaptive control parameter formulas are as follows:
[0234] ,
[0235] Wherein, CF is the adaptive control parameter, with a value range of [0,1]; This represents the current iteration number; This represents the maximum number of iterations.
[0236] The formula for updating local parameters is as follows:
[0237] ,
[0238] in, It is the best individual in the current population;
[0239] Steps 4-6: Pareto non-dominated sorting and optimal parameter selection. The non-dominated solution set is selected to determine the optimal parameters for TVF-FMD. The selection criteria formula is as follows:
[0240] and ,
[0241] in To combine the objective function values;
[0242] The formula for determining the optimal parameters is as follows:
[0243] ,
[0244] in, The optimal parameter combination for TVF-FMD; the Pareto front is the optimal solution set obtained after non-dominated sorting; argmin represents the parameter vector that minimizes the objective function value;
[0245] Steps 4-6: TVF-FMD parameters are fixed, and the optimal parameters are determined. The final TVF-FMD parameter formula is as follows:
[0246] ,
[0247] in, This represents the optimal value for the optimized time-varying filter bandwidth. The optimal value of the control coefficient for the number of modal decomposition layers is obtained after optimization. The optimal value for the fuzzy entropy aggregation threshold is the result of optimization.
[0248] S5. The parameters in the IScaleformer model are optimized using the spatiotemporal correlation weighted high-order moment error (ST-FHME) and multi-scale peak weighted nonlinear error (MS-PLWE-NL) as objective functions.
[0249] The specific steps are as follows:
[0250] Step 5-1: Define the IScaleformer dual optimization objectives;
[0251] Objective function 1 is the higher-order moment error of spatiotemporal correlation weighted fluctuations, as shown in the following formula:
[0252] ,
[0253] The formulas for higher-order moments are as follows:
[0254] ,
[0255] The spatiotemporal weight formula is as follows:
[0256] ,
[0257] in, Here, S represents the ST-FHME value, and T represents the time duration. For spatiotemporal correlation weights, To predict the k-th order central moment of the load, The actual load is represented by the k-th order central moment, where k is the order of higher-order moments, and W is the length of the calculation window. This represents the average actual load within the window. The spatiotemporal correlation coefficient;
[0258] Step 5-2: Objective function 2 is the multi-scale peak-weighted nonlinear error, as shown in the following formula:
[0259] ,
[0260] The peak weight formula is as follows:
[0261] ,
[0262] The formula for the nonlinear penalty term is as follows:
[0263] ,
[0264] in, For MS-PLWE-NL value, For multi-scale layers, Let m be the length of the sequence at scale m. For peak weight, To predict load, This is the actual load. The peak penalty coefficient, Peak threshold This represents the maximum load value at the m-th scale. This is the average load.
[0265] Step 5-3: Determine the parameters to be optimized for IScaleformer:
[0266] ,
[0267] in, For spatial feature dimensions, For multi-scale convolution kernel number, For the number of spatiotemporal attention heads, The spatiotemporal fusion coefficient; This represents the set of parameters that need to be optimized in the IScaleformer model;
[0268] Step 5-4, MOMPA parameter optimization:
[0269] The initialization formula is as follows:
[0270] ,
[0271] in, For the i-th parameter combination;
[0272] Parameter update, global exploration formula is as follows:
[0273] ,
[0274] The formula for partial development is as follows:
[0275] ,
[0276] in, It is a 4-dimensional random vector. Represents element-wise product. As a leading figure in Pareto;
[0277] Step 5-5: Optimal parameter selection and consolidation, the formula is as follows:
[0278] ,
[0279] ,
[0280] in, For the optimal parameter combination, They are respectively , , , The optimal value.
[0281] To verify the effectiveness and superiority of the multi-objective microgrid power load forecasting method model (MSG-GAN-TVF-FMD-IScaleformer-MOMPA) proposed in this invention, the experimental data used in this invention were open-source data provided by the UCSD microgrid dataset, with microgrid power load data from January 2015 to January 2020 as the experimental case. The algorithm program was written in MATLAB, and the prediction model (MSG-GAN-TVF-FMD-IScaleformer-MOMPA) and five sets of control experiments (MSG-GAN-TVF-FMD-IScaleformer, TVF-FMD-IScaleformer, IScaleformer, ELM, and LSTM) of this invention were constructed.
[0282] The proposed model was compared with the five control group models mentioned above. The statistical results of the predictive indicators for each model are shown in Table 1, including RMSE, MAE, MAPR, and R. 2 The value of .
[0283] Table 1. Statistical table of results and performance indicators of the model of this invention and the control group model.
[0284]
[0285] As shown in Table 1, the prediction performance of the IScaleformer model is significantly better than that of the ELM and LSTM models, indicating that the model has better nonlinear fitting ability and robustness, and is more applicable to complex sequence prediction tasks. Compared with the TVF-FMD-IScaleformer model, MSG-GAN-TVF-FMD-IScaleformer achieves better prediction performance, indicating that the multi-scale gradient transfer generative adversarial network MSG-GAN can efficiently complete missing value imputation, and high-quality generated data provides strong support for improving the model's generalization ability. The RMSE of the MSG-GAN-TVF-FMD-IScaleformer-MOMPA prediction model is 3.1826, the MAE is 2.0915, the MAPE is 1.7642%, and the R² is 0.9928. Comparing the MSG-GAN-TVF-FMD-IScaleformer-MOMPA and the MSG-GAN-TVF-FMD-IScaleformer models, it can be seen that the overall prediction performance of the model is significantly improved after introducing the multi-target marine predator algorithm MOMPA. This indicates that MOMPA can effectively search and optimize the model hyperparameters, thereby improving the prediction accuracy.
[0286] This invention is in Figure 4 The text uses bar charts, vertical line charts, vertical histograms, and radar charts to visually present the RMSE, MAE, MAPE, and R² indices for six models. Figure 4 The results show that the MSG-GAN-TVF-FMD-IScaleformer-MOMPA model has the smallest error index and the largest coefficient of determination, exhibiting the best overall performance. The ELM model performs relatively poorly. The IScaleformer model's prediction performance is generally higher than that of the ELM and LSTM models. The MSG-GAN-TVF-FMD-IScaleformer model outperforms the TVF-FMD-IScaleformer model. The performance of the model is further improved after adding the MOMPA algorithm for optimization, verifying the effectiveness and rationality of the present invention.
[0287] This invention provides a multi-objective microgrid power load forecasting method. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A method for predicting power load in a multi-objective microgrid, characterized in that, Includes the following steps: Step 1: Collect power load data of the microgrid, and use multi-scale gradient transfer generative adversarial network MSG-GAN to fill in missing values in the acquired charging load data by real-time focusing to reduce computational costs. Step 2: Decompose the power load data using a method combining time-varying filter eigenmode decomposition (TVF-FMD) and fuzzy entropy aggregation. Step 3: Input the decomposed power load data into the Scaleformer model improved by introducing a multi-objective attention mechanism for prediction, and obtain the final charging load prediction result. Step 4: Optimize the parameters in the TVF-FMD method using the Multi-Objective Marine Predator Algorithm (MOMPA) with the objective function of minimizing local fuzzy entropy and decomposed error; Step 5: Optimize the parameters in the improved Scaleformer model using the spatiotemporal correlation weighted fluctuation higher-order moment error and the multi-scale peak weighted nonlinear error as objective functions.
2. The method according to claim 1, characterized in that, Step 1 includes: Step 1-1: First, analyze and process the obtained basic data. Input the initial data into the generator to obtain the data, and then input it into the discriminator to verify the accuracy of the obtained data and whether it meets the requirements. , The generator outputs noise x to obtain load data G(Z). The obtained data is put into the discriminator D to determine whether the input data comes from the true distribution or the generated distribution. The generator G takes the latent variable z as input and outputs synthetic data G(z). The discriminator D receives the true data x or the generated data G(z) and outputs the probability estimate of the true data. To generate the objective function in an adversarial network that minimizes the generator G and maximizes the discriminator D, Let x be the expected value of the random variable. Steps 1-2: The discriminator simultaneously receives intermediate outputs from different levels of the generator. It then discriminates the obtained data. For the data G(Z) provided by the generator, the discriminator optimizes its own capabilities while the generator minimizes its loss, continuously improving the quality of the generated data. The loss function is expanded to a weighted sum of losses at each scale. , Where S is the scale index. The weights are represented by the discriminator, which simultaneously receives intermediate outputs from different levels of the generator to form gradient information for a multi-scale discrimination mechanism, which is then backpropagated from the discriminator's levels to the generator. The objective function for the multi-scale generative adversarial network MSG-GAN is used to measure the game objective between the generator G and the discriminator D. Represents a random variable that follows the true distribution of data. Take the expected value; The discriminator believes It is a probability estimate of real data; This represents the expected value of a random variable z that follows a noise distribution; This represents the probability estimate that the discriminator considers the generated data to be real data. Steps 1-3, combining signal decomposition methods, involve the following data processing steps: Original load sequence Decomposed into K intrinsic modal components: , in It is a fragment; This represents the k-th intrinsic mode component; The residual term; the MSG-GAN generator output is: , in, This represents the k-th intrinsic mode component output by the generator; This represents the IMF composite data output by the k-th generator; the composite complete payload sequence is then obtained. , in, This represents the complete loading sequence synthesized; The sequence of intrinsic modal components generated by the generator is represented; for missing values in the load data, a mask matrix M' is introduced, and the reconstruction loss is defined. : , Where x represents the actual load data; This represents the synthesized load data output by the generator; where This is element-wise multiplication; Introducing cycle consistency loss : , in, Indicates a filter generator; This indicates a destruction generator.
3. The method according to claim 2, characterized in that, Step 2 includes: Step 2-1: Obtain the signal at time t ; Step 2-2: Calculate the instantaneous amplitude and instantaneous frequency of the input signal using the Hilbert transform: , in, yes Hilbert transform, Indicates Cauchy's principal value; Representing the integral variable Signal strength at the corresponding time point; Representing the integral variable The differential symbol; instantaneous amplitude Represented as: , instantaneous frequency The calculation formula is: ; in, The sign for the differential of the time variable t; Step 2-3: Determine the initial decomposition interval of the characteristic mode function: based on the signal... The time-domain characteristics determine the initial decomposition interval. ,in and These are the start and end times of the signal, respectively. Steps 2-4: Apply time-varying filtering: Within the initial decomposition interval, process the signal... Applying time-varying filtering, the filtering function for: , in Let exp be the standard deviation of the time-varying filter, and let exp represent the natural exponential function. Step 2-5: Perform Eigenmode Decomposition: Perform eigenmode decomposition on the filtered signal to obtain a series of eigenmode functions IMFi(t), with the following expression: , in For the first The number of modes of a characteristic mode function. ϕi represents the modal coefficients. For modal basis functions, , Modal time position; Steps 2-6: Calculate the fuzzy entropy of each characteristic mode function: For each characteristic mode function... Calculate fuzzy entropy : , Where M and N are the number of rows and columns in the fuzzy partition, respectively. For fuzzy membership degree; Step 2-7: Introduce fuzzy entropy aggregation for feature fusion: Aggregate the fuzzy entropies of each feature mode function to obtain the fused feature vector F(t): , Among them, w i The weights are for the i-th feature mode function.
4. The method according to claim 3, characterized in that, Step 3 includes: Step 3-1, TimeNorm (Time Series Normalization Layer): , in, and These represent the learnable scaling factor and offset factor, respectively. It is a constant; Step 3-2, Multi-scale attention mechanism; Trend branches: , Among them, Q T Represents the query matrix. This represents a one-dimensional convolutional layer with 168 output channels; K represents the residual components after decomposition. T V represents the bond matrix. T Represents a value matrix; This indicates the attention output for the trend branch. Indicates the dimensions of the query matrix and the key matrix; Seasonal branches: , , in, The attention mask matrix representing the seasonal branches; This represents the time point of the i-th time step; Indicates the preset peak period; This represents the query matrix after merging multiple branches; This indicates the attention output for trend branches; This represents the attention output for seasonal branches; W represents the attention output of the residual branch. T W S W R It is a learnable weight matrix, and b is the bias term; Step 3-3: Construct a power load forecasting model based on fusion features.
5. The method according to claim 4, characterized in that, Step 3-3 includes: constructing an electricity load prediction model using the fused feature vector F(t), and the output of the prediction model is the predicted electricity load. The calculation formula is: , in, This represents the mapping function of the prediction model.
6. The method according to claim 5, characterized in that, Step 4 includes: Step 4-1: Define the dual optimization objectives of TVF-FMD, setting the minimum local fuzzy entropy and the minimum decomposition error as the dual objective functions; The formula for the local fuzzy entropy of objective function 1 is as follows: , in, is the objective function value of the local fuzzy entropy; t is the center time of the local time window, and the length of the local window L is preset according to the time domain characteristics of the signal. Let be the fuzzy membership degree of the signal feature within the local window at time t in the fuzzy grid at row i and column j. The natural logarithm of fuzzy membership degree; Step 4-2, the decomposition error formula for objective function 2 is as follows: , in, To decompose the objective function value of the error; The original power load sequence at time t; Let be the value of the i-th characteristic mode function at time t; It is an L2 norm; Step 4-3: Determine the parameters and range to be optimized for TVF-FMD, and clarify the core adjustable parameters and preset search range of TVF-FMD, as shown in the following formula: , in, The parameter vector to be optimized for TVF-FMD; The bandwidth is the time-varying filter bandwidth; The modal decomposition layer number control coefficient; The threshold for fuzzy entropy aggregation; Step 4-4: Initialize the MOMPA algorithm, construct the population structure, and set the core operating parameters of MOMPA. The individual encoding formula for the population is as follows: , in , For the i-th individual in the population; Let be the time-varying filter bandwidth parameter value for the i-th individual; Let be the modal decomposition layer control coefficient value for the i-th individual; is the fuzzy entropy aggregation threshold parameter value for the i-th individual; N is the population size; The formulas for setting the core parameters are as follows: , Where T is the maximum number of iterations; P is the predator's step size; and d is the dimension of the parameter to be optimized. Steps 4-5: MOMPA two-stage search optimization, guided by Pareto leading edge leaders, performs a global parameter search. The global exploration formula for the leader stage is as follows: , in, For the updated combination of parameters for the i-th individual in the population; This refers to the i-th individual in the population before the update. It is a 3-dimensional random vector uniformly distributed in the interval [0,1]. This indicates that the corresponding parameter components are multiplied one by one; To be a leading individual at the forefront of Pareto; Steps 4-6, the follower phase, involve fine-tuning local parameters based on the optimal individual in the population. The adaptive control parameter formulas are as follows: , Wherein, CF is the adaptive control parameter, with a value range of [0,1]; This represents the current iteration number; This represents the maximum number of iterations. The formula for updating local parameters is as follows: , in, It is the best individual in the current population; Steps 4-6: Pareto non-dominated sorting and optimal parameter selection. The non-dominated solution set is selected to determine the optimal parameters for TVF-FMD. The selection criteria formula is as follows: and , in To combine the objective function values; The formula for determining the optimal parameters is as follows: , in, The optimal parameter combination for TVF-FMD; the Pareto front is the optimal solution set obtained after non-dominated sorting; argmin represents the parameter vector that minimizes the objective function value; Steps 4-6: TVF-FMD parameters are fixed, and the optimal parameters are determined. The final TVF-FMD parameter formula is as follows: , in, This represents the optimal value for the optimized time-varying filter bandwidth. The optimal value of the control coefficient for the number of modal decomposition layers is obtained after optimization. The optimal value for the fuzzy entropy aggregation threshold is the result of optimization.
7. The method according to claim 6, characterized in that, Step 5 includes: Step 5-1: Define the IScaleformer dual optimization objectives; Objective function 1 is the higher-order moment error of spatiotemporal correlation weighted fluctuations, as shown in the following formula: , The formulas for higher-order moments are as follows: , The spatiotemporal weight formula is as follows: , in, Here, S represents the ST-FHME value, and T represents the time duration. For spatiotemporal correlation weights, To predict the k-th order central moment of the load, The actual load is represented by the k-th order central moment, where k is the order of higher-order moments, and W is the length of the calculation window. This represents the average actual load within the window. The spatiotemporal correlation coefficient; Step 5-2: Objective function 2 is the multi-scale peak-weighted nonlinear error, as shown in the following formula: , The peak weight formula is as follows: , The formula for the nonlinear penalty term is as follows: , in, For MS-PLWE-NL value, For multi-scale layers, Let m be the length of the sequence at scale m. For peak weight, To predict load, This is the actual load. The peak penalty coefficient, Peak threshold This represents the maximum load value at the m-th scale. This is the average load. Step 5-3: Determine the parameters to be optimized for IScaleformer: , in, For spatial feature dimensions, For multi-scale convolution kernel number, For the number of spatiotemporal attention heads, The spatiotemporal fusion coefficient; This represents the set of parameters that need to be optimized in the IScaleformer model; Step 5-4, MOMPA parameter optimization: The initialization formula is as follows: , in, For the i-th parameter combination; Parameter update, global exploration formula is as follows: , The formula for partial development is as follows: , in, It is a 4-dimensional random vector. Represents element-wise product. As a leading figure in Pareto; Step 5-5: Optimal parameter selection and solidification.
8. The method according to claim 7, characterized in that, In step 5-5, the optimal parameters are selected and solidified using the following formula: , , in, For the optimal parameter combination, They are respectively , , , The optimal value.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, It stores a computer program or instructions that, when run on a computer, perform the steps of the method as described in any one of claims 1 to 8.