Power load prediction method and device, electronic equipment and storage medium
By combining variational mode decomposition, self-attention mechanism and quantum particle swarm optimization algorithm, the problem of poor stability in power load forecasting is solved, and deep fusion and precise decoupling of unstructured data are achieved, thereby improving the accuracy and stability of power load forecasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing power load forecasting technologies suffer from poor forecasting stability when faced with complex industrial scenarios, struggle to capture the deep impact of unstructured data on load fluctuations, and lack an understanding of the underlying business logic.
By acquiring historical power load data and related feature data, and after preprocessing, variational mode decomposition technology is used for feature extraction and decoupling. A self-attention mechanism is introduced for multimodal cross-attention fusion, and quantum particle swarm optimization algorithm is used for global hyperparameter optimization to construct a power load prediction model.
It significantly improves the global generalization performance and dynamic response accuracy of power load forecasting in complex environments such as virtual power plants, enhances the model's causal reasoning ability for complex business logic, and improves the stability and accuracy of forecasts.
Smart Images

Figure CN121840580A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for predicting power load. Background Technology
[0002] With the deepening of global energy transition and electricity market reform, the large-scale integration of new energy sources, such as wind and solar power, into the power grid has brought severe challenges to the safe and stable operation of the power system due to their inherent randomness and volatility. As the core foundation for power dispatch, resource optimization, and peak shaving and valley filling, power load forecasting plays an irreplaceable role in ensuring the economical operation of the power grid. Especially in the complex system of virtual power plants (VPPs), which aggregate a large number of heterogeneous loads and distributed power sources, the accuracy of load forecasting is directly related to energy utilization and the scientific nature of dispatch decisions, and there is an urgent need to improve accuracy across all time scales.
[0003] Existing load forecasting techniques have evolved from traditional statistical methods to a deep learning-based combinatorial forecasting stage. Mainstream approaches often employ signal processing techniques such as variational mode decomposition (VMD) to adaptively decompose the original load signal into multiple intrinsic mode functions (EMFs) to achieve frequency domain decoupling. Subsequently, deep learning models are constructed using long short-term memory (LSTM) networks, convolutional neural networks (CNNs), and attention mechanisms to enhance the ability to extract spatial features from power time series data. To further optimize forecasting performance, researchers often introduce particle swarm optimization (PSO) or genetic algorithms to optimize hyperparameters, aiming to improve the model's generalization performance under non-stationary fluctuations.
[0004] However, although existing AI algorithms have improved prediction accuracy to some extent, they still face significant limitations in complex industrial scenarios. First, most existing models are small, task-specific "numerical-driven" models whose inputs heavily rely on structured historical load and meteorological data, severely lacking an understanding of deep business logic. This makes it difficult for these models to capture the profound impact of unstructured data on load fluctuations when faced with complex situations such as production scheduling adjustments or sudden social events, resulting in poor prediction stability. Summary of the Invention
[0005] This invention provides a power load forecasting method, apparatus, electronic device, and storage medium to solve the problem of poor stability in power load forecasting.
[0006] According to one aspect of the present invention, a method for predicting electricity load is provided, the method comprising:
[0007] Historical power load data and related feature data are acquired, and the historical power load data and feature data are preprocessed; the feature data includes environmental feature data, time feature data and text feature data.
[0008] The target features are obtained by performing feature extraction and feature decoupling on the preprocessed historical power load data based on variational mode decomposition technology.
[0009] A self-attention mechanism is introduced to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and a power load prediction model is constructed.
[0010] The quantum particle swarm optimization algorithm is used to optimize the global hyperparameters of the power load prediction model to obtain the target hyperparameters.
[0011] The target hyperparameters are used to train the power load prediction model to obtain the target power load prediction model, and the target power load prediction model is used to predict the power load.
[0012] According to another aspect of the present invention, an electricity load forecasting device is provided, the device comprising:
[0013] The data acquisition and preprocessing module is used to acquire historical power load data and related feature data, and to preprocess the historical power load data and feature data; the feature data includes environmental feature data, time feature data and text feature data.
[0014] The feature extraction and decoupling module is used to extract and decouple features from preprocessed historical power load data based on variational mode decomposition technology to obtain target features.
[0015] The model building module is used to introduce a self-attention mechanism to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and to build a power load prediction model.
[0016] The global optimization module is used to perform global hyperparameter optimization on the power load prediction model using the quantum particle swarm optimization algorithm to obtain the target hyperparameters.
[0017] The power load forecasting module is used to train the power load forecasting model using the target hyperparameters to obtain the target power load forecasting model, and to use the target power load forecasting model to forecast the power load.
[0018] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the power load forecasting method according to any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the power load forecasting method according to any embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the power load forecasting method as described in any embodiment of the present invention.
[0024] The technical solution of this invention involves acquiring historical power load data and related feature data, and preprocessing the historical power load data and feature data. The feature data includes environmental feature data, time feature data, and text feature data. Based on variational mode decomposition (VMD), feature extraction and feature decoupling are performed on the preprocessed historical power load data to obtain target features. A self-attention mechanism is introduced to dynamically fuse the target features and the preprocessed feature data through multimodal cross-attention, and a power load prediction model is constructed. A quantum particle swarm optimization (QPSO) algorithm is used to optimize the global hyperparameters of the power load prediction model to obtain target hyperparameters. The power load prediction model is trained using the target hyperparameters to obtain the target power load prediction model, and the target power load prediction model is used for power load prediction. The technical solution adopted in this invention effectively overcomes the problems of traditional methods' reliance on human experience and severe lack of understanding of deep business logic, making it difficult to capture the deep impact of unstructured data on load fluctuations and resulting in poor prediction stability. By deeply fusing numerical load sequences with unstructured instruction text through a self-attention mechanism, the model's causal reasoning ability for complex business logic is significantly enhanced. This achieves accurate decoupling of multi-scale features of power load and efficient fusion of cross-modal features, thereby significantly improving the global generalization performance, convergence speed, and dynamic response accuracy of power load prediction in complex environments such as virtual power plants, as well as in extreme situations.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of a power load forecasting method provided in Embodiment 1 of the present invention;
[0028] Figure 2 This is a flowchart of a power load forecasting method provided in Embodiment 2 of the present invention;
[0029] Figure 3 This is a flowchart of another power load forecasting method provided in Embodiment 2 of the present invention;
[0030] Figure 4 This is a schematic diagram of the structure of a power load forecasting device according to Embodiment 3 of the present invention;
[0031] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] The acquisition, storage, use, and processing of data in the technical solution of this application all comply with relevant laws and regulations. It should be noted that the terms "first," "second," "target," and "original," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising," "etc.," and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] Example 1
[0035] Figure 1 This document provides a flowchart of a power load forecasting method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where a power load forecasting model is constructed by deeply mining the semantic relationships between unstructured business instructions and numerical loads using a self-attention mechanism, and by introducing a quantum particle swarm optimization algorithm to perform global optimization of the model's internal hyperparameters. This method can be executed by a power load forecasting device, which can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities. Figure 1 As shown, the method includes:
[0036] S110. Obtain historical power load data and related feature data, and preprocess the historical power load data and feature data.
[0037] Among them, power load data can refer to the power demand data of the power system within a specific time area; in this embodiment of the invention, massive amounts of historical power load data are obtained through a power monitoring and acquisition system. The data span is usually 6 to 10 years, and the sampling frequency covers once every 15 minutes (96 sampling points per day) or once per hour (87,600 samples per year).
[0038] Secondly, to determine the accuracy of power load forecasting, this embodiment of the invention also acquires feature data related to power load data; wherein, feature data includes environmental feature data, time feature data, and text feature data. The environmental feature data includes, but is not limited to, feature data such as temperature, humidity, and wind speed; the time feature data includes, but is not limited to, weekdays, holidays, and weekends, etc., and power load data at special time points differs from other normal time points; the text feature data includes, but is not limited to, unstructured text feature data such as production order descriptions, virtual power plant process instructions, and power market policy texts.
[0039] After acquiring historical power load data and related feature data, preprocessing is performed on the historical power load data and feature data; the preprocessing includes, but is not limited to, normalization. Specifically, to eliminate the impact of differences in the dimensions of different features on the convergence speed of the large model, this embodiment of the invention performs min-max normalization on all historical power load data and feature data, mapping the values to the [0, 1] interval, and the calculation formula is as follows:
[0040] ;
[0041] in, The data is normalized, and x represents the historical power load data or characteristic data to be preprocessed. and These represent the minimum and maximum values in the sample sequence, respectively. Normalizing historical power load data and feature data can significantly improve the learning efficiency of the subsequent neural network for feature weights.
[0042] In one optional embodiment of the present invention, during the preprocessing stage, a sliding window technique is used to cyclically divide the acquired historical power load data and feature data according to a 24-hour cycle in order to capture the periodic pattern of power load.
[0043] S120. Based on variational mode decomposition technology, feature extraction and feature decoupling are performed on the preprocessed historical power load data to obtain the target features.
[0044] Variational Mode Decomposition (VMD) is an adaptive signal decomposition method. Its core idea is to iteratively optimize a variational model to decompose the original signal into multiple eigenmode functions with specific frequency characteristics. In this embodiment of the invention, VMD is used to adaptively decompose the preprocessed historical power load data into k eigenmodes, each with its own center frequency. Finite-bandwidth intrinsic mode function (IMF) components Achieve frequency domain decoupling of power load data.
[0045] Feature extraction refers to extracting key information that characterizes the essential patterns of historical power load data from preprocessed historical power load data, mapping high-dimensional data to a low-dimensional feature space. For example, statistical features (such as mean, variance, peak value, and kurtosis), time-domain features (such as autocorrelation coefficient and trend slope), and frequency-domain features (such as center frequency and energy percentage) can be extracted from the IMF components decomposed by VMD. Optionally, when extracting features from historical power load data, a comprehensive feature set can be constructed by combining preprocessed feature data, such as temperature, date type, holidays, and virtual power plant process instructions, to perform feature extraction on the preprocessed historical power load data. For example, the time of maximum daily load and the load fluctuation amplitude can be calculated for the decomposed "daily cycle mode" to reflect the user's habitual electricity consumption characteristics.
[0046] Feature decoupling refers to the process of eliminating redundant correlations or cross-influences between features, allowing each feature to more independently reflect a specific physical meaning or causal relationship. In power load scenarios, power load is affected by multiple coupled factors, and direct modeling can easily lead to feature redundancy or overfitting. In this embodiment of the invention, feature decoupling of power load data can effectively reduce the complexity of power load data.
[0047] The target feature can refer to the feature obtained after feature extraction and feature decoupling of preprocessed historical power load data based on variational mode decomposition technology.
[0048] S130. Introduce a self-attention mechanism to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and construct a power load prediction model.
[0049] The self-attention mechanism refers to a mechanism that allows a model to dynamically focus on the relationships between different positions within a sequence when processing sequential data, used to calculate the correlation weight of each element in the sequence with all other elements. In this embodiment of the invention, the self-attention mechanism is used to capture the deep causal logic between text features and electricity load data, thereby addressing the problem that traditional models cannot understand unstructured data.
[0050] Multimodal cross-attention dynamic fusion refers to a mechanism that allows features from different modalities to "communicate" with each other, enabling the model to dynamically select and integrate information from different data sources. In this embodiment of the invention, it refers to using a self-attention mechanism to determine the deep causal logic between preprocessed feature data and target features obtained based on variational mode decomposition technology, and then fusing them to construct a power load forecasting model for power load prediction.
[0051] S140. The quantum particle swarm optimization algorithm is used to optimize the global hyperparameters of the power load prediction model to obtain the target hyperparameters.
[0052] Quantum Particle Swarm Optimization (QSO) is an improved version of traditional Particle Swarm Optimization (PSO), incorporating principles of quantum mechanics, such as quantum tunneling and wave function probability distribution. QSO breaks through the velocity-position update formula of traditional PSO by describing particle positions through quantum states, enhancing global search capabilities and reducing the risk of getting trapped in local optima. In power load forecasting, QSO is used to optimize the hyperparameters of deep learning models such as LSTM and GRU, finding the optimal parameter combination through swarm intelligence iteration.
[0053] Global hyperparameter optimization refers to the process of searching for the optimal configuration within all possible hyperparameter values using the quantum particle swarm optimization algorithm before model training. These hyperparameters include, but are not limited to, the initial learning rate of the Adam optimizer, the regularization coefficient in the loss function, and the overlap ratio of the sliding window.
[0054] The hyperparameters of power load forecasting models are often strongly coupled. The global search characteristics of the quantum particle swarm optimization algorithm can efficiently traverse parameter combinations. Compared with existing methods such as grid search and random search, it can find an approximate optimal solution in fewer iterations, making it particularly suitable for forecasting tasks involving massive historical power load data and multivariate inputs.
[0055] S150. The target hyperparameters are used to train the power load prediction model to obtain the target power load prediction model, and the target power load prediction model is used to predict the power load.
[0056] In this process, after determining the hyperparameters of the model, the power load prediction model is trained to determine the target power load prediction model in order to predict the power load.
[0057] This invention provides a method for predicting power load. The method involves acquiring historical power load data and related feature data, and preprocessing the historical power load data and feature data. The feature data includes environmental feature data, time feature data, and text feature data. Based on variational mode decomposition (VMD), feature extraction and feature decoupling are performed on the preprocessed historical power load data to obtain target features. A self-attention mechanism is introduced to dynamically fuse the target features and the preprocessed feature data using multimodal cross-attention, and a power load prediction model is constructed. A quantum particle swarm optimization (QPSO) algorithm is used to optimize the global hyperparameters of the power load prediction model to obtain target hyperparameters. The power load prediction model is trained using the target hyperparameters to obtain a target power load prediction model, and the target power load prediction model is then used for power load prediction. The technical solution adopted in this invention effectively overcomes the problems of traditional methods' reliance on human experience and severe lack of understanding of deep business logic, making it difficult to capture the deep impact of unstructured data on load fluctuations and resulting in poor prediction stability. By deeply fusing numerical load sequences with unstructured instruction text through a self-attention mechanism, the model's causal reasoning ability for complex business logic is significantly enhanced. This achieves accurate decoupling of multi-scale features of power load and efficient fusion of cross-modal features, thereby significantly improving the global generalization performance, convergence speed, and dynamic response accuracy of power load prediction in complex environments such as virtual power plants, as well as in extreme situations.
[0058] Example 2
[0059] Figure 2 This is a flowchart of a power load forecasting method provided in Embodiment 2 of the present invention. This embodiment further optimizes the aforementioned embodiments, and can be combined with various optional schemes from one or more of the above embodiments. For example... Figure 2 As shown, the method includes:
[0060] S210. Obtain historical power load data and related feature data, and preprocess the historical power load data and feature data.
[0061] S220. Based on variational mode decomposition technology, feature extraction and feature decoupling are performed on the preprocessed historical power load data to obtain the target features.
[0062] Given the strong non-stationarity, nonlinearity, and multi-scale coupling characteristics of power load data—that is, fluctuations of different frequencies and driving sources superimposed together—directly inputting power load data into a large model may lead to the model's inability to capture local detailed features. Therefore, this embodiment of the invention introduces variational mode decomposition technology as a key pre-feature extraction step.
[0063] As an optional but non-limiting implementation, the feature extraction and feature decoupling processing of the preprocessed historical power load data based on variational mode decomposition technology to obtain target features includes, but is not limited to, steps A1-A2:
[0064] Step A1: Based on variational mode decomposition technology, construct a variational constraint problem based on the preprocessed historical power load data; wherein, the variational constraint problem refers to finding at least two modal components, the sum of the estimated bandwidths of each modal component is minimized, and the sum of all modal components is equal to the preprocessed historical power load data;
[0065] Step A2: By introducing a quadratic penalty factor and Lagrange multipliers, the variational constraint problem is transformed into an unconstrained augmented Lagrange function problem, and the alternating direction multiplier method is used for iterative solution to obtain the target feature after feature decoupling.
[0066] The mathematical essence of VMD is to find k modal components. This minimizes the sum of the estimated bandwidths for each mode while satisfying the constraint that the sum of all modal components equals the original signal. To estimate each mode... To determine the bandwidth, we first calculate the analytic signal using the Hilbert transform to obtain the one-sided spectrum; then, we multiply it by a complex exponential term. The spectral centers of each mode are shifted to the baseband; finally, the bandwidth is estimated by calculating the squared L2 norm of the demodulated signal gradient. The resulting variational constraint problem is shown in the following formula:
[0067] ,
[0068] ;
[0069] in, and These represent the set of k modal components obtained from the decomposition and their corresponding set of center frequencies, respectively. The summation operator represents the accumulation of k preset modal components; The partial derivative operator represents the derivative with respect to time t, used to calculate the gradient of the demodulated signal to estimate the bandwidth; is the Dirac distribution function, used to construct the impulse response of an analytic signal; For convolution operators, it represents the process of convolving a signal with an analytic operator; This is a complex exponential term used to shift the spectral center of each mode to the baseband.
[0070] To solve the aforementioned optimization problem with equality constraints, this invention introduces a quadratic penalty factor. and Lagrange multipliers The above variational constraint problem is transformed into an unconstrained augmented Lagrangian function problem; where the quadratic penalty factor... Lagrange multipliers are used to balance signal reconstruction accuracy and robustness to noise. Used to strictly enforce constraints.
[0071] Subsequently, the alternating direction multiplier method is used to alternately update each modal component in the frequency domain through iterative cycles. Center frequency and Lagrange multipliers The goal is to find the saddle point of the augmented Lagrange function. The specific iterative update formula is as follows:
[0072] 1) Modal components Update:
[0073] The update process of each mode in the frequency domain is essentially a Wiener filtering process, and its update formula is:
[0074] ;
[0075] Where n represents the current iteration number, This represents the frequency domain representation after the Fourier transform. This represents the analytic signal in the frequency domain of the k-th intrinsic mode function (IMF) obtained after the (n+1)-th iteration; This represents the Fourier transform spectrum of the original power load signal after preprocessing. This represents the sum of the spectra of all modes except the k-th mode in the previous iteration in the current iteration; The Lagrange multiplier spectrum at the nth iteration is used to constrain the reconstruction error. It is a continuous frequency variable; This represents the estimated center frequency of the k-th mode in the nth iteration. The formula shows that the spectrum of the k-th mode is calculated by subtracting the spectra of all other modes and multipliers from the original signal spectrum, and then passing the result through a center frequency of... It is obtained from a bandpass filter.
[0076] 2) Center frequency Update:
[0077] The new center frequency is set as the centroid of the current modal power spectrum, denoted as:
[0078] ;
[0079] in, This represents the center frequency of the updated k-th mode, which is the centroid of the power spectrum of that mode. This represents the energy density spectrum (power spectrum) of the kth modal component obtained in the (n+1)th iteration. This represents the integral operation in the positive frequency domain.
[0080] 3) Lagrange multipliers The update (dual ascent) is represented as:
[0081] ;
[0082] in, This represents the updated Lagrange multiplier spectrum; The update step size for the Lagrange multiplier.
[0083] The above iterative process continues until a preset convergence condition is met; for example, the sum of the squares of the relative errors of each modal component between two consecutive iterations is less than a specific threshold. .
[0084] After the above VMD iterative processing, the original complex power load data was successfully decoupled into k IMF component sequences with clear physical meanings, which serve as input features for subsequent large-scale models:
[0085] High-frequency IMF components typically capture details of high-frequency fluctuations caused by transient noise, sudden disturbances, or rapid switching operations in the load.
[0086] Mid-to-low frequency IMF components typically reflect the intraday periodicity, seasonality, and baseline load trends under the influence of macroeconomic factors.
[0087] This multi-scale feature decoupling effectively reduces the complexity of the signal, enabling subsequent deep learning models to learn components with different frequency characteristics separately, thus significantly improving prediction accuracy.
[0088] S230. Introduce a self-attention mechanism to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and construct a power load prediction model.
[0089] In this embodiment of the invention, a self-attention mechanism is used to determine the deep causal logic between the preprocessed feature data and the target features obtained based on variational mode decomposition technology, and then they are fused to construct a power load prediction model for power load prediction.
[0090] As an optional but non-limiting implementation, the introduction of a self-attention mechanism dynamically fuses the target features and preprocessed feature data through multimodal cross-attention to construct a power load prediction model, including but not limited to steps B1-B3:
[0091] Step B1: Build a large model based on the Transformer architecture and use it as a high-level semantic awareness layer;
[0092] Step B2: Introduce a self-attention mechanism to obtain the causal logic between historical power load data and text feature data, and perform multimodal cross-attention dynamic fusion of the target features and preprocessed feature data;
[0093] Step B3: Build an electricity load forecasting model by dynamically adjusting the weight matrix and bias terms of the large model based on the Transformer architecture.
[0094] This invention constructs a composite architecture integrating the feature representation capabilities of large-scale artificial intelligence models with the ability to capture temporal features. The underlying architecture employs a multi-layered Long Short-Term Memory (LSTM) network or a Backpropagation (BP) neural network as the temporal feature extractor. Utilizing the input gate (it), forget gate (ft), and output gate (ot) mechanisms of LSTM, and through the sigmoid activation function and element-wise multiplication, it effectively solves the gradient vanishing and gradient exploding problems in long-sequence training of power load data, thereby accurately capturing long-term dependencies in the sequence. Simultaneously, this invention introduces a large-scale model based on the Transformer architecture as a high-level semantic perception layer. It utilizes a self-attention mechanism to mine the deep causal logic between textual features (such as production plan text) and numerical load data, addressing the problem that traditional models cannot understand unstructured data. This multimodal fusion architecture achieves deep representation of the features of massive high-dimensional power data through dynamic adjustment of the weight matrix W and the bias term b.
[0095] In this embodiment of the invention, a high-level semantic layer of a large model based on a pre-trained Transformer is constructed. The numerical load sequence and unstructured instruction text features are deeply fused through a self-attention mechanism, which significantly enhances the model's ability to make causal inferences about complex business logic.
[0096] S240. The quantum particle swarm optimization algorithm is used to optimize the global hyperparameters of the power load prediction model to obtain the target hyperparameters.
[0097] In response to the pain points of high-dimensional, non-convex parameter spaces in multimodal large models and underlying neural networks, which are prone to getting trapped in local optima, this invention introduces a quantum heuristic optimization mechanism to achieve adaptive global optimization of key model configurations, effectively overcoming the limitation of traditional methods that are prone to getting trapped in local optima.
[0098] As an optional but non-limiting implementation, the method employs quantum particle swarm optimization to perform global hyperparameter optimization on the power load prediction model to obtain the target hyperparameters, including but not limited to steps C1-C2:
[0099] Step C1: The quantum particle swarm optimization algorithm is used to perform global hyperparameter optimization on the power load prediction model, and the state of each particle is iteratively updated through the quantum evolution formula during the search process;
[0100] Step C2: Introduce the quantum excitation mechanism from the dynamic particle swarm optimization algorithm to map the quantum probability state of each particle back to the actual parameter value space. When the population convergence radius is detected to be lower than the preset threshold, quantum perturbation is automatically triggered to determine the globally optimal target hyperparameter that minimizes the model prediction residual.
[0101] In this invention, the embodiment does not rely on the deterministic trajectory update of traditional particle swarm optimization, but instead uses the wave function probability distribution in quantum mechanics to describe the position of the particle in the search space, giving the particle the ability to penetrate the local minimum "barrier".
[0102] This invention utilizes the Quantum Particle Swarm Optimization (QPSO) algorithm, focusing on a global joint search of architecture parameters, fusion parameters, and training hyperparameters. The architecture parameters include the number of multi-head attention heads H in the Transformer encoder, the hidden layer mapping dimension dmodel, and the number of hidden units in the underlying temporal feature extraction module. The fusion parameters include the initial weight distribution range of the query matrix Q, key matrix K, and value matrix V in the multimodal cross-attention fusion layer. The training hyperparameters include the initial learning rate of the Adam optimizer. The regularization coefficient in the loss function and the overlap rate of the sliding window.
[0103] During the search process, the state of each particle is iteratively updated using the following quantum evolution formula:
[0104] ,
[0105] ;
[0106] Where t is the number of iterations. and These are controlled quantum mechanical parameters used to adjust the balance between global exploration and local convergence in the algorithm; pbest represents the historical best position found by the particle itself.
[0107] To further enhance the population diversity of the model under extreme power conditions, this embodiment of the invention introduces a quantum excitation mechanism from Dynamic Particle Swarm Optimization (DPSO). The Delta operator maps the quantum probability states of particles back to the actual parameter value space. When the population convergence radius is detected to be below a preset threshold, a quantum perturbation is automatically triggered, ensuring that the algorithm can escape local minima and ultimately lock in the globally optimal hyperparameter combination that minimizes the model's prediction residuals.
[0108] The embodiments of this invention introduce an optimization algorithm with quantum excitation mechanism (QPSO / DPSO) to perform synchronous global optimization on the key structural parameters of variational mode decomposition and the internal hyperparameters of large models. This effectively overcomes the limitations of traditional methods, such as reliance on human experience and the tendency to get trapped in local optima in high-dimensional non-convex parameter spaces.
[0109] S250. The target hyperparameters are used to train the power load prediction model to obtain the target power load prediction model, and the target power load prediction model is used to predict the power load.
[0110] In this process, after determining the target hyperparameters, the model is trained to obtain the target power load prediction model used for power load prediction.
[0111] As an optional but non-limiting implementation, the step of training the power load prediction model using the target hyperparameters to obtain the target power load prediction model, and then using the target power load prediction model to perform power load prediction, includes, but is not limited to, steps D1-D3:
[0112] Step D1: Train the power load prediction model using the target hyperparameters and utilize the Adam optimizer or an adaptive learning rate adjustment mechanism;
[0113] Step D2: Use mean squared error as the loss function and update the gradient in real time based on the mean squared error of each iteration;
[0114] Step D3: Determine the target power load prediction model that satisfies the preset convergence conditions, and use the target power load prediction model to predict the power load; wherein, the preset convergence conditions include the mean square error being stable or increasing for a preset number of consecutive times and the change in the position of the particles being less than a preset change threshold.
[0115] In this embodiment of the invention, the optimal weights and biases are searched and used as the initial parameters of the large model to initiate end-to-end fine-tuning training. During training, the mean squared error (MSE) is used as the loss function, and the Adam optimizer or adaptive learning rate adjustment mechanism is used to update the gradient in real time based on the error of each iteration. To ensure the model's generalization ability, the training set, validation set, and test set are divided in a proportional ratio (e.g., 7:3 or 8:2). This embodiment of the invention establishes multiple convergence stopping conditions: first, the MSE on the validation set does not decrease for more than 10 consecutive iterations; second, the change in particle position is less than a threshold of 0.001, indicating that the algorithm has reached equilibrium. Simultaneously, the root mean square error (RMSE) is monitored in real time to perform closed-loop evaluation of the model performance. The root mean square error is expressed as:
[0116] ;
[0117] in, These are the model's predicted values. This represents the actual load value.
[0118] This invention significantly improves the global optimization capability and prediction accuracy of power load forecasting by deeply integrating a large artificial intelligence model with an optimization algorithm with a quantum excitation mechanism. It uses quantum mechanical probability distribution to describe particle positions, enabling the model to effectively escape local optima and overcome common problems in traditional neural networks such as overfitting and slow convergence speed. Its actual convergence speed is significantly improved compared to traditional models.
[0119] As an optional but non-limiting implementation, the method further includes the following steps before employing the target power load forecasting model for power load forecasting:
[0120] The target power load prediction model is tested under preset scenarios to determine the average absolute percentage error of the model under different preset scenarios in order to evaluate the model performance. The preset scenarios refer to extreme power environments, including peak load scenarios caused by high temperature or extreme cold, low-valley stable load scenarios at night or on holidays, and sudden load fluctuation scenarios caused by production task switching or failure.
[0121] To verify the robustness of the embodiments of the present invention under complex power environments, tests and analyses will be conducted under three core scenarios: first, peak load scenarios caused by high temperatures or extreme cold; second, stable load scenarios during nighttime or holidays; and third, sudden load fluctuation scenarios caused by production task switching or faults. Experimental results show that, through the quantum dynamic adjustment mechanism, the present invention achieves a mean absolute percentage error (MAPE) as low as 1.3% under peak load, maintains a MAPE below 1.5% during off-peak periods, and consistently maintains a response error below 3% under sudden fluctuations, effectively solving the prediction delay or distortion problems commonly found in traditional models. The core evaluation index, MAPE, is calculated using the following formula:
[0122] .
[0123] This invention introduces a quantum excitation mechanism of Dynamic Particle Swarm Optimization (DPSO). When a sudden load change is detected (such as a process switch in a virtual power plant), the model automatically jumps out of the local optimum through quantum perturbation, ensuring the prediction accuracy under peak and sudden fluctuations and ensuring the reliability of the model in real-time application under diverse load conditions such as virtual power plants.
[0124] When processing massive amounts of power load data, the embodiments of the present invention can not only significantly reduce the root mean square error (RMSE) and mean absolute percentage error (MAPE), but also demonstrate excellent robustness under extreme situations such as peak, off-peak, and sudden fluctuations. The peak load prediction error is lower than that of traditional models, thus providing reliable technical support and decision-making basis for the stable operation, precise scheduling, and optimal allocation of resources of the power system.
[0125] S260. Perform inverse normalization on the power load predicted using the target power load prediction model to determine the actual dimensions of the predicted power load.
[0126] After completing the forward prediction derivation of the large model, the normalized output value is obtained. Then, a final denormalization procedure is performed to restore the actual dimensions (MW) of the power load.
[0127] The maximum value saved during the preprocessing stage in this embodiment of the invention and minimum value The parameters, which map the prediction results back to the actual load value space using the inverse operation formula, can be expressed as:
[0128] .
[0129] The resulting actual load forecast sequence accurately tracks the changing trends of the power grid load. This result provides power dispatchers with high-precision real-time data support for peak shaving and valley filling, production planning and scheduling, and virtual power plant resource allocation decisions, significantly improving the resource utilization and operational stability of the power system.
[0130] Example 3
[0131] Figure 4 This is a schematic diagram of the structure of a power load forecasting device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes:
[0132] The data acquisition and preprocessing module 410 is used to acquire historical power load data and feature data related to the power load data, and to preprocess the historical power load data and feature data; wherein, the feature data includes environmental feature data, time feature data and text feature data.
[0133] The feature extraction and decoupling module 420 is used to perform feature extraction and feature decoupling on the preprocessed historical power load data based on variational mode decomposition technology to obtain target features.
[0134] The model building module 430 is used to introduce a self-attention mechanism to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and to build a power load prediction model.
[0135] The global optimization module 440 is used to perform global hyperparameter optimization on the power load prediction model using the quantum particle swarm optimization algorithm to obtain the target hyperparameters.
[0136] The power load prediction module 450 is used to train the power load prediction model using the target hyperparameters to obtain the target power load prediction model, and to use the target power load prediction model to predict the power load.
[0137] Optional feature extraction and decoupling module, specifically used for:
[0138] Based on variational mode decomposition technology, a variational constraint problem is constructed based on preprocessed historical power load data. The variational constraint problem refers to finding at least two modal components, where the sum of the estimated bandwidths of each modal component is minimized, and the sum of all modal components is equal to the preprocessed historical power load data.
[0139] By introducing a quadratic penalty factor and Lagrange multipliers, the variational constraint problem is transformed into an unconstrained augmented Lagrange function problem, and the alternating direction multiplier method is used for iterative solution to obtain the target feature after feature decoupling.
[0140] Optional, the model building module, specifically used for:
[0141] Build a large model based on the Transformer architecture and use it as a high-level semantic awareness layer;
[0142] A self-attention mechanism is introduced to obtain the causal logic between historical power load data and text feature data, and the target features and preprocessed feature data are dynamically fused through multimodal cross-attention.
[0143] A power load forecasting model is constructed by dynamically adjusting the weight matrix and bias terms of a large model based on the Transformer architecture.
[0144] Optional, global optimization module, specifically used for:
[0145] The quantum particle swarm optimization algorithm is used to optimize the global hyperparameters of the power load prediction model, and the state of each particle is iteratively updated through the quantum evolution formula during the search process.
[0146] The quantum excitation mechanism in the dynamic particle swarm optimization algorithm is introduced to map the quantum probability state of each particle back to the actual parameter value space. When the population convergence radius is detected to be lower than the preset threshold, quantum perturbation is automatically triggered to determine the globally optimal target hyperparameter that minimizes the model prediction residual.
[0147] Optional, power load forecasting module, specifically used for:
[0148] The target hyperparameters are used to train the power load prediction model, and the Adam optimizer or adaptive learning rate adjustment mechanism is used.
[0149] Mean squared error is used as the loss function, and the gradient is updated in real time based on the mean squared error of each iteration;
[0150] A target power load prediction model is determined when the model satisfies the preset convergence conditions, and the target power load prediction model is used to predict the power load; wherein, the preset convergence conditions include the mean square error being stable or increasing for a preset number of consecutive preset times and the change in the position of the particles being less than a preset change threshold.
[0151] Optionally, before performing power load forecasting using the target power load forecasting model, the device further includes a performance evaluation module, specifically used for:
[0152] The target power load prediction model is tested under preset scenarios to determine the average absolute percentage error of the model under different preset scenarios in order to evaluate the model performance. The preset scenarios refer to extreme power environments, including peak load scenarios caused by high temperature or extreme cold, low-valley stable load scenarios at night or on holidays, and sudden load fluctuation scenarios caused by production task switching or failure.
[0153] Optionally, after training the power load prediction model using the target hyperparameters to obtain the target power load prediction model, and then using the target power load prediction model to predict power load, the device further includes an inverse normalization module, specifically used for:
[0154] The power load predicted by the target power load prediction model is inversely normalized to determine the actual dimensions of the predicted power load.
[0155] The power load forecasting device provided in the embodiments of the present invention can execute the power load forecasting method provided in any of the embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the power load forecasting method. For details, please refer to the relevant operations of the power load forecasting method in the foregoing embodiments.
[0156] Example 4
[0157] Figure 5 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0158] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0159] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0160] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as power load forecasting methods.
[0161] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.
[0162] In some embodiments, the power load forecasting method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the power load forecasting method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the power load forecasting method by any other suitable means (e.g., by means of firmware).
[0163] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0164] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0165] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0167] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0168] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0169] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0170] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for predicting electricity load, characterized in that, The method includes: Historical power load data and related feature data are acquired, and the historical power load data and feature data are preprocessed; the feature data includes environmental feature data, time feature data and text feature data. The target features are obtained by performing feature extraction and feature decoupling on the preprocessed historical power load data based on variational mode decomposition technology. A self-attention mechanism is introduced to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and a power load prediction model is constructed. The quantum particle swarm optimization algorithm is used to optimize the global hyperparameters of the power load prediction model to obtain the target hyperparameters. The target hyperparameters are used to train the power load prediction model to obtain the target power load prediction model, and the target power load prediction model is used to predict the power load.
2. The method according to claim 1, characterized in that, The variational mode decomposition technique is used to extract and decouple features from the preprocessed historical power load data to obtain target features, including: Based on variational mode decomposition technology, a variational constraint problem is constructed based on preprocessed historical power load data. The variational constraint problem refers to finding at least two modal components, where the sum of the estimated bandwidths of each modal component is minimized, and the sum of all modal components is equal to the preprocessed historical power load data. By introducing a quadratic penalty factor and Lagrange multipliers, the variational constraint problem is transformed into an unconstrained augmented Lagrange function problem, and the alternating direction multiplier method is used for iterative solution to obtain the target feature after feature decoupling.
3. The method according to claim 1, characterized in that, The introduction of a self-attention mechanism dynamically fuses the target features and preprocessed feature data through multimodal cross-attention, and constructs a power load prediction model, including: Build a large model based on the Transformer architecture and use it as a high-level semantic awareness layer; A self-attention mechanism is introduced to obtain the causal logic between historical power load data and text feature data, and the target features and preprocessed feature data are dynamically fused through multimodal cross-attention. A power load forecasting model is constructed by dynamically adjusting the weight matrix and bias terms of a large model based on the Transformer architecture.
4. The method according to claim 1, characterized in that, The quantum particle swarm optimization algorithm is used to perform global hyperparameter optimization on the power load prediction model to obtain the target hyperparameters, including: The quantum particle swarm optimization algorithm is used to optimize the global hyperparameters of the power load prediction model, and the state of each particle is iteratively updated through the quantum evolution formula during the search process. The quantum excitation mechanism in the dynamic particle swarm optimization algorithm is introduced to map the quantum probability state of each particle back to the actual parameter value space. When the population convergence radius is detected to be lower than the preset threshold, quantum perturbation is automatically triggered to determine the globally optimal target hyperparameter that minimizes the model prediction residual.
5. The method according to claim 1, characterized in that, The process of training the power load prediction model using the target hyperparameters to obtain the target power load prediction model, and then using the target power load prediction model to predict power load, includes: The target hyperparameters are used to train the power load prediction model, and the Adam optimizer or adaptive learning rate adjustment mechanism is used. Mean squared error is used as the loss function, and the gradient is updated in real time based on the mean squared error of each iteration; A target power load prediction model is determined when the model satisfies the preset convergence conditions, and the target power load prediction model is used to predict the power load; wherein, the preset convergence conditions include the mean square error being stable or increasing for a preset number of consecutive preset times and the change in the position of the particles being less than a preset change threshold.
6. The method according to claim 1, characterized in that, Before using the target power load forecasting model for power load forecasting, the method further includes: The target power load prediction model is tested under preset scenarios to determine the average absolute percentage error of the model under different preset scenarios in order to evaluate the model performance. The preset scenarios refer to extreme power environments, including peak load scenarios caused by high temperature or extreme cold, low-valley stable load scenarios at night or on holidays, and sudden load fluctuation scenarios caused by production task switching or failure.
7. The method according to claim 1, characterized in that, After training the power load prediction model using the target hyperparameters to obtain the target power load prediction model, and then using the target power load prediction model to predict power load, the method further includes: The power load predicted by the target power load prediction model is inversely normalized to determine the actual dimensions of the predicted power load.
8. A power load forecasting device, characterized in that, The device includes: The data acquisition and preprocessing module is used to acquire historical power load data and related feature data, and to preprocess the historical power load data and feature data; the feature data includes environmental feature data, time feature data and text feature data. The feature extraction and decoupling module is used to extract and decouple features from preprocessed historical power load data based on variational mode decomposition technology to obtain target features. The model building module is used to introduce a self-attention mechanism to dynamically fuse the target features and preprocessed feature data through multimodal cross-attention, and to build a power load prediction model. The global optimization module is used to perform global hyperparameter optimization on the power load prediction model using the quantum particle swarm optimization algorithm to obtain the target hyperparameters. The power load forecasting module is used to train the power load forecasting model using the target hyperparameters to obtain the target power load forecasting model, and to use the target power load forecasting model to forecast the power load.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the power load forecasting method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the power load forecasting method according to any one of claims 1-7.