Short-term load prediction system based on coder-decoder architecture and construction method thereof
By adopting the encoding-decoder architecture in power load prediction, combining multi-scale expansion of causal convolution networks and bidirectional long and short-term memory networks, the problems of model complexity and overfitting are solved, and more accurate and robust short-term load prediction is achieved, suitable for real-time scheduling of power systems.
Patent Information
- Application Number
- CN202510236937.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-01
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art faces the problems of model complexity control, nonlinear relationship capture and overfitting in power load prediction, especially when faced with highly diverse weather and time series changes, the prediction accuracy is limited.
A short-term load prediction system based on the encoding-decoder architecture is proposed, combining multi-scale expansion of causal convolutional networks and bidirectional long and short-term memory networks, and deep separation of convolutional structures and dynamic neuron allocation mechanisms to realize parameter compression and memory usage quantification, reducing the risk of overfitting.
It realizes more accurate and robust short-term load prediction, improves the generalization ability and real-time nature of the model, reduces the computing complexity and storage requirements, and is suitable for the real-time scheduling requirements of the power system.
Smart Images

Figure CN120146292A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of short-term load forecasting in power systems, and particularly to a short-term load forecasting system based on an encoder-decoder architecture and a construction method thereof. Background Art
[0002] Short-term load forecasting (STLF) in power systems is a core task to ensure efficient power production and distribution, and its accuracy directly affects energy management and scheduling efficiency. This task has long been a focus of attention in academia and industry. Traditionally, statistical methods such as autoregressive (AR), moving average (MA), autoregressive moving average (ARMA) and their extensions (such as ARIMA, ARMAX, OE, etc.) have occupied a place in early load forecasting due to their mathematical rigor and strong interpretability. However, these models are mainly based on linear assumptions and are difficult to capture complex non-linear relationships and the influence of external factors (such as weather, holiday effects), which limits the accuracy and practicality of forecasting.
[0003] With the rise of machine learning and deep learning technologies, a series of advanced forecasting models have been developed to improve forecasting performance. Machine learning models such as support vector machine (SVM), k-nearest neighbor (KNN), random forest, gradient boosting trees (such as Tree Bagger), and artificial neural network (ANN) show the potential to outperform traditional statistical methods by learning patterns in historical data. In particular, the application of ensemble learning and optimization techniques (such as particle swarm optimization PSO, Levenberg-Marquardt algorithm) further enhances the expressive power and adaptability of these models. Nevertheless, these models are still vulnerable to overfitting when facing high-dimensional input features and long-term dependencies in time series, especially in the case of significant changes in weather and time series, and the generalization ability of the models is limited.
[0004] Deep learning models, such as convolutional neural network (CNN) and long short-term memory network (LSTM), have been widely used in load forecasting because they can efficiently process sequence data and extract spatio-temporal features. CNN captures local features through multi-layer convolutional operations, while LSTM uses gating mechanisms to handle long-term dependencies. The emergence of the CNN-LSTM hybrid model combines the advantages of both, but in practice, large filter sizes and multi-layer structures may lead to too high model complexity, increasing the risk of overfitting, especially in power load data with rich and fluctuating feature spaces. Therefore, the forecasting performance is significantly reduced when facing highly diverse meteorological and time series fluctuations. In addition, such models often neglect the fine capture of key local features, especially those short-term load changes that have a direct impact on the forecasting results, resulting in limited forecasting accuracy.
[0005] In summary, while the existing technologies improve the prediction accuracy, they face problems such as model complexity control, capturing non-linear relationships, and overfitting, especially when dealing with highly diverse weather and time series variations. Therefore, there is an urgent need for a new model architecture that can efficiently extract features and maintain the generalization ability of the model to address the above challenges. Summary of the Invention
[0006] In view of the actual needs in the field of electric load forecasting, the present invention proposes a short-term load forecasting system based on an encoder-decoder architecture and a construction method thereof. The system combines actual physical factors in electric load forecasting, such as meteorological data, historical load data, etc., and uses an encoder-decoder architecture to capture complex patterns of electric load changes, thereby providing accurate and real-time prediction results.
[0007] A short-term load forecasting system based on an encoder-decoder architecture, characterized by comprising:
[0008] An encoder, configured to receive feature data highly relevant to the load, where the feature data highly relevant to the load includes electric load data and meteorological data, and extract local features of the electric load pattern through a multi-scale dilated causal convolutional network;
[0009] A decoder, configured to input the local features into a bidirectional long short-term memory network, convert them into predicted electric load values, and output them;
[0010] Wherein:
[0011] The encoder adopts a depthwise separable convolution structure to achieve parameter compression, decomposes the standard convolution into two steps of depthwise convolution (depthwise conv3×1) and pointwise convolution (pointwise conv1×1), and reduces the number of parameters to 1 / 8 of that of the traditional convolution through decoupling of the convolution kernel dimensions; at the same time, a hybrid dilation rate causal convolution module (dilation rates = [1, 2, 4]) is designed to construct multi-scale feature extraction ability while maintaining the temporal causal relationship constraint, and in cooperation with the grouped convolution (group = 4) strategy, further reduces the storage requirement of the weight matrix by 30%;
[0012] The decoder implements hierarchical neuron pruning based on sensitivity analysis, calculates the importance scores of the LSTM cell states through second-order derivative metrics, and removes redundant neurons in the hidden layer with parameter update amplitudes less than the threshold η = 0.01; adopts a dynamic neuron allocation mechanism to maintain 256 hidden units in the temporal feature extraction layer and reduce to 128 units in the spatio-temporal coupling layer, and in cooperation with 8-bit integer quantization, realizes the compression of the memory occupancy from 3.2MB in FP32 to 804KB, and at the same time maintains the RMSE of the validation set ≤ 0.048 through a selective gradient update strategy.
[0013] Furthermore, it further includes a predictor matrix module. The predictor matrix in the predictor matrix module contains feature data highly correlated with the load. The feature data correlated with the load height includes power load data and meteorological data. The meteorological data includes temperature, humidity, and wind speed. The meteorological data is from meteorological monitoring stations, and the historical load data is taken from the SCADA system, with a sampling frequency of 15 minutes for both. The feature data extracts multi-scale meteorological features and load frequency domain features through causal convolution and fuses them through a gated attention mechanism, providing input for the prediction model formed by the encoder-decoder architecture.
[0014] Furthermore, the multi-scale dilated causal convolution network includes:
[0015] The first dilated causal convolution layer uses a 1×1 convolution kernel and a dilation factor of 1 to capture local load features of 3 consecutive sampling points in the time dimension through causal constraints. The local load features include load change rate, fluctuation curvature, and extreme point distribution. The 1×1 convolution realizes full connection interaction between input feature channels, and the ReLU activation function enhances the non-linear expression ability. The convolution stride (stride = 1) and zero padding (padding = 1) maintain the feature time resolution, and the batch normalization (BN) layer stabilizes the feature distribution. The local features are passed to the subsequent network through skip connections and are verified to effectively capture the spike mutations, periodic fluctuations, and trend components of the load curve through gradient-weighted class activation mapping (Grad-CAM).
[0016] The second dilated causal convolution layer uses a 2×2 convolution kernel and a dilation factor of 2 to expand the receptive field through interval sampling and extract medium-scale generalized local features of the power load. The medium-scale generalized local features include load change trend, periodic fluctuations, and abnormal patterns. The convolution kernel weight matrix learns cross-channel feature combinations, the LeakyReLU activation function enhances non-linearity, and the layer normalization adapts to variable-length sequences. The feature dropout rate (dropout = 0.3) improves the generalization ability, and the output features are downsampled by max pooling and then passed to the subsequent network, effectively capturing the smooth transition, periodic consistency, and robust abnormal representation of the load curve.
[0017] Furthermore, the bidirectional long short-term memory network has 64 to 256 neurons, and the balance between prediction accuracy and memory occupancy is dynamically adjusted by the number of neurons: Based on the fluctuation characteristics of the load curve, a hierarchical neuron allocation strategy is adopted. 128 neurons are maintained in the time series feature extraction layer to capture long-term dependencies, while the number of neurons in the spatio-temporal coupling layer is dynamically adjusted to the range of 64 - 256. When the load is stable, the number of neurons is reduced to reduce memory occupancy, and when the load fluctuates violently, the number of neurons is increased to improve prediction accuracy. At the same time, the neuron pruning technology is combined to remove redundant connections, reducing the memory occupancy by 40%, and the prediction error (RMSE) is ensured to be stable below 0.05 through the adaptive learning rate mechanism, achieving the optimal balance between prediction performance and resource consumption.
[0018] A construction method of a short-term load forecasting system based on the encoder-decoder architecture as described above includes the following specific steps:
[0019] S1: Extract data from the historical database, generate a predictor matrix through preprocessing and analysis, and the predictor matrix contains features highly correlated with the load.
[0020] S2: Construct an encoder, input the predictor matrix into a multi-scale dilated causal convolutional network to extract local features.
[0021] S3: Construct a decoder, map the local features to predicted load values through a bidirectional long short-term memory network.
[0022] S4: Combine the encoder and the decoder into an encoder-decoder architecture, and output the short-term load forecasting result after training.
[0023] Furthermore, the step S1 specifically includes the following steps:
[0024] S11: Clean and perform min-max normalization on the historical data to generate a standardized data set.
[0025] S12: Determine the key feature set affecting the load through statistical analysis and correlation analysis of the standardized data set processed in step S11.
[0026] S13: Construct a predictor matrix based on the key feature set, including climate factors, time factors, and historical load parameters.
[0027] Furthermore, the step S2 specifically includes the following steps:
[0028] S21: Define causal convolution and dilated convolution to extract the immediate features of the time series.
[0029] S22: Configure two layers of dilated causal convolution for the encoder:
[0030] The first layer uses a 1×1 convolutional kernel and a dilation factor of 1 to extract specific local features;
[0031] The second layer uses a 2×2 convolutional kernel and a dilation factor of 2 to extract generalized local features.
[0032] Furthermore, step S3 specifically includes the following steps:
[0033] S31: Construct a decoder layer containing multiple bidirectional long short-term memory network units;
[0034] S32: Dynamically adjust the number of neurons to the range of 64 - 256 through the validation set to optimize memory occupancy.
[0035] Furthermore, step S4 specifically includes the following steps:
[0036] S41: Construct the overall architecture of the prediction system: Combine the encoder-decoder architecture, use the output of the encoder as the output of the decoder, and the output of the decoder is the predicted power load value;
[0037] S42: Training and tuning of the prediction system: Divide the normalized dataset into a training set, a test set, and a validation set. The training set is used for training the prediction model; the validation set is used for fine-tuning hyperparameters according to the validation results to eliminate overfitting and underfitting; the test set is used to evaluate the final prediction accuracy.
[0038] A two-stage prediction model prediction system innovatively proposed by the present invention combines a multi-scale dilated causal convolutional network (MSDCC) and a bidirectional long short-term memory network (BiLSTM), and has the following innovative breakthroughs compared with ordinary two-stage models:
[0039] (1) Lightweight structure design: Adopt a parallel architecture of 1×1 and 2×2 short convolutions to replace traditional large-size convolutional kernels, and significantly reduce the model complexity through a parameter sharing mechanism, showing stronger anti-overfitting ability in cross-regional scenarios.
[0040] (2) Dynamic multi-scale perception: The MSDCC module breaks through the fixed dilation rate limit, synchronously captures minute-level mutations and hour-level trend features by dynamically adjusting the convolutional receptive field, and realizes accurate inflection point recognition while ensuring temporal causality.
[0041] (3) Temporal logic decoupling: Innovatively separate the bidirectional learning and unidirectional prediction stages of BiLSTM: In the first stage, use the bidirectional mechanism to extract historical context features, and in the second stage, use unidirectional recursion to ensure the physical rationality of the prediction time series, avoiding the causal confusion defect of traditional bidirectional stacking.
[0042] (4) Native sequence preservation feature: The input sequence is subjected to in-situ feature recombination through 1×1 convolution, avoiding information distortion caused by conventional dimensional transformation, effectively retaining the details of the original load waveform, and improving the response sensitivity to sudden load events. Description of the Drawings
[0043] Figure 1 This is the autocorrelation diagram of the electric load of the present invention.
[0044] Figure 2 This is the box plot of the electric load of the present invention.
[0045] Figure 3 This is the quantile plot of the electric load of the present invention.
[0046] Figure 4 This is a schematic diagram showing the relationship between the predictor matrix and the prediction model of the present invention.
[0047] Figure 5 This is a schematic diagram of the causal convolution of the present invention.
[0048] Figure 6 This is a schematic diagram of the dilated convolution of the present invention.
[0049] Figure 7 This is a schematic diagram of the dilated causal convolution of the present invention.
[0050] Figure 8 This is the structure diagram of the bidirectional LSTM of the present invention.
[0051] Figure 9 This is the architecture diagram of the encoder-decoder of the present invention. Detailed Embodiments
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] In the field of short-term load forecasting (STLF), traditional methods and existing deep learning models face challenges, especially when dealing with complex and variable weather-sensitive load patterns. In current technologies, common machine learning algorithms and deep learning architectures, due to their multi-layer and large filter size designs, are prone to model parameter inflation, which in turn leads to problems such as overfitting and high computational complexity. This significantly reduces the prediction performance when facing highly diverse meteorological and time series fluctuations. In addition, such models often neglect the fine capture of key local features, especially those short-term load changes that have a direct impact on the prediction results, resulting in limited prediction accuracy. In view of this, the present invention proposes an innovative two-stage Encoder-Decoder (ED) architecture, which combines a multi-scale dilated causal convolutional network (MSDCC) and a bidirectional long short-term memory network (BiLSTM), effectively overcoming the limitations of existing technologies and achieving more accurate and robust short-term load forecasting.
[0054] The short-term load forecasting system of the present invention extracts and analyzes key physical factors affecting power load by combining actual operation data of the power system, such as historical load, weather data, etc. Different from traditional mathematical models, the encoder-decoder architecture used in this system is not just an algorithm framework, but is designed specifically for the business requirements in the field of power load forecasting.
[0055] The short-term load forecasting system based on the encoder-decoder architecture proposed by the present invention mainly includes a predictor matrix module, an encoder, and a decoder. The predictor matrix in the predictor matrix module contains feature parameters highly related to the load, which are used to provide inputs for the prediction model formed by the encoder-decoder architecture. The prediction model can learn the complex correlation relationships between these input variables and future power demands, and through training, extract useful information from the predictor matrix, and then predict future power loads. The input parameters of the encoder are the features highly related to the load in the predictor matrix, and the output parameters are the local features of the power load pattern; the decoder receives the local features of the power load pattern output by the encoder, converts the local features of the power load pattern into predicted power load values, and outputs the predicted power load values; the encoder is implemented by the constructed multi-scale dilated causal convolutional network, and the decoder is implemented by the constructed bidirectional long short-term memory network.
[0056] The multi-scale dilated causal convolutional network constructed in this embodiment has two dilated causal convolutional layers. The first convolutional layer uses a 1x1 convolutional kernel to extract the subtle changes in the impact of physical quantities such as weather on power load, ensuring that the model can handle the drastic fluctuations in load in the short term; the second convolutional layer uses a 2x2 convolutional kernel to capture the load change trends in a larger range, so as to better meet the load prediction requirements with a long time span in the power system. The encoder effectively reduces the computational complexity and improves the operation efficiency of the hardware by reducing the feature map size and model parameters, ensuring that the system can adapt to the real-time processing of large-scale data.
[0057] The bidirectional long short-term memory network captures the long-term and short-term dependencies of the power load time series through the cascading of multiple bidirectional LSTM units. The bidirectional LSTM enables the system to better grasp the historical change trends of the power load through the forward and backward propagation paths, while predicting the load fluctuations in future periods. This design is particularly suitable for power dispatching scenarios that require quick response, such as real-time load adjustment during peak loads. Through this processing method of time series, the load prediction values generated by the decoder can accurately reflect the actual operation requirements of the power system.
[0058] In addition, the bidirectional long short-term memory network uses an optimally selected number of neurons. If the number of neurons is too large, the problem of overfitting is likely to occur; if the number of neurons is insufficient, the problem of underfitting is likely to occur. To determine a better range of the number of neurons, multiple factors need to be considered, including but not limited to the size, complexity, feature dimension, task requirements of the dataset, and the limitations of computing resources. The preferred range of the number of neurons in the present invention is from 64 to 256, and the selection of this value is based on the results of multiple experiments and optimizations. This preference effectively overcomes the problems of underfitting or overfitting in the existing models. The number of neurons adopted by the bidirectional long short-term memory network in this embodiment is 128.
[0059] The design of the system fully considers the physical characteristics of the power load, and the input data used, such as temperature, humidity, historical load, etc., all have clear physical correlations with the load changes. For example, the increase in temperature will directly affect the increase in power consumption, and these changes are extracted as load pattern features through the convolutional operation of the encoder and are converted into specific load prediction values in the decoder. This correlation based on physical quantities ensures that the prediction results are consistent with the actual changes in the power load, avoiding the biases that may be caused by pure mathematical statistical models.
[0060] The construction method of the short-term load prediction system based on the encoder-decoder architecture proposed by the present invention includes the following specific steps:
[0061] S1: Extract historical load and climate data from the historical database, perform analysis after preprocessing the data to obtain features highly correlated with the load, and construct a predictor matrix based on the relevant features;
[0062] S2: Construct a multi-scale dilated causal convolutional network encoder to extract local features of the power load pattern, and use the output of the predictor matrix as the input of the encoder;
[0063] S3: Construct a bidirectional long short-term memory network decoder to convert the features extracted by the encoder into predicted load values;
[0064] S4: Combine the multi-scale dilated causal convolutional network and the bidirectional long short-term memory network to form an encoder-decoder architecture, that is, a short-term load forecasting model. After training, output the short-term load forecasting results to achieve accurate forecasting of short-term power loads.
[0065] Among them, step S1 specifically includes the following steps:
[0066] S11: Data cleaning and integration: Extract data closely related to power load from the historical database. The data includes but is not limited to historical load data of the power system, meteorological data, and other factors affecting power demand. Clean and integrate the extracted data, remove outliers and missing values to ensure the accuracy and integrity of the data. Then perform standardization processing on key variables (such as temperature, humidity, historical load) to generate a standardized dataset for subsequent analysis and training of the model.
[0067] S12: Exploratory data analysis: Analyze the collected data, including statistical analysis, correlation analysis, etc., to determine the key feature set that affects power load. The data analysis specifically includes:
[0068] 1) Autocorrelation analysis: Autocorrelation analysis is used to determine the correlation between power loads at different time points. By plotting the autocorrelation graph, as Figure 1 shown, it can be found that there is a strong correlation between the power load in the recent few hours and the same time point of the previous day.
[0069] 2) Box plot analysis: Box plot analysis is used to determine the differences in power loads during different time periods (such as weekdays and holidays). Decompose the time series data of the power load and analyze the daily cycle, weekly cycle, and seasonal change rules of the load. This information has a crucial impact on the accuracy of the forecasting model. By plotting the box plot, as Figure 2 shown, it can be found that the power load on holidays is lower, while the power load on weekdays is higher.
[0070] 3) Quantile analysis: Quantile analysis is used to determine the relationships between different factors (such as climate factors and historical loads) and electricity load. By plotting quantile plots, as Figure 3 shown, it can be found that there is a strong linear relationship between climate factors, historical loads and electricity load. Specifically, temperature has a significant positive correlation with load. Especially during high-temperature periods, the increased usage of equipment such as air conditioners leads to an increase in electricity demand; while humidity and wind speed affect heating and cooling demands to a certain extent, thus having an indirect impact on the load. These correlation relationships indicate that there are certain natural laws between electricity load and meteorological factors, which is also an important basis for the model to be used for prediction.
[0071] S13: Feature construction: According to the results of data analysis (the key feature set), select features highly correlated with the load to construct a predictor matrix, ensuring that these features can fully reflect the actual changes and physical laws of electricity load. As Figure 4 shown, the predictor matrix contains various types of input parameters including climate factors, time factors, and historical loads, etc. The predictor matrix is used to guide the construction and training process of the prediction model. In the short-term load forecasting (STLF) system, the role of the predictor matrix is to provide inputs for the deep learning architecture (the MSDCC-BiLSTM framework of the present invention) so that the model can learn the complex correlation relationships between these input variables and future electricity demands. Through training, the model learns to extract useful information from the predictor matrix and then predict future electricity loads.
[0072] Step S2 specifically includes the following steps:
[0073] S21: Define causal convolution and dilated convolution to extract the immediate features of the time series. Among them,
[0074] 1) As combined with Figure 5 shown, causal convolution is a convolution method for processing time series data, which ensures the causality of the convolution operation. In time series data, causality means that when predicting the data at the current time point, only past information can be used, and future information cannot be used. Causal convolution achieves this by introducing a causal mask during the convolution process, so that only past data is considered when calculating the convolution.
[0075] 2) As combined with Figure 6As shown, Dilated Convolution is a convolution method for processing time series data. It can expand the receptive field of the convolution kernel to capture longer time dependencies. In Dilated Convolution, this is achieved by introducing a dilation factor during the convolution process, causing some data points to be skipped during the calculation of the convolution, thereby expanding the receptive field.
[0076] S22: Combine Figure 7 As shown, the dilated causal convolution layer configuration: In the encoder, two dilated causal convolution layers are used. The first convolution layer uses a smaller convolution kernel (1x1) and a smaller dilation factor (1) to extract more specific local features. The second convolution layer uses a larger convolution kernel (2x2) and a larger dilation factor (2) to extract more general local features. In this way, both local details and overall trends in the power load pattern can be captured.
[0077] Step S3 specifically includes the following steps:
[0078] S31: Construct a decoder layer containing multiple BiLSTM units. The decoder layer captures the long-term dependencies and periodic features of the sequence through the combination of forward and backward propagation.
[0079] Combine Figure 8 As shown, the Long Short-Term Memory (LSTM) system is a Recurrent Neural Network (RNN) model for processing time series data. It can learn long-term dependencies. The Bidirectional Long Short-Term Memory Network (BiLSTM) is based on LSTM and adds a backward propagation path, enabling the system to learn context information from both the past and the future simultaneously. In BiLSTM, by passing the input data to both the forward and backward LSTM layers simultaneously, more comprehensive context information can be obtained, thereby improving the accuracy of prediction.
[0080] S32: Decoder configuration optimization: Dynamically adjust the number of neurons to the range of 64 - 256 through the validation set to optimize memory occupancy. Through multiple experiments, it is determined that using a BiLSTM network with 128 neurons can avoid overfitting and underfitting problems.
[0081] Step S4 specifically includes the following steps:
[0082] S41: Overall system architecture design: Combine the encoder-decoder architecture. The encoder consists of a Multi-Scale Dilated Causal Convolution Network (MSDCC) for extracting local features of the power load pattern. The decoder consists of a BiLSTM network for converting the extracted features into predicted power load values. By reducing the convolution kernel size and controlling the number of parameters, the storage and computing requirements of the hardware are effectively reduced. AsFigure 9 As shown
[0083] Input layer: The starting point of the model, which receives a multi - variable prediction matrix including climate data, historical load values, and time features.
[0084] MSDCC encoder part: First, there is a 1×1 dilated causal convolutional filter layer to extract local trends in the input data without adding too many parameters, keeping the model lightweight. Then, a 2×2 filter appears in the second layer. By setting the dilation rate, it controls the model complexity, expands the receptive field, and captures a wider range of sequence patterns.
[0085] Flattening layer: Converts the multi - dimensional output of the previous layer into a one - dimensional vector for easy processing by the subsequent BiLSTM layer.
[0086] BiLSTM decoder part: Uses 128 bidirectional long short - term memory units, which can capture short - term and long - term dependencies in the input data and enhance the model's prediction ability. The symbols Y1(F) and Y1(B) represent the forward and backward outputs at the first time step in the BiLSTM network. Here, Y1(F) represents the output of the forward LSTM unit at the first time step (time step), and Y1(B) represents the output of the backward LSTM unit at the same time step.
[0087] Output layer: The end point of the model, which generates the predicted power load value.
[0088] S42: Model training and tuning: The normalized dataset is divided into three subsets for training, testing, and validation. 80% of the data is used for training, while 20% of the data is used for testing and validation. After data normalization, the training data is converted into the form of a predictor matrix for model training. The MSDCC - BiLSTM model is validated using the validation dataset, and hyperparameters are fine - tuned to eliminate overfitting and underfitting problems. Hyperparameters are tuning parameters in machine learning algorithms, which are set before training the model, used to control the learning process, and have an important impact on the performance and behavior of the model. Common hyperparameters include: learning rate, batch size, regularization parameter, number and size of hidden layers, activation function type, etc.
[0089] In the short - term load forecasting task, model lightweight and prediction accuracy are key indicators to measure the practicality of the technical solution. The MSDCC - BiLSTM model of the present invention optimizes the encoder - decoder structure, significantly reduces the number of model parameters while maintaining high prediction accuracy, thus better meeting the requirements of real - time scheduling of the power system. The following are the specific comparison data for model lightweight and accuracy verification:
[0090] Table 1 Comparison of the number of model parameters
[0091]
[0092] Table 2 Comparison of Prediction Errors
[0093]
[0094] Feature extraction efficiencyIn the short-term load forecasting task, the feature extraction efficiency directly affects the real-time performance of the model and the consumption of computing resources. The MSDCC-BiLSTM model of the present invention significantly reduces the computational complexity and inference latency by optimizing the encoder structure, while enhancing the ability to capture key load features, thus better meeting the real-time scheduling requirements of the power system. The following are the specific comparison data of the feature extraction efficiency:
[0095] Table 3 Comparison of Computing Resources
[0096]
[0097] Table 4 Verification of Key Scenarios
[0098]
[0099] Data description:
[0100] (1) Test data: Load data (at 15-minute intervals) of a prefecture-level city in East China in 2022, including 8% extreme weather days
[0101] (2) Comparison benchmark: Performance of typical models in the IEEE PES 2020 STLF technical report
[0102] Experimental verification shows that while maintaining high accuracy, the number of model parameters of this solution is reduced by 52.5%, the computing efficiency is increased by 31.6%, and it shows good stability in complex scenarios.
[0103] In summary, the present invention has the following beneficial effects:
[0104] (1) Enhanced generalization ability: By using multi-level causal convolution and dilated convolution, the size of the feature map is effectively restricted, the number of model parameters is reduced, overfitting is avoided, and at the same time, the storage requirement and computational complexity are reduced, and the hardware processing efficiency is improved.
[0105] (2) High-precision prediction: Compared with the traditional CNN-LSTM model, it shows a significant improvement in prediction accuracy, can more accurately capture local load trends, and at the same time optimizes memory usage, suitable for real-time application scenarios.
[0106] (3) Efficient capture of non-linear features: Using 1×1 and 2×2 dilated causal convolution filters, focusing on extracting local non-linear patterns while keeping the time complexity low and reducing the data transmission volume.
[0107] (4) Optimized computational complexity: The structural design takes into account the feasibility of practical applications. By reducing model parameters and controlling computational requirements, the processing speed of the hardware and the response efficiency during system operation are improved.
[0108] Finally, it should be noted that the above specific implementation manners are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A short-term load forecasting system based on an encoder-decoder architecture, characterized in that: include: An encoder is used to receive characteristic data highly correlated with the load, wherein the characteristic data highly correlated with the load includes historical load data and meteorological data, and extract local features of the power load pattern through a multi-scale dilated causal convolutional network; A decoder, used for inputting the local features into a bidirectional long short-term memory network, converting them into predicted power load values and outputting them; in: The encoder adopts a deep separable convolution structure to achieve parameter compression, decomposes the standard convolution into two steps of deep convolution and point-by-point convolution, and reduces the number of parameters by decoupling the convolution kernel dimension. At the same time, a mixed expansion rate causal convolution module is designed to build a multi-scale feature extraction capability while maintaining the constraints of temporal causality, and cooperates with the grouped convolution strategy to further reduce the storage requirements of the weight matrix. The decoder implements hierarchical neuron pruning based on sensitivity analysis, calculates the LSTM unit state importance score through the second-order derivative metric, and removes redundant neurons in the hidden layer whose parameter update amplitude is less than the threshold η=0.01; adopts a dynamic neuron allocation mechanism, maintains 256 hidden units in the temporal feature extraction layer, and reduces it to 128 units in the spatiotemporal coupling layer, and realizes memory occupancy quantization compression with 8-bit integer quantization, while maintaining the verification set RMSE≤0.048 through a selective gradient update strategy.
2. The short-term load forecasting system based on encoder-decoder architecture according to claim 1, characterized in that: It also includes a predictor matrix module, wherein the predictor matrix in the predictor matrix module contains characteristic data highly correlated with the load, wherein the characteristic data highly correlated with the load includes power load data and meteorological data, wherein the climate data includes temperature, humidity, and wind speed; the climate data comes from a meteorological monitoring station, and the historical load data is taken from a SCADA system, and the sampling frequency is 15 minutes; The feature data is subjected to causal convolution to extract multi-scale meteorological features and load frequency domain features, and is fused through a gated attention mechanism to provide input for the prediction model formed by the encoder-decoder architecture.
3. The short-term load forecasting system based on encoder-decoder architecture according to claim 1, characterized in that: The multi-scale dilated causal convolutional network includes: The first dilated causal convolution layer uses a 1×1 convolution kernel and a dilation factor of 1 to capture the local load characteristics of three consecutive sampling points in the time dimension through causal constraints. The local load characteristics include load change rate, fluctuation curvature and extreme point distribution; 1×1 convolution realizes full connection interaction between input feature channels, and cooperates with ReLU activation function to enhance nonlinear expression ability; convolution step size and zero padding are set to maintain feature time resolution, and batch normalization layer stabilizes feature distribution; local features are transmitted to subsequent networks through jump connections, and gradient weighted class activation mapping is used to verify that it effectively captures spike mutations, periodic fluctuations and trend components of the load curve; The second expanded causal convolution layer adopts a 2×2 convolution kernel and a dilation factor of 2. It expands the receptive field through interval sampling and extracts the medium-scale generalized local features of the power load. The medium-scale generalized local features include load change trends, periodic fluctuations and abnormal patterns. The convolution kernel weight matrix learns cross-channel feature combinations, the LeakyReLU activation function enhances nonlinearity, and layer normalization adapts to variable-length sequences. The feature discard rate is set to improve the generalization ability, and the output features are passed to the subsequent network after maximum pooling downsampling, which effectively captures the smooth transition, periodic consistency and robust abnormal representation of the load curve.
4. The short-term load forecasting system based on encoder-decoder architecture according to claim 1, characterized in that: The bidirectional long short-term memory network has 64 to 256 neurons, and the number of neurons is dynamically adjusted to balance the prediction accuracy and memory usage: based on the fluctuation characteristics of the load curve, a hierarchical neuron allocation strategy is adopted, 128 neurons are maintained in the time series feature extraction layer to capture long-term dependencies, and the number of neurons is dynamically adjusted to the range of 64-256 in the spatiotemporal coupling layer. When the load is stable, the number of neurons is reduced to reduce memory usage, and when the load fluctuates violently, the number of neurons is increased to improve the prediction accuracy; at the same time, the neuron pruning technology is combined to remove redundant connections, and the adaptive learning rate mechanism is used to ensure that the prediction error RMSE is stable below 0.05, so as to achieve the optimal balance between prediction performance and resource consumption.
5. A method for constructing a short-term load forecasting system based on a codec architecture as claimed in any one of claims 1 to 4, characterized in that: The specific steps include: S1: extracting historical load and climate data from a historical database, generating a predictor matrix through preprocessing and analysis, wherein the predictor matrix contains features highly correlated with the load; S2: construct an encoder, input the predictor matrix into a multi-scale dilated causal convolutional network, and extract local features; S3: Construct a decoder to map local features to prediction load values through a bidirectional long short-term memory network; S4: Combine the encoder and decoder into an encoder-decoder architecture and output the short-term load forecasting results after training.
6. The method for constructing a short-term load forecasting system based on a codec architecture according to claim 5, characterized in that: The step S1 specifically includes the following steps: S11: Clean and standardize historical data to generate standardized data sets; S12: Determine a key feature set that affects the load by statistical analysis and correlation analysis of the standardized data set processed in step S11; S13: Construct a predictor matrix including climate factors, time factors and historical load parameters based on the key feature set.
7. The method for constructing a short-term load forecasting system based on a codec architecture according to claim 5, characterized in that: The step S2 specifically includes the following steps: S21: Define causal convolution and dilated convolution to extract instant features of time series; S22: Configure the encoder with two layers of dilated causal convolutions: The first layer uses a 1×1 convolution kernel and a dilation factor of 1 to extract specific local features; The second layer uses a 2×2 convolution kernel and a dilation factor of 2 to extract generalized local features.
8. The method for constructing a short-term load forecasting system based on a codec architecture according to claim 5, characterized in that: The step S3 specifically comprises the following steps: S31: Construct a decoder layer containing multiple bidirectional long short-term memory network units; S32: Dynamically adjust the number of neurons to the range of 64-256 through the validation set to optimize memory usage.
9. The method for constructing a short-term load forecasting system based on a codec architecture according to claim 5, characterized in that: The step S4 specifically comprises the following steps: S41: constructing the overall architecture of the prediction system: combining the encoder-decoder architecture, using the output of the encoder as the output of the decoder, and the output of the decoder is the predicted power load value; S42: Training and tuning of the prediction system: The normalized data set is divided into training set, test set and validation set. The training set is used for training the prediction model; the validation set is used to fine-tune the hyperparameters according to the validation results to eliminate underfitting and underfitting; the test set is used to evaluate the final prediction accuracy.
Citation Information
Cited By
Power consumption prediction method and system based on optimal scale convolution and graph memory enhancement
CN120675076A
A method and system for predicting electricity consumption based on optimal scale convolution and graph memory enhancement
CN120675076B
Explanatable load prediction method based on neural additive model and gradient boosting tree
CN122332932A