A Dynamic Power Demand Prediction Method Based on a Multi-Scale Hybrid Architecture

By combining the hybrid architecture of TCN and LSTM and phased training strategies, nonlinear coupled modeling and multi-scale feature extraction problems in power demand prediction are solved, and prediction accuracy and interpretability are improved, which is suitable for intelligent prediction of power systems.

CN119990477BActive Publication Date: 2025-07-22WUXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510468974.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the existing power demand prediction technology, there are problems such as insufficient nonlinear dynamic coupling modeling of external variables and loads, insufficient extraction of multi-scale timing feature and insufficient interpretability of prediction results, especially in extreme weather and holiday scenarios.

Method used

A hybrid architecture based on time domain convolution network (TCN) and long and short-term memory network (LSTM) is adopted, combining temperature sensitivity segmented nonlinear coding and holiday multi-dimensional dynamic attenuation model, and a dynamic prediction model is built through multi-source data feature fusion, time convolution network TCN module, bidirectional long and short-term memory network BLSTM module, spatiotemporal attention mechanism and feature fusion layer, and a phased dynamic training strategy is used for model training.

Benefits of technology

It significantly improves the prediction accuracy in extreme temperature scenarios, reduces high-temperature daily load and holiday prediction errors, improves the interpretability of prediction results, supports grid scheduling decision optimization, and reduces computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990477B_ABST
    Figure CN119990477B_ABST
Patent Text Reader

Abstract

The present invention provides a dynamic power demand prediction method based on a multi-scale hybrid architecture, comprising the following steps: Step 1, obtaining multi-source data and performing preprocessing; Step 2, constructing dynamic feature engineering; Step 3, constructing a hybrid prediction model; Step 4, implementing a phased dynamic training strategy; Step 5, validating the model and deploying and applying the model. By introducing temperature sensitivity piecewise non-linear coding and a multi-dimensional dynamic decay model for holidays, the present invention significantly improves the prediction accuracy in extreme temperature scenarios, reduces the prediction error of high-temperature day loads and the error during the Spring Festival holiday. In addition, through lightweight design and phased dynamic strategies to gradually expand the input length in phases, the training convergence speed is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent prediction of power systems, and particularly relates to a dynamic power demand prediction method based on a multi-scale hybrid architecture. Background Technique

[0002] Power demand prediction is the core link of power system optimal dispatching, which directly affects the formulation of power generation plans and the safe operation of power grids. Traditional statistical models (such as ARIMA, Prophet) rely on manual setting of time series parameters and are difficult to handle the non-linear coupling relationships of multi-source heterogeneous data (such as temperature, holidays). With the development of deep learning technology, recurrent neural networks such as Long Short-Term Memory (LSTM) have gradually become the mainstream solutions, but they still face significant bottlenecks. For example, Shi Min et al. used LSTM for short-term prediction of photovoltaic power generation, and the prediction results showed high adaptability and optimization potential, but there were still problems such as insufficient feature coupling, multi-scale conflicts, and relatively weak local feature extraction ability, and it was easy to ignore the details in short-term dependence relationships. Different from traditional recurrent neural networks, the Temporal Convolutional Network (TCN) is a brand-new deep learning structure. By integrating multiple mechanisms such as dilated convolution, causal convolution, and residual modules, it can effectively extract the correlation between data and better handle time series prediction problems. And Wang et al. and Li Feifei's research shows that the dilated convolution in the TCN architecture expands the receptive field, but due to the lack of construction of a dynamic association model, the prediction variance will increase significantly in the scenario of policy mutation, and the ability to model daily / weekly cycles is weak, and the prediction error of holiday load attenuation is large.

[0003] In addition, there are inherent contradictions in the existing methods in terms of feature engineering, model architecture, and training strategies. The non-linear relationship between temperature and load needs to be modeled in segments (such as thresholds of 10°C and 26°C), but traditional linear normalization leads to blurred features in key intervals; it is difficult for LSTM and TCN to be used alone to balance local fluctuations and long-term trends; complex models (such as Transformer) rely on high-end hardware and cannot be adapted to ordinary equipment of power grid enterprises. The core difficulty in solving these problems lies in: how to design a lightweight hybrid architecture to balance the efficiency of multi-scale feature extraction, while constructing a dynamic coupling model to quantify the influence of external variables and ensure training stability under limited computing power. Summary of the Invention

[0004] Objective of the Invention: To effectively improve the prediction performance and accuracy of traditional power quality prediction methods, the present invention provides a dynamic power demand prediction method based on a multi-scale hybrid architecture, specifically a power quality prediction method based on a Temporal Convolutional Network (TCN) and a Long Short-Term Memory Network (LSTM). The TCN effectively processes time series data through dilated convolution and causal convolution, and can capture local features at different time scales, while the LSTM is good at dealing with long-term dependencies in time series. The combination of the two can retain short-term features in the data and capture long-term trends when dealing with complex power quality data. The present invention aims to solve the following problems existing in the existing power demand prediction technology:

[0005] (1) Insufficient non-linear dynamic coupling modeling of external variables and load: The correlation mechanism between external factors such as temperature and holidays and power load has not been accurately quantified. In particular, the temperature threshold effect (such as the sudden change response of air conditioning load at 26°C) lacks effective characterization, resulting in a significant increase in prediction error in extreme weather scenarios;

[0006] (2) Insufficient extraction of multi-scale time series features: A single model architecture is difficult to simultaneously capture hourly fluctuation features (such as sudden changes in morning peak load) and weekly / monthly long-term trends (such as seasonal electricity consumption growth). Feature conflicts at different time scales lead to a decline in prediction stability;

[0007] (3) Insufficient interpretability of prediction results: Existing deep learning models lack a feature contribution quantification mechanism and cannot analyze the specific impact of key factors such as temperature and holidays on prediction results, making it difficult to support the optimization of power grid dispatching decisions.

[0008] The method of the present invention includes the following steps:

[0009] Step 1, obtain multi-source data and perform preprocessing: Obtain historical power load data, perform outlier detection and dynamic correction, fill in missing values, and standardize and normalize the data; construct time series samples, and divide the historical power load data into a training set, a validation set, and a test set according to a certain proportion;

[0010] Step 2, construct dynamic feature engineering: Perform piecewise non-linear encoding of temperature sensitivity, encode the temperature data through a piecewise function to form new data, and normalize it to form temperature feature data; perform multi-dimensional dynamic decay encoding of holidays to obtain holiday encoding based on the impact of holidays;

[0011] Step 3, construct a hybrid prediction model; the hybrid prediction model includes a multi-source data feature fusion module, a Temporal Convolutional Network (TCN) module, a Bidirectional Long Short-Term Memory (BLSTM) module, a spatio-temporal attention mechanism, a feature fusion layer, and an output layer; the multi-source data feature fusion module is used for multi-source data feature fusion, the TCN module is used for dilated convolution, the BLSTM module is used to output the concatenated hidden states, the spatio-temporal attention mechanism is used for temporal attention and feature attention calculations, the feature fusion layer is used to complete feature concatenation, and the output layer is used to output the power load prediction result;

[0012] Step 4, execute a phased dynamic training strategy: in each phase, train the hybrid model based on the training set, monitor the validation loss, and after each round of training, calculate the Mean Absolute Error (MAE) of the validation set, and select the weights with the smallest MAE for test set evaluation;

[0013] Step 5, verify the model and deploy the model: based on the weights obtained in Step 4, calculate the MAE and Weighted Mean Absolute Percentage Error (WMAPE) based on the test set, complete the verification of the model, and deploy and apply the hybrid prediction model.

[0014] Step 1 includes:

[0015] Step 1.1, obtain historical power load data with a time resolution of 1 hour, and the data fields include load values, area codes, and timestamps;

[0016] Obtain hourly meteorological temperature data;

[0017] Obtain and integrate legal holidays and regional holidays to form a structured holiday time label table;

[0018] Step 1.2, outlier detection and dynamic correction;

[0019] Adopt a sliding window quantile detection method to dynamically correct abnormal power load values, and dynamically calculate the 0.05 quantile and 0.95 quantile at each time point. If the daily power load value exceeds the , range, then is replaced with the median within the window , and the calculation formula is:

[0020] ,

[0021] Among them, represents the corrected power load value, which is used to replace the outlier;

[0022] Because the closer the data time is, the higher the weight should be, so the window mean is calculated by the moving average algorithm , and the weights of each point in the window are distributed according to exponential decay. The weight of the i-th (where i ranges from 0 to 167) point in the window is:

[0023] ,

[0024] Among them, is an intermediate parameter, generally set to , and taking λ as 0.01 is an empirical value, which can effectively balance the importance of recent data and the retention of historical trends. In hourly load forecasting, this value can highlight the latest data without overly ignoring historical information, and it is commonly used and has stable effects in practical applications. e is the natural constant;

[0025] The window mean The calculation formula is:

[0026] ;

[0027] Step 1.3, Missing value filling strategy;

[0028] For missing data, the cubic spline interpolation algorithm with natural boundary conditions is adopted. For each interval , a cubic polynomial is defined:

[0029] ,

[0030] Among them, ; , , , are the coefficients of the cubic polynomial, and is the timestamp of the known data point. By the continuity and boundary conditions, the spline coefficients , , , are solved, and finally the cubic polynomial on each interval is obtained;

[0031] Filling missing values in the spatial dimension using the mean of physically adjacent regions: Regions are divided based on the physical connection relationships in the power grid topology (such as the power supply scope of substations and line connection relationships), rather than geographically adjacent regions, which solves the filling deviation caused by uneven power flow distribution in the power grid. If the data of a certain moment in region A is missing, the average power load of physically adjacent regions B, C, and D is taken as the filling value:

[0032] ;

[0033] where represents the filling result of the missing value in region A, represents the average power load of region B, represents the average power load of region C, represents the average power load of region D;

[0034] Step 1.4, Data standardization and normalization:

[0035] To eliminate the influence of outliers and make the data distribution more stable and more suitable for neural network training. It is necessary to apply RobustScaler to the processed power load data for standardization to obtain the standard data , and the calculation formula is:

[0036] ,

[0037] where median( ) is the median of the set of processed power load data , and IQR( ) is the interquartile range, and the calculation formula is:

[0038] IQR( ) = - ,

[0039] where represents the 0.75 quantile of each time point, represents the 0.25 quantile of each time point;

[0040] Step 1.5, Construct time series samples;

[0041] The historical power load data is cut with a sample length of 720 consecutive hours and a sliding step of 24 hours to form the input sequence , where N is the number of samples and R represents the real number space;

[0042] When cutting the samples, the time sliding window method is used to ensure the temporal continuity of the sequence;

[0043] Finally, the historical power load data is divided into a training set, a validation set, and a test set according to a ratio. The training set: 70% (used for model parameter update and feature learning), the validation set: 20% (used for hyperparameter tuning and early stopping decision), and the test set: 10% (used for final performance evaluation).

[0044] Step 2 includes:

[0045] Step 2.1, piecewise non-linear encoding of temperature sensitivity;

[0046] Different from traditional linear or single non-linear encoding (such as exponential function), the present invention defines 10°C and 26°C as key inflection points and constructs a double-critical temperature piecewise function. Among them, according to the low-temperature section (Temp < 10°C): simulating the gentle growth of heating load; the medium-temperature section (10°C to 26°C): reflecting the square relationship of the basic load; the high-temperature section (Temp > 26°C): capturing the sudden change response of air-conditioning load (linear + exponential term). Therefore, a piecewise non-linear function of temperature sensitivity is constructed :

[0047] ,

[0048] The temperature data Temp is encoded through the piecewise function to form data , and then further normalized to the interval [0, 1] to form temperature feature data , and the formula is:

[0049] ;

[0050] Step 2.2, multi-dimensional dynamic decay encoding for holidays.

[0051] Step 2.2 includes: quantifying the immediate impact of holidays on load (such as the first day of the Spring Festival), the decay effect (such as returning to work after the festival), and historical patterns (such as year-on-year changes). The holiday impact includes type weight, time decay factor and historical comparison coefficient γ(t);

[0052] For the type weight, national holidays are assigned a value of 1.0, regional holidays are assigned a value of 0.7, and weekdays are 0;

[0053] The time decay factor The calculation formula is:

[0054] ,

[0055] where d is the number of days from the holiday;

[0056] The calculation formula of the historical comparison coefficient γ(t) is:

[0057] ,

[0058] where represents the average power load in the same period of the past three years;

[0059] The holiday code is calculated using the following formula :

[0060] ,

[0061] where Y1 represents the type weight.

[0062] Step 3 specifically includes the following steps:

[0063] Step 3.1, the multi-source data feature fusion module performs the following operations:

[0064] Existing methods mostly use simple splicing or static weight fusion. In the present invention, the standardized power load data , temperature code and holiday code are concatenated along the feature axis to form an input tensor . The feature fusion layer (including two fully connected layers, where a fully connected layer refers to a hierarchical structure that connects each neuron to all neurons in the previous layer) dynamically adjusts the weights of each feature through a fully connected network (a network structure stacked by multiple fully connected layers) to achieve scene adaptive coupling of multi-source signals. The formula is:

[0065] ,

[0066] where is the hidden representation after feature fusion, containing the dynamically adjusted multi-source feature information, is the learnable weight matrix, D is the dimension of the hidden layer, is the bias term of the fully connected layer, and ReLU is the activation function.

[0067] Step 3.2, establish a temporal convolutional network TCN module;

[0068] For the multi-scale spike characteristics of power load (such as instantaneous overload), progressive dilation rate design is adopted to gradually expand the receptive field and avoid the loss of local features caused by a single dilation rate. The time convolutional network TCN module contains 3 layers of dilated convolutions. Each layer of dilated convolution structure is the same, and each includes a dilated convolution layer, a ReLU activation function, and a residual connection. The number of convolutional kernels in each layer of dilated convolution layer is 32, the convolutional kernel size is 3, the stride is 1, and the padding method is causal convolution. The dilation rates of the 3 layers of dilated convolution layers are 2, 4, and 8 in sequence. The output of each layer of dilated convolution is added element-wise to the input through the residual connection, gradually expanding the receptive field to 168 hours. The residual connection formula is:

[0069] ,

[0070] The input sequence of the dilated convolution is a tensor obtained by concatenating the standardized load data, temperature encoding, and holiday encoding , with a dimension of [B, 720, 3], where B is the batch size; is the input of the th layer of dilated convolution (the input of the first layer is the original data, and the input of subsequent layers is the output of the previous layer); The th layer of dilated convolution output, with the final output dimension of [B, 720, 32]; Conv1D is a one-dimensional dilated convolution operation;

[0071] Step 3.3, establish a bidirectional long short-term memory network BLSTM module;

[0072] The bidirectional long short-term memory network BLSTM module includes a forward long short-term memory module and a backward long short-term memory module. 64 hidden units are set in the bidirectional long short-term memory network BLSTM module. The output of the time convolutional network TCN module is intercepted for the last 72 hours as the input of the forward long short-term memory module. The forward long short-term memory module extracts the forward trend, and the backward long short-term memory module captures the reverse dependence. The hidden state update formula is:

[0073] ,

[0074] ,

[0075] ,

[0076] where t is the time step, h t is the long short-term memory hidden state, is the forward long short-term memory hidden state, is the backward long short-term memory hidden state, and LSTM is the model function;

[0077] The hidden state with a dimension of [B, 72, 128] is output after splicing by the bidirectional long short-term memory network BLSTM module;

[0078] Step 3.4, establish a spatio-temporal attention mechanism, including temporal attention and feature attention;

[0079] The temporal attention calculates the weights of each time step through the Softmax function (the higher the weight, the more important the moment is for prediction):

[0080] ,

[0081] where w t represents the weight at time step t, is the output of the temporal convolutional network TCN module at time step t, v, , are learnable parameters, ;

[0082] is the key matrix of the attention mechanism, which maps the hidden state of the BiLSTM to the attention space, is the value matrix, which maps the TCN features to the attention space; TP represents transpose; Softmax is a normalized exponential function, whose role is to convert the input vector into a probability distribution, indicating the importance of each time step; tanh is a hyperbolic tangent function, whose role is to perform a non-linear transformation on the input features to enhance the non-linear expression ability;

[0083] The feature attention uses a squeeze-and-excitation (SE) module to perform channel weighting on the temperature and holiday encoding. The squeeze-and-excitation SE module is mainly used to enhance the interdependence between channels in image classification tasks. Transferring the squeeze-and-excitation SE module from the image domain to power demand prediction and performing channel weighting on time series features (temperature, holiday encoding) is an improved application of existing technologies. Its advantages include enhancing the sensitivity of the model to key features, suppressing the influence of noise or irrelevant features, and enhancing the interpretability of the model. The squeeze-and-excitation SE module performs the following steps:

[0084] Step 3.4.1, adopt global average pooling to compress the features along the time dimension to obtain the pooled global feature vector s:

[0085] ,

[0086] where, The input feature at time step t, with a dimension of [B, T, C] (B is the batch size, T is the time step length, and C is the number of channels). In the present invention: C = 3 (electric load, temperature encoding, holiday encoding), T: the time step length of the input sequence (e.g., T = 720 hours is set);

[0087] Step 3.4.2: Obtain the channel weight vector through a two-layer fully connected network :

[0088] ,

[0089] where, is the weight matrix of the first fully connected layer, is the bias term of the first fully connected layer, is the weight matrix of the second fully connected layer, is the bias term of the second fully connected layer, ReLU is the activation function, and σ is the Sigmoid function;

[0090] Step 3.4.3: Perform channel weighting to obtain the weighted output feature :

[0091] ,

[0092] where, , x is the original input feature (load, temperature encoding, holiday encoding, here x is different from before pooling, and before pooling is a single time step feature, while here x is the complete time series), represents element-wise multiplication across channels;

[0093] Step 3.5: Concatenate the feature fusion layer and the output layer.

[0094] Step 3.5 includes: Implementing the feature fusion layer to concatenate the local feature output by the time convolutional network TCN module and the global feature output by the bidirectional long short-term memory network BLSTM module, and inputting it into a two-layer fully connected network containing 64 neurons and the Swish activation function:

[0095] ,

[0096] where, is the output layer weight for mapping the 64-dimensional hidden layer to the predicted value, is the weight matrix of the fully connected network that maps the concatenated features to the 64-dimensional hidden layer, is the bias term of the hidden layer in the feature fusion layer, is the bias term of the output layer, The future 6-hour power load prediction result output by the output layer.

[0097] Existing deep learning models often use a fixed input length and a unified activation function, making it difficult to distinguish the coupled interference of hourly mutations, daily cycles, and long-term trends in power demand. The present invention realizes feature decoupling through a phased strategy. Step 4 includes:

[0098] Step 4.1, in each stage, the hybrid model is trained for 10 epochs based on the training set. Each epoch includes a complete traversal of all training samples, and the parameters are dynamically adjusted in stages;

[0099] The first stage is from 1 to 10 epochs: the model is trained completely 10 times (i.e., Y1 epochs) using the training set, and all training data is traversed each time.

[0100] Input length: 6 hours;

[0101] Activation function: ReLU function (Rectified Linear Unit), which is defined as: when the input value is greater than or equal to 0, the output value is equal to the input value; when the input value is less than 0, the output is 0. By suppressing negative inputs, reducing the number of neuron activations, and enhancing the sparse expression ability of the model, it is suitable for quickly extracting local mutation patterns (such as morning rush hours);

[0102] Learning rate: initial value 1×10 -3 , using a linear warm-up strategy, gradually increasing from 1×10 -4 to 1×10 -3 ;

[0103] Objective: Learn hourly load fluctuation characteristics and suppress long-term noise interference;

[0104] The second stage is from 11 to 20 epochs: continue training for 10 epochs, but adjust the input length and hyperparameters.

[0105] Input length: extended to 24 hours;

[0106] Activation function: switched to LeakyReLU function (Leaky Rectified Linear Unit), which is defined as: when the input value is non-negative, the original value is output, and when it is negative, the product of the input value and the leakage coefficient (default 0.1) is output. It allows weak activation of negative inputs, alleviates the vanishing gradient, and avoids the problem of neuron inactivation, making it suitable for fusing medium-complexity daily cycle features (such as temperature-sensitive intervals, holiday attenuation):

[0107] ,

[0108] Learning rate: decreased to 5×10 -4 , and the cosine decay strategy is adopted, decaying to 50% of the initial value every 10 rounds;

[0109] Objective: Optimize daily feature fusion;

[0110] The third stage is from round 21 to 30: Retrain for another 10 rounds (epochs) to further optimize long-term trend prediction.

[0111] Input length: Set to 72 hours;

[0112] Activation function: The Swish function (self-gated activation function) is adopted. It has continuous differentiable non-linearity near x = 0, avoiding the hard truncation effect of ReLU, and dynamically adjusts the activation threshold through the β parameter to adapt to different data distributions. The smooth non-linear characteristics of the Swish function can effectively capture more complex long-term dependencies (such as seasonal changes, holiday decay effects), and its self-adaptability helps to stabilize the distribution of attention weights:

[0113] ,

[0114] where β is a learnable parameter (default β = 1);

[0115] Learning rate: adjusted to 1×10 -4 , and the exponential decay strategy is adopted to calculate the dynamically adjusted learning rate , which gradually decreases as the number of training rounds increases:

[0116] ,

[0117] where, is the initial learning rate (1×10 in the third stage -4 ), decay_rate is the decay rate (controlling the decay speed of the learning rate, the smaller the value, the faster the decay), is the number of training rounds (from round 21 to round 30, the relative round count starting from the third stage, round 21 is ), and decay_steps is the decay step size (if decay_steps = 5, then decay every 5 rounds);

[0118] Objective: Strengthen long-term trend prediction and stabilize the distribution of attention weights;

[0119] Step 4.2, Implement the early stopping mechanism and save the model: Monitor the validation loss. After each round of training, calculate the mean absolute error (MAE) of the validation set. If the validation loss does not decrease for 5 consecutive rounds, terminate the training and roll back to the historical best weights. The model weights are saved once per round, and finally, the weights with the minimum mean absolute error MAE are selected for test set evaluation. The test set is isolated throughout and only used in the final evaluation to ensure unbiased performance metrics.

[0120] Step 5 includes:

[0121] Step 5.1, Load the best weights saved in Step 4.2 and calculate the following metrics based on the test set:

[0122] Mean absolute error MAE:

[0123] ,

[0124] Weighted mean absolute percentage error WMAPE:

[0125] ,

[0126] where, is the actual load value (test set data); is the model prediction value; N is the number of test set samples.

[0127] Step 5.2, Conduct interpretability analysis:

[0128] Traditional interpretability analysis is mostly limited to feature importance ranking, while the present invention innovatively realizes the organic combination of attention weights (attention weight visualization) and spatio-temporal contribution degree (SHAP value analysis). In this method, the validation set is used for interpretability analysis to analyze the contribution degrees of features such as temperature and holidays.

[0129] Attention weight visualization: Through forms such as heatmap, display the weights assigned by the attention mechanism in the neural network, intuitively reflecting the attention degree of the model to different time steps or features during prediction.

[0130] SHAP is a model interpretability method based on game theory, which is widely used to interpret machine learning models (including deep learning). DeepExplainer is an interpretation tool specifically designed for deep learning models in the SHAP library. Using the SHAP (SHapley Additive exPlanations) method, through the DeepExplainer tool, calculate the contribution degrees of each input feature in the model to the prediction result, and quantify the impact of each feature on the final prediction.

[0131] Step 5.3, Lightweight deployment and application;

[0132] Traditional model deployment often relies on high-performance computing resources. However, in the present invention, by retaining the hourly feature extraction layer (the first two layers of TCN) and only retaining the long-term trend layer (BiLSTM) in stage 3, the computational redundancy during inference is reduced.

[0133] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method described above.

[0134] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction runs on a computer, the steps of the method described above are executed.

[0135] The present invention has the following technical solution advantages:

[0136] Multi-scale feature fusion: Dilated convolutions in TCN capture hourly spikes, bidirectional LSTM models weekly / monthly trends, and spatio-temporal attention focuses on key time periods and features;

[0137] Dynamic interpretability: SHAP values and attention weights analyze the prediction logic to support the optimization of power grid dispatching decisions;

[0138] Progressive training: Compared with the training strategy with a fixed input length, the phased dynamic strategy improves the convergence speed while reducing the prediction error in extreme weather scenarios.

[0139] Beneficial effects: By introducing piecewise non-linear coding of temperature sensitivity and a multi-dimensional dynamic decay model for holidays, the present invention significantly improves the prediction accuracy in extreme temperature scenarios, reduces the prediction error of daily load on high-temperature days and the error during the Spring Festival holiday. In addition, through lightweight design and the phased dynamic strategy to gradually expand the input length, the training convergence speed is significantly improved. In terms of interpretability, the time attention weight focuses on the morning and evening peak hours, and the SHAP value quantification shows that the temperature contribution degree on high-temperature days is relatively high, providing a decision-making basis for power grid dispatching. Economically, combined with the time-of-use electricity price strategy, the daily dispatching cost of a single region, the prediction variance in extreme weather scenarios, and the fault loss decrease significantly, having significant social application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0140] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0141] Figure 1 The flowchart of the power demand dynamic prediction method based on a multi-scale hybrid architecture provided for the embodiment shows the complete process from multi-source data collection, preprocessing, dynamic feature engineering, model training to prediction result output.

[0142] Figure 2 The structural diagram of the hybrid prediction model provided for the embodiment includes the dilated convolution layer (dilation rates 2 / 4 / 8) of the TCN module, the temporal dependence extraction of the bidirectional LSTM module, the dynamic weight allocation of the spatio-temporal attention mechanism, and the multi-scale information integration of the feature fusion layer.

[0143] Figure 3 The model performance comparison diagram provided for the embodiment uses MAE and WMAPE as indicators to compare the prediction accuracies of ARIMA, BiLSTM, CNN-LSTM, and the present invention, and marks the error reduction amplitude in extreme weather scenarios.

[0144] Figure 4 To mark the focusing effect (proportion > 60%) of the temporal attention weight during the morning and evening rush hours.

[0145] Figure 5 For SHAP feature contribution analysis, mainly showing the SHAP value distributions of the load / temperature / holidays in the previous 24 hours. Detailed implementation manners

[0146] As Figure 1 shown, the embodiment of the present invention provides a dynamic power demand prediction method based on a multi-scale hybrid architecture, including:

[0147] Step 1, obtaining multi-source data and performing preprocessing;

[0148] Step 1.1, data source and collection specification;

[0149] The data collection module of this embodiment obtains historical load data from the provincial company of the State Grid, with a time resolution of 1 hour. The data fields include load value (unit: megawatt, MW), area code (identifying the power grid partition), and timestamp (accurate to the hour). The area code adopts a three-level structure, such as "East China - Jiangsu - Nanjing", to support subsequent spatial correlation analysis. The meteorological data obtains the hourly temperature (°C) data from the China Meteorological Administration. The holiday data integrates legal holidays and regional holidays announced by the provincial government to form a structured time tag table. All data is synchronously transferred to the distributed time series database (InfluxDB) through the API interface in real time to ensure the efficient reading and writing of data and the time window query ability.

[0150] Step 1.2, outlier detection and dynamic correction;

[0151] The sliding window quantile detection method is used to dynamically correct the abnormal load values. The window length is set to 168 hours (7 days), and the 0.05 quantile ( ) and 0.95 quantile ( ) of each time point are calculated dynamically. If the current power load value Beyond the range, it is replaced by the median within the window , and the specific calculation formula is:

[0152] ,

[0153] Window mean Calculated by the moving average algorithm, the weights of each point within the window are distributed with exponential decay, and the weights of recent data are higher. For example, the weight of the i-th point within the window is:

[0154] ,

[0155] This method can eliminate abnormal fluctuations caused by extreme weather (such as typhoons) or equipment failures, while retaining the continuity of the normal load trend.

[0156] Step 1.3, Missing value filling strategy;

[0157] For missing data, the cubic spline interpolation algorithm with natural boundary conditions is adopted. The basic idea of its interpolation formula is that every two adjacent data points are represented by a cubic polynomial, so that these polynomials are continuous and smooth at adjacent nodes (including the continuity of the first derivative and the second derivative). For each interval , define a cubic polynomial:

[0158] ,

[0159] where ;

[0160] By continuity and boundary conditions, the spline coefficients are solved, and finally the cubic polynomials on each interval are obtained, which not only ensures the continuity of the function values, but also ensures the smoothness of the data curve.

[0161] For missing values in the spatial dimension, the mean value of adjacent regions is used for filling. The region division is based on the physical connection relationship in the power grid topology, such as the coverage area of the same substation. If the data at a certain moment in region A1 is missing, the load mean values of adjacent regions A2, A3, and A4 are taken as the filling values:

[0162] ;

[0163] This method ensures the smoothness and consistency of spatial data.

[0164] Step 1.4, Data standardization and normalization;

[0165] Apply RobustScaler to the processed load data for standardization, and the calculation formula is:

[0166] ,

[0167] where IQR( ) is the interquartile range , which can eliminate the influence of outliers.

[0168] After the temperature data is encoded by a piecewise function, it is further normalized to the interval [0, 1]. The formula is:

[0169] ;

[0170] The normalized data distribution better meets the input requirements of the neural network and accelerates model convergence.

[0171] Step 1.5, construct time series samples;

[0172] The complete data is cut with a sample length of 720 consecutive hours (30 days) and a sliding step of 24 hours to form an input sequence (N is the number of samples, and 3 is the feature dimension: load, temperature encoding, holiday encoding). When cutting the samples, the time sliding window method is used to ensure the temporal continuity of the sequence. For example, the first sample covers hours 1 to 720, the second sample covers hours 25 to 744, and so on. The final dataset is divided into a training set, a validation set, and a test set in a 7:2:1 ratio. The validation set is used for hyperparameter tuning, and the test set is used for final performance evaluation. During the division, time overlap is strictly avoided to prevent future information from leaking into the training process.

[0173] Step 2, construct dynamic feature engineering;

[0174] Step 2.1, piecewise non-linear encoding of temperature sensitivity;

[0175] Based on the correlation analysis of load and temperature in a certain provincial power grid in recent years, 10°C and 26°C are defined as key inflection points, and a piecewise non-linear function of temperature sensitivity is constructed. The function form is:

[0176] ,

[0177] Low - temperature section (Temp < 10°C): A combination of a linear term (0.8 Temp) and a quadratic term (0.1Temp²) is used to simulate the gentle growth of heating load. For example, when Temp = 5°C, the coded value is 0.8×5 + 0.1×5² = 4 + 2.5 = 6.5. Medium - temperature section (10°C ≤ Temp ≤ 26°C): The base load is reflected by a square relationship (Temp²). When Temp = 20°C, the coded value is 400. High - temperature section (Temp > 26°C): A linear offset term (1.5(Temp−26)) and an exponential term (approximately 67.3) are introduced to capture the abrupt response of air - conditioning load. For example, when Temp = 30°C, the coded value is 1.5×(30−26) + 67.3 = 6 + 67.3 = 73.3. The coded temperature data is normalized to the [0,1] interval by Min - Max normalization to eliminate the dimension difference.

[0178] Step 2.2, Multi - dimensional dynamic decay coding for holidays;

[0179] The holiday impact Holiday_impact consists of three parts:

[0180] (1) Type weight: A value of 1.0 is assigned to national holidays (such as the Spring Festival), 0.7 to regional holidays (such as temple fairs), and 0 to weekdays.

[0181] (2) Time decay factor ( ) :

[0182] ,

[0183] where d is the number of days from the holiday. For example, if a certain day is the 3rd day of the Spring Festival (d = 0), then β(0)=1; the next day (d = 1) β(1)=0.5; the previous day (d = - 1) β(-1)=0.5.

[0184] (3) Historical comparison coefficient (γ(t)):

[0185] ,

[0186] Calculate the load y on the current day t and the average load ȳ hist over the same period in the past three years. For example, if the current day is May 1st, then ȳ hist is the average load of May 1st from 2020 - 2024.

[0187] The final coding is a weighted sum:

[0188] ,

[0189] Among them, Y1 represents the type weight;

[0190] This encoding quantifies the immediate impact (60%), decay effect (30%), and historical pattern (10%) of holidays. For example, the encoding value on the first day of the Spring Festival (Y1 = 1, β(0) = 1, γ(t) = 1.2) is 0.6×1 + 0.3×1 + 0.1×1.2 = 0.6 + 0.3 + 0.12 = 1.02.

[0191] Step 3, construct a hybrid prediction model;

[0192] As Figure 2 shown, the structure of the hybrid prediction model includes the following core modules: Build a hybrid architecture consisting of a TCN module, a bidirectional LSTM module, a spatio-temporal attention mechanism, and a feature fusion layer.

[0193] Step 3.1, multi-source feature data fusion;

[0194] Concatenate the standardized load data, temperature encoding, and holiday encoding along the feature axis to form an input tensor . The feature fusion layer dynamically adjusts the weights of each feature through a fully connected network:

[0195] ,

[0196] where is a learnable weight matrix, and d is the dimension of the hidden layer (default d = 64). This operation maps the original 3D features to a high-dimensional space, enhancing the model's ability to capture feature interactions.

[0197] Step 3.2, design the TCN module;

[0198] The TCN module contains 3 layers of dilated convolutions. The number of convolutional kernels in each layer is 32, and the dilation rates are 2, 4, and 8 in sequence. The activation function uses ReLU. The output of each layer is added element-wise to the input through a residual connection, gradually expanding the receptive field to 168 hours (covering a one-week cycle). The structure of each layer is as follows:

[0199] (1) Dilated convolution layer: The convolutional kernel size is 3, the stride is 1, and the padding method is causal convolution (CausalPadding) to ensure that the output only depends on the current and historical inputs.

[0200] (2) Activation function: Use the ReLU activation function, and the formula is:

[0201] ,

[0202] (3) Residual connection: The output of each layer is added element-wise to the input to alleviate the problem of gradient disappearance:

[0203] ,

[0204] Step 3.3, Design of Bidirectional LSTM Module

[0205] The bidirectional LSTM is set with 64 hidden units, and the input sequence length is 72 hours (3 days). The forward LSTM extracts the forward trend, and the backward LSTM captures the reverse dependence. The hidden state update formula is:

[0206] ,

[0207] ,

[0208] ,

[0209] After concatenating the outputs, the dimension is [B, 72, 128] (B is the batch size), where 128 = 64 (forward) + 64 (backward).

[0210] Step 3.4, Spatiotemporal Attention Mechanism

[0211] (1) Temporal Attention: Calculate the weights of each time step through Softmax to focus on key periods (such as morning and evening rush hours):

[0212] ,

[0213] where h t is the LSTM hidden state, is the output of the TCN, is the learnable parameter, , , , d is the hidden dimension, d = 64; T represents transpose;

[0214] (2) Feature Attention: Use the SE (Squeeze-and-Excitation) module to perform channel weighting on the temperature and holiday encoding:

[0215] Global Average Pooling (GAP): Compress the features along the time dimension:

[0216] ,

[0217] Fully connected layer:

[0218] ,

[0219] where , , σ is the Sigmoid function.

[0220] Channel weighting:

[0221] ,

[0222] where represents channel-wise multiplication.

[0223] Step 3.5, concatenate the feature fusion layer and the output layer;

[0224] The feature fusion layer concatenates the local features (dimension [B, T, 32]) output by the TCN and the global features (dimension [B, T, 128]) output by the LSTM, and inputs them into a two-layer fully connected network with 64 neurons and the Swish activation function:

[0225] ,

[0226] where the Swish function is defined as:

[0227] ,

[0228] where ;

[0229] Output the load prediction results for the next 6 hours, and set the Dropout rate to 0.2 to prevent overfitting.

[0230] Step 4, implement a phased dynamic training strategy;

[0231] Step 4.1, dynamically adjust the parameters in phases;

[0232] The first phase (rounds 1 - 10):

[0233] Input length: 6 hours, covering intraday volatility patterns (such as morning rush hours and midday lows).

[0234] Activation function: ReLU, to accelerate initial convergence.

[0235] Learning rate: initial value 1×10 -3 , using a linear warm-up strategy, gradually increasing from 1×10 -4 to 1×10 -3 in the first 5 rounds.

[0236] Objective: Learn hourly load volatility features and initially fit the short-term impacts of temperature and holidays.

[0237] The second phase (rounds 11 - 20):

[0238] Input length: Extended to 24 hours, covering a full day cycle;

[0239] Activation function: Switched to LeakyReLU (α = 0.1) to alleviate the problem of neuron death:

[0240] ,

[0241] Learning rate: reduced to 5×10 -4 , and the cosine decay strategy is adopted, decaying to 50% of the initial value every 10 rounds.

[0242] Objective: Optimize medium-term (daily-level) feature fusion and enhance the model's response ability to temperature-sensitive intervals.

[0243] The third stage (rounds 21 - 30):

[0244] Input length: set to 72 hours, covering a three-day cycle (including weekend effects).

[0245] Activation function: Use Swish to balance nonlinearity and gradient stability:

[0246] Learning rate: further adjusted to 1×10 -4 , and the exponential decay strategy is adopted:

[0247] ,

[0248] Objective: Strengthen long-term trend prediction and stabilize the distribution of attention weights.

[0249] Step 4.2, implement the early stopping mechanism and save the model;

[0250] Monitor the validation loss. When the validation loss does not decrease for 5 consecutive rounds, terminate the training and roll back to the best weights. The loss function uses MAE. The model weights are saved every round, and finally the weights with the minimum validation loss are selected for test set evaluation.

[0251] Step 5, model validation and deployment;

[0252] Step 5.1, calculation of evaluation metrics;

[0253] (1) MAE (Mean Absolute Error):

[0254] ,

[0255] (2) WMAPE (Weighted Mean Absolute Percentage Error):

[0256] ,

[0257] Step 5.2, interpretability analysis:

[0258] (1) Visualization of attention weights: As Figure 4 shown, the weight proportion of time attention in the morning peak (8:00 - 10:00) and the evening peak (18:00 - 20:00) exceeds 60%.

[0259] (2)SHAP value analysis: Calculate the feature contribution degree through DeepExplainer, as Figure 5 shown.

[0260] Step 6, Experimental results and benefit analysis;

[0261] Step 6.1, Comparison of prediction accuracy;

[0262] The MAE of the present invention is reduced by 3.33% compared with the ARIMA benchmark model; the WMAPE is decreased by 1.50% compared with the ARIMA benchmark model. The specific comparison is as Figure 3 shown.

[0263] The present invention significantly improves the prediction accuracy and stability of power demand through multi-scale feature fusion, dynamic coding mechanism and progressive training strategy, while taking into account interpretability and resource efficiency. Experiments show that the prediction error of this method decreases significantly in extreme scenarios, and the economic benefit improves significantly, providing reliable technical support for smart grid scheduling. In the future, the fusion of multi-modal data (such as satellite images, social media) can be further explored to cope with more complex prediction scenarios.

[0264] The present invention provides a dynamic power demand prediction method based on a multi-scale hybrid architecture. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation mode of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the existing technology.

Claims

1. A dynamic prediction method for power demand based on a multi-scale hybrid architecture, characterized in that, It includes the following steps: Step 1, obtaining multi-source data and preprocessing: obtaining historical power load data, performing outlier detection and dynamic correction, filling in missing values, and normalizing and standardizing the data; constructing time series samples, and dividing the historical power load data into a training set, a validation set, and a test set according to a certain proportion; Step 2, constructing dynamic feature engineering: performing piecewise non-linear encoding of temperature sensitivity, encoding the temperature data through a piecewise function to form new data, and normalizing it to form temperature feature data; performing multi-dimensional dynamic decay encoding of holidays, and obtaining holiday encoding based on the influence of holidays; Step 3, constructing a hybrid prediction model; the hybrid prediction model includes a multi-source data feature fusion module, a temporal convolutional network (TCN) module, a bidirectional long short-term memory network (BLSTM) module, a spatio-temporal attention mechanism, a feature fusion layer, and an output layer; the multi-source data feature fusion module is used for multi-source data feature fusion, the TCN module is used for dilated convolution, the BLSTM module is used for outputting the concatenated hidden states, the spatio-temporal attention mechanism is used for calculating temporal attention and feature attention, the feature fusion layer is used for completing feature concatenation, and the output layer is used for outputting the power load prediction result; Step 4, implementing a phased dynamic training strategy: training the hybrid model based on the training set in each phase, monitoring the validation loss, calculating the mean absolute error (MAE) of the validation set after each round of training, and selecting the weight with the smallest MAE for test set evaluation; Step 5, validating the model and deploying and applying the model: based on the weight obtained in Step 4, calculating the MAE and weighted mean absolute percentage error (WMAPE) based on the test set, completing the validation of the model, and deploying and applying the hybrid prediction model; Step 1 includes: Step 1.1, obtaining historical power load data with a time resolution of 1 hour, and the data fields include load values, area codes, and timestamps; Obtaining hourly meteorological temperature data; Obtaining and integrating legal holidays and regional holidays to form a structured holiday time label table; Step 1.2, outlier detection and dynamic correction; Dynamically correct the abnormal power load value by using the sliding window quantile detection method, and dynamically calculate the 0.05 quantile Q at each time point 0.05 and the 0.95 quantile Q 0.95 . If the daily power load value x′ load exceeds the range of [Q 0.05 , Q 0.95 , then replace x′ load with the median x median within the window. The calculation formula is: where x load represents the corrected power load value for replacing the outlier; Calculate the window mean through the moving average algorithm The weights of each point within the window are distributed with exponential decay. The weight w of the i-th point within the window i is as follows: w i = e -λ(168-i) , where λ is an intermediate parameter and e is the natural constant; Window mean The calculation formula is as follows: Step 1.3, missing value filling strategy; For missing data, the cubic spline interpolation algorithm with natural boundary conditions is adopted. For each interval [x m , x m+1 , a cubic polynomial S m (x load ) is defined as follows: S m (x load ) = b0 + b1(x load - x m ) + b2(x load - x m ) 2 + b3(x load - x m ) 3 , where m = 0, 1, …, n-1; b0, b1, b2, b3 are the coefficients of the cubic polynomial, and x i is the timestamp of the known data point; by using the continuity and boundary conditions, the spline coefficients b0, b1, b2, b3 are solved, and finally the cubic polynomials on each interval are obtained. Using the mean value of physically adjacent regions to fill in the missing values in the spatial dimension: if the data of a certain moment in region A1 is missing, then take the mean value of the power loads of physically adjacent regions A2, A3, and A4 as the filling value; Among them represents the missing value filling result of area A1, represents the average power load of area A2, represents the average power load of area A3, represents the average power load of area A4; Step 1.4, data standardization and normalization: For the processed power load data x load Apply RobustScaler for standardization to obtain the standardized data The calculation formula is as follows: where median(X load ) is the median of the processed power load data set X load , and IQR(X load ) is the interquartile range, and the calculation formula is: IQR(X load ) = Q 0.75 - Q 0.25 , where Q 0.75 represents the 0.75 quantile at each time point, and Q 0.25 represents the 0.25 quantile at each time point; Step 1.5, constructing time series samples; The historical power load data is cut with a sample length of 720 consecutive hours and a sliding step of 24 hours to form an input sequence where N is the number of samples and R represents the real number space; When cutting the samples, the time sliding window method is used to ensure the temporal continuity of the sequence; Finally, the historical power load data is divided into a training set, a validation set, and a test set according to a certain proportion; Step 2 includes: Step 2.1, piecewise non-linear encoding of temperature sensitivity; Defining 10°C and 26°C as key inflection points, and constructing the piecewise non-linear function of temperature sensitivity α(Temp): The temperature data Temp is encoded by the piecewise function α(Temp) to form the data Temp encoded , and then further normalized to the interval [0, 1] to form the temperature feature data x temp , and the formula is: Step 2.2, multi-dimensional dynamic decay encoding of holidays; Step 2.2 includes: holiday impact x holiday including a type weight, a time decay factor β(d), and a historical comparison coefficient γ(t); The calculation formula of the time decay factor β(d) is: where d is the number of days from the holiday; The calculation formula for the historical comparison coefficient γ(t) is as follows: Among them represents the average power load in the same period of the recent three years; The holiday code x is calculated using the following formula holiday :[[]]END]] x holiday = 0.6 × Y1 + 0.3 × β(d) + 0.1 × γ(t), where Y1 represents the type weight; Step 3 includes the following steps: Step 3.1, the multi-source data feature fusion module performs the following operations: The standardized power load data Temperature encoding x temp And holiday encoding x holiday Are concatenated along the feature axis to form an input tensor The feature fusion layer dynamically adjusts the weights of each feature through a fully connected network to achieve scene adaptive coupling of multi-source signals. The formula is: where X in ∈R N×720×D is the hidden representation after feature fusion, W f ∈R 3×D is the learnable weight matrix, D is the hidden layer dimension, b f is the bias term of the fully connected layer, and ReLU is the activation function; Step 3.2, establish a Temporal Convolutional Network (TCN) module; The Temporal Convolutional Network (TCN) module contains 3 layers of dilated convolutions. Each layer of the dilated convolution structure is the same and includes a dilated convolutional layer, a ReLU activation function, and a residual connection. The number of convolutional kernels in each layer of the dilated convolutional layer is 32, the convolutional kernel size is 3, the stride is 1, and the padding method is causal convolution. The dilation rates of the 3 layers of dilated convolutional layers are 2, 4, and 8 in sequence. The output of each layer of dilated convolution is added element-wise to the input through the residual connection, gradually expanding the receptive field to 168 hours. The formula for the residual connection is: The input sequence of the dilated convolution is the tensor X spliced by the standardized load data, temperature encoding, and holiday encoding in , with the dimension of [B, 720, 3], where B is the batch size; is the input of the l-th layer of dilated convolution; The output of the l-th layer of dilated convolution, and the final output dimension is [B, 720, 32]; Conv1D is a one-dimensional dilated convolution operation; Step 3.3, establish a Bidirectional Long Short-Term Memory (BLSTM) module; The bidirectional long short-term memory network BLSTM module includes a forward long short-term memory module and a backward long short-term memory module. 64 hidden units are set in the bidirectional long short-term memory network BLSTM module. The output of the temporal convolutional network TCN module The last 72 hours are intercepted as the input of the forward long short-term memory module. The forward long short-term memory module extracts the forward trend, and the backward long short-term memory module captures the reverse dependence. The hidden state update formula is: where t is the time step, h t is the long short-term memory hidden state, is the forward long short-term memory hidden state, is the backward long short-term memory hidden state, and LSTM is the model function; The output of the Bidirectional Long Short-Term Memory (BLSTM) module is concatenated to obtain a hidden state with a dimension of [B, 72, 128]; Step 3.4, establish a spatio-temporal attention mechanism, including temporal attention and feature attention; The temporal attention calculates the weights of each time step through the Softmax function: w t = Softmax(v TP tanh(W h h t + W x x TCN,t )) where w t represents the weight at time step t, x TCN,t is the output of the temporal convolutional network TCN module at time step t, v, W h , W x are learnable parameters, v ∈ R D , W h ∈ R D×128 is the key matrix of the attention mechanism, mapping the hidden state of the BiLSTM to the attention space, W x ∈ R D×32 is the value matrix, mapping the TCN features to the attention space; TP represents transpose; Softmax is the normalized exponential function; tanh is the hyperbolic tangent function; The feature attention uses a Squeeze-and-Excitation (SE) module to perform channel weighting on the temperature and holiday encoding. The Squeeze-and-Excitation (SE) module performs the following steps: Step 3.4.1, use global average pooling to compress the features along the time dimension to obtain the pooled global feature vector s: where x t is the input feature at time step t, with dimensions [B, T, C]; Step 3.4.2, obtain the channel weight vector s' through two fully connected networks: s' = σ(W2ReLU(W1s + b1) + b2), where, W1 ∈ R 32×16 is the weight matrix of the first fully connected layer, b1 is the bias term of the first fully connected layer, W2 ∈ R 16 ×3 is the weight matrix of the second fully connected layer, b2 is the bias term of the second fully connected layer, ReLU is the activation function, and σ is the Sigmoid function; Step 3.4.3, perform channel weighting to obtain the weighted output feature x': where x ∈ R B×T×C , x is the original input feature represents channel-wise multiplication Step 3.5, splice the feature fusion layer and the output layer; Step 3.5 includes: implementing a feature fusion layer to concatenate the local feature x output by the Temporal Convolutional Network (TCN) module TCN with the global feature x output by the Bidirectional Long Short-Term Memory (BLSTM) module LSTM and inputting the result into a two-layer fully connected network with 64 neurons and a Swish activation function: x pred = W0·Swish(W fuse ·[x TCN ; x LSTM +b h )+b0, where, W0 ∈ R 64×1 is the weight of the output layer, and W fuse ∈ R 160×64 is the weight matrix of the fully connected network, b h is the bias term of the hidden layer in the feature fusion layer, b0 is the bias term of the output layer, and x pred is the predicted result of the future 6-hour power load output by the output layer; Step 4 includes: Step 4.1, each stage is based on the training set for 10 rounds of hybrid model training. Each round includes a complete traversal of all training samples, and the parameters are dynamically adjusted in stages; The first stage is from round 1 to Y1: Use the training set to train the model 10 times completely. Each training traverses all training data; Input length: 6 hours; Activation function: ReLU function; Learning rate: initial value 1×10 -3 , using a linear warm-up strategy, gradually increasing from 1×10 -4 to 1×10 -3 in the first 5 rounds; Objective: Learn the hourly load fluctuation characteristics and suppress long-term noise interference; The second stage is from round 11 to 20: Continue training for 10 rounds and adjust the input length and hyperparameters; Input length: Extended to 24 hours; Activation function: Switched to the LeakyReLU function; Learning rate: decreased to 5×10 -4 , using the cosine annealing strategy, decaying to 50% of the initial value every 10 epochs; Objective: Optimize the daily feature fusion; The third stage is from round 21 to 30: Retrain for 10 rounds to further optimize the long-term trend prediction; Input length: Set to 72 hours; Activation function: Adopt the Swish function; Learning rate: Adjusted to 1×10 -4 , and an exponential decay strategy is adopted to calculate the dynamically adjusted learning rate η(t), which gradually decreases as the number of training rounds increases: Among them, η0 is the initial learning rate, decay_rate is the decay rate, t epoch is the number of training rounds, and decay_steps is the decay step size; Objective: Strengthen the long-term trend prediction and stabilize the attention weight distribution; Step 4.2, execute the early stopping mechanism and save the model: Monitor the validation loss. After each round of training, calculate the mean absolute error (MAE) of the validation set. If the validation loss does not decrease for 5 consecutive rounds, terminate the training and roll back to the historical best weights; the model weights are saved once per round, and finally, the weights with the minimum mean absolute error (MAE) are selected for test set evaluation; the test set is isolated throughout and only used in the final evaluation. Step 5 includes: Step 5.1, load the best weights saved in Step 4.2 and calculate the following metrics based on the test set: Mean absolute error (MAE): Weighted mean absolute percentage error (WMAPE): Among them, y i is the actual load value; is the model prediction value; N is the number of samples in the test set; Step 5.2, conduct interpretability analysis; Step 5.3, perform lightweight deployment and application.

2. An electronic device, characterized in that, Comprising a processor and a memory, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method according to claim 1.

3. A storage medium, characterized in that, Stored with a computer program or instructions, when the computer program or instructions are run on a computer, the steps of the method according to claim 1 are executed.

Citation Information

Patent Citations

  • Urban rail transit passenger flow prediction method based on similar day selection

    CN115169638A

  • CNN-BiLSTM load prediction method based on isolated forest

    CN118821991A