Power demand dynamic prediction method based on multi-scale hybrid architecture

By adopting a multi-scale hybrid architecture of TCN and LSTM in power demand forecasting, combining multi-source data feature fusion and dynamic feature engineering, the problems of insufficient nonlinear coupling modeling of external variables and loads and insufficient multi-scale feature extraction are solved, and more accurate and stable power demand forecasting is achieved, especially in extreme weather and holiday scenarios.

CN119990477AActive Publication Date: 2025-05-13WUXI UNIV

Patent Information

Application Number
CN202510468974.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the existing power demand prediction technology, the nonlinear dynamic coupling modeling of external variables and loads is insufficient, the multi-scale timing feature extraction is insufficient, and the prediction results are insufficient, resulting in a significant increase in prediction error in extreme weather scenarios and a large prediction error during holiday load decay.

Method used

A multi-scale hybrid architecture based on time-domain convolutional network (TCN) and long and short-term memory network (LSTM) is adopted, combining multi-source data feature fusion, dynamic feature engineering and phased dynamic training strategies, a dynamic coupled model is built to quantify the impact of external variables, and the multi-scale feature extraction ability and interpretability of prediction are improved through spatiotemporal attention mechanism and feature fusion layer.

Benefits of technology

It significantly improves the accuracy and stability of power demand forecasting, especially in extreme weather and holiday scenarios, reduces prediction errors, and improves the interpretability of prediction results, supporting grid scheduling decision optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990477A_ABST
    Figure CN119990477A_ABST
Patent Text Reader

Abstract

The invention provides a power demand dynamic prediction method based on a multi-scale hybrid architecture. The power demand dynamic prediction method comprises the following steps: step 1, acquiring multi-source data and preprocessing the multi-source data; step 2, constructing a dynamic feature project; step 3, constructing a hybrid prediction model; 4, executing a staged dynamic training strategy; and 5, verifying the model, and deploying and applying the model. According to the invention, by introducing the temperature sensitivity piecewise nonlinear coding and the holiday multi-dimensional dynamic attenuation model, the prediction precision in an extreme temperature scene is significantly improved, and the high-temperature daily load prediction error and the spring festival holiday error are reduced. In addition, through lightweight design and staged dynamic strategy staged extension of the input length, the training convergence speed is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent prediction of power systems, and specifically relates to a method for dynamic prediction of power demand based on a multi-scale hybrid architecture. Background Art

[0002] Electricity demand forecasting is the core link of power system optimization and dispatching, which directly affects the formulation of power generation plans and the safe operation of power grids. Traditional statistical models (such as ARIMA and Prophet) rely on manually set time series parameters and are difficult to handle the nonlinear coupling relationship of multi-source heterogeneous data (such as temperature and holidays). With the development of deep learning technology, recurrent neural networks such as Long Short-Term Memory (LSTM) have gradually become mainstream solutions, but they still face significant bottlenecks. For example, Shi Min et al. used LSTM to make short-term predictions of photovoltaic power generation. The prediction results showed high adaptability and optimization potential, but there were still problems such as insufficient feature coupling, multi-scale conflicts, and relatively weak local feature extraction capabilities, and it was easy to ignore the details in short-term dependencies. Different from traditional recurrent neural networks, Temporal Convolutional Network (TCN) is a new deep learning structure. By integrating multiple mechanisms such as dilated convolution, causal convolution, and residual modules, it can effectively extract the correlation between data and better handle time series prediction problems. Wang et al. and Li Feifei believe that the atrous convolution in the TCN architecture expands the receptive field, but because a dynamic association model is not built, the prediction variance will increase significantly under policy mutation scenarios, and the ability to model daily / weekly cycles is weak, resulting in large errors in holiday load attenuation predictions.

[0003] In addition, existing methods have inherent contradictions in feature engineering, model architecture, and training strategies. The nonlinear relationship between temperature and load needs to be modeled in segments (such as 10°C and 26°C thresholds), but traditional linear normalization leads to blurred features in key intervals; LSTM and TCN alone are difficult to take into account both local fluctuations and long-term trends; complex models (such as Transformer) rely on high-end hardware and cannot be adapted to ordinary equipment of power grid companies. The core difficulty in solving these problems lies in: how to design a lightweight hybrid architecture to balance the efficiency of multi-scale feature extraction, while building a dynamic coupling model to quantify the impact of external variables and ensure training stability under limited computing power. Summary of the invention

[0004] Purpose of the invention: In order to effectively improve the prediction performance and accuracy of traditional power quality prediction methods, the present invention provides a method for dynamic prediction of power demand based on a multi-scale hybrid architecture, specifically a method for power quality prediction based on a time-domain convolutional network (TCN) and a long short-term memory network (LSTM). TCN effectively processes time series data through dilated convolution and causal convolution, and can capture local features at different time scales, while LSTM is good at processing long-term dependencies in time series. The combination of the two can retain short-term features in the data and capture long-term trends when processing complex power quality data. The present invention aims to solve the following problems existing in existing power demand forecasting technologies: (1) Insufficient modeling of nonlinear dynamic coupling between external variables and loads: The association mechanism between external factors such as temperature and holidays and power loads has not been accurately quantified. In particular, the temperature threshold effect (such as the sudden change response of air conditioning load at 26°C) lacks effective characterization, resulting in a significant increase in prediction errors under extreme weather scenarios. (2) Insufficient extraction of multi-scale time series features: A single model architecture is difficult to capture both hourly fluctuation characteristics (such as sudden changes in morning peak load) and weekly / monthly long-term trends (such as seasonal electricity consumption growth). The conflict between features at different time scales leads to decreased prediction stability. (3) Insufficient interpretability of prediction results: Existing deep learning models lack a quantification mechanism for feature contribution and are unable to analyze the specific impact of key factors such as temperature and holidays on prediction results, making it difficult to support the optimization of power grid dispatching decisions.

[0005] The method of the present invention comprises the following steps: Step 1, obtain multi-source data and perform preprocessing: obtain historical power load data, perform outlier detection and dynamic correction, fill missing values, and standardize and normalize the data; construct a time series sample, and divide the historical power load data into training set, validation set, and test set in proportion; Step 2: construct dynamic feature engineering: perform piecewise nonlinear coding of temperature sensitivity, encode the temperature data through piecewise functions to form new data, and normalize it to form temperature feature data; perform multi-dimensional dynamic attenuation coding of holidays, and obtain holiday codes based on holiday impacts; Step 3, constructing a hybrid prediction model; the hybrid prediction model includes a multi-source data feature fusion module, a temporal convolutional network TCN (Temporal Convolutional Network, TCN) module, a bidirectional long short-term memory network BLSTM module, a spatiotemporal attention mechanism, a feature fusion layer and an output layer; the multi-source data feature fusion module is used to perform multi-source data feature fusion, the temporal convolutional network TCN module is used for hole convolution, the bidirectional long short-term memory network BLSTM (Bidirectional Long Short-Term Memory, BLSTM) module is used to output the spliced ​​hidden state, the spatiotemporal attention mechanism is used for time attention and feature attention calculation, the feature fusion layer is used to complete feature splicing, and the output layer is used to output the power load prediction result; Step 4: Execute a phased dynamic training strategy: perform hybrid model training based on the training set in each phase, monitor the validation loss, calculate the mean absolute error (MAE) of the validation set after each round of training, and select the weight with the smallest MAE for the test set evaluation; Step 5, verify the model and deploy and apply the model: Based on the weights obtained in step 4, calculate the mean absolute error MAE and the weighted mean absolute percentage error WMAPE based on the test set, complete the model verification, and deploy and apply the hybrid prediction model.

[0006] Step 1 includes: Step 1.1, obtain historical power load data with a time resolution of 1 hour. The data fields include load value, area code and timestamp; Get hourly meteorological temperature data; Obtain and integrate statutory holidays and regional holidays to form a structured holiday time label table; Step 1.2, outlier detection and dynamic correction; The sliding window quantile detection method is used to dynamically correct abnormal power load values ​​and dynamically calculate the 0.05 quantile at each time point. and 0.95 quantile If the power load value on that day is Beyond , ] range, then Replace with the median within the window , the calculation formula is: , in It represents the corrected power load value, which is used to replace the abnormal value; Because the more recent the data is, the higher the weight should be, so the window mean is calculated using the moving average algorithm , the weight of each point in the window is allocated according to exponential decay, and the weight of the i-th point in the window (with a value of 0~167) for: , in is an intermediate parameter, generally set to , λ is 0.01, which is an empirical value and can effectively balance the importance of recent data and the retention of historical trends. In hourly load forecasting, this value can highlight the latest data without excessively ignoring historical information. It is commonly used in practical applications and has stable effects. e is a natural constant; Window mean The calculation formula is: ; Step 1.3, missing value filling strategy; For missing data, the cubic spline interpolation algorithm with natural boundary conditions is used. , define a cubic polynomial : , in ; , , , are cubic polynomial coefficients, is the timestamp of the known data point. By using continuity and boundary conditions, solve the spline coefficients , , , , and finally obtain the cubic polynomial on each interval; Use the mean of physically adjacent areas to fill in missing values ​​in the spatial dimension: The areas are divided according to the physical connection relationship in the power grid topology (such as the power supply range of substations and line connection relationships) rather than the geographically adjacent areas, which solves the filling deviation caused by the uneven distribution of power grid trends; if the data of area A is missing at a certain moment, the mean of the power load of physically adjacent areas B, C, and D is taken as the filling value: ; in Indicates the missing value filling result of area A. represents the average power load in area B, represents the average power load in area C, represents the average power load of area D; Step 1.4, data standardization and normalization: In order to eliminate the influence of outliers and make the data distribution more stable and more suitable for neural network training, it is necessary to process the power load data Apply Robust Scaler to standardize and get standard data , the calculation formula is: , Where median( ) is the processed power load data set The median, IQR( ) is the interquartile range, and the calculation formula is: IQR( )= - , in represents the 0.75 quantile at each time point, represents the 0.25 quantile at each time point; Step 1.5, construct time series samples; The historical power load data is cut into samples of 720 consecutive hours and a sliding step of 24 hours to form an input sequence , where N is the number of samples and R represents the real number space; When cutting samples, the time sliding window method is used to ensure the temporal continuity of the sequence; Finally, the historical power load data is divided into training set, validation set and test set in proportion: training set: 70% (for model parameter update and feature learning), validation set: 20% (for hyperparameter tuning and early stopping decision), test set: 10% (for final performance evaluation).

[0007] Step 2 includes: Step 2.1, temperature sensitivity piecewise nonlinear coding; Different from traditional linear or single nonlinear coding (such as exponential function), the present invention defines 10℃ and 26℃ as key inflection points and constructs a double critical temperature piecewise function. Among them, according to the low temperature segment (Temp < 10℃): simulate the gentle growth of heating load; medium temperature segment (10℃~26℃): reflect the square relationship of basic load; high temperature segment (Temp > 26℃): capture the sudden change response of air conditioning load (linear + exponential term). Therefore, a temperature sensitivity piecewise nonlinear function is constructed. : , Temperature data Temp through piecewise function After encoding, the data is formed , and then further normalized to the [0,1] interval to form temperature characteristic data , the formula is: ; Step 2.2, multi-dimensional dynamic attenuation encoding of holidays.

[0008] Step 2.2 includes: quantifying the immediate impact of holidays on load (such as the first day of the Spring Festival), attenuation effects (such as resuming work after the holidays), and historical patterns (such as like-for-like changes). Including type weight, time decay factor and historical comparison coefficient γ(t); For type weights, national holidays are assigned a value of 1.0, regional holidays are assigned a value of 0.7, and weekdays are assigned a value of 0; Time decay factor The calculation formula is: , Where d is the number of days until the holiday; The calculation formula of historical comparison coefficient γ(t) is: , in It indicates the average power load in the same period of the past three years; The holiday code is calculated using the following formula : , Where Y1 represents the type weight.

[0009] Step 3 specifically includes the following steps: Step 3.1, the multi-source data feature fusion module performs the following operations: Existing methods mostly use simple splicing or static weight fusion. The present invention converts the standardized power load data into , Temperature coding and holiday codes Concatenate along the feature axis to form the input tensor , the feature fusion layer (including two fully connected layers, where the fully connected layer refers to a hierarchical structure that connects each neuron to all neurons in the previous layer) dynamically adjusts the weights of each feature through a fully connected network (a network structure composed of multiple fully connected layers stacked together) to achieve scene-adaptive coupling of multi-source signals. The formula is: , in, It is the hidden representation after feature fusion, including the dynamically adjusted multi-source feature information. is the learnable weight matrix, D is the hidden layer dimension, is the bias term of the fully connected layer, and ReLU is the activation function.

[0010] Step 3.2, establish a temporal convolutional network TCN module; In view of the multi-scale peak characteristics of power load (such as instantaneous overload), a progressive expansion rate design is adopted to gradually expand the receptive field and avoid the loss of local features caused by a single expansion rate. The temporal convolutional network TCN module contains 3 layers of dilated convolutions. Each layer of dilated convolution has the same structure, including a dilated convolution layer, a ReLU activation function and a residual connection. The number of convolution kernels in each dilated convolution layer is 32, the convolution kernel size is 3, the step size is 1, and the filling method is causal convolution. The expansion rates of the 3 dilated convolution layers are 2, 4, and 8 respectively. The output and input of each dilated convolution layer are added element by element through a residual connection, gradually expanding the receptive field to 168 hours; the residual connection formula is: , The input sequence of the dilated convolution is the tensor of the normalized load data, temperature code, and holiday code concatenation. , the dimension is [B,720,3], B is the batch size; For the The input of the dilated convolution layer (the first layer input is the original data, and the subsequent layer input is the output of the previous layer); No. The output of the layer dilated convolution, the final output dimension is [B, 720, 32]; Conv1D is a one-dimensional dilated convolution operation; Step 3.3, establish a bidirectional long short-term memory network BLSTM module; The bidirectional long short-term memory network BLSTM module includes a forward long short-term memory module and a backward long short-term memory module. The bidirectional long short-term memory network BLSTM module is provided with 64 hidden units. The temporal convolution network TCN module outputs The last 72 hours are taken as the input of the forward long short-term memory module. The forward long short-term memory module extracts positive trends, and the backward long short-term memory module captures reverse dependencies. The hidden state update formula is: , , , Where t is the time step, h t is the long short-term memory hidden state, is the forward long short-term memory hidden state, is the backward long short-term memory hidden state, and LSTM is the model function; The bidirectional long short-term memory network BLSTM module outputs a hidden state with a concatenated dimension of [B, 72, 128]; Step 3.4, establish the spatiotemporal attention mechanism, including temporal attention and feature attention; The temporal attention calculates the weight of each time step through the Softmax function (the higher the weight, the more important the moment is for prediction): , where w t represents the weight of time step t, is the output of the temporal convolutional network TCN module at time step t, v, , is a learnable parameter, ; is the key matrix of the attention mechanism, mapping the hidden state of BiLSTM to the attention space. is the value matrix, which maps the TCN features to the attention space; TP stands for transpose; Softmax is a normalized exponential function, which converts the input vector into a probability distribution to indicate the importance of each time step; tanh is a hyperbolic tangent function, which performs a nonlinear transformation on the input features to enhance the nonlinear expression capability; The feature attention uses the Squeeze-and-Excitation (SE) module to perform channel weighting on temperature and holiday codes. The Squeeze-and-Excitation (SE) module is mainly used to enhance inter-channel dependencies in image classification tasks. The Squeeze-and-Excitation (SE) module is migrated from the image field to power demand forecasting to perform channel weighting on time series features (temperature, holiday codes). This is an improved application of the existing technology, and its advantages include improving the model's sensitivity to key features, suppressing the impact of noise or irrelevant features, and enhancing the model's interpretability. The Squeeze-and-Excitation (SE) module performs the following steps: Step 3.4.1, use global average pooling to compress features along the time dimension to obtain the pooled global feature vector s: , in, is the input feature of time step t, with dimension [B, T, C] (B is the batch size, T is the time step, and C is the number of channels). In the present invention, C=3 (power load, temperature coding, holiday coding), T: the time step of the input sequence (such as setting T=720 hours); Step 3.4.2, through the two-layer fully connected network, get the channel weight vector : , in, is the weight matrix of the first fully connected layer, is the bias term of the first fully connected layer, is the weight matrix of the second fully connected layer, The bias term of the second fully connected layer, ReLU is the activation function, and σ is the Sigmoid function; Step 3.4.3, perform channel weighting to obtain weighted output features : , in, , x is the original input feature (load, temperature code, holiday code, where x is the same as before pooling Different, before pooling is a single time step feature, where x is a complete time series), represents channel-by-channel multiplication; Step 3.5, the feature fusion layer is concatenated with the output layer.

[0011] Step 3.5 includes: implementing a feature fusion layer to transform the local features output by the temporal convolutional network TCN module The global features output by the bidirectional long short-term memory network BLSTM module Concatenate and input into a two-layer fully connected network with 64 neurons and Swish activation function: , in, is the output layer weight, which is used to map the 64-dimensional hidden layer to the predicted value, is the weight matrix of the fully connected network, which maps the concatenated features to the 64-dimensional hidden layer. is the bias term of the hidden layer in the feature fusion layer, is the bias term of the output layer, The output layer outputs the power load forecast results for the next 6 hours.

[0012] Existing deep learning models often use fixed input length and unified activation function, which makes it difficult to distinguish the coupled interference of hourly mutations, daily cycles and long-term trends in power demand. The present invention achieves feature decoupling through a phased strategy. Step 4 includes: Step 4.1: In each stage, 10 epochs of hybrid model training are performed based on the training set. Each epoch includes a complete traversal of all training samples, and the parameters are dynamically adjusted in stages. The first stage is 1 to 10 rounds: the model is trained 10 times (i.e., Y1 epochs) using the training set, and all training data are traversed each time.

[0013] Input length: 6 hours; Activation function: ReLU function (Rectified Linear Unit), which is defined as: when the input value is greater than or equal to 0, the output value is equal to the input value; when the input value is less than 0, the output is 0. By suppressing negative inputs, reducing the number of neuron activations, and enhancing the sparse expression ability of the model, it is suitable for quickly extracting local mutation patterns (such as morning peaks); Learning rate: initial value 1×10 -3 , using a linear warm-up strategy, the first 5 rounds gradually increase from 1×10 -4 Increase to 1×10 -3 ; Objective: Learn the hourly load fluctuation characteristics and suppress long-term noise interference; The second stage is 11~20 rounds: continue training for 10 rounds (epochs), but adjust the input length and hyperparameters.

[0014] Input length: extended to 24 hours; Activation function: Switch to LeakyReLU function (Leaky Rectified Linear Unit), which is defined as: when the input value is non-negative, the original value is output, and when it is negative, the product of the input value and the leakage coefficient (default 0.1) is output. It allows weak activation of negative inputs to alleviate the gradient disappearance and avoid neuron inactivation problems. It is suitable for the fusion of medium-complexity daily periodic features (such as temperature sensitive intervals and holiday attenuation): , Learning rate: reduced to 5×10 -4 , using the cosine decay strategy, decaying to 50% of the initial value every 10 rounds; Objective: Optimize day-level feature fusion; The third stage is 21~30 rounds: train for another 10 rounds (epochs) to further optimize the long-term trend prediction.

[0015] Input length: set to 72 hours; Activation function: The Swish function (self-gated activation function) is used, which has a continuous and differentiable nonlinearity near x=0, avoiding the hard truncation effect of ReLU, and dynamically adjusts the activation threshold through the β parameter to adapt to different data distributions. The smooth nonlinear characteristics of the Swish function can effectively capture more complex long-term dependencies (such as seasonal changes and holiday attenuation effects), and its adaptability helps stabilize the distribution of attention weights: , Among them, β is a learnable parameter (default β=1); Learning rate: adjusted to 1×10 -4, using an exponential decay strategy to calculate the dynamically adjusted learning rate , which gradually decreases with the increase of training rounds: , in, is the initial learning rate (1×10 -4 ), decay_rate is the decay rate (controls the decay speed of the learning rate, the smaller the value, the faster the decay), is the training round (the 21st to 30th rounds are relative rounds counted from the third stage, the 21st round is ), decay_steps is the decay step length (if decay_steps=5, it decays every 5 rounds); Goal: Strengthen long-term trend prediction and stabilize attention weight distribution; Step 4.2, execute the early stopping mechanism and save the model: monitor the validation loss, calculate the mean absolute error (MAE) of the validation set after each round of training, and terminate the training and roll back to the historical best weight if the validation loss does not decrease for 5 consecutive rounds. The model weight is saved once per round, and the weight with the smallest mean absolute error MAE is finally selected for the test set evaluation. The test set is isolated throughout the process and is only used in the final evaluation to ensure that the performance indicators are unbiased.

[0016] Step 5 includes: Step 5.1, load the best weights saved in step 4.2, and calculate the following indicators based on the test set: Mean absolute error (MAE): , Weighted Mean Absolute Percent Error WMAPE: , in, is the actual load value (test set data); is the model prediction value; N is the number of test set samples.

[0017] Step 5.2, perform interpretability analysis: Traditional interpretability analysis is mostly limited to feature importance ranking, while this invention innovatively realizes the organic combination of attention weight (attention weight visualization) and spatiotemporal contribution (SHAP value analysis). In this method, the validation set is used for interpretability analysis to analyze the contribution of features such as temperature and holidays.

[0018] Attention weight visualization: The weights assigned by the attention mechanism in the neural network are displayed in the form of heatmaps, which intuitively reflects the degree of attention the model pays to different time steps or features during prediction.

[0019] SHAP is a model interpretability method based on game theory, which is widely used to explain machine learning models (including deep learning). DeepExplainer is an explanation tool in the SHAP library specifically designed for deep learning models. Using the SHAP (SHapley Additive exPlanations) method, the DeepExplainer tool calculates the contribution of each input feature in the model to the prediction result and quantifies the impact of each feature on the final prediction.

[0020] Step 5.3, lightweight deployment application; Traditional model deployment often relies on high-performance computing resources. However, the present invention reduces computational redundancy during inference by retaining the hourly feature extraction layer (the first two layers of TCN) and only retaining the long-term trend layer (BiLSTM) in stage 3.

[0021] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.

[0022] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.

[0023] The present invention has the following technical advantages: Multi-scale feature fusion: TCN dilated convolution captures hourly spikes, bidirectional LSTM models weekly / monthly trends, and spatiotemporal attention focuses on key time periods and features; Dynamic interpretability: SHAP value and attention weight analysis predictive logic to support grid dispatch decision optimization; Progressive training: Compared with the training strategy with fixed input length, the phased dynamic strategy improves the convergence speed while reducing the prediction error of extreme weather scenarios.

[0024] Beneficial effects: The present invention significantly improves the prediction accuracy in extreme temperature scenarios by introducing temperature sensitivity piecewise nonlinear coding and a multi-dimensional dynamic attenuation model for holidays, and reduces the load prediction error on high temperature days and the error during the Spring Festival holiday. In addition, the training convergence speed is significantly improved by expanding the input length in stages through lightweight design and phased dynamic strategies. In terms of interpretability, the time attention weight focuses on the morning and evening peak hours, and the SHAP value quantification shows that the temperature contribution on high temperature days is relatively high, providing a decision-making basis for power grid dispatching. In terms of economic benefits, combined with the time-of-use electricity price strategy, the daily dispatching cost of a single region, the prediction variance of extreme weather scenarios, and the fault loss are significantly reduced, which has significant social application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.

[0026] Figure 1 A flow chart of a method for dynamic prediction of power demand based on a multi-scale hybrid architecture provided in an embodiment shows the complete process from multi-source data collection, preprocessing, dynamic feature engineering, model training to prediction result output.

[0027] Figure 2 The hybrid prediction model structure diagram provided for the embodiment includes a dilated convolution layer of a TCN module (expansion rate 2 / 4 / 8), temporal dependency extraction of a bidirectional LSTM module, dynamic weight allocation of a spatiotemporal attention mechanism, and multi-scale information integration of a feature fusion layer.

[0028] Figure 3 The model performance comparison chart provided in the embodiment uses MAE and WMAPE as indicators to compare the prediction accuracy of ARIMA, BiLSTM, CNN-LSTM and the present invention, and annotates the error reduction in extreme weather scenarios.

[0029] Figure 4 To mark the focusing effect of time attention weight during the morning and evening peak hours (accounting for >60%).

[0030] Figure 5 This is a SHAP feature contribution analysis, which mainly displays the SHAP value distribution of load / temperature / holidays in the previous 24 hours. DETAILED DESCRIPTION

[0031] like Figure 1 As shown, an embodiment of the present invention provides a method for dynamic prediction of power demand based on a multi-scale hybrid architecture, comprising: Step 1, obtain multi-source data and preprocess it; Step 1.1, data source and collection specifications; The data acquisition module of this embodiment obtains historical load data from the provincial companies of the State Grid with a time resolution of 1 hour. The data fields include load value (unit: megawatt, MW), regional code (identifying the grid partition), and timestamp (accurate to the hour). The regional code adopts a three-level structure, such as "East China-Jiangsu-Nanjing", to support subsequent spatial correlation analysis. Meteorological data obtains hourly temperature (℃) data from the China Meteorological Administration. Holiday data integrates statutory holidays and regional holidays announced by provincial governments to form a structured time tag table. All data is synchronized to the distributed time series database (InfluxDB) in real time through the API interface to ensure efficient reading and writing of data and time window query capabilities.

[0032] Step 1.2, outlier detection and dynamic correction; The sliding window quantile detection method is used to dynamically correct abnormal load values. The window length is set to 168 hours (7 days), and the 0.05 quantile ( ) and the 0.95 quantile ( ). If the current power load value Beyond range, it is replaced by the median within the window , the specific calculation formula is: , Window mean The weight of each point in the window is calculated by the moving average algorithm, and the weight of recent data is higher. For example, the weight of the i-th point in the window is: , This approach eliminates abnormal fluctuations caused by extreme weather (such as typhoons) or equipment failures while retaining the continuity of normal load trends.

[0033] Step 1.3, missing value filling strategy; For missing data, the cubic spline interpolation algorithm with natural boundary conditions is used. The basic idea of ​​the interpolation formula is to represent every two adjacent data points with a cubic polynomial, so that these polynomials are continuous and smooth at adjacent nodes (including the continuity of the first-order derivative and the second-order derivative). For each interval , define a cubic polynomial: , in ; Through continuity and boundary conditions, the spline coefficients are solved and finally the cubic polynomial on each interval is obtained, which not only ensures the continuity of the function value but also ensures the smoothness of the data curve.

[0034] The missing values ​​in the spatial dimension are filled with the mean value of the adjacent regions. The regional division is based on the physical connection relationship in the power grid topology, such as the coverage area of ​​the same substation. If the data of region A1 is missing at a certain moment, the load mean value of the adjacent regions A2, A3, and A4 is taken as the filling value: ; This method ensures the smoothness and consistency of spatial data.

[0035] Step 1.4, data standardization and normalization; The processed load data is standardized by RobustScaler, and the calculation formula is: , Where IQR( ) is the interquartile range , which can eliminate the influence of outliers.

[0036] After the temperature data is encoded by the piecewise function, it is further normalized to the [0,1] interval. The formula is: ; The normalized data distribution is more in line with the neural network input requirements and accelerates model convergence.

[0037] Step 1.5, construct time series samples; The complete data is cut into samples of 720 consecutive hours (30 days) and a sliding step of 24 hours to form an input sequence (N is the number of samples, 3 is the feature dimension: load, temperature coding, holiday coding). When cutting samples, the time sliding window method is used to ensure the temporal continuity of the sequence. For example, the first sample covers the 1st to 720th hour, the second sample covers the 25th to 744th hour, and so on. The final data set is divided into training set, validation set and test set in a ratio of 7:2:1. The validation set is used for hyperparameter tuning, and the test set is used for final performance evaluation. When dividing, time overlap is strictly avoided to prevent future information from leaking into the training process.

[0038] Step 2: Build dynamic feature engineering; Step 2.1, temperature sensitivity piecewise nonlinear coding; Based on the load and temperature correlation analysis of a provincial power grid in recent years, 10℃ and 26℃ are defined as key inflection points, and a piecewise nonlinear function of temperature sensitivity is constructed. The function form is: , Low temperature range (Temp< 10℃): Linear term (0.8 Temp) and quadratic term (0.1Temp²) are combined to simulate the gentle increase of heating load. For example, when Temp = 5℃, the encoding value is 0.8×5 + 0.1×5² = 4 + 2.5 = 6.5. Medium temperature range (10℃ ≤ Temp ≤ 26℃): The base load is reflected by the square relationship (Temp²). When Temp = 20℃, the encoding value is 400. High temperature range (Temp > 26℃): Linear offset term (1.5(Temp−26)) and exponential term (about 67.3) are introduced to capture the sudden change response of air conditioning load. For example, when Temp = 30℃, the encoding value is 1.5×(30−26) + 67.3 = 6 + 67.3 = 73.3. The encoded temperature data is normalized to the [0,1] interval through Min-Max to eliminate the dimension difference.

[0039] Step 2.2, multi-dimensional dynamic attenuation coding of holidays; Holiday impact Holiday_impact consists of three parts: (1) Type Weight: National holidays (such as the Spring Festival) are assigned a value of 1.0, regional holidays (such as temple fairs) are assigned a value of 0.7, and weekdays are assigned a value of 0.

[0040] (2) Time decay factor ( ): , Where d is the number of days until the holiday. For example, if a day is the third day of the Spring Festival (d=0), then β(0)=1; the next day (d=1) β(1)=0.5; the day before (d=−1) β(−1)=0.5.

[0041] (3) Historical comparison coefficient (γ(t)): , Calculate the load y for the day t Compared with the average load in the past three years hist For example, if the current day is May 1, then hist It is the average load from May 1, 2020 to 2024.

[0042] The final encoding is a weighted sum: , Where Y1 represents the type weight; This code quantifies the immediate impact of holidays (60%), the attenuation effect (30%), and the historical pattern (10%). For example, the code value of the first day of the Spring Festival (Y1=1, β(0)=1, γ(t)=1.2) is 0.6×1 + 0.3×1 + 0.1×1.2 = 0.6 +0.3 + 0.12 = 1.02.

[0043] Step 3, constructing a hybrid prediction model; like Figure 2 As shown in the figure, the structure of the hybrid prediction model includes the following core modules: building a hybrid architecture consisting of a TCN module, a bidirectional LSTM module, a spatiotemporal attention mechanism, and a feature fusion layer.

[0044] Step 3.1, multi-source feature data fusion; Concatenate the standardized load data, temperature code, and holiday code along the feature axis to form the input tensor The feature fusion layer dynamically adjusts the weights of each feature through a fully connected network: , in is a learnable weight matrix, and d is the hidden layer dimension (default d=64). This operation maps the original 3D features to a high-dimensional space, enhancing the model's ability to capture feature interactions.

[0045] Step 3.2, design the TCN module; The TCN module contains 3 layers of dilated convolutions, each with 32 convolution kernels, expansion rates of 2, 4, and 8, respectively. The activation function uses ReLU, and the output and input of each layer are added element by element through residual connections, gradually expanding the receptive field to 168 hours (covering a week). The structure of each layer is as follows: (1) Dilated convolution layer: The convolution kernel size is 3, the stride is 1, and the padding method is causal padding, which ensures that the output depends only on the current and historical inputs.

[0046] (2) Activation function: ReLU activation function is used, the formula is: , (3) Residual connection: The output and input of each layer are added element by element to alleviate the gradient vanishing problem: , Step 3.3, bidirectional LSTM module design; The bidirectional LSTM is set with 64 hidden units and the input sequence length is 72 hours (3 days). The forward LSTM extracts positive trends and the backward LSTM captures reverse dependencies. The hidden state update formula is: , , , The output dimension after concatenation is [B, 72, 128] (B is the batch size), where 128=64 (forward) + 64 (backward).

[0047] Step 3.4, spatiotemporal attention mechanism; (1) Temporal attention: Calculate the weight of each time step through Softmax and focus on key time periods (such as morning and evening peaks): , where h t is the LSTM hidden state, is the TCN output, is a learnable parameter, , , , d is the hidden dimension, d=64; T represents transposition; (2) Feature attention: The SE (Squeeze-and-Excitation) module is used to perform channel weighting on temperature and holiday codes: Global Average Pooling (GAP): compresses features along the time dimension: , Fully connected layer: , in , , σ is the Sigmoid function.

[0048] Channel weighting: , in Represents channel-by-channel multiplication.

[0049] Step 3.5, the feature fusion layer is concatenated with the output layer; The feature fusion layer concatenates the local features (dimension [B, T, 32]) output by TCN with the global features (dimension [B, T, 128]) output by LSTM and inputs them into a two-layer fully connected network with 64 neurons and Swish activation function: , The Swish function is defined as: , in ; Output the load forecast results for the next 6 hours, and the Dropout rate is set to 0.2 to prevent overfitting.

[0050] Step 4, executing a phased dynamic training strategy; Step 4.1, dynamically adjust the parameters in stages; Phase 1 (Rounds 1-10): Input length: 6 hours, covering intraday volatility patterns (such as morning peaks and midday troughs).

[0051] Activation function: ReLU, to accelerate initial convergence.

[0052] Learning rate: initial value 1×10 -3 , using a linear warm-up strategy, the first 5 rounds gradually increase from 1×10 -4 Increase to 1×10 -3 .

[0053] Objective: To learn the hourly load fluctuation characteristics and preliminarily fit the short-term impact of temperature and holidays.

[0054] The second stage (rounds 11-20): Input length: extended to 24 hours, covering the full daily cycle; Activation function: Switch to LeakyReLU (α=0.1) to alleviate the problem of neuron death: , Learning rate: reduced to 5×10 -4 , using the cosine decay strategy, decaying to 50% of the initial value every 10 rounds.

[0055] Objective: Optimize medium-period (daily) feature fusion and enhance the model's responsiveness to temperature-sensitive ranges.

[0056] The third stage (rounds 21-30): Input Length: Set to 72 hours, covering a three-day period (including weekend effects).

[0057] Activation function: Swish is used to balance nonlinearity and gradient stability: Learning rate: further adjusted to 1×10 -4 , using an exponential decay strategy: , Goal: Strengthen long-term trend prediction and stabilize attention weight distribution.

[0058] Step 4.2, execute the early stopping mechanism and save the model; Monitor the validation loss. When the validation loss does not decrease for 5 consecutive rounds, terminate the training and roll back to the best weight. The loss function uses MAE. The model weight is saved once per round, and the weight with the smallest validation loss is finally selected for test set evaluation.

[0059] Step 5: Model verification and application deployment; Step 5.1, evaluation index calculation; (1) MAE (mean absolute error): , (2) WMAPE (weighted mean absolute percentage error): , Step 5.2, Interpretability Analysis: (1) Attention weight visualization: Figure 4 As shown in the figure, the weight of time attention during the morning peak (8:00~10:00) and evening peak (18:00~20:00) accounts for more than 60%.

[0060] (2) SHAP value analysis: DeepExplainer is used to calculate the feature contribution, such as Figure 5 shown.

[0061] Step 6: Experimental results and benefit analysis; Step 6.1, prediction accuracy comparison; The MAE of the present invention is 3.33% lower than that of the ARIMA benchmark model; the WMAPE is 1.50% lower than that of the ARIMA benchmark model. Figure 3 shown.

[0062] This invention significantly improves the accuracy and stability of power demand forecasting through multi-scale feature fusion, dynamic coding mechanism and progressive training strategy, while taking into account interpretability and resource efficiency. Experiments show that the prediction error of this method in extreme scenarios is significantly reduced, and the economic benefits are significantly improved, providing reliable technical support for smart grid scheduling. In the future, we can further explore the fusion of multimodal data (such as satellite images and social media) to cope with more complex prediction scenarios.

[0063] The present invention provides a method for dynamic prediction of power demand based on a multi-scale hybrid architecture. There are many methods and approaches to implement the technical solution. The above is only a preferred implementation of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented using existing technologies.

Claims

1. A method for dynamic prediction of power demand based on a multi-scale hybrid architecture, characterized in that: The following steps are involved: Step 1, obtain multi-source data and perform preprocessing: obtain historical power load data, perform outlier detection and dynamic correction, fill missing values, and standardize and normalize the data; construct a time series sample, and divide the historical power load data into training set, validation set, and test set in proportion; Step 2: construct dynamic feature engineering: perform piecewise nonlinear coding of temperature sensitivity, encode the temperature data through piecewise functions to form new data, and normalize it to form temperature feature data; perform multi-dimensional dynamic attenuation coding of holidays, and obtain holiday codes based on holiday impacts; Step 3, constructing a hybrid prediction model; the hybrid prediction model includes a multi-source data feature fusion module, a temporal convolutional network TCN module, a bidirectional long short-term memory network BLSTM module, a spatiotemporal attention mechanism, a feature fusion layer and an output layer; the multi-source data feature fusion module is used to perform multi-source data feature fusion, the temporal convolutional network TCN module is used for hole convolution, the bidirectional long short-term memory network BLSTM module is used to output the spliced ​​hidden state, the spatiotemporal attention mechanism is used for time attention and feature attention calculation, the feature fusion layer is used to complete feature splicing, and the output layer is used to output the power load prediction result; Step 4: Execute a phased dynamic training strategy: perform hybrid model training based on the training set in each phase, monitor the validation loss, calculate the mean absolute error (MAE) of the validation set after each round of training, and select the weight with the smallest MAE for the test set evaluation; Step 5, verify the model and deploy and apply the model: Based on the weights obtained in step 4, calculate the mean absolute error MAE and the weighted mean absolute percentage error WMAPE based on the test set, complete the model verification, and deploy and apply the hybrid prediction model.

2. The method according to claim 1, characterized in that Step 1 includes: Step 1.1, obtain historical power load data with a time resolution of 1 hour. The data fields include load value, area code and timestamp; Get hourly meteorological temperature data; Obtain and integrate statutory holidays and regional holidays to form a structured holiday time label table; Step 1.2, outlier detection and dynamic correction; The sliding window quantile detection method is used to dynamically correct abnormal power load values ​​and dynamically calculate the 0.05 quantile at each time point. and 0.95 quantile If the power load value on that day is Beyond , ] range, then Replace with the median within the window , the calculation formula is: , in It represents the corrected power load value, which is used to replace the abnormal value; Calculate the window mean using the moving average algorithm , the weight of each point in the window is allocated according to exponential decay, and the weight of the i-th point in the window for: , in is an intermediate parameter, e is a natural constant; Window mean The calculation formula is: ; Step 1.3, missing value filling strategy; For missing data, the cubic spline interpolation algorithm with natural boundary conditions is used. , define a cubic polynomial : , in ; , , , are cubic polynomial coefficients, is the timestamp of the known data point; solve the spline coefficients through continuity and boundary conditions , , , , and finally obtain the cubic polynomial on each interval; Use the mean of the physically adjacent areas to fill in the missing values ​​in the spatial dimension: If the data of area A1 is missing at a certain moment, the mean of the power load of the physically adjacent areas A2, A3, and A4 is taken as the filling value: ; in It represents the missing value filling result of area A1. represents the average power load in area A2, represents the average power load in area A3, represents the average power load of area A4; Step 1.4, data standardization and normalization: After processing, the power load data Apply robust standardization RobustScaler to standardize and obtain standard data , the calculation formula is: , Where median( ) is the processed power load data set The median, IQR( ) is the interquartile range, and the calculation formula is: IQR( )= - , in represents the 0.75 quantile at each time point, represents the 0.25 quantile at each time point; Step 1.5, construct time series samples; The historical power load data is cut into samples of 720 consecutive hours and a sliding step of 24 hours to form an input sequence , where N is the number of samples and R represents the real number space; When cutting samples, the time sliding window method is used to ensure the temporal continuity of the sequence; Finally, the historical power load data is divided into training set, validation set and test set in proportion.

3. The method according to claim 2, characterized in that Step 2 includes: Step 2.1, temperature sensitivity piecewise nonlinear coding; Define 10℃ and 26℃ as key inflection points and construct a piecewise nonlinear function of temperature sensitivity : , Temperature data Temp through piecewise function After encoding, the data is formed , and then further normalized to the [0,1] interval to form temperature characteristic data , the formula is: ; Step 2.2, multi-dimensional dynamic attenuation encoding of holidays.

4. The method according to claim 3, characterized in that Step 2.2 includes: Holiday impact includes type weight and time decay factor and historical comparison coefficient γ(t); Time decay factor The calculation formula is: , Where d is the number of days until the holiday; The calculation formula of historical comparison coefficient γ(t) is: , in It indicates the average power load in the same period of the past three years; The holiday code is calculated using the following formula : , Where Y1 represents the type weight.

5. The method according to claim 4, characterized in that Step 3 includes the following steps: Step 3.1, the multi-source data feature fusion module performs the following operations: The standardized power load data , Temperature coding and holiday codes Concatenate along the feature axis to form the input tensor , the feature fusion layer dynamically adjusts the weights of each feature through a fully connected network to achieve scene adaptive coupling of multi-source signals. The formula is: , in, is the hidden representation after feature fusion, is the learnable weight matrix, D is the hidden layer dimension, is the bias term of the fully connected layer, and ReLU is the activation function; Step 3.2, establish a temporal convolutional network TCN module; The temporal convolutional network TCN module includes 3 layers of dilated convolutions. Each layer of dilated convolutions has the same structure, including a dilated convolution layer, a ReLU activation function, and a residual connection. The number of convolution kernels in each dilated convolution layer is 32, the convolution kernel size is 3, the step size is 1, and the filling method is causal convolution. The expansion rates of the three dilated convolution layers are 2, 4, and 8 respectively. The output and input of each dilated convolution layer are added element by element through a residual connection, and the receptive field is gradually expanded to 168 hours. The residual connection formula is: , The input sequence of the dilated convolution is the tensor of the normalized load data, temperature code, and holiday code concatenation. , the dimension is [B,720,3], B is the batch size; For the Input of the layer dilated convolution; No. The output of the layer dilated convolution, the final output dimension is [B, 720, 32]; Conv1D is a one-dimensional dilated convolution operation; Step 3.3, establish a bidirectional long short-term memory network BLSTM module; The bidirectional long short-term memory network BLSTM module includes a forward long short-term memory module and a backward long short-term memory module. The bidirectional long short-term memory network BLSTM module is provided with 64 hidden units. The temporal convolution network TCN module outputs The last 72 hours are taken as the input of the forward long short-term memory module. The forward long short-term memory module extracts positive trends, and the backward long short-term memory module captures reverse dependencies. The hidden state update formula is: , , , Where t is the time step, h t is the long short-term memory hidden state, is the forward long short-term memory hidden state, is the backward long short-term memory hidden state, and LSTM is the model function; The bidirectional long short-term memory network BLSTM module outputs a hidden state with a concatenated dimension of [B, 72, 128]; Step 3.4, establish the spatiotemporal attention mechanism, including temporal attention and feature attention; The temporal attention calculates the weight of each time step through the Softmax function: , where w t represents the weight of time step t, is the output of the temporal convolutional network TCN module at time step t, v, , is a learnable parameter, ; is the key matrix of the attention mechanism, is the Value matrix; TP means transpose; Softmax is the normalized exponential function; tanh is the hyperbolic tangent function; The feature attention uses a compression excitation SE module to perform channel weighting on temperature and holiday codes. The compression excitation SE module performs the following steps: Step 3.4.1, use global average pooling to compress features along the time dimension to obtain the pooled global feature vector s: , in, is the input feature at time step t, with dimension [B, T, C]; Step 3.4.2, through the two-layer fully connected network, get the channel weight vector : , in, is the weight matrix of the first fully connected layer, is the bias term of the first fully connected layer, is the weight matrix of the second fully connected layer, The bias term of the second fully connected layer, ReLU is the activation function, and σ is the Sigmoid function; Step 3.4.3, perform channel weighting to obtain weighted output features : , in, , x is the original input feature, represents channel-by-channel multiplication; Step 3.5, the feature fusion layer is concatenated with the output layer.

6. The method according to claim 5, characterized in that Step 3.5 includes: implementing a feature fusion layer to transform the local features output by the temporal convolutional network TCN module The global features output by the bidirectional long short-term memory network BLSTM module Concatenate and input into a two-layer fully connected network with 64 neurons and Swish activation function: , in, is the output layer weight, is the weight matrix of the fully connected network, is the bias term of the hidden layer in the feature fusion layer, is the bias term of the output layer, The output layer outputs the power load forecast results for the next 6 hours.

7. The method according to claim 6, characterized in that Step 4 includes: Step 4.1: In each stage, 10 rounds of hybrid model training are performed based on the training set. Each round includes a complete traversal of all training samples, and the parameters are dynamically adjusted in stages. The first stage is rounds 1 to Y1: the model is fully trained 10 times using the training set, and all training data are traversed each time; Input length: 6 hours; Activation function: ReLU function; Learning rate: initial value 1×10 -3 , using a linear warm-up strategy, the first 5 rounds gradually increase from 1×10 -4 Increase to 1×10 -3 ; Objective: Learn the hourly load fluctuation characteristics and suppress long-term noise interference; The second stage is 11-20 rounds: continue training for 10 rounds and adjust the input length and hyperparameters; Input length: extended to 24 hours; Activation function: switch to LeakyReLU function; Learning rate: reduced to 5×10 -4 , using the cosine decay strategy, decaying to 50% of the initial value every 10 rounds; Objective: Optimize day-level feature fusion; The third stage is 21~30 rounds: 10 more rounds of training to further optimize long-term trend predictions; Input length: set to 72 hours; Activation function: Swish function is used; Learning rate: adjusted to 1×10 -4 , using an exponential decay strategy to calculate the dynamically adjusted learning rate , which gradually decreases with the increase of training rounds: , in, is the initial learning rate, decay_rate is the decay rate, is the training round, decay_steps is the decay step size; Goal: Strengthen long-term trend prediction and stabilize attention weight distribution; Step 4.2, execute the early stopping mechanism and save the model: monitor the validation loss, calculate the mean absolute error (MAE) of the validation set after each round of training, and terminate the training and roll back to the historical best weight if the validation loss does not decrease for 5 consecutive rounds; save the model weight once per round, and finally select the weight with the smallest mean absolute error (MAE) for the test set evaluation; isolate the test set throughout the process and use it only in the final evaluation.

8. The method according to claim 7, characterized in that Step 5 includes: Step 5.1, load the best weights saved in step 4.2, and calculate the following indicators based on the test set: Mean absolute error (MAE): , Weighted Mean Absolute Percent Error WMAPE: , in, is the actual load value; is the model prediction value; Step 5.2, perform interpretability analysis; Step 5.3: Lightweight deployment of applications.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.

Citation Information

Patent Citations

  • Solar photovoltaic power generation prediction method based on TCN-LSTM

    CN110909926A

  • Prophet-LSTNet combination model-based power supply cost analysis and prediction method

    CN113256036A

  • Urban rail transit passenger flow prediction method based on similar day selection

    CN115169638A

  • Time sequence prediction method applied to optical storage and charging field

    CN118395255A

  • Household electricity short-term load prediction method and system

    CN118586559A

Cited By

  • Method for measuring effluent ammonia nitrogen concentration based on neural network and Bayesian optimization

    CN120144969A

  • Flight height prediction method based on neural network

    CN120430348A

  • Flight altitude prediction method based on neural network

    CN120430348B

  • Regional power utilization scheduling method and system for power equipment

    CN120433206A

  • A regional power dispatching method and system for power equipment

    CN120433206B