NST-IRN combined prediction model based on time feature analysis

By constructing an NST-IRN combined forecasting model, and combining time feature analysis, improved residual neural networks, and DS evidence theory, the robustness and uncertainty problems of existing short-term load forecasting methods are solved, and higher accuracy power load forecasting is achieved.

CN120911645APending Publication Date: 2025-11-07XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510698323.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing short-term load forecasting methods are not robust enough when dealing with non-adaptive regular data, and the prediction results of a single forecasting model are uncertain and one-sided, making it difficult to accurately predict changes in power load.

Method used

A combined NST-IRN prediction model based on time feature analysis is constructed. By deeply exploring the trend and periodicity of load power, and combining an improved residual neural network model and DS evidence theory for weight fusion, the advantages of the two models are comprehensively utilized to improve prediction accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and robustness of short-term load forecasting, and can more accurately capture the multi-timescale characteristics and daily type effects of load power, reducing the uncertainty of single-model forecasting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911645A_ABST
    Figure CN120911645A_ABST
Patent Text Reader

Abstract

The invention discloses an NST-IRN combined prediction model based on time feature analysis. The method comprises the following steps: firstly, deeply mining the time characteristic change of load power, decomposing the time characteristic change into a trend component and a cyclic component, and constructing a new time series (NTS) model; secondly, considering the influence of a multi-time scale input feature and a day type on the load power, and constructing an improved residual neural network (IRN) model containing a feature input structure and a deep learning structure; and finally, performing weight fusion on prediction results of the NTS model and the IRN model by using a D-S evidence theory to obtain a final load prediction result. A simulation experiment is carried out by using real load data of ISO New Engine, and a result shows that the provided model has relatively high prediction precision and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power supply load prediction, and particularly relates to an NST-IRN combined prediction model based on time feature analysis. BACKGROUND

[0002] Load prediction is of great significance to promote power supply and demand balance, ensure power supply safety and promote the refinement of power dispatch. Short-term load prediction can flexibly adapt to the fluctuations of power demand, effectively avoid power shortage and power curtailment, thereby optimizing resource allocation and promoting the safe and economic operation of the power grid, and is a key technical means to ensure stable power supply. Common short-term load prediction methods mainly include time series prediction methods and machine learning prediction methods.

[0003] Time series prediction methods include grey model method, exponential smoothing method, cumulative autoregressive moving average method, etc. Time series prediction methods have fast calculation speed, but perform poorly when dealing with non-adaptive regular data, lack robustness, are easily disturbed by external factors, and often need to be combined with other algorithms or models to enhance the effect. In addition, the above methods have shortcomings in revealing the trend and periodic characteristics of load power.

[0004] Machine learning prediction method is the current trend in the field of load prediction, including BP neural network, Transformer model, long short-term memory network, residual neural network, etc. The short-term power load prediction based on multi-branch gated residual convolutional neural network by Fan Jiangchuan, Yu Haozheng, Liu Huiting, etc. uses a multi-branch gated residual convolutional neural network to extract multi-time scale features of historical load, and uses an attention mechanism to reasonably allocate weights to calculate the prediction result, but does not consider the influence of day type and season on load power, and the neural network structure is relatively simple, limiting the improvement of its learning ability.

[0005] Since time series and machine learning prediction methods have their own unique advantages and limitations, and the prediction results of a single prediction model are uncertain and one-sided, the advantages of the two prediction methods can be combined by using a combination prediction method. Common combination algorithm methods include weighted average, Bayesian model, and D-S evidence theory, etc. The existing weighted average algorithm constructs a weight matrix for weighting according to the prediction error covariance, which is simple to operate and efficient to calculate, but does not consider the weight allocation deviation caused by the difference in data quality, which affects the prediction accuracy. The support vector machine prediction method based on Bayesian fusion predicts the power by fusing power, numerical weather prediction and remote sensing data, and the Bayesian processing has strong ability to deal with uncertain and incomplete data, high prediction accuracy and good robustness, but the calculation complexity is high, and sufficient data and reasonable prior knowledge are needed to construct the probability model. The short-term wind power prediction model based on D-S evidence theory takes the weight of a single model as evidence, constructs a combination model by multiple fusion of the confidence function, has unique advantages in dealing with incomplete information and conflicting information, effectively integrates information by a reasonable conflict allocation strategy, and improves the robustness of information fusion and prediction accuracy.

[0006] In view of the above problems, the present application provides a NST-IRN combination prediction model based on time feature analysis. A new time series (NTS) model is constructed to deeply mine the trend and periodicity of load power, an improved residual neural network (IRN) model is used to deeply study the influence of multi-time scale and day type on load power, and D-S evidence theory is used to weight the prediction results of the two models, so that the advantages of the two models are comprehensively utilized. The effectiveness of the method is verified by experiments on ISO New England real load data. SUMMARY

[0007] To solve the problems existing in the prior art, the purpose of the present application is to provide a NST-IRN combination prediction model based on time feature analysis.

[0008] To achieve the above purpose, the technical scheme of the present application is:

[0009] A NST-IRN combined prediction model based on time characteristic analysis, firstly, deeply mining the time characteristic change of load power, decomposing it into trend component and cycle component, constructing a new time series (NTS) model; secondly, considering the influence of multi-time scale input features and day type on load power, constructing an improved residual neural network (IRN) model containing feature input structure and deep learning structure; finally, using D-S evidence theory to weight fuse the prediction results of NTS and IRN models to obtain the final load prediction result; the real load data of ISO New England is used for simulation experiment, and the results show that the proposed model has high prediction accuracy and robustness.

[0010] Further, a NST-IRN combined prediction model based on time characteristic analysis, composed of the following steps:

[0011] Step one, construction of the combined prediction model

[0012] Step 1.1 Construction of a new time series model

[0013] Considering the inherent time characteristics of load power, the NTS model is constructed, and the load power can be decomposed into trend, cycle and random error components; let denote the load power at time t, and its additive model is:

[0014]

[0015] In the formula: trend component at time t, MW; cycle component at time t, MW; random error component at time t, MW; the first two components are deterministic, and the random error component contains changes that cannot be explained by the prediction model; the parameter model of the two deterministic parts and are established respectively as trend model and cycle model;

[0016] The structure and important parameters of the trend model are determined by ordinary least squares (OLS) using the training data set, and the trend after removing the trend from the training set is obtained, and the cycle model structure and its important parameters are determined again using OLS; in the sliding window, the test set is used to and The coefficients in the structure are subjected to dummy variable regression and combined with input features such as temperature and day type to form a complete NTS prediction model and predict the load power at the next moment. The sliding window is moved and the above steps are repeated to make predictions for the next moment.

[0017] 1.1.1 Estimation of Trend Components

[0018] Trend function Represents load power The slow changes in electricity consumption, time, temperature, and day type are important factors explaining the changes in electricity consumption. Day type is represented by a dummy variable, and a trend function with a linear correlation between time, temperature, and the dummy variable is proposed:

[0019]

[0020] Where: t—time exponent from 1 to the window size, h; T t —Temperature value at time t, in °C; , —Dummy variables, with values ​​between 0 and 1, indicate whether a corresponding regression exists; —The number of different categories; parameters a, b, c and —The coefficients of the regression model are evaluated using linear regression with OLS;

[0021] 1.1.2 Cyclic Component Estimation

[0022] Circulating section captures load power The cyclic behavior; let p represent the periodicity of the trendless function; it can write all samples: The average value over the entire period is zero. Let the detrending sequence be:

[0023]

[0024] (1) Stability test

[0025] Since the sum of the sine and cosine functions can completely describe the dynamics of any stationary periodic signal, Fourier component regression is used to estimate the cyclic component; this is to determine the detrended sequence. To determine stationarity, a combined Kwiatkowski-Phillips-Schmidt-Shin (KPSS) and Augmented Dickey-Fuller (ADF) test is used. If both stationarity tests confirm stationarity... The stationarity of the sliding window of any size indicates that the designed trend model is sufficient to represent the load power The dynamics of slow variation can be estimated by

[0026] (2) Smoothed periodogram

[0027] The smoothed periodogram can find the main harmonic frequencies of the cyclic components, which represent the fast variation of the data. Let n be the size of the stationary data set, the periodogram is defined as:

[0028]

[0029] where: — the detrended series — the fundamental frequency of the discrete Fourier transform (DFT) of the detrended series; — the normalized real variable of the DFT; — the normalized imaginary variable of the DFT;

[0030]

[0031]

[0032] where: — the mean of the detrended series

[0033] Equation (6) shows that the sum of squares can be split into the harmonic components represented by the amplitudes of the periodogram at frequencies The frequencies corresponding to the peaks in a periodogram can explain the variation of the data, i.e. in a periodogram, the peak values represent the harmonic components;

[0034] The periodogram has a large variance at given frequencies, the solution is to use the smoothed periodogram; assume that the spectral density is fairly constant within the considered frequency band, and that adjacent frequencies have asymptotically independent values, and for any sample size n, the periodogram is the sum of squares of real and imaginary variables; a smoothing filter of size L = 2m + 1 « n, centered at frequency B:

[0035]

[0036] The Daniell kernel is a symmetric positive weight, centered at the estimated frequency, and the sum of all weights is 1;

[0037] ​​​

[0038] The smoothed periodic chart is as follows:

[0039]

[0040] (3) Cyclic model with Fourier components

[0041] Based on frequency set The cyclic component of the set is:

[0042]

[0043] Where: parameters , , and —The coefficients of the regression model are evaluated using linear regression with ordinary least squares (OLS) method;

[0044] 1.1.3 NTS Prediction Model

[0045] The NTS forecasting model consists of a trend estimation model and a cyclic component model. The complete NTS forecasting model can be expressed as:

[0046]

[0047] Step 1.2 Construction of the Improved Residual Neural Network Model

[0048] 1.2.1 RN Model Structure

[0049] A residual neural network is composed of multiple residual blocks stacked together. A residual block can be represented as:

[0050]

[0051] In the formula: —x to Mapping; —A set of weights and biases associated with the residual block. The presence of the residual block allows the input to bypass some of the neural network, making the gradient easier to propagate downwards and reducing the problem of gradient explosion.

[0052] The activation function for the residual block is SELU, and its expression is:

[0053]

[0054] In the formula: λ, α — two adjustable parameters; when and , and the output of the full connection layer will be close to the standard normal distribution when the input follows the standard normal distribution, which helps the network to prevent the problems of gradient vanishing and explosion; all full connection layers of the residual neural network designed by the application adopt SELU as the activation function;

[0055] Loss function of RN model Error from prediction Hyper-range penalty term for accelerating the training process ;

[0056]

[0057]

[0058]

[0059] In the formula: Model output of the hth hour, MW; Actual normalized load, MW; N-number of data samples; H-number of loads per hour in a day (H=24), h;

[0060] When the predicted daily load is out of the range of the actual load curve, The model is penalized, prompting the model to converge faster to a more accurate prediction range; when the model exhibits high prediction accuracy, The accuracy of peak and trough prediction will be emphasized more;

[0061] 1.2.2 IRN model structure

[0062] In order to make the residual neural network fully learn the time characteristics such as multi-time scale and day type, the IRN model increases the multi-time scale feature input structure and the deep learning structure to improve the learning ability on the basis of the residual neural network model, analyzes the time characteristics of load power, and further improves the model prediction performance;

[0063] (1) Feature input structure

[0064] In order to further analyze the time characteristics of load power, the load power of different time scales is selected as the input feature, and the load power data of different time scales and the time-related one-hot code are fused to enhance the ability of the model to capture load dynamics and abnormal patterns, and the input features are as follows:

[0065]

[0066] Connect, and pass through a fully connected layer, the second layer is denoted as ; Connect , , Connect, and connect with three independent fully connected layers, then connect the three fully connected layers to another fully connected layer ; Connect S, W and H to produce two identical fully connected layers as the part of the input of and ; Connect and with fully connected layers, and connect to the output layer of the input structure or the input layer of the IRN model; Replace with , the corresponding value, so as to realize the feature input at continuous time points;

[0067] Through the designed feature input structure, the IRN model can analyze the input features in a hierarchical manner, so that the residual neural network can capture the time characteristics of the load power to the greatest extent, reduce information loss, and improve the prediction accuracy;

[0068] (2) Deep learning structure

[0069] The deep learning structure is composed of main residual blocks, bilateral residual blocks and shortcut connections; the output of each main residual block and the output of the bilateral residual block at the same level are averaged through shortcut connection, and the obtained average value is further transmitted to all main residual blocks in the subsequent layer, and is combined with the input layer of the network to form the input of the main residual block;

[0070] By introducing additional bilateral residual blocks and dense shortcut connections, the representation ability of the deep learning network can be significantly improved, and the complex features in the data can be effectively captured. At the same time, this structure optimizes the error backpropagation path, improves the efficiency and stability of the training process, and thus helps to improve the prediction accuracy and generalization ability of the model;

[0071] Step 1.3 Weight fusion of D-S evidence theory

[0072] When combining models for weight fusion, models exhibiting higher prediction accuracy are assigned larger weights, while models with lower prediction accuracy receive smaller weights. This weight allocation strategy is similar to the basic credibility allocation mechanism in Dempster's evidence theory. The prediction results of a single model and its corresponding weights can be regarded as "evidence" of the true load value, while the model weights represent the credibility of the model's prediction results. Therefore, the Dempster composition rule of evidence theory can be used to fuse weights, making full use of the advantages of each model, reducing the uncertainty of single model predictions, and thus improving overall prediction accuracy. When the prediction results are highly correlated, there is a problem of reduced effectiveness of Dempster's theory. To solve this problem, different model structures and different feature sets are required. The prediction models selected in this invention meet the above requirements and introduce a weight mechanism to reflect their reliability and importance.

[0073] 1.3.1 Model Weight Allocation Method

[0074] The variance-covariance weighting method is used to allocate weights for a single prediction model. Assuming there are m prediction models and n predictions are made, the actual value of the j-th (j=1,2,...,m) prediction for the i-th (i=1,2,...,m) model is... The predicted value is The prediction error is The model's prediction error variance is Combined prediction results Combined prediction error for:

[0075]

[0076] In the formula: —The weights of the i-th prediction model;

[0077] Typically, the prediction data for the same object are independent of each other, with all covariances being 0. The variance of the combined prediction result is:

[0078]

[0079] Introducing Lagrange pairs Finding the minimum value, we can obtain the model weights as follows:

[0080]

[0081] Model weights Similar to the basic belief value in evidence theory, the weight fusion can be carried out by using Dempster combination rule to obtain the combined model load power prediction value at the next moment;

[0082] 1.3.2 D-S combination model prediction process

[0083] The application adopts the variance covariance weight method and the D-S evidence theory to distribute and fuse the prediction results of the NTS and the IRN model, and a flow chart is as shown in Figure 6 , and the specific process is as follows:

[0084] Step 1: training the NTS and IRN prediction models by using the training set, predicting the load power on the two models by using the test set, obtaining the prediction values of the two prediction models and the prediction errors;

[0085] Step 2: determining the prediction time length [t0, t f ] and the prediction moment t of the D-S evidence theory combination model, t starting from the starting moment t0 and t f ending at the next moment of t

[0086] Step 3: obtaining the prediction value at t moment and the prediction value and prediction error of the previous N hours of t moment, and calculating the combination weight of the two prediction models, and the fusion effect is best when N=3 after multiple comparison experiments;

[0087] Step 4: constructing the basic belief distribution on the recognition framework, fusing the weights by using the Dempster rule, obtaining the weight distribution of the two prediction models at t moment, and calculating the prediction value of the D-S evidence theory combination model at t moment according to the prediction values of the two models at t moment obtained in Step 2;

[0088] Step 5: advancing 1 h from the prediction moment t, judging whether the moment t exceeds the prediction time length, if not, jumping to Step 3, otherwise ending the prediction process;

[0089] Step 2 simulation example

[0090] 2.1 Data source

[0091] The application uses the power load data provided by ISO New England from January 1, 2023 to December 31, 2023, with a resolution of 1 h, and uses the load data from January 1 to September 30 as the training set, and uses the load data from October 1 to December 31, 2023 as the test set, for the convenience of observing the prediction results, only the prediction results from October 11 to 20 in the test set are selected to draw curves, and the calculation of the evaluation index uses the complete test set;

[0092] 2.2 Model parameter selection

[0093] 2.2.1 NTS model parameters

[0094] In the process of constructing the NTS model, it is particularly important to determine the size of the sliding window of the trend component, the detrended order stationarity of the cyclic component estimation and the main harmonic frequency; the size of the sliding window is selected by the MAE standard of the trend prediction results based on different window sizes, and the minimum unit is one week, in which different day types can be found;

[0095] The optimal duration of the sliding window size is 3 weeks, once the size of the sliding window is determined, the detrended order is calculated according to formula (3) , and KPSS test and ADF test are performed, the test statistics of KPSS and ADF test are less than the critical value at 10%, 5% and 1% significant level, and the P value is less than 0.1, proving is stationary; the smooth periodogram using Daniell kernel (7,7) filter is plotted on it, as shown in formula (20), which represents the periods of 24 h, 12 h, 6 h, 5 h, etc.;

[0096]

[0097] 2.2.2 IRN model parameters

[0098] In the feature input structure of the IRN model, , , and The fully connected layer has 10 hidden nodes, while The fully connected layer has 5 hidden nodes, , and the fully connected layer before it has 10 hidden nodes; in the deep learning structure, each residual block has a hidden layer with 20 hidden nodes, the input and output size of the residual block is 24, the residual block is divided into main residual block and bilateral residual block, and the average value is taken by using shortcut connection, the main and bilateral each stack 30 residual blocks, forming a left-middle-right three-side 60-layer residual network.

[0099] Compared with the prior art, the beneficial effects of the present application are:

[0100] A NST-IRN combined prediction model based on time feature analysis,

[0101] (1) The present application deeply mines the time feature of load power change, and constructs the NTS prediction model by decomposing the trend and cyclic component according to time, compared with the ARIMA prediction model, the prediction performance of the model is improved;

[0102] (2) In order to perceive the time characteristics of load power data on different time scales, based on the RN model, the feature input structure and deep learning structure are introduced, and the IRN prediction model is proposed. The experimental results show that the model can capture the time characteristics of multi-time scale input features, reduce the loss of information, and further improve the prediction accuracy;

[0103] (3) In view of the uncertainty and one-sidedness of the prediction results of single prediction model, the D-S combination prediction model is proposed. The model integrates the prediction results of NTS and IRN, effectively reduces the uncertainty of the prediction results of single model, and the experiment shows that the D-S combination prediction model significantly improves the overall prediction accuracy and robustness. BRIEF DESCRIPTION OF DRAWINGS

[0104] Figure 1 Combination prediction model framework

[0105] Figure 2 NTS prediction model construction process

[0106] Figure 3 Residual neural network structure

[0107] Figure 4 Feature input structure of IRN model

[0108] Figure 5 Deep learning structure of IRN model

[0109] Figure 6 D-S combination model load power prediction process

[0110] Figure 7 MAE standard calculated according to different sliding window size

[0111] Figure 8 Best trend curve of NTS

[0112] Figure 9 Smooth periodogram of detrended series

[0113] Figure 10 Prediction curve of NTS and ARIMA model

[0114] Figure 11 Prediction curve of RN and LSTM, BP, Transformer model

[0115] Figure 12 Prediction curve of IRN and RN_input, RN model

[0116] Figure 13 Prediction curve of three combination prediction models and NTS, IRN model. DETAILED DESCRIPTION

[0117] The technical solutions of the present application will be described in further detail below in combination with the drawings and specific embodiments:

[0118] As shown in the specific embodiments of the present application, Figures 1-13 a NST-IRN combined prediction model based on time characteristic analysis,

[0119] 1 Load prediction idea

[0120] The combined prediction model framework proposed by the present application is shown in the specific embodiments of the present application, Figure 1 The load power has strong time characteristics, and the NTS model is constructed by decomposing the trend and cycle components of the load power according to time; considering the influence of multiple time scales and day types and other time characteristics, the IRN model is constructed to further analyze the time characteristics of the load power; since the prediction results of a single prediction model have uncertainty and one-sidedness, the D-S combined prediction model is constructed to make full use of the advantages of the two models, improve the prediction accuracy, and enhance the robustness of the model.

[0121] 1) Preprocessing of load data, including screening of abnormal load data, replacing missing and abnormal values with historical data or mean value, data normalization, etc.

[0122] 2) Introducing time as an explanatory variable, decomposing the load data into trend components and cycle components and performing parameter estimation, constructing an NTS model for load power training and prediction.

[0123] 3) Selecting load power of different time scales as input features, constructing a multi-dimensional time series feature input structure and an IRN prediction of deep learning structure, fully analyzing the time characteristics of the load power and performing load power prediction.

[0124] 4) Using D-S evidence theory to weight fuse the prediction results of the two prediction models, complementing the advantages of both sides and further improving the prediction accuracy.

[0125] 2 Construction of combined prediction model

[0126] 2.1 Construction of new time series model

[0127] Considering the inherent time characteristics of the load power, an NTS model is constructed, and the load power can be decomposed into trend, cycle and random error components. Let denote the load power at time t, and its additive model is:

[0128]

[0129] In the formula, trend component at time t, MW; — the cyclical component at time t, MW; — the random error component at time t, MW. The first two components are deterministic, while the random error component contains the variation that the prediction model cannot explain. The parameter model of the two deterministic parts and are called the trend model and the cycle model, respectively. The NTS prediction model construction process is shown in Figure 2 .

[0130] As shown in Figure 2 , the structure of the trend model and its important parameters are determined by ordinary least squares (OLS) using the training data set, and the trend is removed from the training set to obtain the detrended sequence, and the cycle model structure and its important parameters are determined again using OLS; within the sliding window, use the test set to perform dummy variable regression on the coefficients in and , and combine them with temperature, day type, and other input features to form a complete NTS prediction model and predict the next time load power. Move the sliding window and repeat the above steps to predict the next time.

[0131] 2.1.1 Trend component estimation

[0132] The trend function represents the slow changes in load power , time, temperature, and day type are important factors to explain the change in electricity consumption, and day type is represented by a dummy variable, so a trend function is proposed that is linearly related to time, temperature, and dummy variables:

[0133]

[0134] In the formula: t—time index from 1 to window size, h; T t — temperature value at time t, ℃; , — dummy variable, value between 0 and 1, indicating whether the corresponding regression exists; — the number of different categories; parameters a, b, c, and — regression model coefficients, linear regression using OLS to evaluate the coefficients.

[0135] 2.1.2 Cycle component estimation

[0136] The cycle part captures the cyclical behavior of load power . Let p represent the periodicity without the trend function. It can be written for all samples: The average value over the entire period is zero. Let the detrending sequence be:

[0137]

[0138] (1) Stability test

[0139] Since the sum of the sine and cosine functions can completely describe the dynamics of any stationary periodic signal, Fourier component regression is used to estimate the cyclic component. This is to determine the detrended sequence. To determine stationarity, a combined Kwiatkowski-Phillips-Schmidt-Shin (KPSS) and Augmented Dickey-Fuller (ADF) test is used. If both stationarity tests confirm stationarity... If the trend model remains stable across any sliding window size, it indicates that the designed trend model is sufficient to represent the load power. Slowly changing dynamics can be used to... Perform cyclic component estimation.

[0140] (2) Smooth periodic chart

[0141] A smooth periodogram can identify the dominant harmonic frequencies of cyclic components, representing rapid changes in the data. Let n be the size of the stationary dataset; this periodogram is defined as:

[0142]

[0143] In the formula: --Detrending sequence The fundamental frequency of the Discrete Fourier Transform (DFT); —Normalized real variables of the DFT; —Normalized dummy variables of the DFT.

[0144]

[0145]

[0146] In the formula: ; --Detrending sequence The average value.

[0147] Formula (6) indicates that the sum of squares can be divided into frequencies of The amplitude of the harmonic components of a periodogram, i.e. the frequencies corresponding to the peaks in a periodogram, can explain the variation in the data.

[0148] Periodograms typically have large variance at given frequencies, and the solution is to use a smoothed periodogram. Assuming that the spectral density is fairly constant over the frequency band considered, and that adjacent frequencies have asymptotically independent values, and that for any sample size n, the periodogram is the sum of the squares of real and imaginary variables. A smoothing filter of size L = 2m + 1 « n, centered at frequency B:

[0149]

[0150] The Daniell kernel is a symmetric positive weight centered at the estimated frequency, with all weights summing to 1.

[0151]

[0152] The smoothed periodogram is:

[0153]

[0154] (3) Cyclical model with Fourier components

[0155] Based on a set of frequencies The cyclical components of the set are:

[0156]

[0157] where: the parameters , , and are the coefficients of the regression model, evaluated using ordinary least squares (OLS) method for linear regression.

[0158] 2.1.3 NTS prediction model

[0159] As shown in Figure 2 , the NTS prediction model is composed of a trend estimation model and a cyclical component model, and the complete NTS prediction model can be expressed as:

[0160]

[0161] 2.2 Construction of improved residual neural network model

[0162] 2.2.1 RN model structure

[0163] AsFigure 3 As shown, the residual neural network is composed of multiple residual blocks stacked together, and a residual block can be represented as:

[0164]

[0165] In the formula: The mapping of x to in the neural network; A set of weights and biases associated with the residual block, the presence of the residual block allows the input to bypass some of the neural network, making it easier for the gradient to pass down, reducing the problem of gradient explosion.

[0166] The activation function of the residual block is selected as SELU, and its expression is:

[0167]

[0168] In the formula: λ, α - two adjustable parameters. When and , and the input follows the standard normal distribution, the output of the full connection layer will be close to the standard normal distribution, which helps the network to prevent the problem of gradient disappearance and explosion. All full connection layers of the residual neural network designed in the application use SELU as the activation function.

[0169] The loss function of the RN model is composed of the error of the prediction and the hyper-range penalty term for accelerating the training process .

[0170]

[0171]

[0172]

[0173] In the formula: The output of the model in the h hour, MW; The actual normalized load, MW; N - the number of data samples; H - the number of loads per hour in a day (H = 24), h.

[0174] When the predicted daily load exceeds the range of the actual load curve, the model is penalized, prompting the model to converge faster to a more accurate prediction range. When the model exhibits high prediction accuracy, the accuracy of peak and trough prediction will be emphasized.

[0175] 2.2.2 IRN Model Structure

[0176] To enable the residual neural network to fully learn time features such as multiple time scales and day types, a multi-time scale feature input structure and a deep learning structure to improve learning ability are added to the IRN model based on the residual neural network model. This further improves the model's predictive performance while analyzing the time characteristics of load power.

[0177] (1) Feature input structure

[0178] To further analyze the temporal characteristics of load power, load power at different time scales was selected as input features. By fusing load power data at different time scales and time-dependent one-hot codes, the model's ability to capture load dynamics and abnormal patterns was enhanced. The input features are shown in Table 1, and the input structure of the IRN model is as follows. Figure 4 As shown.

[0179] Table 1 Input characteristics of load forecasting

[0180]

[0181] like Figure 4 It can be seen that, Connect and move forward through a fully connected layer, the second layer being denoted as... ;Will , , Connect the layers, and then connect them using three separate fully connected layers, and then connect the three fully connected layers to another fully connected layer. Connecting S, W, and H creates two identical fully connected layers, which serve as... and Partial input. (The rest of the text is missing.) and Connect them using fully connected layers, and then connect them to the output layer of the input structure. Or the input layer of an IRN model. The corresponding value is replaced with This allows for feature input at continuous time points.

[0182] Through the designed feature input structure, the IRN model can analyze the input features hierarchically, enabling the residual neural network to capture the temporal characteristics of load power to the greatest extent, reduce information loss, and improve prediction accuracy.

[0183] (2) Deep learning structure

[0184] Deep learning architectures such as Figure 5The deep learning structure is composed of main residual blocks, bilateral residual blocks and shortcut connections. The output of each main residual block and the output of the bilateral residual block at the same level are averaged by shortcut connection, and the obtained average value is further transmitted to all main residual blocks in the subsequent layer, and is combined with the input layer of the network to form the input of the main residual block.

[0185] By introducing additional bilateral residual blocks and dense shortcut connections, the representation ability of the deep learning network can be significantly improved, effectively capturing complex features in the data. At the same time, this structure optimizes the error backpropagation path, improves the efficiency and stability of the training process, and thus helps to improve the prediction accuracy and generalization ability of the model.

[0186] 2.3 Weight fusion of D-S evidence theory

[0187] When the combined model performs weight fusion, a model showing higher prediction accuracy is assigned a larger weight, and a model with lower prediction accuracy obtains a smaller weight. This weight allocation strategy is similar to the basic credibility allocation mechanism in D-S evidence theory. The prediction result of a single model and its corresponding weight can be regarded as a kind of "evidence" of the true load value, and the model weight represents the credibility of the prediction result of the model. Therefore, the Dempster combination rule of evidence theory can be used to fuse the weights, fully utilize the advantages of each model, reduce the uncertainty of single model prediction, and thus improve the overall prediction accuracy. When there is a high correlation between the prediction results, there is a problem of reduced effectiveness of D-S theory. To solve this problem, the prediction model selected by the present application meets the above requirements and introduces a weight mechanism to reflect its reliability and importance.

[0188] 2.3.1 Model weight allocation method

[0189] The variance covariance weight allocation method is selected to allocate the weight of a single prediction model. Assuming that there are m prediction models and n predictions are performed, the actual value of the jth(j=1, 2,..., n) prediction of the ith(i=1, 2,..., m) model is , the predicted value is , the prediction error is , the prediction error variance of the model is , the combined prediction result and the combined prediction error are:

[0190]

[0191] In the formula: The weight of the ith prediction model.

[0192] Generally, each group of prediction data of the same object is independent of each other, all covariances are 0, and the error variance of the combined prediction result is:

[0193]

[0194] Introducing Lagrange pair Taking the minimum value, the model weight is:

[0195]

[0196] Model weight Similar to the basic belief value in the evidence theory, the weight fusion can be carried out by using the Dempster combination rule to obtain the combined model load power prediction value at the next moment.

[0197] 2.3.2 D-S combined model prediction process

[0198] The present application adopts the variance covariance weight method and the D-S evidence theory to distribute and fuse the prediction results of the NTS and IRN models, and the flow chart is as shown in Figure 6 , and the specific process is as follows:

[0199] Step 1: training the NTS and IRN prediction models by using the training set, predicting the load power on the two models by using the test set, obtaining the prediction values of the two prediction models and the prediction errors thereof;

[0200] Step 2: determining the prediction time length [t0, t f ] and the prediction moment t of the D-S evidence theory combined model, t starting from the starting moment t0, and t f next moment ending;

[0201] Step 3: obtaining the prediction value at t moment and the prediction value and prediction error of N hours before t moment, and calculating the combined weight of the two prediction models, and through multiple comparison experiments, the fusion effect is best when N=3;

[0202] Step 4: constructing the basic belief allocation on the recognition framework, fusing the weights by using the Dempster rule, obtaining the weight distribution of the two prediction models at t moment, and calculating the prediction value of the D-S evidence theory combined model at t moment according to the prediction values of the two models at t moment obtained in Step 2;

[0203] Step 5: advancing 1 h from the prediction moment t, judging whether the moment t exceeds the prediction time length, if not, jumping to Step 3, otherwise ending the prediction process.

[0204] 3 Simulation Examples

[0205] 3.1 Data Source

[0206] This invention uses electricity load data from January 1 to December 31, 2023, provided by ISO New England, with a resolution of 1 hour. The load data from January 1 to September 30 is used as the training set, and the load data from October 1 to December 31, 2023 is used as the test set. To facilitate observation of the prediction results, only the prediction results from October 11 to 20 in the test set are selected to plot the curve, while the evaluation index is calculated using the complete test set.

[0207] 3.2 Model Parameter Selection

[0208] 3.2.1 NTS Model Parameters

[0209] In the construction of the NTS model, determining the sliding window size for the trend component, the detrending series stationarity of the cyclic component estimation, and the main harmonic frequencies is particularly important. The sliding window size is selected using the MAE criterion based on trend prediction results with different window sizes, with a minimum unit of one week, within which at least different daily types can be found. Figure 7 The MAE standard calculated for different sliding window sizes is shown.

[0210] Depend on Figure 7 It can be seen that the optimal duration of the sliding window is 3 weeks. Once the sliding window size is determined, the optimal trend curve is also determined. Figure 8 As shown, the detrending series is calculated according to formula (3). The results of the KPSS and ADF tests are shown in Table 2.

[0211] Table 2 Results of Detrended Series Stationarity Test

[0212]

[0213] As shown in Table 2, the test statistics of both the KPSS and ADF tests are less than the critical values ​​at the 10%, 5%, and 1% significance levels, and the p-values ​​are all less than 0.1, proving that... It is stable. A smoothed periodogram using a Daniell kernel (7,7) filter is plotted on it. Figure 9 For the smoothed periodic plot of the detrended training dataset, the dashed lines indicate the main harmonic frequency set, as shown in Equation (20), which represents the periods of 24 h, 12 h, 6 h, 5 h, etc.

[0214]

[0215] 3.2.2 IRN model parameters

[0216] In the feature input structure of the IRN model, , , and The full connection layer of has 10 hidden nodes, while The full connection layer of has 5 hidden nodes, , and the full connection layer before it has 10 hidden nodes. In the deep learning structure, each residual block has a hidden layer with 20 hidden nodes, the input and output size of the residual block are both 24, the residual block is divided into a main residual block and a bilateral residual block, and the average value is taken by using a shortcut connection, the main and bilateral residual blocks are each stacked with 30 residual blocks, forming a left-middle-right three-side 60-layer residual network.

[0217] 3.3 Experimental result analysis

[0218] 3.3.1 Comparison of time series models

[0219] In order to verify the effectiveness of the NTS model, a simulation experiment is carried out by comparing with the classical ARIMA time series model, the prediction curve of the simulation result is as shown in Figure 10 The experimental evaluation indexes adopt the mean absolute error MAE, the root mean square error RMSE and the mean absolute percentage error MAPE, and the calculation results are as shown in Table 3.

[0220] As can be seen from Figure 10 , compared with ARIMA, the predicted value of NTS is closer to the true value of the load power, showing a better fitting effect, but the prediction effect of the part with sharp changes at the peak is decreased, according to formula (1), the main reason is the influence of the random error component.

[0221] Table 3 Evaluation indexes of prediction results of NTS and ARIMA models

[0222]

[0223] As can be seen from Table 3, in the NTS and ARIMA models, the prediction effect of the NTS prediction model is better, the MAE, RMSE and MAPE of NTS compared with ARIMA are reduced by 57.5%, 61.9% and 57.9% respectively, indicating that the deviation degree of the predicted value of the model from the true value is smaller, and this result verifies the effectiveness of the model.

[0224] 3.3.2 Comparison of neural network models

[0225] To verify the selected residual neural network prediction model RN has certain superiority in predicting load power, it is compared with long short-term memory neural network model LSTM, BP neural network and Transformer neural network for simulation experiment, and the experimental results are shown in Figure 11 Table 4.

[0226] As shown in Figure 11 , the predicted values of RN, LSTM, BP and Transformer models at the peak of load power, especially at the time when the fluctuation amplitude is large (such as 60-70, 80-90 and 220-240), have certain deviation from the true values. The main reason is that the demand for electricity at this stage is complex and there is a change in the type of day, and it is difficult for traditional RN, LSTM, BP and Transformer prediction models to learn complex dynamic characteristics. Compared with LSTM, BP and Transformer, the predicted values of RN at the time when the fluctuation amplitude is large have large deviation from the true values, but the predicted values at the climbing and downhill stages are closer to the true values, and RN is better than other prediction models as a whole. To solve the above-mentioned problems, RN model needs to be further optimized.

[0227] Table 4. Evaluation indexes of prediction results of RN, LSTM, BP and Transformer models

[0228]

[0229] As shown in Table 4, compared with LSTM model, the MAE, RMSE and MAPE of BP model are reduced by 10.3%, 4.4% and 6.6% respectively; compared with BP model, the MAE, RMSE and MAPE of Transformer model are reduced by 11.5%, 17.0% and 8.6% respectively; compared with Transformer model, the MAE, RMSE and MAPE of RN model are reduced by 14.8%, 6.9% and 25.3% respectively. Among the four neural network prediction models, the prediction effect of RN is the best, the prediction effect of LSTM is the worst, and RN model has certain superiority.

[0230] 3.3.3 Comparison of residual neural network models

[0231] To verify the prediction performance of the improved residual neural network, the classical residual neural network RN, the residual neural network with feature input structure RN_input and the improved residual neural network IRN are compared for simulation experiment, and the experimental results are shown in Figure 12 Table 5.

[0232] As shown in Figure 12It can be seen that, regardless of the trough, peak or climbing stage, the fitting effect of the IRN prediction model is better than that of the RN and RN_input prediction models, because the IRN model not only analyzes the multi-scale time input features, but also shortens the connection between the input layer and the output layer without reducing the network depth, reduces the loss of information, and enables the model to be more in-depth, more accurate and more effective for training.

[0233] Table 5 Evaluation indexes of prediction results of IRN, RN_input and RN models

[0234]

[0235] As shown in Table 5, the IRN model has better prediction performance. Compared with the RN model, the MAE, RMSE and MAPE of the RN_input model are reduced by 23.9%, 26.5% and 21.9% respectively; compared with the RN_input model, the MAE, RMSE and MAPE of the IRN model are reduced by 34.2%, 26.0% and 37.3% respectively. Thus, it is shown that the feature input structure and deep learning structure designed by the present application can further improve the prediction performance of the residual neural network model.

[0236] 3.3.4 Comparison of combined prediction models

[0237] In order to verify the effectiveness of the D-S combined prediction model in fusing the respective advantages of the NTS and IRN models and the improvement of the model accuracy, the NTS, IRN, weighted average combination, Bayesian combination and D-S combined prediction models are compared and simulated, and the test results are shown in Table 6. Figure 13

[0238] As shown in Table 6, the test results are calculated by using the above experimental evaluation indexes, and the results are shown in Table 6. Figure 13 It can be seen that the combined prediction model can select the best prediction model according to the fitting effect of the NTS and IRN prediction models at each time, and calculate the best prediction result in the time period when the prediction effect of both is poor (such as 130-140, 150-160), especially the D-S combined prediction model shows higher prediction performance.

[0239] Table 6 Evaluation indexes of prediction results of three combined models and NTS, IRN models

[0240]

[0241] ​From table 6, the combined prediction model has better prediction performance, wherein the D-S combination model has the highest prediction precision.Compared with the NTS model, the MAE, RMSE and MAPE of the three combined models are reduced by 32.4%-52.9%, 22.2%-37.8% and 37.2%-55.6% respectively; compared with the IRN model, the MAE, RMSE and MAPE are reduced by 0%-30.4%, 5.4%-24.3% and 3.2%-31.4% respectively; compared with the weighted average and Bayesian combination methods, the MAE of the D-S method is reduced by 30.4% and 20.0% respectively, the RMSE is reduced by 20.0% and 17.6% respectively, and the MAPE is reduced by 33.5% and 27.7% respectively.It can be seen that the combined algorithm can play the advantages of the model and make up for the shortcomings of the model, wherein the D-S evidence theory method has the best effect on the combination of the IRN and NTS models, and can significantly improve the prediction accuracy of the model.

[0242] 3.3.5 Robustness comparison of D-S combined prediction model

[0243] During the collection and storage process, the load power may be abnormal due to equipment failure and other factors.In order to verify the robustness of the proposed model, white noise with a standard deviation of 0.2 and a noise data ratio of 5%, 10% and 20% is added to the normalized load power, and the experimental evaluation indexes are calculated, and the results are shown in table 7.

[0244] Table 7 Evaluation indexes of D-S combined model prediction results under different noise ratios

[0245]

[0246] As shown in table 7, after adding a small amount of white noise, the prediction accuracy of the model is not affected much, and after adding 20% white noise, the accuracy of the model prediction result is decreased, but the prediction performance is still better than that of the NTS model without white noise, and is similar to that of the IRN model without white noise, and the experimental results verify the robustness of the combined model.

[0247] CONCLUSION

[0248] The present application aims at the problem of large short-term load power fluctuation and low prediction accuracy, and proposes a combined prediction model based on time feature analysis, and the experimental conclusions are as follows:

[0249] (1) The present application deeply mines the time feature of load power change, decomposes the trend and cyclic component according to time, constructs an NTS prediction model, and improves the prediction performance of the model compared with the ARIMA prediction model;

[0250] (2) In order to perceive the time characteristics of the load power data on different time scales, the IRN prediction model is proposed based on the RN model, the feature input structure and the deep learning structure. The experimental results show that the model can capture the time characteristics of the multi-time scale input features, reduce the loss of information, and further improve the prediction accuracy.

[0251] (3) In view of the uncertainty and one-sidedness of the prediction results of a single prediction model, the D-S combination prediction model is proposed. The model integrates the prediction results of NTS and IRN to effectively reduce the uncertainty of the prediction results of a single model. The experiment shows that the D-S combination prediction model significantly improves the overall prediction accuracy and robustness.

[0252] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions without creative labor should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be limited by the protection scope defined in the claims.

Claims

1. A combined NST-IRN prediction model based on time feature analysis, characterized in that, First, the temporal characteristics of load power are deeply analyzed and decomposed into trend components and cyclical components to construct a novel time series model. Second, considering the impact of multi-timescale input features and daily types on load power, an improved residual neural network model with feature input structure and deep learning structure is constructed. Finally, the prediction results of the NTS and IRN models are weighted and fused using DS evidence theory to obtain the final load prediction result. Simulation experiments were conducted using real load data from ISO New England, and the results show that the proposed model has high prediction accuracy and robustness.

2. The prediction model according to claim 1, characterized in that, It consists of the following steps: Step 1: Construction of the combined prediction model Step 1.1 Construction of a novel time series model Considering the inherent time characteristics of load power, an NTS model is constructed, in which load power can be decomposed into trend, cyclic, and random error components; let... Let the load power at time t be represented, and its additivity model be: ; In the formula: —The trend component at time t, MW; —The cyclic component at time t, MW; —The random error component at time t, MW; The first two components are deterministic, while the random error component contains variations that the prediction model cannot explain; a parametric model is established for the two deterministic components. and These are respectively called trend model and cycle model; Determine the trend model using ordinary least squares (OLS) with the training dataset. The structure and key parameters of the sequence are determined, and the sequence is removed from the training set to obtain the detrended sequence. OLS is then used again to determine the recurrent model. Structure and its key parameters; Within the sliding window, use the test set pair and The coefficients in the structure are subjected to dummy variable regression and combined with input features such as temperature and day type to form a complete NTS prediction model and predict the load power at the next moment. The sliding window is moved and the above steps are repeated to make predictions for the next moment. Step 1.2 Construction of the Improved Residual Neural Network Model 1.2.1 RN Model Structure A residual neural network is composed of multiple residual blocks stacked together. A residual block can be represented as: ; In the formula: — x to Mapping; —A set of weights and biases associated with the residual block. The presence of the residual block allows the input to bypass some of the neural network, making the gradient easier to propagate downwards and reducing the problem of gradient explosion. The activation function for the residual block is SELU, and its expression is: ; In the formula: λ, α — two adjustable parameters; when and Furthermore, when the input follows a standard normal distribution, the output of the fully connected layer will be close to a standard normal distribution, which helps the network prevent gradient vanishing and exploding problems; all fully connected layers of the residual neural network designed in this invention use SELU as the activation function; Loss function of RN model Due to the error in prediction Over-range penalty term for accelerated training process constitute; ; ; ; In the formula: —Model output at hour h, MW; —Actual normalized load, MW; N —Number of data samples; H —Load per hour in a day (H=24), h; When the predicted daily load exceeds the range of the actual load curve, Penalizing the model encourages it to converge more quickly to a more accurate prediction range; when the model exhibits high prediction accuracy, Greater emphasis will be placed on the accuracy of peak and trough predictions; 1.2.2 IRN Model Structure To enable the residual neural network to fully learn time features such as multiple time scales and day types, a multi-time scale feature input structure and a deep learning structure to improve learning ability are added to the IRN model based on the residual neural network model. This further improves the model's predictive performance while analyzing the time characteristics of load power. (1) Feature input structure To further analyze the temporal characteristics of load power, load power at different time scales was selected as input features. By fusing load power data at different time scales and time-dependent one-hot codes, the model's ability to capture load dynamics and abnormal patterns was enhanced. The input features are as follows: ; Connect and move forward through a fully connected layer, the second layer being denoted as... ;Will , , Connect the layers, and then connect them using three separate fully connected layers, and then connect the three fully connected layers to another fully connected layer. Connecting S, W, and H creates two identical fully connected layers, which serve as... and Partial input; will and Connect them using fully connected layers, and then connect them to the output layer of the input structure. Or the input layer of an IRN model; The corresponding value is replaced with This allows for feature input at continuous time points; Through the designed feature input structure, the IRN model can analyze the input features hierarchically, enabling the residual neural network to capture the time characteristics of load power to the greatest extent, reduce information loss, and improve prediction accuracy. (2) Deep learning structure The deep learning architecture consists of main residual blocks, bilateral residual blocks, and shortcut connections. The output of each main residual block is averaged with the output of the bilateral residual blocks at the same level through shortcut connections. The average value is then passed to all main residual blocks in subsequent layers and together with the input layer of the network, it constitutes the input of the main residual blocks. By introducing additional two-sided residual blocks and dense shortcut connections, the representational ability of deep learning networks can be significantly improved, effectively capturing complex features in the data. At the same time, this structure optimizes the error backpropagation path, improves the efficiency and stability of the training process, and thus helps to improve the prediction accuracy and generalization ability of the model. Step 1.3 Weighting of DS Evidence Theory When combining models for weight fusion, models exhibiting higher prediction accuracy are assigned larger weights, while models with lower prediction accuracy receive smaller weights. This weight allocation strategy is similar to the basic credibility allocation mechanism in Dempster's evidence theory. The prediction results of a single model and its corresponding weights can be regarded as "evidence" of the true load value, while the model weights represent the credibility of the model's prediction results. Therefore, the Dempster composition rule of evidence theory can be used to fuse weights, making full use of the advantages of each model, reducing the uncertainty of single model predictions, and thus improving overall prediction accuracy. When the prediction results are highly correlated, there is a problem of reduced effectiveness of Dempster's theory. To solve this problem, different model structures and different feature sets are required. The prediction models selected in this invention meet the above requirements, and a weight mechanism is introduced to reflect their reliability and importance. 1.3.1 Model Weight Allocation Method The variance-covariance weighting method is used to allocate weights for a single prediction model. Assuming there are m prediction models and n predictions are made, the actual value of the j-th (j=1,2,...,m) prediction for the i-th (i=1,2,...,m) model is... The predicted value is The prediction error is The model's prediction error variance is Combined prediction results Combined prediction error for: ; In the formula: —The weights of the i-th prediction model; Typically, the prediction data for the same object are independent of each other, with all covariances being 0. The variance of the combined prediction result is: ; Introducing Lagrange pairs Finding the minimum value, we can obtain the model weights as follows: ; Model weights Similar to the basic confidence value in evidence theory, the Dempster composition rule can be used to perform weighted fusion to obtain the load power prediction value of the combined model at the next time step. 1.3.2 DS Combined Model Prediction Process This invention employs the variance-covariance weighting method and DS evidence theory to weight and fuse the prediction results of NTS and IRN models. The flowchart is shown in Figure 6, and the specific process is as follows: Step 1: Train the NTS and IRN prediction models using the training set, and use the test set to predict the load power on both models to obtain the predicted values ​​and prediction errors of the two prediction models. Step 2: Determine the prediction duration [t0, t] of the DS evidence theory combination model. f And the predicted time t, t starts from the initial time t0, t f The next moment ends; Step 3: Obtain the predicted value at time t and the predicted value and prediction error N hours before time t, and calculate the combined weight of the two prediction models. After multiple comparative experiments, the fusion effect is the best when N=3. Step 4: Construct a basic credibility assignment on the identification framework, use Dempster's rule to perform weight fusion, obtain the weight assignment of the two prediction models at time t, and calculate the prediction value of the DS evidence theory combined model at time t based on the prediction values ​​of the two models obtained in Step 2. Step 5: Advance the predicted time t by 1 hour, and determine whether time t exceeds the prediction duration. If it does not exceed the prediction duration, jump to Step 3; otherwise, end the prediction process. Step 2 Simulation Example 2.1 Data Source This invention uses electricity load data from January 1 to December 31, 2023, provided by ISO New England, with a resolution of 1 hour. The load data from January 1 to September 30 is used as the training set, and the load data from October 1 to December 31, 2023 is used as the test set. To facilitate observation of the prediction results, only the prediction results from October 11 to 20 in the test set are selected to plot the curve, while the evaluation index is calculated using the complete test set. 2.2 Model Parameter Selection 2.2.1 NTS Model Parameters In the process of constructing the NTS model, it is particularly important to determine the sliding window size of the trend component, the detrending series stationarity of the cyclic component estimation, and the main harmonic frequencies; the size of the sliding window is selected by the MAE criterion based on the trend prediction results for different window sizes, with the minimum unit being one week, within which at least different daily types can be found; The optimal duration for the sliding window size is 3 weeks. Once the sliding window size is determined, the detrending series is calculated according to formula (3). The KPSS and ADF tests were performed, and the test statistics of both tests were less than the critical values ​​at the 10%, 5%, and 1% significance levels, and the p-values ​​were all less than 0.1, proving that... It is stable; a smooth period plot is drawn on it using the Daniell kernel (7,7) filter, as shown in Equation (20), which represents the periods of 24 h, 12 h, 6 h, 5 h, etc. ; 2.2.2 IRN model parameters In the feature input structure of the IRN model , , and The fully connected layer has 10 hidden nodes, while The fully connected layer has 5 hidden nodes. , The preceding fully connected layer has 10 hidden nodes; in the deep learning architecture, each residual block has a hidden layer with 20 hidden nodes. The input and output size of the residual block is 24. The residual block is divided into a main residual block and bilateral residual blocks, and the average value is taken by shortcut connection. 30 residual blocks are stacked on each of the main and bilateral sides to form a 60-layer residual network with three sides on the left, middle and right.

3. The prediction model according to claim 2, wherein: The construction of the new time series model in step 1.1 is specifically as follows: 1.1.1 Trend component estimation Trend function Represents load power Y t The slow changes in electricity consumption, time, temperature, and day type are important factors explaining the changes in electricity consumption. Day type is represented by a dummy variable, and a trend function with a linear correlation between time, temperature, and the dummy variable is proposed: Where: t—time exponent from 1 to the window size, h; T t —Temperature value at time t, in °C; D α α = 1, ..., κ-1 — dummy variables, with values ​​between 0 and 1, indicating whether a corresponding regression exists; κ — the number of distinct categories; parameters a, b, c, and γ α —The coefficients of the regression model are evaluated using linear regression with OLS; 1.1.2 Cyclic component estimation Cyclic section captures load power Y t cyclical behavior; Let p represent the periodicity of the detrended function; It can write all samples: S t+p =S t The average value over the entire period is zero. Let the detrending sequence be: (1) Stationarity test Since the sum of the sine and cosine functions can completely describe the dynamics of any stationary periodic signal, Fourier component regression is used to estimate the cyclic component; to determine the detrended sequence W t To determine stationarity, a combined Kvitkowski-Phillips-Schmidt-Hine test and an enhanced Dickie-Fuller test are used. If both stationarity tests confirm W... t If the trend model remains stationary across any sliding window size, it indicates that the designed trend model is sufficient to represent the load power Y. t Slowly changing dynamics can be applied to W. t Perform cyclic component estimation; (2) Smoothed periodogram The smoothed periodogram can find the main harmonic frequencies of the cyclic component, which represents the rapid changes in the data; let n be the size of the stationary data set, and the periodogram is defined as: In the formula: ν k =k / n, k = {0, 1, ..., n-1} — Detrended sequence W t The fundamental frequency of the Discrete Fourier Transform (DFT); d c (ν k — Normalized real variables of the DFT; d s (ν k — Normalized dummy variables of the DFT; In the formula: m = (n-1) / 2; —Detrending sequence W t The average value; Formula (6) indicates that the sum of squares can be divided into components with frequency ν. k The amplitude of a time periodogram represents the harmonic components; that is, in a time periodogram, the frequency corresponding to the peak value can explain the variation in the data. The periodogram usually has a large variance at a given frequency, and the solution is to use the smoothed periodogram; assuming that the spectral density is fairly constant within the considered frequency band, and adjacent frequencies have asymptotically independent values, and for any sample size n, the periodogram is the sum of the squares of the real and imaginary variables; a smoothing filter of size L = 2m + 1 << n, centered at frequency B: The Daniell kernel is a symmetric positive weight, centered at the estimated frequency, and the sum of all weights is 1; The smoothed periodogram is: (3) Cyclic model with Fourier components Based on frequency set The cyclic component of the set is: In the formula: parameter c i s i c i,α and s i,α —The coefficients of the regression model are evaluated using linear regression with ordinary least squares (OLS) method; 1.1.3 NTS prediction model The NTS prediction model consists of a trend estimation model and a cyclic component model, and the complete NTS prediction model can be expressed as: