Short-term power load prediction method based on spatial-temporal feature fusion

By introducing spatiotemporal feature fusion and deep learning model (TFT) in short-term power load prediction, combining clustering and signal decomposition techniques, nonlinear and long-term dependency problems are solved, and more accurate and stable load prediction is achieved.

CN119944670AInactive Publication Date: 2025-05-06NANJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413593.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Short-term power load forecasting faces nonlinear and long-term dependency problems, and existing technologies are difficult to effectively solve these challenges.

Method used

Using a method based on spatiotemporal feature fusion, combined with Canopy coarse clustering, KFCM algorithm and CEEMDAN signal decomposition, a Temporal Fusion Transformer (TFT) model was constructed, and the long-term dependence between time steps was captured through self-attention mechanism and gated residual network.

Benefits of technology

It significantly improves the accuracy and robustness of short-term power load prediction, reduces noise interference, and enhances the adaptability and prediction capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944670A_ABST
    Figure CN119944670A_ABST
Patent Text Reader

Abstract

The invention discloses a short-term power load prediction method based on spatial-temporal feature fusion. The method comprises the following steps: constructing a feature database; identifying and eliminating abnormal values, filling missing values, and performing standardization processing on the data by adopting a Z-score method; a Canopy coarse clustering algorithm is adopted to determine a clustering number C and a cluster S, and the clustering number C and the cluster S serve as initial parameters of a KFCM algorithm; a membership matrix, noise points and a clustering center are calculated through C-KFCM clustering, and data clustering is completed; decomposing the clustered data through a CEEMDAN algorithm, and sequentially extracting IMF components and residual signals until the residual signals cannot be further decomposed; converting the static features and the time-varying features into vectors through an embedded layer; dynamically selecting key features by using a GRN (Gated Residual Network) and a variable selection network; the decoder is combined with the encoder to output a future known feature prediction load value; the method provides powerful support for optimal dispatching of the power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electric power technology, and in particular to a short-term electric power load forecasting method based on spatiotemporal feature fusion. Background Art

[0002] Power system load forecasting analyzes the intrinsic relationship between power load and historical power load data, social, economic, meteorological and other factors; explores the changing laws of power load; and accurately and scientifically predicts the power load usage in a certain period of time in the future under the premise of ensuring a certain degree of accuracy. Load forecasting is of great significance to the management, design and analysis of power grid systems. With the development of smart microgrids, load forecasting has gradually become an important part of the energy management system. Accurate power load forecasting provides an important reference for the operation of the power grid. According to the length of the forecast time, power load forecasting can be divided into long-term, medium-term and short-term forecasting. This paper mainly focuses on short-term power load forecasting. Accurate short-term load forecasting helps to reasonably arrange the maintenance of various equipment in the power system, optimize the dispatch of power resources, and help save energy such as coal and oil, which is highly consistent with the current goal of building an energy-saving society. Therefore, fast, accurate and stable load forecasting is crucial to the daily operation of the power system and is the basis for building a new power system. Summary of the invention

[0003] Purpose of the invention: The purpose of the present invention is to provide a short-term power load forecasting method based on the fusion of spatiotemporal features, which solves the nonlinear and long-term dependency problems in power load forecasting by introducing clustering and signal decomposition methods and combining them with a Transformer-based deep learning model (i.e., TFT model).

[0004] Technical solution: The short-term power load forecasting method based on spatiotemporal feature fusion described in the present invention comprises the following steps: S1. Obtain the historical load data of the target power system and the historical data of related factors affecting load changes, and build a feature database; identify and eliminate outliers, fill in missing values, and use the Z-score method to standardize the data; S2, using the Canopy rough clustering algorithm to determine the number of clusters C and clusters S as the initial parameters of the KFCM algorithm; using C-KFCM to calculate the membership matrix, noise points and cluster centers to complete data clustering; S3, decomposing the clustered data by CEEMDAN algorithm, extracting IMF components and residual signals in sequence, until the residual signal cannot be further decomposed; S4. Construct a TFT hybrid model: convert static features and time-varying features into vectors through the embedding layer; use the gated residual network GRN and the variable selection network to dynamically select key features; the encoder captures long-term dependencies between time steps through the self-attention mechanism; the decoder combines the encoder output and future known features to predict the load value.

[0005] Further, step S1 includes the following steps: (11) Use the Laida criterion to identify outliers, set confidence limits, and eliminate data that exceed the confidence limits; (12) Non-continuous missing data were filled with the average of the previous and next data, and continuous missing data were deleted; (13) The cleaned data were normalized using the Z-score.

[0006] Furthermore, the C-KFCM clustering algorithm in step S2 includes: (21) Set distance thresholds T1 and T2 (T1>T2) and randomly select data points as initial cluster centers; (22) Calculate the distance d between the data point and the cluster center, and assign it to the corresponding cluster or create a new cluster based on the relationship between d and T1 and T2; (23) Iterate until the data set is clear, and output the number of clusters C and cluster S; (24) Use the KFCM algorithm to initialize the membership matrix, the formula is: The formula is: ; in, It is a sample For The membership degree of the cluster centers is is the kernel function, is the fuzzification parameter, is the set of cluster centers; Calculate the noise points in the membership matrix, the formula is as follows: ; in, is the modified membership of noise point i to the jth cluster center, E is the neighborhood of noise point i, is the number of samples in domain E, is the membership of the qth sample in the neighborhood E to the jth cluster center; Calculate the cluster center of the data, the calculation formula is as follows: ; in, is the fuzzy weighted membership, is the kernel function value; Until the convergence condition is met ;in, and is the cluster center of two adjacent iterations, is the set threshold.

[0007] Furthermore, the CEEMDAN decomposition in step S3 specifically includes: (31) To the original sequence Add adaptive Gaussian white noise to generate a noisy sequence ; (32) Yes Perform EMD decomposition and calculate the first IMF component and residual ; (33) Residual Perform EMD decomposition to obtain the second IMF component; (34) Iteratively calculate the i-th residual and IMF component; (35) Repeat until the residual cannot be decomposed, and the final residual is: .

[0008] Furthermore, the TFT model construction in step S4 includes: (41) Convert static features, time-varying known features, time-varying unknown features, and timestamp features into vectors through an embedding layer; (42) Generate feature selection weights using gated residual network GRN; (43) The encoder uses a multi-head self-attention mechanism to process historical data; (44) The decoder combines the encoder output, static context information, and known features of future time steps to generate future load forecasts.

[0009] Furthermore, the embedding formula of the static feature in step (41) is as follows: ; in, is a static feature, The function is an embedding representation function, which is converted into a low-dimensional representation after embedding. ; The known feature embedding formula at time step t is as follows: ; in, is the known dynamic feature at time step t; The unknown feature embedding representation formula at time step t is as follows: ; in, is the historical observation data at time step t; The embedding formula of time step information is as follows: ; in, is the timestamp of time step t.

[0010] Furthermore, in step (42), the GRN generates the feature selection weight formula as follows: ; ; in, is the gated recurrent network function, is a static input feature, is the dynamic input feature at time step t.

[0011] Furthermore, the formula of step (43) is as follows: ; in, are the query, key, and value obtained from the input through linear transformation, is the dimension of the key; It is a function that normalizes the score to a weight; the output is fed through a feedforward network and normalized to obtain the encoding result .

[0012] Furthermore, in step (44), the decoder output formula is as follows: ; in, is the time step The predicted output is is the decoder module, is the final hidden state output by the encoder, is a static context, is the time step known features of; The decoder generates predictions for all future time steps, and the output is a time series as follows: ; in, It’s the future The prediction sequence of time steps is is the total number of future time steps predicted.

[0013] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: The present invention introduces a combination of Canopy coarse clustering and KFCM algorithms, which effectively reduces data complexity and improves the recognition ability of key features. Traditional methods usually rely on static features or original time series modeling, while the present invention uses EMD and CEEMDAN algorithms to perform multi-level signal decomposition on data, which can better extract the intrinsic pattern of the load series and reduce noise interference, thereby improving the prediction accuracy. Unlike the single time series processing of traditional models, the TFT model of the present invention combines the advantages of Transformer and LSTM, accurately captures long-term dependencies through the self-attention mechanism, dynamically selects features and optimizes information flow, and significantly improves the prediction ability and robustness. Traditional methods often rely on artificial feature design, while the present invention automatically selects key features related to the task through the gated residual network (GRN) and the variable selection network (VSN), reduces human intervention, and enhances the adaptability of the model. In general, the present invention provides a more accurate, stable and efficient short-term power load forecasting method, which provides strong support for the optimization and dispatching of power systems and the construction of smart grids. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 This is the architecture diagram of the Temporal Fusion Transformer. DETAILED DESCRIPTION

[0015] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0016] like Figure 1 As shown, an embodiment of the present invention provides a short-term power load forecasting method based on spatiotemporal feature fusion, comprising the following steps: (1) Obtain the historical load data of the target power system and the historical data of factors affecting load changes (such as temperature, humidity, power consumption mode, etc.), establish a feature database, identify outliers in the original data and remove them, fill in missing values ​​in the original data, and standardize the data using Z-score; (2) Use the Canopy coarse clustering algorithm to calculate the number of clusters C and clusters S of the original data set, use C and S as the initial input of the KFCM algorithm, and set the corresponding parameters; use random numbers [0,1] to perform initial operations on the membership matrix, calculate the noise points in the membership matrix, and calculate the cluster centers of the data; (3) Obtain the first IMF component through EMD decomposition, calculate CEEMDAN decomposition to obtain a unique residual signal, and obtain the second IMF component through the formula, and repeat the above steps to calculate the nth residual signal; repeat the above steps until the residual signal can no longer be decomposed by EMD, and finally obtain the decomposed signal; (4) Combine Transformer and LSTM to establish a Temporal Fusion Transformer (TFT) hybrid model; the model converts static features and time-varying features into vector representations suitable for processing through an embedding layer. Then, using the gated residual network (GRN) and variable selection network, the model can dynamically select features related to the prediction task and adjust the information flow through the gating mechanism to capture the complex nonlinear relationships in the input data. Next, the encoder of the model processes historical data through the Transformer network and uses the self-attention mechanism to model the dependencies between time steps to fully explore the long-term dependencies in the sequence. Finally, the decoder predicts future load values ​​based on the contextual information generated by the encoder; The specific contents are as follows: (1) Use the Laida criterion method to eliminate outliers. The specific operation is as follows: Assume that the measured value is measured with equal precision and independently obtain , calculate its arithmetic mean and residual error , and calculate the standard deviation , if a measurement value The residual error , satisfying the following formula: ; It is believed that Should be removed.

[0017] Then, the average method is used to fill in the missing data. The specific operation is as follows: for non-continuous missing data, the average method is used to fill in; for multiple continuous missing data, the direct deletion method is used. The formula of the average method is as follows: ; in, is the i-th missing data, and is the data before and after the missing value. If If it is a missing value, then ; Then perform data standardization: ; (2) Use the combination of the Canopy rough clustering algorithm and the KFCM algorithm to perform clustering calculations on the data; use the Canopy rough clustering algorithm and the KFCM algorithm to cluster the historical load data of the target power system to reduce the complexity of data features and improve the recognition ability of the prediction model for data features. The specific steps of the C-KFCM algorithm for clustering are as follows: Use the Canopy rough clustering algorithm to process the sample data to determine the number of clusters C and the cluster S, where the number of clusters C and the cluster S are used as the initial input parameters of the KFCM algorithm; at the same time, set the fuzzy parameter and the kernel function condition , the convergence boundary value and the number of iterations to ensure the effective execution of the algorithm; select two distance thresholds T1 and T2 (where T1>T2), randomly select a point P from the data set, define it as the center of the first Canopy, and remove this point from the sample data set. Continue to select data points from the sample data set and calculate the distance d between it and all the generated cluster centers; perform corresponding operations according to the size relationship of d, T1, and T2: when d<T2, assign this point to the current cluster and remove it from the original data set; when T2<d<T1, assign this point to the current cluster but keep it in the original data; when d>T1, this point is set as the new cluster center and removed from the original data set; repeat the above steps until the original data set is emptied; each Canopy consists of a center point and all points within T1 of this center point. The number of clusters C is the number of Canopies formed, and each Canopy corresponds to a cluster center. The cluster S consists of all points assigned to each Canopy, and each Canopy contains several points. Use a random number in [0,1] to perform an initial operation on the membership matrix and calculate the membership matrix. The calculation formula is as follows: ; where, is the sample 's membership degree to the th cluster center, is the kernel function, is the fuzzification parameter; is the set of cluster centers; Calculate the noise points in the membership matrix. The formula is as follows: ; where, is the corrected membership degree of the noise point i to the jth cluster center, E is the neighborhood of the noise point i, is the number of samples in the neighborhood E, is the membership degree of the qth sample in the neighborhood E to the jth cluster center; Calculate the cluster center of the data, the calculation formula is as follows: ; in, is the fuzzy weighted membership, is the kernel function value; Until the convergence condition is met ;in, and is the cluster center of two adjacent iterations, is the threshold value set; (3) Use the CEEMDAN algorithm to decompose the preprocessed data into multiple intrinsic mode functions (IMFs) and a final residual signal. The specific steps are as follows: In the original sequence Add a finite number of adaptive Gaussian white noises to , the formula is as follows: ; In the formula It is The sequence obtained after adding Gaussian white noise is is the fixed noise figure, is a Gaussian white noise sequence; For the first IMF component of The EMD decomposes with the mean as follows, and the formula is as follows: ; In the formula, For the The first IMF obtained after the Gaussian white noise decomposition, M is the number of noise additions; From the original sequence Subtract IMF1 from the original to get the first residual sequence. The formula is as follows: ; right Perform EMD decomposition to obtain the second IMF component, the formula is as follows: ; In the formula It represents the first IMF components, is the noise factor, is a Gaussian white noise sequence; Repeat the following steps to obtain the remaining IMF components, the formula is as follows: ; Where k is the total number of IMF components; All IMF components are obtained through the above operations, and the final sequence of residuals is calculated according to the IMF components. The formula is as follows: ; in, is the final residual, is the original signal, It is An intrinsic mode function.

[0018] (4) If Figure 2 As shown in the figure, the TFT hybrid model is constructed. First, the data is converted into a vector representation that the model can handle through the input embedding layer to support different types of input data, including static features, time-dynamic known features, and time-dynamic unknown features. Then, the variable selection network is used to dynamically select input features related to the prediction task. In the encoder, the long-term dependencies between historical time steps are captured, and global static features are extracted through the static context network. These static features serve as conditional information in the decoding stage. Finally, the decoder combines historical information, future known features, and static context to generate the final prediction value.

[0019] Static features are global features that are independent of time. For example, the area number of the power grid, equipment category, load type, etc. Static features are converted into low-dimensional vectors through the embedding layer to facilitate model processing. The embedding representation formula of static features is as follows: ; in, is a static feature, The function is an embedding representation function, which is converted into a low-dimensional representation after embedding. ; The time-dynamic known features are processed. Time-dynamic known features refer to features that can be obtained in advance during prediction, such as weather forecasts, holidays, seasonal changes, etc. The features of each time step are mapped through the embedding layer. The embedding representation formula of the known features at time step t is as follows: ; in, is the known dynamic feature at time step t; For the processing of time-dynamic unknown features, the time-dynamic unknown features are past observation data, such as historical load, historical meteorological data, etc. The embedding formula for the unknown features at time step t is as follows: ; in, is the historical observation data at time step t;

[0020] Time step information (such as hours, days of the week, months, etc.) can also be encoded as vectors through the embedding layer as additional input to the model; the embedding representation formula for time step information is as follows: ; in, is the timestamp of time step t.

[0021] The variable selection network of TFT dynamically selects the most relevant input features. This process is implemented by the Gated Residual Network (GRN). GRN performs nonlinear transformations on input features and adjusts the information flow through a gating mechanism. It uses skip connections and gated layers to control the flow of features. The output formula of the features after transformation by GRN is as follows: ; in, It is a nonlinear transformation consisting of a fully connected layer and an activation function (such as ReLU). It is a gating layer based on the Sigmoid activation function, which is used to control the flow of information.

[0022] For both static and dynamic features, the output generated by GRN is used as the selection weight of the feature. The dynamic features (known and unknown features at time step t) are dynamically selected by GRN at each time step to be the most relevant features.

[0023] The dynamic feature selection formula is as follows: ; ; The TFT model uses a Transformer encoder to capture long-term dependencies in time series. The encoder calculates the correlation between different time steps based on a multi-head attention mechanism and further extracts features through a feed-forward network (FFN).

[0024] The multi-head attention mechanism captures the dependencies between time steps in time series data through the relationship between query, key, and value. Each head performs attention calculation independently and then concatenates the results.

[0025] The calculation formula of the attention mechanism is as follows: ; in, are the query vector, key vector, and value vector obtained from the input through linear transformation, is the dimension of the key vector, is a function that normalizes the scores to weights.

[0026] The dynamic feature input of each time step is processed by a multi-head attention mechanism to obtain the encoder output. Subsequently, the output is further processed by a feed-forward network and finally normalized and residually connected to obtain the final output of the encoder.

[0027] The encoder output formula is as follows: ; ; ; in is the dynamic embedding vector, is the multi-head attention output, is the output of the feedforward network, is the final output of the encoder; The decoder combines the encoder output, static context information, and known features of future time steps to generate future load forecasts. The decoder inputs include: encoder output , static context , known features of future time steps ,The decoder uses a multi-head attention mechanism to capture the dependency between historical data and future predictions, and then generates prediction results; The decoder output formula is as follows: ; in, is the time step The predicted output is is the decoder module, is the final hidden state output by the encoder, is a static context, is the time step known features.

[0028] The decoder generates predictions for all future time steps, and the output is a time series as follows: ; in, It’s the future The prediction sequence of time steps is is the total number of future time steps predicted.

Claims

1. A short-term power load forecasting method based on spatiotemporal feature fusion, characterized in that: The following steps are involved: S1. Obtain historical load data of the target power system and historical data of related factors affecting load changes, and build a feature database; Identify and remove outliers, fill in missing values, and use the Z-score method to standardize the data; S2, using the Canopy coarse clustering algorithm to determine the number of clusters C and clusters S as the initial parameters of the KFCM algorithm; The membership matrix, noise points and cluster centers are calculated by C-KFCM clustering algorithm to complete data clustering; S3, decomposing the clustered data by CEEMDAN algorithm, extracting IMF components and residual signals in sequence, until the residual signal cannot be further decomposed; S4. Construct a TFT hybrid model: convert static features and time-varying features into vectors through the embedding layer; dynamically select key features using the gated residual network GRN and the variable selection network; the encoder captures long-term dependencies between time steps through the self-attention mechanism; The decoder combines the encoder output with future known features to predict the load value.

2. According to claim 1, a short-term power load forecasting method based on spatiotemporal feature fusion is characterized in that: Step S1 includes the following steps: (11) Use the Laida criterion to identify outliers, set confidence limits, and eliminate data that exceed the confidence limits; (12) Non-continuous missing data were filled with the average of the previous and next data, and continuous missing data were deleted; (13) The cleaned data were normalized using the Z-score.

3. The short-term power load forecasting method based on spatiotemporal feature fusion according to claim 1 is characterized in that: The C-KFCM clustering algorithm in step S2 includes: (21) Set distance thresholds T1 and T2, T1 > T2, and randomly select data points as initial cluster centers; (22) Calculate the distance d between the data point and the cluster center, and assign it to the corresponding cluster or create a new cluster based on the relationship between d and T1 and T2; (23) Iterate until the data set is clear, and output the number of clusters C and cluster S; (24) Use the KFCM algorithm to initialize the membership matrix, the formula is: ; in, It is a sample For The membership degree of the cluster centers is is the kernel function, is the fuzzification parameter, is the set of cluster centers; Calculate the noise points in the membership matrix, the formula is as follows: ; in, is the modified membership of noise point i to the jth cluster center, E is the neighborhood of noise point i, is the number of samples in domain E, is the membership of the qth sample in the neighborhood E to the jth cluster center; Calculate the cluster center of the data, the calculation formula is as follows: ; in, is the fuzzy weighted membership, is the kernel function value; Until the convergence condition is met ;in, and is the cluster center of two adjacent iterations, is the set threshold.

4. The short-term power load forecasting method based on spatiotemporal feature fusion according to claim 1 is characterized in that: The CEEMDAN decomposition in step S3 specifically includes: (31) To the original sequence Add adaptive Gaussian white noise to generate a noisy sequence ; (32) Yes Perform EMD decomposition and calculate the first IMF component and residual ; (33) Residual Perform EMD decomposition to obtain the second IMF component; (34) Iteratively calculate the i-th residual and IMF component; (35) Repeat until the residual cannot be decomposed, and the final residual is: .

5. The short-term power load forecasting method based on spatiotemporal feature fusion according to claim 1 is characterized in that: The TFT model construction in step S4 includes: (41) Convert static features, time-varying known features, time-varying unknown features, and timestamp features into vectors through an embedding layer; (42) Generate feature selection weights using gated residual network GRN; (43) The encoder uses a multi-head self-attention mechanism to process historical data; (44) The decoder combines the encoder output, static context information, and known features of future time steps to generate future load forecasts.

6. A short-term power load forecasting method based on spatiotemporal feature fusion according to claim 5, characterized in that: Step (41) The embedding formula of the static feature is as follows: ; in, is a static feature, The function is an embedding representation function, which is converted into a low-dimensional representation after embedding. ; The known feature embedding formula at time step t is as follows: ; in, is the known dynamic feature at time step t; The unknown feature embedding representation formula at time step t is as follows: ; in, is the historical observation data at time step t; The embedding formula of time step information is as follows: ; in, is the timestamp of time step t.

7. The short-term power load forecasting method based on spatiotemporal feature fusion according to claim 5 is characterized in that: In step (42), the GRN generates the feature selection weight formula as follows: ; ; in, is the gated recurrent network function, is a static input feature, is the dynamic input feature at time step t.

8. The short-term power load forecasting method based on spatiotemporal feature fusion according to claim 5 is characterized in that: The formula of step (43) is as follows: ; in, are the query, key, and value obtained from the input through linear transformation, is the dimension of the key; It is a function that normalizes the score to a weight; the output is fed through a feedforward network and normalized to obtain the encoding result .

9. The short-term power load forecasting method based on spatiotemporal feature fusion according to claim 5 is characterized in that: In step (44), the decoder output formula is as follows: ; in, is the time step The predicted output is is the decoder module, is the final hidden state output by the encoder, is a static context, is the time step known features of; The decoder generates predictions for all future time steps, and the output is a time series as follows: ; in, It’s the future The prediction sequence of time steps is is the total number of future time steps predicted.

Citation Information

Patent Citations

  • Table sequence identification method and system based on contextual relationship attention mechanism

    CN114241497A

  • Ultra-short-term solar irradiance prediction method based on multi-modal feature fusion

    CN118470486A