Power system short-term load forecasting method and power system short-term load forecasting device

CN116404637BActive Publication Date: 2026-09-25TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310326869.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2026-09-25
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

然而,与系统级的负荷相比,低层级,尤其是居民配电变压器这类层级的负荷,其负荷非线性特性更加明显,目前已提及的方法预测精度受到限制,导致预测精度降低

Benefits of technology

[0100]之后根据目标应用场景的不同,建立基于Transformer的不同负荷预测模型,并利用短期负荷历史曲线和已知特征,对每个类别中的负荷预测模型进行模型训练和评估,也即实现提取通用特征,从而得到每个类别各自的性能最佳模型;对每个类别各自的性能最佳模型进行模型迁移,以使得每个性能最佳模型各自应用于类间其它应用场景中,进行短期负荷预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116404637B_ABST
    Figure CN116404637B_ABST
Patent Text Reader

Abstract

The application provides a kind of power system short-term load prediction and power system short-term load prediction device, it is related to electric power technical field, including: based on each prediction scene respective short-term load history curve, the time sequence and distribution similarity of each short-term load history curve are combined clustering, obtain optimal clustering result;According to the difference of target application scene, establish different load prediction model based on Transformer, and utilize short-term load history curve and known feature, the model training and evaluation of load prediction model in each category are carried out to obtain the performance best model of each category respectively;The performance best model of each category is migrated to model, so that each performance best model is applied to other application scenarios between categories for short-term load prediction.The application simply carries out short-term load prediction for any application scene.Improve low-level load prediction accuracy, small amount of operation, relatively fast operation speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power technology, and in particular to a power system short-term load forecasting and power system short-term load forecasting device. Background Technology

[0002] Power data analysis has entered the era of big data. The vast amounts of real-time, fine-grained energy consumption data collected through advanced metering infrastructure (AMI) provide more reliable information for power supply and demand balance analysis. Accurate short-term load forecasting (STLF) can provide references for power sellers, dispatchers, and users, enabling them to formulate more reasonable power sales, dispatching, and consumption plans. Therefore, in numerous similar load forecasting scenarios, researching ways to reduce the training time cost of forecasting models while maintaining a certain level of forecasting accuracy, as well as improving load forecasting under small sample sizes, is of great significance.

[0003] Over the past few decades, researchers have designed a variety of short-term load forecasting models (SLFMs), such as autoregressive models (autoregressive moving averages, autoregressive integral moving averages), and support vector machines. A large body of research literature has demonstrated their effectiveness in the field of short-term load forecasting (SLFM). However, compared to system-level loads, lower-level loads, especially those at the level of residential distribution transformers, exhibit more pronounced load nonlinearity, limiting the prediction accuracy of currently mentioned methods and leading to reduced prediction precision. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a power system short-term load forecasting and a power system short-term load forecasting device.

[0005] This invention provides a method for short-term load forecasting of a power system, the method comprising:

[0006] Based on the short-term load history curves of each prediction scenario, clustering is performed by combining the temporal and distributional similarities of each short-term load history curve to obtain the optimal clustering result. The optimal clustering result includes: multiple optimal clustering categories.

[0007] Based on different target application scenarios, different load prediction models based on Transformer are established. Using the short-term load history curves and known features, the load prediction models in each category are trained and evaluated to obtain the best-performing model for each category.

[0008] For each category, the best performing model is transferred to other application scenarios across categories for short-term load forecasting.

[0009] Optionally, the different load prediction models may include: a single encoder and multiple decoders, or include: a single encoder and a single decoder;

[0010] Establish different load forecasting models based on Transformer, including:

[0011] Each encoder receives known historical features and known forecast features, performs encoding operations using the characteristics of Transformer, obtains output features, and transmits them to a single decoder or multiple decoders. Each decoder corresponds to a prediction scenario.

[0012] Each decoder performs decoding operations on the output features using the characteristics of the Transformer to obtain a short-term load prediction curve corresponding to the prediction scenario;

[0013] Multiple decoders perform decoding operations on the output features using the characteristics of the Transformer to obtain short-term load prediction curves for their respective prediction scenarios.

[0014] Optionally, the encoder receives known historical features and known forecast features, performs encoding operations using the characteristics of the Transformer, and obtains output features, including:

[0015] The known historical feature is preset to be X. h ∈R h×n The known forecast feature is X. p ∈R p×n Where h represents the known historical feature time length, p represents the known forecast feature time length, and n represents the number of features used for prediction;

[0016] The known historical features and the known predicted features are each input into a multi-head attention layer. Through linear mapping, dot product, and normalization, the output of each attention layer is obtained. Multiple attention layers are stacked to obtain the output Multihead(Q, K, V), represented as follows:

[0017] Multihead(Q,K,V)=concat(head1,...,head m W O

[0018]

[0019] In the above formula, m represents the number of attention heads, and WO This represents the fusion of multi-head attention and the linear mapping weights to an appropriate dimension, Q = XW. Q K = XW K V = XW V X represents the input data X of the attention layer. h and X p , They represent the linear mapping weights, and Q, K, and V represent the value matrix, key matrix, and query matrix, respectively.

[0020] The output Multihead(Q, K, V) of the multihead attention layer is added to the input X of the attention layer, and then layer normalization is performed to obtain Norm. out1 , is represented as:

[0021] Norm out1 =Norm(X+Multihead(Q,K,V))

[0022] The Norm out1 The input is fed into a feedforward neural network to obtain the output features of the feedforward neural network, and then the Norm is... out1 The Norm is obtained by adding the output features of the feedforward neural network and performing layer normalization. out2 , is represented as:

[0023] Norm out2 =Norm(Norm) out1 +FC(Norm out1 ))

[0024] In the above formula, FC(·) represents a fully connected neural network;

[0025] The output features of the known historical features and the output features of the known forecast features are stacked to obtain the Encoder. out , is represented as:

[0026]

[0027] In the above formula, the known historical feature X h The output features are The known forecast feature X p The output features are

[0028] Optionally, based on the short-term load history curves for each prediction scenario, clustering is performed by combining the temporal and distributional similarities of each short-term load history curve to obtain the optimal clustering result, including:

[0029] Z-SCORE standardization is performed on each short-term load history curve to obtain the standardized curve;

[0030] Set the peak height and peak width, and extract the sequence peak and valley points for each standardized curve;

[0031] The peak and valley points are stretched on the horizontal axis, and density clustering is performed on the vertical axis using DBSCAN to align the peak and valley points, extract the sequence key points of each standardized curve, and ignore outliers.

[0032] Using Euclidean distance as a measure of temporal similarity, we perform similarity metric calculations, as well as hierarchical clustering and similarity index calculations.

[0033] Based on the similarity metric calculation results, as well as the hierarchical clustering and similarity index calculation results, and combined with the preset distribution similarity threshold, the optimal clustering result is obtained.

[0034] Optionally, Euclidean distance is used as a measure of temporal similarity to perform similarity metric calculation, as well as hierarchical clustering and similarity index calculation, including:

[0035] Calculate the Euclidean distance matrix between keypoints in different sequences: D E ∈R m×m ;

[0036] Kernel density is used to estimate the probability distribution of short-term load forecast curves, and KL divergence is used as a measure of sequence distribution similarity to calculate the KL divergence matrix between different short-term load forecast curves.

[0037] Calculate the result of hierarchical clustering with n clusters: C = {c1, ..., c2} n} and the corresponding distribution similarity index: Sim dis As shown in the following formula:

[0038]

[0039] In the above formula, Represents the divergence between classes. x represents the sum of divergences between different classes. p It is any class C i The load sequence in x q It is any class C j The load sequence in.

[0040] Optionally, depending on the target application scenario, different load forecasting models based on Transformer are established. Using the short-term load history curves and known features, the load forecasting models for each category are trained and evaluated to obtain the best-performing model for each category, including:

[0041] The target application scenario is determined to be either a large-sample scenario or a small-sample scenario;

[0042] When the target application scenario is the large sample scenario, a first load prediction model based on Transformer is constructed;

[0043] When the target application scenario is the small sample scenario, a second load prediction model based on Transformer is constructed;

[0044] Based on the optimal clustering results, using the short-term load history curve and the known features, the first load prediction model or the second load prediction model is divided into training and testing sets, and the model is trained and evaluated to obtain the best-performing model for each category.

[0045] Optionally, the known features include: known historical features and known forecast features;

[0046] Based on the optimal clustering results, using the short-term load history curves and known features, the first load prediction model or the second load prediction model is divided into training and testing sets, and the model is trained and evaluated to obtain the best-performing model for each category, including:

[0047] Based on the optimal clustering result C opt In the target application scenario corresponding to the first load prediction model, any scenario r is selected as the reference scenario;

[0048] The known historical features, known forecast features, and short-term load history curves of the reference scenario are divided into training sets and test sets corresponding to the first load prediction model according to preset conditions.

[0049] Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set via gradient backpropagation, and then evaluated using the test set to obtain the best-performing model for each category; or,

[0050] Based on the optimal clustering result C opt In the target application scenario corresponding to the second load prediction model, select any type c i ;

[0051] Class c i The known historical characteristics, known forecast characteristics, and short-term load history curves are divided into training and test sets corresponding to the second load prediction model according to preset conditions.

[0052] Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set through the gradient backpropagation algorithm, and the second load prediction model is evaluated using the test set to obtain the best performing model for each category.

[0053] Optionally, model transfer is performed on the best-performing model for each category, so that each best-performing model can be applied to other application scenarios across categories for short-term load forecasting, including:

[0054] Based on the optimal clustering result C opt The best-performing model in each category is then applied to other target scenarios within that category.

[0055] Based on the other target scenarios, a third load prediction model based on Transformer is built, and the known historical features, known forecast features, and short-term load history curves of the other target scenarios are divided into training sets and test sets corresponding to the third load prediction model according to preset conditions. The output layer of the best performing model in the category of other target scenarios is fine-tuned.

[0056] The output layer parameters of the third load prediction model are solved using the least squares method with 2-norm constraints in the training set, and the third load prediction model is evaluated using the test set, so that the third load prediction model can be applied to the other target scenarios for short-term load prediction.

[0057] Optionally, model transfer is performed on the best-performing model for each category, so that each best-performing model can be applied to other application scenarios across categories for short-term load forecasting, including:

[0058] Based on the optimal clustering result C opt The best-performing model in each category is then applied to other target scenarios within that category.

[0059] Calculate the temporal and distributional similarity between other target scenes and each predicted scene, and assign the other target scenes to the closest predicted scene;

[0060] Based on the closest predicted scenario, a fourth load prediction model based on Transformer is built, and the known historical features, known forecast features, and short-term load history curves of the other target scenarios are divided into training sets and test sets corresponding to the fourth load prediction model according to preset conditions.

[0061] Using the minimum mean squared error as the objective function, the decoder of the fourth load prediction model is trained using the training set through the backpropagation algorithm, and the fourth load prediction model after decoder training is evaluated using the test set, so that the fourth load prediction model after decoder training can be applied to other target scenarios for short-term load prediction.

[0062] This invention also provides a short-term load forecasting device for power systems, the short-term load forecasting device for power systems comprising:

[0063] The clustering module is used to cluster based on the short-term load history curve of each prediction scenario, combined with the temporal and distributional similarity of each short-term load history curve, to obtain the optimal clustering result. The optimal clustering result includes: multiple optimal clustering categories.

[0064] The modeling training and evaluation module is used to establish different load forecasting models based on Transformer according to different target application scenarios, and to train and evaluate the load forecasting models in each category using the short-term load history curves and known features to obtain the best-performing model for each category.

[0065] The migration module is used to migrate the best-performing model for each category, so that each best-performing model can be applied to other application scenarios across categories for short-term load forecasting.

[0066] Optionally, the clustering module includes:

[0067] The standardized unit is used to perform Z-SCORE standardization on each short-term load history curve to obtain the standardized curve.

[0068] The extraction unit is used to set the peak height and peak width, and to extract the peak and valley points of the sequence for each standardized curve.

[0069] Alignment unit is used to stretch the peak and valley points on the horizontal axis and perform density clustering on the vertical axis using DBSCAN to align the peak and valley points, extract the sequence key points of each standardized curve, and ignore outliers;

[0070] The computing unit is used to perform similarity measurement calculations, hierarchical clustering, and similarity index calculations, using Euclidean distance as a measure of temporal similarity.

[0071] Clustering units are used to obtain the optimal clustering result based on the similarity metric calculation results, as well as the hierarchical clustering and similarity index calculation results, combined with a preset distribution similarity threshold.

[0072] Optionally, the computing unit is specifically used for:

[0073] Calculate the Euclidean distance matrix between keypoints in different sequences: D E ∈R m×m ;

[0074] Kernel density is used to estimate the probability distribution of short-term load forecast curves, and KL divergence is used as a measure of sequence distribution similarity to calculate the KL divergence matrix between different short-term load forecast curves.

[0075] Calculate the result of hierarchical clustering with n clusters: C = {c1, ..., c2} n} and the corresponding distribution similarity index: Sim dis As shown in the following formula:

[0076]

[0077] In the above formula, Represents the sum of divergences between classes. x represents the sum of divergences between different classes. p It is any class C i The load sequence in x q It is any class C j The load sequence in.

[0078] Optionally, the modeling training and evaluation module includes:

[0079] A scenario unit is used to determine whether the target application scenario is a large-sample scenario or a small-sample scenario;

[0080] The first modeling unit is used to construct a first load prediction model based on Transformer when the target application scenario is the large sample scenario.

[0081] The second modeling unit is used to construct a second load prediction model based on Transformer when the target application scenario is the small sample scenario.

[0082] The training and evaluation unit is used to divide the first load prediction model or the second load prediction model into training and testing sets based on the optimal clustering results, using the short-term load history curve and the known features, and to train and evaluate the model to obtain the best-performing model for each category.

[0083] Optionally, the known features include: known historical features and known forecast features; the training and evaluation unit is specifically used for:

[0084] Based on the optimal clustering result C opt In the target application scenario corresponding to the first load prediction model, any scenario r is selected as the reference scenario;

[0085] The known historical features, known forecast features, and short-term load history curves of the reference scenario are divided into training sets and test sets corresponding to the first load prediction model according to preset conditions.

[0086] Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set via gradient backpropagation, and then evaluated using the test set to obtain the best-performing model for each category; or,

[0087] Based on the optimal clustering result C opt In the target application scenario corresponding to the second load prediction model, select any type c i ;

[0088] Class c i The known historical characteristics, known forecast characteristics, and short-term load history curves are divided into training and test sets corresponding to the second load prediction model according to preset conditions.

[0089] Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set through the gradient backpropagation algorithm, and the second load prediction model is evaluated using the test set to obtain the best performing model for each category.

[0090] Optionally, the migration module includes:

[0091] The first application unit is used to determine the optimal clustering result C. opt The best-performing model in each category is then applied to other target scenarios within that category.

[0092] The modeling and fine-tuning unit is used to build a third load prediction model based on Transformer based on the other target scenarios, and divide the known historical features, known forecast features and short-term load history curves of the other target scenarios into the training set and test set corresponding to the third load prediction model according to preset conditions, and fine-tunes the output layer of the best performing model in the category of other target scenarios.

[0093] The application unit is used to solve the output layer parameters of the third load prediction model in the training set using the least squares method with 2-norm constraints, and to evaluate the third load prediction model using the test set, so that the third load prediction model can be applied to the other target scenarios for short-term load prediction.

[0094] Optionally, the migration module further includes:

[0095] The second application unit is used to determine the optimal clustering result C. opt The best-performing model in each category is then applied to other target scenarios within that category.

[0096] The classification unit is used to calculate the temporal and distributional similarity between other target scenes and each predicted scene, and classify the other target scenes to the closest predicted scene.

[0097] The modeling and partitioning unit is used to build a fourth load prediction model based on the closest prediction scenario, and to divide the known historical features, known forecast features, and short-term load history curves of the other target scenarios into the training set and test set corresponding to the fourth load prediction model according to preset conditions.

[0098] The training application unit is used to train the decoder of the fourth load prediction model using the training set with the minimum mean square error as the objective function, and to evaluate the fourth load prediction model after decoder training using the test set, so that the fourth load prediction model after decoder training can be applied to other target scenarios for short-term load prediction.

[0099] The short-term load forecasting method for power systems provided by this invention first performs clustering based on the short-term load history curves of each forecast scenario and the temporal and distributional similarity of each short-term load history curve to obtain the optimal clustering result.

[0100] Then, based on different target application scenarios, different load forecasting models based on Transformer are established. Using short-term load history curves and known features, the load forecasting models in each category are trained and evaluated, that is, common features are extracted to obtain the best performing model for each category. The best performing model for each category is then transferred to other application scenarios between categories to perform short-term load forecasting.

[0101] This invention targets low-level loads, particularly those at the level of residential distribution transformers. First, it obtains the short-term load history curves for each prediction scenario. Since the characteristics of these loads are relatively concentrated, clustering can be performed to obtain a certain number of clusters. Then, common features are extracted for each practical application scenario, resulting in the optimal performance model for each category. Finally, the optimal performance model is transferred to other application scenarios across the categories, enabling simple and convenient short-term load prediction for any application scenario. This significantly improves the prediction accuracy of low-level loads, while requiring less computation and operating at a faster speed, making it highly practical. Attached Figure Description

[0102] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0103] Figure 1 This is a flowchart of a short-term load forecasting method for a power system according to an embodiment of the present invention;

[0104] Figure 2 This is a schematic diagram of the encoder and multiple decoders in an embodiment of the present invention;

[0105] Figure 3 This is a block diagram of a short-term load forecasting device for a power system according to an embodiment of the present invention. Detailed Implementation

[0106] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention, and are only some, not all, embodiments of the present invention, and are not intended to limit the present invention.

[0107] The inventors discovered that although various short-term forecasting models have been designed, the nonlinear characteristics of low-level loads, especially those at the level of residential distribution transformers, are more pronounced compared to system-level loads. This limits the forecasting accuracy of the currently mentioned short-term forecasting models, resulting in reduced forecasting precision.

[0108] Further research by the inventors revealed that in recent years, with the rise of artificial intelligence, numerous methods based on artificial neural networks (ANNs) have been combined with STLF for application, such as fully connected neural networks (FCNs), convolutional neural networks (CNNs), and long short-term memory networks (LSTMs). LSTMs, in particular, have demonstrated excellent performance in STLF. However, LSTMs are recursive along the temporal direction of the sequence; in other words, they are limited by time complexity, making their application in STLF relatively cumbersome and computationally intensive.

[0109] In the field of natural language processing, Transformers are generally used to capture contextual information and utilize attention mechanisms to allocate limited computational resources to focus on important parts, which can greatly improve the interpretability of various models. Furthermore, the network models built using Transformers are feedforward neural networks, which can be computed in parallel, significantly reducing the training cost of various models.

[0110] Furthermore, the inventors discovered that the model-based transfer learning hypothesis, through proper training, allows the source domain model to learn a wealth of structural knowledge from the data. Therefore, reusing the model learned from the source domain avoids the need to re-extract training data or reason about relationships in complex data representations, making model-based transfer learning more effective and enabling the mastery of high-level knowledge from the source domain. Applying model-based transfer learning methods to natural language processing, by pre-training with large amounts of text data and then fine-tuning the model's output layer on downstream tasks, often yields excellent results.

[0111] Based on the aforementioned innovative research findings, the inventors creatively combined clustering, Transformer model construction, model transfer, and other technologies in STLF, proposing the power system short-term load forecasting and power system short-term load forecasting device of this invention. The following provides a detailed explanation and description of the power system short-term load forecasting and power system short-term load forecasting device proposed in this invention.

[0112] Reference Figure 1 The flowchart illustrates a short-term load forecasting method for a power system according to an embodiment of the present invention. The method includes:

[0113] Step 101: Based on the short-term load history curve of each prediction scenario, cluster the data by combining the temporal and distribution similarity of each short-term load history curve to obtain the optimal clustering result. The optimal clustering result includes the categories of multiple optimal clusters.

[0114] For each scenario in a power system, especially for lower-level equipment, the characteristics of load data changes are highly similar. Therefore, clustering can be performed to obtain one or more most representative load data to cover other load data, thereby reducing the amount of data required for subsequent calculations and improving computational efficiency. Based on this consideration, a method is proposed that clustering be performed based on the short-term load history curves of each prediction scenario, combined with the temporal and distributional similarities of each short-term load history curve, to obtain the optimal clustering result. Generally, the optimal clustering result can yield one or more clusters, each clustering result being distinct from the others. Therefore, classification can be based on the number of clusters, with one clustering result corresponding to one category.

[0115] In a preferred embodiment, the specific method of clustering includes:

[0116] First, Z-SCORE standardization is performed on each short-term load history curve to obtain a standardized curve. Then, the peak height and peak width are set, and the sequence peak and valley points are extracted for each standardized curve. Next, the horizontal axis of the peak and valley points is stretched, and density clustering is performed on the vertical axis using DBSCAN to align the peak and valley points, extract the sequence key points of each standardized curve, and ignore outliers.

[0117] Then, Euclidean distance is used as a measure of temporal similarity to perform similarity measurement calculation, as well as hierarchical clustering and similarity index calculation. Finally, based on the similarity measurement calculation results, hierarchical clustering and similarity index calculation results, combined with a preset distribution similarity threshold, the optimal clustering result is obtained.

[0118] This includes using Euclidean distance as a measure of temporal similarity, calculating similarity metrics, and performing hierarchical clustering and similarity index calculations, including:

[0119] First, calculate the Euclidean distance matrix between keypoints in different sequences: D E ∈R m×m Then, the probability distribution of the short-term load forecast curve is estimated using kernel density, and the KL (Kullback-Leibler) divergence is used as a measure of the similarity of the sequence distribution to calculate the KL divergence matrix between different short-term load forecast curves.

[0120] Assuming the number of clusters is n, the result of hierarchical clustering with n clusters is calculated as: C = {c1, ..., c2} n} and the corresponding distribution similarity index: Sim dis As shown in the following formula:

[0121]

[0122] In the above formula, Represents the sum of divergences between classes. x represents the sum of divergences between different classes. p It is any class C i The load sequence in x q It is any class C j The load sequence in the equation. It should be noted that this function is a monotonically decreasing function, so the smaller the value, the better the clustering effect. However, this formula has no extreme points, so a threshold needs to be set manually, i.e., a preset distribution similarity threshold needs to be set.

[0123] Step 102: Based on different target application scenarios, establish different load forecasting models based on Transformer, and use short-term load history curves and known features to train and evaluate the load forecasting models in each category to obtain the best performing model for each category.

[0124] After obtaining the clustering results, different load forecasting models based on Transformer can be established according to different target application scenarios. Using short-term load history curves and known features, the load forecasting models in each category are trained and evaluated, thus extracting common features and obtaining the best-performing model for each category. Generally, application scenarios in practice include large-sample scenarios and small-sample scenarios. Large-sample scenarios refer to scenarios with relatively large amounts of data, while small-sample scenarios refer to scenarios with relatively small amounts of data. Taking residential distribution transformers as an example, a large-sample scenario involves a large amount of load data for the transformer, storing load data from the current day to one year prior. A small-sample scenario, on the other hand, involves a small amount of load data, storing only load data from one or two months prior. Therefore, a small-sample scenario cannot accurately reflect the true load data of the residential distribution transformer.

[0125] There are two types of load forecasting models: one can be defined as a multi-objective load forecasting model, which includes a single encoder and multiple decoders; the other can be defined as a single-objective load forecasting model, which includes a single encoder and a single decoder.

[0126] First, based on the characteristics of the Transformer, a single or multi-objective load forecasting model is established. Then, this model is used to process known historical and forecast characteristics to obtain short-term load forecast curves for each forecast scenario. This is because the single or multi-objective load forecasting model is essentially a recursive model, reflecting the relationship between known historical characteristics, known forecast characteristics, and the short-term load forecast curves for each scenario.

[0127] In a preferred embodiment, an encoder receives known historical features and known forecast features, performs encoding operations using the characteristics of a Transformer, obtains output features, and transmits them to multiple decoders. A single decoder performs decoding operations on the output features using the Transformer's characteristics to obtain a short-term load forecast curve corresponding to a given forecast scenario; multiple decoders perform decoding operations on the output features using the Transformer's characteristics to obtain their respective short-term load forecast curves for their respective forecast scenarios. Since one decoder obtains a short-term load forecast curve for one forecast scenario, naturally multiple decoders obtain their respective short-term load forecast curves for their respective forecast scenarios.

[0128] For the structure of the encoder and multiple decoders mentioned above, please refer to... Figure 2 The structural diagram shown provides a clearer understanding. The known historical features are assumed to be X. h ∈R h×n The known forecast feature is X p ∈R p×n , where h represents the known historical feature time length, p represents the known forecast feature time length, and n represents the number of features used for prediction; these two features serve as the input to the encoder.

[0129] The known historical features and the known forecast features are each input into the multi-head attention layer. Figure 2 In this context, Multi-HeadAttention (represented by this concept) is used to obtain the output of each attention layer through linear mapping, dot product, and normalization. Multiple attention layers are stacked to obtain the output of a multi-head attention layer, Multihead(Q, K, V), as shown below:

[0130] Multihead(Q,K,V)=concat(head1,...,head m W O

[0131]

[0132] In the above formula, m represents the number of attention heads, and W O This represents the fusion of multi-head attention and the linear mapping weights to an appropriate dimension, Q = XW. Q K = XW K V = XW V X represents the input data X of the attention layer. h and X p , They represent the linear mapping weights, and Q, K, and V represent the value matrix, key matrix, and query matrix, respectively.

[0133] Then, the output Multihead(Q, K, V) of the multihead attention layer is combined with the input X (including X) of the attention layer. h and X p Add them together and perform layer normalization. Figure 2 In Chinese, we use Add&Norm to obtain the Norm. out1 , is represented as:

[0134] Norm out1 =Norm(X+Multihead(Q,K,V))

[0135] Norm out1 Input to feedforward neural network ( Figure 2The Norm (represented by Feed Forward) is used to obtain the output features of the feedforward neural network, and then the Norm is... out1 Add the output features of the feedforward neural network and perform layer normalization. Figure 2 (Using "Add & Norm" above FeedForward to get the Norm) out2 , is represented as:

[0136] Norm out2 =Norm(Norm) out1 +FC(Norm out1 ))

[0137] In the above formula, FC(·) represents a fully connected neural network.

[0138] Finally, the output features of known historical features and the output features of known forecast features are stacked. Figure 2 (Using Concate to represent this) to obtain the Encoder out , is represented as:

[0139]

[0140] In the above formula, the historical feature X is known. h The output features are Known forecast feature X p The output features are

[0141] Multiple decoders will receive the Encoder out Decode the data and output the corresponding short-term load forecast curves y1, y2, ... y3 for each scenario. m The specific decoding operations of the decoder can be understood by referring to the principles of the encoder and known decoders, and will not be elaborated further. It should be noted that the so-called features in this invention can be understood as features corresponding to information other than load data or load curves, such as: temperature, rainfall, air pressure, wind speed, humidity, sunshine duration, etc.

[0142] In a preferred embodiment, step 102 specifically includes:

[0143] Determine whether the target application scenario is a large-sample scenario or a small-sample scenario; depending on whether the target application scenario to be predicted is a large-sample scenario or a small-sample scenario, a prediction model needs to be built accordingly. That is:

[0144] When the target application scenario is a large-sample scenario, a first load prediction model based on Transformer is constructed; when the target application scenario is a small-sample scenario, a second load prediction model based on Transformer is constructed.

[0145] Finally, based on the optimal clustering results, using short-term load history curves and known features, the first load prediction model or the second load prediction model is divided into training and test sets, and the model is trained and evaluated to obtain the best-performing model for each category.

[0146] Assume the optimal clustering result is defined as C opt For large sample scenarios, then:

[0147] Based on the optimal clustering result C opt In the target application scenario corresponding to the first load forecasting model, any scenario r is selected as a reference scenario. Then, the known historical features, known forecast features, and short-term load history curves of the reference scenario r are divided into a training set and a test set corresponding to the single-objective load forecasting model according to preset conditions. The preset conditions in this embodiment can be divided according to actual needs. For example, 80% of all known historical features, 80% of all known forecast features, and 80% of all short-term load history curves of the reference scenario r are selected as the training set, and the remaining 20% ​​is selected as the test set.

[0148] Finally, using the minimum mean squared error as the objective function, the first load prediction model was trained multiple times using the training set through the gradient backpropagation algorithm, and evaluated using the test set to obtain the best-performing model {M1, ..., M} for each category. k}; where M1 represents the best-performing model for the first category, M k This indicates that the model with the best performance in the k-th category is .

[0149] For example, during training: the input consists of known historical features and forecast features relative to these historical features, and the output is the actual short-term load. During prediction, the unknown load is predicted. For instance, assuming we have features and load data from the past week, during the first training iteration, the input known historical features are those from Monday and Tuesday, and the known forecast features relative to these historical features are those from Wednesday, resulting in the output being the actual short-term load for Wednesday. During the second training iteration, the input known historical features are those from Tuesday and Wednesday, and the known forecast features relative to these historical features are those from Thursday, resulting in the output being the actual short-term load for Thursday. This process continues, with repeated training and final evaluation. Using the evaluated load prediction model, load prediction for the following Monday can be performed by Sunday. In this case, the input known historical features are those from Saturday and Sunday, and the known forecast features are those predicted for Monday. The load prediction model can then accurately predict the unknown short-term load for Monday.

[0150] It should be noted that because the data volume for large-sample scenarios is relatively large, the corresponding scenario for each category is fixed. Therefore, the first load forecasting model is actually a single-objective load forecasting model. However, the data volume for small-sample scenarios is relatively small, and it cannot be determined which forecasting scenario's load characteristics are more similar to theirs. Therefore, the second load forecasting model needs to be constructed as a multi-objective load forecasting model to make short-term load forecasts for small-sample scenarios more accurate. For small-sample scenarios:

[0151] Based on the optimal clustering result C opt In the target application scenario corresponding to the second load prediction model, select any type c i The selected class c i Known historical characteristics, known forecast characteristics, and short-term load history curves are used to divide the data into training and test sets for the second load prediction model according to preset conditions. Finally, using the minimum mean squared error as the objective function, the second load prediction model is trained multiple times using the training set through gradient backpropagation, and evaluated using the test set. This yields the best-performing model for each category, which can also be represented by {M}. 1, ..., M k Let M1 represent the best performing model for the first category, and M... k This indicates that the model with the best performance in the k-th category is .

[0152] Step 103: Perform model migration on the best performing model for each category so that each best performing model can be applied to other application scenarios across categories for short-term load forecasting.

[0153] After obtaining the best-performing model for each category, since the best-performing model is obtained by extracting general characteristics from the clustering results, it cannot accurately reflect the situation of other scenarios in a category. Therefore, model transfer is still required so that the transferred model can more accurately reflect the situation of other scenarios in its category, and thus more accurately predict short-term load.

[0154] For large-sample scenarios, since the category to which the target application scenario belongs is clearly known, the optimal clustering result C is used. opt The best-performing model from each category is applied to other target scenarios within that category. The resulting Transformer-based third load forecasting model, built upon these other target scenarios, is essentially a single-objective load forecasting model. It is understandable that for small-sample scenarios, based on the optimal clustering result C... optThe best-performing model in each category is applied to other target scenarios within each category. The fourth load prediction model based on Transformer, built on these other target scenarios, is essentially a multi-objective load prediction model.

[0155] Based on the above principles, for large-sample scenarios, it is necessary to divide the known historical features, known forecast features, and short-term load history curves of other target scenarios into training and test sets corresponding to the third load prediction model according to the aforementioned preset conditions. Simultaneously, it is also necessary to fine-tune the output layer of the best-performing model in the category of other target scenarios. For example, a distribution transformer A and another distribution transformer B belong to the same category C. The best-performing model for category C is obtained based on distribution transformer A. When distribution transformer B is the target application scenario, its known historical features, known forecast features, and short-term load history curves are divided into training and test sets corresponding to the third load prediction model according to the aforementioned preset conditions. Furthermore, it is necessary to fine-tune the output layer of the best-performing model for category C.

[0156] After fine-tuning, the output layer parameters of the third load forecasting model are solved using the least squares method with 2-norm constraints in the training set, and the third load forecasting model is evaluated using the test set. This allows the third load forecasting model to be applied to other target scenarios for short-term load forecasting. Continuing with the previous example: after fine-tuning, the output layer parameters of the third load forecasting model in category C are solved using the least squares method with 2-norm constraints in the training set, and the third load forecasting model in category C is evaluated using the test set. This allows the third load forecasting model in category C to be applied to distribution transformer B, thereby accurately forecasting the short-term load of distribution transformer B.

[0157] For small sample scenarios, the differences are slightly different from those for large sample scenarios:

[0158] First, it is necessary to calculate the temporal and distributional similarity between other target scenes and each predicted scene, and then classify the other target scenes into the closest predicted scene; that is, first determine which category the target scene should be classified into.

[0159] Then, based on the closest prediction scenario, a fourth load prediction model based on Transformer is built. The known historical features, known forecast features, and short-term load history curves of other target scenarios are divided into training and test sets corresponding to the fourth load prediction model according to preset conditions. Finally, with the minimum mean squared error as the objective function, the decoder of the fourth load prediction model is trained using the BP algorithm (backpropagation algorithm) with the training set, and the fourth load prediction model after decoder training is evaluated using the test set. This allows the fourth load prediction model after decoder training to be applied to other target scenarios, thereby accurately predicting the short-term load of the target scenario.

[0160] In this embodiment of the invention, based on the above-described method for short-term load forecasting of power systems, a short-term load forecasting device for power systems is also proposed, with reference to... Figure 3 A block diagram of a short-term load forecasting device for a power system is shown, which includes:

[0161] Clustering module 310 is used to perform clustering based on the short-term load history curve of each prediction scenario and the temporal and distributional similarity of each short-term load history curve to obtain the optimal clustering result. The optimal clustering result includes: multiple optimal clustering categories.

[0162] The modeling training and evaluation module 320 is used to establish different load prediction models based on Transformer according to different target application scenarios, and to train and evaluate the load prediction models in each category using the short-term load history curves and known features to obtain the best performing model for each category.

[0163] The migration module 330 is used to migrate the best performing model for each category, so that each best performing model can be applied to other application scenarios between categories for short-term load forecasting.

[0164] Optionally, the clustering module 310 includes:

[0165] The standardized unit is used to perform Z-SCORE standardization on each short-term load history curve to obtain the standardized curve.

[0166] The extraction unit is used to set the peak height and peak width, and to extract the peak and valley points of the sequence for each standardized curve.

[0167] Alignment unit is used to stretch the peak and valley points on the horizontal axis and perform density clustering on the vertical axis using DBSCAN to align the peak and valley points, extract the sequence key points of each standardized curve, and ignore outliers;

[0168] The computing unit is used to perform similarity measurement calculations, hierarchical clustering, and similarity index calculations, using Euclidean distance as a measure of temporal similarity.

[0169] Clustering units are used to obtain the optimal clustering result based on the similarity metric calculation results, as well as the hierarchical clustering and similarity index calculation results, combined with a preset distribution similarity threshold.

[0170] Optionally, the computing unit is specifically used for:

[0171] Calculate the Euclidean distance matrix between keypoints in different sequences: D E ∈R m×m ;

[0172] Kernel density is used to estimate the probability distribution of short-term load forecast curves, and KL divergence is used as a measure of sequence distribution similarity to calculate the KL divergence matrix between different short-term load forecast curves.

[0173] Calculate the result of hierarchical clustering with n clusters: C = {c1, ..., c2} n} and the corresponding distribution similarity index: Sim dis As shown in the following formula:

[0174]

[0175] In the above formula, Represents the sum of divergences between classes. x represents the sum of divergences between different classes. p It is any class C i The load sequence in x q It is any class C j The load sequence in.

[0176] Optionally, the modeling training and evaluation module 320 includes:

[0177] A scenario unit is used to determine whether the target application scenario is a large-sample scenario or a small-sample scenario;

[0178] The first modeling unit is used to construct a first load prediction model based on Transformer when the target application scenario is the large sample scenario.

[0179] The second modeling unit is used to construct a second load prediction model based on Transformer when the target application scenario is the small sample scenario.

[0180] The training and evaluation unit is used to divide the first load prediction model or the second load prediction model into training and testing sets based on the optimal clustering results, using the short-term load history curve and the known features, and to train and evaluate the model to obtain the best-performing model for each category.

[0181] Optionally, the known features include: known historical features and known forecast features; the training and evaluation unit is specifically used for:

[0182] Based on the optimal clustering result C opt In the target application scenario corresponding to the first load prediction model, any scenario r is selected as the reference scenario;

[0183] The known historical features, known forecast features, and short-term load history curves of the reference scenario are divided into training sets and test sets corresponding to the first load prediction model according to preset conditions.

[0184] Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set via gradient backpropagation, and then evaluated using the test set to obtain the best-performing model for each category; or,

[0185] Based on the optimal clustering result C opt In the target application scenario corresponding to the second load prediction model, select any type c i ;

[0186] Class c i The known historical characteristics, known forecast characteristics, and short-term load history curves are divided into training and test sets corresponding to the second load prediction model according to preset conditions.

[0187] Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set through the gradient backpropagation algorithm, and the second load prediction model is evaluated using the test set to obtain the best performing model for each category.

[0188] Optionally, the migration module 330 includes:

[0189] The first application unit is used to determine the optimal clustering result C. opt The best-performing model in each category is then applied to other target scenarios within that category.

[0190] The modeling and fine-tuning unit is used to build a third load prediction model based on Transformer based on the other target scenarios, and divide the known historical features, known forecast features and short-term load history curves of the other target scenarios into the training set and test set corresponding to the third load prediction model according to preset conditions, and fine-tunes the output layer of the best performing model in the category of other target scenarios.

[0191] The application unit is used to solve the output layer parameters of the third load prediction model in the training set using the least squares method with 2-norm constraints, and to evaluate the third load prediction model using the test set, so that the third load prediction model can be applied to the other target scenarios for short-term load prediction.

[0192] Optionally, the migration module 330 further includes:

[0193] The second application unit is used to determine the optimal clustering result C. opt The best-performing model in each category is then applied to other target scenarios within that category.

[0194] The classification unit is used to calculate the temporal and distributional similarity between other target scenes and each predicted scene, and classify the other target scenes to the closest predicted scene.

[0195] The modeling and partitioning unit is used to build a fourth load prediction model based on the closest prediction scenario, and to divide the known historical features, known forecast features, and short-term load history curves of the other target scenarios into the training set and test set corresponding to the fourth load prediction model according to preset conditions.

[0196] The training application unit is used to train the decoder of the fourth load prediction model using the training set with the minimum mean square error as the objective function, and to evaluate the fourth load prediction model after decoder training using the test set, so that the fourth load prediction model after decoder training can be applied to other target scenarios for short-term load prediction.

[0197] In summary, the power system short-term load forecasting method of the present invention first performs clustering based on the short-term load historical curves of each forecast scenario and the temporal and distributional similarity of each short-term load historical curve to obtain the optimal clustering result.

[0198] Then, based on different target application scenarios, different load forecasting models based on Transformer are established. Using short-term load history curves and known features, the load forecasting models in each category are trained and evaluated, that is, common features are extracted to obtain the best performing model for each category. The best performing model for each category is then transferred to other application scenarios between categories to perform short-term load forecasting.

[0199] This invention targets low-level loads, particularly those at the level of residential distribution transformers. First, it obtains the short-term load history curves for each prediction scenario. Since the characteristics of these loads are relatively concentrated, clustering can be performed to obtain a certain number of clusters. Then, common features are extracted for each practical application scenario, resulting in the optimal performance model for each category. Finally, the optimal performance model is transferred to other application scenarios across the categories, enabling simple and convenient short-term load prediction for any application scenario. This significantly improves the prediction accuracy of low-level loads, while requiring less computation and operating at a faster speed, making it highly practical.

[0200] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0201] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0202] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for short-term load forecasting of a power system, characterized in that, The power system short-term load forecasting method includes: Based on the short-term load history curves of each prediction scenario, clustering is performed by combining the temporal and distributional similarities of each short-term load history curve to obtain the optimal clustering result. The optimal clustering result includes: multiple optimal clustering categories. Based on different target application scenarios, different load prediction models based on Transformer are established. Using the short-term load history curves and known features, the load prediction models in each category are trained and evaluated to obtain the best-performing model for each category. For each category, the best performing model is transferred to other application scenarios across categories for short-term load forecasting. The different load forecasting models include: a single encoder and multiple decoders, or a single encoder and a single decoder; depending on the target application scenario, different load forecasting models based on Transformer are established, and the short-term load history curves and known features are used to train and evaluate the load forecasting model for each category, obtaining the best-performing model for each category, including: A single encoder receives known historical features and known forecast features, performs encoding operations using the characteristics of a Transformer, obtains output features, and transmits them to a single decoder or multiple decoders, with each decoder corresponding to a prediction scenario; a single decoder performs decoding operations on the output features using the characteristics of a Transformer to obtain a short-term load prediction curve corresponding to a prediction scenario; multiple decoders perform decoding operations on the output features using the characteristics of a Transformer to obtain short-term load prediction curves corresponding to their respective prediction scenarios. The target application scenario is determined to be a large-sample scenario or a small-sample scenario; if the target application scenario is the large-sample scenario, a first load prediction model based on Transformer is constructed; the first load prediction model includes: a single encoder and a single decoder; When the target application scenario is the small sample scenario, a second load prediction model based on Transformer is constructed; the second load prediction model includes: a single encoder and multiple decoders; Based on the optimal clustering results, using the short-term load history curve and the known features, the first load prediction model or the second load prediction model is divided into training and testing sets, and the model is trained and evaluated to obtain the best-performing model for each category.

2. The short-term load forecasting method for power systems according to claim 1, characterized in that, The encoder receives known historical features and known forecast features, performs encoding operations using the characteristics of the Transformer, and obtains output features, including: The known historical features are presupposed to be The known forecast features are ,in Indicates the length of time of known historical characteristics. Indicates the known forecast characteristic time length. Indicates the number of features used for prediction; The known historical features and the known predicted features are each input into a multi-head attention layer. Through linear mapping, dot product, and normalization, the output of each attention layer is obtained. Multiple attention layers are stacked to obtain the output of the multi-head attention layer. , means as follows: In the above formula, m represents the number of attention heads. This represents the fusion of multi-head attention and the linear mapping weights to an appropriate dimension. , X represents the input data of the attention layer. and , They represent the linear mapping weights, and Q, K, and V represent the value matrix, key matrix, and query matrix, respectively. The output of the multi-head attention layer Add the input X of the attention layer and perform layer normalization to obtain , represented as: The The input is fed into a feedforward neural network to obtain the output features of the feedforward neural network, and then the aforementioned... The output features of the feedforward neural network are added together and then normalized by layer to obtain the result. , is represented as: In the above formula, This represents a fully connected neural network; The output features of the known historical features and the output features of the known forecast features are obtained by stacking them. , represented as: In the above formula, the known historical features The output features are The known forecast features The output features are .

3. The short-term load forecasting method for power systems according to claim 1, characterized in that, Based on the short-term load history curves for each prediction scenario, clustering is performed using the temporal and distributional similarities of each short-term load history curve to obtain the optimal clustering results, including: Z-SCORE standardization is performed on each short-term load history curve to obtain the standardized curve; Set the peak height and peak width, and extract the sequence peak and valley points for each standardized curve; The peak and valley points are stretched on the horizontal axis, and density clustering is performed on the vertical axis using DBSCAN to align the peak and valley points, extract the sequence key points of each standardized curve, and ignore outliers. Using Euclidean distance as a measure of temporal similarity, we perform similarity metric calculations, as well as hierarchical clustering and similarity index calculations. Based on the similarity metric calculation results, as well as the hierarchical clustering and similarity index calculation results, and combined with the preset distribution similarity threshold, the optimal clustering result is obtained.

4. The short-term load forecasting method for power systems according to claim 3, characterized in that, Using Euclidean distance as a measure of temporal similarity, similarity metrics are calculated, along with hierarchical clustering and similarity index calculations, including: Calculate the Euclidean distance matrix between keypoints in different sequences: ; Kernel density is used to estimate the probability distribution of short-term load forecast curves, and KL divergence is used as a measure of sequence distribution similarity to calculate the KL divergence matrix between different short-term load forecast curves. Calculate the number of clusters Below are the results of hierarchical clustering: and the corresponding distribution similarity index: As shown in the following formula: In the above formula, , representing the divergence between classes, , representing the sum of divergences between different classes, x p It is any type C i The load sequence in x q It is any type C j The load sequence in.

5. The short-term load forecasting method for power systems according to claim 1, characterized in that, The known features include: known historical features and known forecast features; Based on the optimal clustering results, using the short-term load history curves and known features, the first load prediction model or the second load prediction model is divided into training and testing sets, and the model is trained and evaluated to obtain the best-performing model for each category, including: Based on the optimal clustering results In the target application scenario corresponding to the first load prediction model, select any scenario. As a reference scenario; The known historical features, known forecast features, and short-term load history curves of the reference scenario are divided into training sets and test sets corresponding to the first load prediction model according to preset conditions. Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set via gradient backpropagation, and then evaluated using the test set to obtain the best-performing model for each category; or, Based on the optimal clustering results In the target application scenario corresponding to the second load prediction model, select any type ; The class The known historical characteristics, known forecast characteristics, and short-term load history curves are divided into training and test sets corresponding to the second load prediction model according to preset conditions. Using the minimum mean squared error as the objective function, the first load prediction model is trained multiple times using the training set through the gradient backpropagation algorithm, and the second load prediction model is evaluated using the test set to obtain the best performing model for each category.

6. The short-term load forecasting method for power systems according to claim 1, characterized in that, For each category, the best-performing model is transferred to other application scenarios across categories for short-term load forecasting, including: Based on the optimal clustering results The best-performing model in each category is then applied to other target scenarios within that category. Based on the other target scenarios, a third load prediction model based on Transformer is built, and the known historical features, known forecast features, and short-term load history curves of the other target scenarios are divided into training sets and test sets corresponding to the third load prediction model according to preset conditions. The output layer of the best performing model in the category of other target scenarios is fine-tuned. The output layer parameters of the third load prediction model are solved using the least squares method with 2-norm constraints in the training set, and the third load prediction model is evaluated using the test set, so that the third load prediction model can be applied to the other target scenarios for short-term load prediction.

7. The short-term load forecasting method for power systems according to claim 1, characterized in that, For each category, the best-performing model is transferred to other application scenarios across categories for short-term load forecasting, including: Based on the optimal clustering results The best-performing model in each category is then applied to other target scenarios within that category. Calculate the temporal and distributional similarity between other target scenes and each predicted scene, and assign the other target scenes to the closest predicted scene; Based on the closest predicted scenario, a fourth load prediction model based on Transformer is built, and the known historical features, known forecast features, and short-term load history curves of the other target scenarios are divided into training sets and test sets corresponding to the fourth load prediction model according to preset conditions. Using the minimum mean squared error as the objective function, the decoder of the fourth load prediction model is trained using the training set through the backpropagation algorithm, and the fourth load prediction model after decoder training is evaluated using the test set, so that the fourth load prediction model after decoder training can be applied to other target scenarios for short-term load prediction.

8. A short-term load forecasting device for a power system, characterized in that, The power system short-term load forecasting device includes: The clustering module is used to cluster based on the short-term load history curve of each prediction scenario, combined with the temporal and distributional similarity of each short-term load history curve, to obtain the optimal clustering result. The optimal clustering result includes: multiple optimal clustering categories. The modeling training and evaluation module is used to establish different load forecasting models based on Transformer according to different target application scenarios, and to train and evaluate the load forecasting models in each category using the short-term load history curves and known features to obtain the best-performing model for each category. The migration module is used to migrate the best performing model for each category, so that each best performing model can be applied to other application scenarios between categories for short-term load forecasting. The different load forecasting models established by the modeling training and evaluation module include: a single encoder and multiple decoders, or a single encoder and a single decoder; the single encoder receives known historical features and known forecast features, performs encoding operations using the characteristics of Transformer, obtains output features, and transmits them to a single decoder or multiple decoders, with one decoder corresponding to one forecasting scenario; the single decoder performs decoding operations on the output features using the characteristics of Transformer to obtain a short-term load forecasting curve corresponding to a forecasting scenario; the multiple decoders perform decoding operations on the output features using the characteristics of Transformer to obtain short-term load forecasting curves corresponding to their respective forecasting scenarios; The modeling training and evaluation module includes: A scenario unit is used to determine whether the target application scenario is a large-sample scenario or a small-sample scenario; The first modeling unit is used to construct a first load prediction model based on Transformer when the target application scenario is the large sample scenario. The second modeling unit is used to construct a second load prediction model based on Transformer when the target application scenario is the small sample scenario. The training and evaluation unit is used to divide the first load prediction model or the second load prediction model into training and testing sets based on the optimal clustering results, using the short-term load history curve and the known features, and to train and evaluate the model to obtain the best-performing model for each category.

Citation Information

Patent Citations

  • Short-term power energy load parallel-prediction method applied to power quality comprehensive-governance scene and system

    CN108734355A

  • Power system short-term load prediction method and device based on feature selection

    CN114881343A