A small-sample learning method for optimal power flow prediction in power systems
Through the stack noise reduction automatic encoder network pre-training and knowledge distillation technology, the accuracy problem of small and medium-sized sample learning of optimal current prediction of DC power system is solved, and rapid deployment and online computing are realized on the newly built system.
Patent Information
- Application Number
- CN202211101634.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-09-09
AI Technical Summary
The existing deep learning methods are difficult to quickly deploy and apply on new or expanded systems due to insufficient data in the prediction of optimal current flow of DC power systems.
The stack noise reduction automatic encoder (SDAE) network is used for pre-training, combining knowledge distillation and task decomposition, and the voltage phase angle and power generation prediction model is constructed, finite annotation samples are used for training, and learning accuracy is improved through improved loss function and teacher model fine-tuning.
It realizes efficient optimal trend prediction on small sample data sets, reduces the demand for labeled samples, improves model training speed and accuracy, and is suitable for rapid deployment and online computing of new systems.
Smart Images

Figure CN115907000B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting optimal power flow in a power system, and more particularly to a small sample learning method for optimal power flow in a power system. Background Art
[0002] In order to quickly respond to changes in network load demand state and break through the bottleneck of computational efficiency, a solution using data-driven technology to predict optimal power flow (OPF) has emerged in recent years. Deep learning-based methods have brought significant efficiency improvements to OPF [1], [2]. It uses a large amount of historical data to approximate the relationship between variables and achieve real-time response. Compared with traditional solvers, deep learning methods have increased the computational speed of DC-OPF by 200 times and AC OPF (AC-OPF) by 35 times [3][4]. In addition, deep learning technology provides a feasible solution for solving OPF in online settings and state combinations. In order to solve the online efficiency problem of OPF learning, several methods based on active constraints [5][6], hot start point prediction [7][8], etc. have been studied. However, the high data requirements of these data-intensive methods limit their application [9]. However, most existing methods are highly dependent on large amounts of data. This limits their training speed and restricts their deployment and application in new or expanded systems.
[0003] To address the problem of training data requirements, a currently used method is a simplified iterative method driven by a hybrid of digital and analog, using Lagrangian reinforcement learning to accelerate convergence and achieve optimality during the iteration process
[10] . It is no longer a simple end-to-end deep learning training, but a method that uses training technology to accelerate the original OPF solution process. Another type of method relies on the prior information of the physical model. Using constraints as prior elements, machine learning methods can predict AC-OPF
[11] , the Lagrangian dual problem of optimal power flow
[12] , and implement pre-classification based on constraints
[13] . Existing deep learning methods in OPF are either data-intensive or require knowledge.
[0004] Since power systems often undergo topological changes due to line failures or dispatch operations, models often need to be retrained frequently. However, retraining models from scratch takes too long, and the limited amount of sample data accumulated in a short period of time restricts the accuracy of training
[14] , which limits its application in practical systems. Therefore, the development of high-accuracy small-sample learning technology is an urgent need for the development of artificial intelligence
[15] .
[0005] References
[0006] [1] O.A. Alimi, K. Ouahada, and A.M. Abu-Mahfouz, “A Review of Machine Learning Approaches to Power System Security and Stability”, IEEE Access, vol. 8, pp. 113512–113531, 2020.
[0007] [2] Y. Zhao and B. Zhang, “Deep learning in power systems,” in Advanced Data Analytics for Power Systems, A. Tajer, S.M. Perlaza, and H.V. Poor, Eds. Cambridge, U.K.: Cambridge Univ. Press, May 2021, pp. 52–71.
[0008] [3] T. Zhao, X. Pan, M. Chen, A. Venzke, et al, “DeepOPF: A deep neural network approach for DC optimal power flow for ensuring feasibility,” in Proc. IEEE Int. Conf. Smart Grid Commun., Tempe, AZ, USA, 2020, pp. 1–6.
[0009] [4] X. Pan, M. Chen, T. Zhao, and S.H. Low, “Deep OPF: A feasibility optimized deep neural network approach for AC optimal power flow problems,” 2020. [Online]. Available: https: / / arxiv.org / abs / 2007.01002.
[0010] [[ID=—12]][5] D. Deka and S.M. Misra, “Learning for DC-OPF: Classifying active sets using neural nets,” in Proc. IEEE Milan Power Tech., 2019, pp. 1–6.
[0011] It should be noted that there seems to be a small error in the original text where "[[ID=—12]]" should probably be "". This has been corrected in the translation as well.[6] F. Hasan, A. Kargarian, and J. Mohammadi, “Hybrid learning aided inactive constraints filtering algorithm to enhance AC OPF solution time,” IEEE Trans. Ind. Appl., vol. 57, no. 2, pp. 1325–1334, Mar. 2021
[0012] [7] K. Baker, “Learning warm start points for Ac optimal power flow,” 2019, arXiv:1905.08860. [Online]. Available: https: / / doi.org / 10.48550 / arXiv.1905.08860.
[0013] [8] L. Chen and J. E. Tate, “Hot-starting the AC power flow with convolutional neural networks,” 2020, arXiv:2004.09342. [Online]. Available: https: / / doi.org / 10.48550 / arXiv.2004.09342.
[0014] [9] M. Chatzos, T. W. K. Mak, and P. V. Hentenryck, “Spatial Network Decomposition for Fast and Scalable AC-OPF Learning”, IEEE Trans. Power Syst., vol. 37, no. 4, pp. 2601–2612, 2022.
[0015]
[10] Z. Yan and Y. Xu, “Real-Time Optimal Power Flow: A Lagrangian Based Deep Reinforcement Learning Approach”, IEEE Trans. Power Syst., vol. 35, no. 4, pp. 3270–3273, 2020.
[0016]
[11] M. Chatzos, F. Fioretto, T. W. K. Mak, and P. V. Hentenryck, “High-fidelity machine learning approximations of large-scale optimal power flow,” 2020, arXiv:2006. Art. no. 16356.
[0017]
[12] F. Fioretto, T. W. Mak, and P. Van Hentenryck, “Predicting AC optimal power flows: Combining deep learning and Lagrangian dual methods,” in Proc. AAAI Conf. Artif. Intell., 2020, pp. 630–637.
[0018]
[13] X. Lei, Z. Yang, J. Yu, J. Zhao, Q. Gao, and H. Yu, “Data-Driven Optimal Power Flow: A Physics-Informed Machine Learning Approach”, IEEE Trans. Power Syst., vol. 36, no. 1, pp. 346–354, 2021.
[0019]
[14] Y. Chen, S. Lakshminarayana, C. Maple, and H. V. Poor, “A Meta-Learning Approach to the Optimal Power Flow Problem Under Topology Reconfigurations”, IEEE Open Access J. Power & Energy, vol. 9, pp. 109–120, 2022.
[0020]
[15] M. K. Singh, V. Kekatos, and G. B. Giannakis, “Learning to Solve the AC-OPF Using Sensitivity-Informed Deep Neural Networks”, IEEE Trans Power Syst., vol. 37, no. 4, pp. 2833–2846, 2022. Summary of the Invention
[0021] The technical problem to be solved by the present invention is to provide a small-sample learning method for optimal power flow prediction of a DC power system, which can avoid the influence of insufficient data of a small number of samples on the accuracy of deep learning in order to overcome the deficiencies of the existing technology.
[0022] The technical solution adopted by the present invention is as follows: A small-sample learning method for optimal power flow prediction of a DC power system includes the following steps:
[0023] 1) Input historical data or simulation data of different states of the power system power flow as training samples;
[0024] 2) Divide the training samples into two categories: labeled data and unlabeled data;
[0025] 3) Construct a voltage phase angle prediction model and a power generation prediction model based on a stacked denoising autoencoder neural network, which are expressed in the following mathematical forms.
[0026] Y l = h l (h l-1 (h l-2 (···h 1 (X)))) (1)
[0027] h k (X) = s(W k X k + b k ) (2)
[0028]
[0029] Z = g l (g l-1 (g l-2 (···g 1 (Y l )))) (4)
[0030] g k (Y k ) = s(W k TY k + b′ k ) (5)
[0031] Among them, X is the input feature vector; Y l is the latent feature vector of the network and also the output vector of the lth encoding layer, which is obtained by continuously encoding the input vector from the first layer to the lth layer; h k () is the encoding function of the kth layer; s() represents the activation function; W k and b krespectively represent the weights and biases in the k-th hidden layer; X k is the input vector of the k-th layer. The input of the first encoding layer is the input feature vector of the network, and the input of each encoding layer from the second layer to the l-th layer is the output of the previous encoding layer; l is the number of encoding layers in the stacked denoising autoencoder neural network; Z is the reconstructed feature output by decoding, and the length of this vector is the same as the original input X; g k () represents the decoding transformation function of the k-th layer; Y k is the output vector of the l-th encoding layer, which is equal to the input vector of the l-th decoding layer; W k T and b k ’ represent the weights and biases in the k-th decoding layer;
[0032] Among them, when the input feature vector X is load demand level data and the latent feature vector Y l of the network is voltage phase angle data, equations (1)-(5) constitute a voltage phase angle prediction model; when the input feature vector X is load demand level data and the latent feature vector Y l of the network is power generation data, equations (1)-(5) constitute a power generation prediction model;
[0033] 4) Model pre-training, the specific process is: input the unlabeled data into the voltage phase angle prediction model and the power generation prediction model, calculate the loss function according to the reconstructed features output by the voltage phase angle prediction model and the power generation prediction model according to the following equation (6), and use the gradient descent method to minimize the loss function to determine the encoding layer parameters;
[0034] L H (X, Z l ) = ||X - Z l ||2 2 (6)
[0035] Among them, L H represents the loss function in the pre-training stage; ||X - Z l ||2 is the Euclidean norm of the residual vector formed by the input feature and the reconstructed feature.
[0036] 5) Decompose the DC optimal power flow calculation task into two subtasks of predicting voltage phase angle and predicting power generation, and the two subtasks are respectively undertaken by the voltage phase angle prediction model and the power generation prediction model;
[0037] 6) Use the phase angle label data to fine-tune and train the voltage phase angle prediction model f θ,t ;
[0038] 7) Use the power generation label data to fine-tune and train the power generation prediction model f G,t ;
[0039] 8) Determine the teacher model and the student model required for knowledge distillation, specifically: establish a deep neural network with the same scale as the voltage phase angle prediction model for the student model to simultaneously predict the voltage phase angle and the power generation, and use the encoding layer parameters obtained in step 4) to initialize the hidden layer parameters of the deep neural network, and represent the deep neural network as f (θ,G),s (P D →V θ ,P G ); Use the voltage phase angle prediction model f G,t and the power generation prediction model f G,t obtained through fine-tuning training as the teacher model, then the loss functions of the voltage phase angle prediction model and the power generation prediction model constitute the loss function of the teacher model; use the deep neural network f (θ,G),s (P D →V θ ,P G ) that simultaneously predicts voltage and phase angle as the student model;
[0040] 10) Set the maximum number of iterations for knowledge distillation training, and initialize the current number of iterations to 1;
[0041] 11) Calculate and measure the difference between the prediction values of the student model and the teacher model using the mean square error:
[0042]
[0043]
[0044] where, L θs,θt and L Gs,Gt respectively represent the differences between the teacher model and the student model for the phase angle prediction value and the power generation prediction value, and the subscripts s and t represent the student model and the teacher model respectively; V θs,i and P Gs,j are the predictions of the student model for the phase angle of the i-th node and the power generation of the j-th generator; V θt,i and P Gt,j are the prediction values of the teacher model for the i-th phase angle and the power generation of the j-th generator unit.
[0045] 12) Calculate the loss function value of the student model, and the calculation formula is expressed as,
[0046]
[0047]
[0048] where, L θs and L Gs respectively represent the loss function values of the phase angle and the power generation;
[0049] 13) Based on the loss function value of the student model and that of the teacher model, compare the accuracies of the student model and the teacher model according to the following formula. If the student model is more accurate, the weight λ of the teacher model is zero; otherwise, the weight λ of the teacher model linearly decreases as the student model is fine-tuned and trained:
[0050]
[0051]
[0052] where λ θ and λ G are the weights of the teacher model; e and e max are the current iteration number and the maximum iteration number of fine-tuning respectively; λ is the weight of the teacher model under the dynamic annealing mechanism, which linearly increases with the iteration number; L θ,t represents the loss function of the voltage phase angle prediction model f θ,t ; L Gt is the loss function of the power generation prediction model f G,t ;
[0053] 14) Calculate the weighted sum of the teacher loss function and the student loss function and use it for parameter update, which is calculated by formula (21):
[0054] L = λ θ L θs,θt + (1 - λ θ )L θs + λ G L Gs,Gt + (1 - λ G )L Gs (21)
[0055] where L is the comprehensive loss function used for knowledge distillation, which is composed of L θs,θt , L θs , L Gs,Gt and L Gs ;
[0056] 15) Repeat steps 11) to 14) until the iteration number reaches the upper limit, indicating that the small-sample learning of the optimal power flow is completed.
[0057] A small-sample learning method for optimal power flow prediction in a DC power system according to the present invention helps a deep neural network learn the optimal power flow of the power system using limited labeled samples. First, a pre-training strategy is adopted in a stacked denoising autoencoder (SDAE) network. By transferring the work to an unsupervised pre-training stage, the demand for labeled data is reduced. Second, a DC-OPF task decomposition strategy is combined with knowledge distillation to reduce the learning complexity. The knowledge distillation learning is improved through a teacher annealing strategy to improve the accuracy. In addition, the loss function is improved based on focal loss during the training stage, enhancing the training effect without adding extra samples. The technology proposed in the present invention can not only achieve the accuracy of small-sample learning, but also improve the model training speed by reusing the pre-training results and reducing the scale of the training set. The present invention improves the deep learning training process, reduces the dependence of model accuracy on the sample scale, and avoids the impact of insufficient data of minority-class samples on the accuracy of deep learning, thereby realizing the application for a power system with limited samples and helping it build and deploy an online calculation model of the optimal power flow faster.
[0058] The small-sample learning method proposed in the present invention is scalable and can be flexibly combined with other data-driven models, which is conducive to the combination of various data-driven technologies and power system applications. The present invention reduces the demand for labeled samples by the deep network, thus saving a large amount of computing resources for data annotation of training samples and facilitating practical applications. Traditional deep network training methods are difficult to establish on a newly built system because they require a large number of labeled samples. The present invention breaks through the limitation of data-driven technology in the application of small-sample data sets, fills the application gap of deep networks in the early stage of a newly built system lacking sufficient samples, provides an online calculation method for the optimal power flow of a newly built network, and provides operation guidance and decision-making reference for the initial construction stage lacking sufficient operation experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a flowchart of a small-sample learning method for optimal power flow prediction in a DC power system according to the present invention;
[0060] Figure 2 is a schematic diagram of a neural network based on a stacked denoising autoencoder established by the method of the present invention;
[0061] Figure 3 is a schematic diagram of a deep neural network established by the method of the present invention;
[0062] Figure 4 is a schematic diagram of the device used in the method of the present invention;
[0063] Figure 5 is a comparison diagram of the voltage phase angle prediction accuracy of the embodiments of the method of the present invention and the direct training method varying with the scale of label samples;
[0064] Figure 6 It is a comparison chart of the power generation prediction accuracy of the method of the present invention and the direct training method for the embodiments varying with the scale of the labeled samples. Detailed implementation manners
[0065] The following will make a detailed description of a small sample learning method for optimal power flow prediction in a DC power system according to the present invention in combination with embodiments and drawings.
[0066] As Figure 1 shown, a small sample learning method for optimal power flow prediction in a DC power system according to the present invention includes the following steps:
[0067] 1) Input historical data or simulation data of different states of the power system power flow as training samples;
[0068] The historical data includes: the load demand levels of each node recorded in several states during the operation history of the power system, the actual active power output values of each generator, and the phase angles of the voltage phasors of each node.
[0069] The simulation data includes: the load demand levels of each node solved by the power flow equation, the actual active power output values of each generator, and the phase angles of the voltage phasors of each node.
[0070] 2) Divide the training samples into two categories: labeled data and unlabeled data. Among them,
[0071] if there is only load data in the training samples, it is unlabeled data; if there are both load data, power generation data, and phase angle data in the training samples, it is labeled data; the labeled data is divided into phase angle labels and power generation labels according to the phase angle variable and the power generation variable.
[0072] 3) Construct a voltage phase angle prediction model and a power generation prediction model based on the stacked denoising autoencoder (SDAE) neural network as Figure 2 shown, which is expressed in the following mathematical form.
[0073] Y l = h l (h l-1 (h l-2 (···h 1 (X)))) (1)
[0074] h k (X) = s(W k X k + b k ) (2)
[0075]
[0076] Z = g l (g l-1 (g l-2 (···g 1 (Y l )))) (4)
[0077] g k (Y k ) = s(W k T Y k + b′ k ) (5)
[0078] Wherein, X is the input feature vector; Y l is the latent feature vector of the network and also the output vector of the l-th encoding layer, which is obtained by continuously encoding the input vector from the first layer to the l-th layer; h k () is the encoding function of the k-th layer; s() represents the activation function; W k and b k respectively represent the weights and biases in the k-th hidden layer; X k is the input vector of the k-th layer. The input of the first encoding layer is the input feature vector of the network, and the input of each encoding layer from the second layer to the l-th layer is the output of the previous encoding layer; l is the number of encoding layers in the stacked denoising autoencoder neural network; Z is the reconstructed feature output by decoding, and the length of this vector is the same as the original input X; g k () represents the decoding transformation function of the k-th layer; Y k is the output vector of the l-th encoding layer and is equal to the input vector of the l-th decoding layer; W k T and b k ’ represent the weights and biases in the k-th decoding layer;
[0079] Wherein, when the input feature vector X is the load demand level data and the latent feature vector Y l of the network is the voltage phase angle data, formulas (1)-(5) constitute a voltage phase angle prediction model; when the input feature vector X is the load demand level data and the latent feature vector Y l of the network is the power generation data, formulas (1)-(5) constitute a power generation prediction model;
[0080] 4) Model pre-training. The specific process is as follows: Input the unlabeled data into the voltage phase angle prediction model and the power generation prediction model, calculate the loss function according to the reconstructed features output by the voltage phase angle prediction model and the power generation prediction model according to the following formula (6), and minimize the loss function using the gradient descent method to determine the encoding layer parameters;
[0081] L H(X, Z l ) = ||X - Z l ||² 2 (6)
[0082] Among them, L H represents the loss function in the pre-training stage; ||X - Z l ||² is the Euclidean norm of the residual vector formed by the input feature and the reconstructed feature.
[0083] The loss function used is the mean square error function (MSE), which is calculated using the following formula:
[0084] L H (X, Z l ) = ||X - Z l ||² 2 (8)
[0085] Among them, L H is the loss function in the pre-training stage; ||X - Z l ||² is the Euclidean norm of the residual vector formed by the input feature and the reconstructed feature.
[0086] 5) Decompose the DC optimal power flow calculation task into two subtasks of predicting the voltage phase angle and predicting the power generation. The two subtasks are respectively undertaken by the voltage phase angle prediction model and the power generation prediction model; among them,
[0087] The input vectors and prediction vectors of the voltage phase angle prediction model and the power generation prediction model are represented by the following formula:
[0088]
[0089] Among them, P D represents the variable composed of the load demands of each node, V θ represents the vector composed of the voltage phase angles of each node, f θ,t is the voltage phase angle prediction model used to predict the phase angle V θ ; P G represents the power generation vector of the generator, f G,t is the power generation prediction model used to predict the power generation P G ; f θ,t (P D →V θ ) represents that the knowledge learned by the voltage phase angle prediction model f θ,t is the correlation relationship from the node load to the node phase angle, and f G,t (P D →P G ) then represents the power generation prediction model f G,tThe knowledge learned is the correlation between node load and power generation; it can be seen that the input feature vectors are all node load vectors, that is, the load demand data is shared by the two models.
[0090] 6) Use the phase angle label data to fine-tune and train the voltage phase angle prediction model f θ,t ; The loss function involved in fine-tuning and training the voltage phase angle prediction model f θ,t uses an improved dynamic scaling loss function, which is expressed as follows.
[0091]
[0092] Among them, n θ is the number of samples of the phase angle label; a i is the attention weight related to the phase angle label data of the i-th node; and p i is the attention weight related to the phase angle prediction value of the i-th node; Ω N is the set of nodes in the power system; V θt,i is the predicted value of the phase angle of the i-th node in the power system; V θ,i represents the phase angle value of the i-th node in the power system in the label data; L θ,t represents the loss function of the voltage phase angle prediction model f θ,t ;
[0093] The mathematical relationship between the attention weight a i and the label data is
[0094]
[0095] Among them, P[V θ,i is the proportion of the category corresponding to the numerical size of V θ,i , which is determined by dividing the value range of all V θ,i into 20 intervals and counting the number of phase angle labels falling into the same interval;
[0096] From the perspective of deep learning, data whose output is often equal to 0 or the maximum value is easier to learn. For these categories, the corresponding parameter p i should have a lower value. Therefore, the attention weight p i of the i-th node is determined based on the power operation result of the predicted value,
[0097] p i = V θt,i r (1 - V θt,i ) r + 0.5 (11)
[0098] Among them, r is set to 1.
[0099] 7) Use the power generation label data to fine-tune and train the power generation prediction model f G,t Perform fine-tuning training on the power generation prediction model f G,t The loss function involved in the fine-tuning training of the power generation prediction model f uses an improved dynamic scaling loss function, which is expressed as follows:
[0100]
[0101] where n G is the number of samples of the power generation label; a j is the attention weight of the jth generator related to the power generation label; and p j is the attention weight of the jth generator related to the power generation prediction value; Ω G is the set of power generation equipment in the power system; P Gt,j is the predicted value of the power generation of the jth generator in the power system by the power generation prediction model f G,t ; P G,j represents the power generation label of the jth generator in the power system; L Gt is the loss function of the power generation prediction model f G,t ;
[0102] The mathematical relationship between a j and the power generation label is
[0103]
[0104] where P[P G,j is the proportion of the category where the numerical value of P G,j is located, which is determined by dividing the value range of all P G,j into 20 intervals and counting the number of power generation labels falling into the same interval;
[0105] From the perspective of deep learning, data with outputs often equal to 0 or the maximum value are easier to learn. For these categories, the corresponding parameter p j should have a lower value. Therefore, the attention weight p j of the jth generator is determined based on the power operation result of the predicted value,
[0106] p j = P Gt,j r (1 - P Gt,j ) r + 0.5 (14)
[0107] where r is set to 1.
[0108] 8) Determine the teacher model and the student model required for knowledge distillation, specifically: establish a deep neural network for the student model with the same scale as the voltage phase angle prediction model as shown in Figure 3 to simultaneously predict the voltage phase angle and the generated power, and use the encoding layer parameters obtained in step 4) to initialize the hidden layer parameters of the deep neural network, and represent the deep neural network as f (θ,G),s (P D →V θ ,P G );Take the voltage phase angle prediction model f G,t and the generated power prediction model f G,t obtained through fine-tuning training as the teacher model, then the loss functions of the voltage phase angle prediction model and the generated power prediction model constitute the loss function of the teacher model; take the deep neural network f (θ,G),s (P D →V θ ,P G ) that is used to simultaneously predict the voltage and the phase angle as the student model;
[0109] 10) Set the maximum number of iterations for knowledge distillation training, and initialize the current number of iterations to 1;
[0110] 11) Calculate and measure the difference between the prediction values of the student model and the teacher model using the mean square error:
[0111]
[0112]
[0113] where L θs,θt and L Gs,Gt respectively represent the differences between the teacher model and the student model for the phase angle prediction value and the generated power prediction value, and the subscripts s and t represent the student model and the teacher model respectively; V θs,i and P Gs,j are the predictions of the student model for the phase angle of the i-th node and the generated power of the j-th generator; V θt,i and P Gt,j are the prediction values of the teacher model for the i-th phase angle and the generated power of the j-th generator unit.
[0114] 12) Calculate the loss function value of the student model, and the calculation formula is expressed as,
[0115]
[0116]
[0117] where L θs and L Gs respectively represent the loss function values of the phase angle and the generated power;
[0118] 13) Based on the loss function values of the student model and the teacher model, compare the accuracies of the student model and the teacher model according to the following formula. If the student model is more accurate, the weight λ of the teacher model is zero; otherwise, the weight λ of the teacher model linearly decreases as the student model is fine-tuned:
[0119]
[0120]
[0121] where λ θ and λ G are the weights of the teacher model; e and e max are the current iteration number and the maximum iteration number of fine-tuning respectively; λ is the weight of the teacher model under the dynamic annealing mechanism, which linearly increases with the iteration number; L θ,t represents the loss function of the voltage phase angle prediction model f θ,t ; L Gt is the loss function of the power generation prediction model f G,t ;
[0122] 14) Calculate the weighted sum of the teacher loss function and the student loss function and use it for parameter update, which is calculated by formula (21):
[0123] L = λ θ L θs,θt + (1 - λ θ )L θs + λ G L Gs,Gt + (1 - λ G )L Gs (21)
[0124] where L is the comprehensive loss function used for knowledge distillation, which is composed of L θs,θt , L θs , L Gs,Gt and L Gs ;
[0125] As can be seen from formulas (19) and (20), with the change of the teacher model weight, the knowledge distillation process is divided into two stages. In the early stage, the student model learns from the teacher models f θ,t and f G,t ; in the later stage, as the fine-tuning period increases, the student model gradually transitions into a supervised learning process based on the original labels.
[0126] 15) Repeat steps 11) to 14) until the iteration number reaches the upper limit, indicating the completion of the optimal power flow small sample learning.
[0127] A small-sample learning method for optimal power flow prediction in a DC power system according to the present invention is tested on the IEEE-RTS79 system. The system includes 24 nodes, 32 generating units, 38 branches, and the peak load is 2850 MW. Among them, the 38 branches consist of 5 transformer branches, 1 cable branch, and 32 transmission branches.
[0128] Table 1 Calculation Performance Analysis
[0129]
[0130] As shown in Table 1, by reusing the pre-trained model, the proposed training method using the knowledge distillation strategy can achieve high-precision results within 1 minute. Table 1 and Figure 5 、 Figure 6 show the application results of the method of the present invention on a small-sample data set. It shows that the pre-training strategy helps to improve the accuracy of a small sample size. This is because a large amount of unlabeled data can be used to train the shallow layer of the SDAE network during the pre-training stage, thus reducing the learning burden in the supervised stage. Compared with other methods, the proposed method maintains a higher accuracy level, demonstrating its feasibility in the small-sample state. For example, M2 requires 15,000 samples to roughly achieve the accuracy achieved by the proposed method (M6) with 200 samples.
[0131] An on-line calculation device for the operation risk of a power system with integrated knowledge transfer is established according to the method proposed by the present invention, including: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the processor to execute any link of the method described in the proposed method.
[0132] The computer hardware configuration of the embodiment of the present invention includes an Intel Core i5-6500 CPU, 8G of memory, the operating system is Windows 10, the simulation software is MATLAB 2020a, and the traditional OPF calculates the optimal load shedding using the matpower toolbox.
[0133] [[ID=SS]]The hyperparameters of each model in the embodiment of the present invention are: the number of hidden layers of each neural network is 3. The size of the teacher encoder for each layer is (150, 200, 250), and the student model is correspondingly half the size. For the pre-training stage, the SGD optimizer is used, with an initial learning rate of 0.1 and a momentum of 0.8. For the fine-tuning stage, the parameters β involved in the Adam optimizer are (0.7, 0.92), and the learning rate is 0.0005. The total number of epochs for pre-training is 40, and the total number of epochs for model fine-tuning is 50. The batch sizes for pre-training and fine-tuning are 256 and 128 respectively.
Claims
1. A small-sample learning method for optimal power flow prediction in a power system, characterized in that, It includes the following steps: 1) Input historical data or simulation data of different states of the power system power flow as training samples; 2) Divide the training samples into two categories: labeled data and unlabeled data; 3) Construct a voltage phase angle prediction model and a power generation prediction model based on the stack denoising autoencoder neural network, expressed in the following mathematical form; Y l = h l (h l-1 (h l-2 (···h 1 (X)))) (1) h k (X) = s(W k X k + b k ) (2) Z = g l (g l-1 (g l-2 (···g 1 (Y l )))) (4) g k (Y k ) = s(W k T Y k + b′ k ) (5) Among them, X is the input feature vector; Y l is the latent feature vector of the network and also the output vector of the l-th encoding layer, which is obtained by continuously encoding the input vector from the 1st layer to the l-th layer; h k () is the encoding function of the k-th layer; s() represents the activation function; W k and b k respectively represent the weight and bias in the hidden layer of the k-th layer; X k is the input vector of the k-th layer. The input of the first encoding layer is the input feature vector of the network, and the input of each encoding layer from the 2nd layer to the l-th layer is the output of the previous encoding layer; l is the number of encoding layers in the stacked denoising autoencoder neural network; Z is the reconstructed feature of the decoded output, and the length of this vector is the same as the original input X; g k () represents the decoding transformation function of the k-th layer; Y k is the output vector of the l-th encoding layer and is equal to the input vector of the l-th decoding layer; W k T and b k ’ represent the weight and bias in the k-th decoding layer; Among them, when the input feature vector X is load demand level data and the latent feature vector Y of the network l is voltage phase angle data, equations (1)-(5) constitute a voltage phase angle prediction model; when the input feature vector X is load demand level data and the latent feature vector Y of the network l is power generation data, equations (1)-(5) constitute a power generation prediction model; 4) Model pre-training, the specific process is: input the unlabeled data into the voltage phase angle prediction model and the power generation prediction model, calculate the loss function according to the reconstructed features output by the voltage phase angle prediction model and the power generation prediction model according to the following formula (6), and use the gradient descent method to minimize the loss function, so as to determine the coding layer parameters; L H (X,Z l )=||X-Z l ||2 2 (6) Among them, L H represents the loss function in the pre-training stage; ||X - Z l ||2 is the Euclidean norm of the residual vector formed by the input feature and the reconstructed feature; 5) Decompose the DC optimal power flow calculation task into two subtasks of predicting the voltage phase angle and predicting the power generation, and the two subtasks are respectively undertaken by the voltage phase angle prediction model and the power generation prediction model; 6) Use the phase angle tag data described above to fine-tune and train the voltage phase angle prediction model f θ,t ; 7) Use the generated power label data to fine-tune and train the generated power prediction model f G,t for fine-tuning training; 8) Determine the teacher model and the student model required for knowledge distillation, specifically: establish a deep neural network with the same scale as the voltage phase angle prediction model for the student model to simultaneously predict the voltage phase angle and the power generation, and use the encoding layer parameters obtained in step 4) to initialize the hidden layer parameters of the deep neural network, and represent the deep neural network as f (θ,G),s (P D →V θ ,P G ); Use the voltage phase angle prediction model f G,t and the power generation prediction model f G,t obtained through fine-tuning training as the teacher model, then the loss function of the voltage phase angle prediction model and the loss function of the power generation prediction model constitute the loss function of the teacher model; use the deep neural network f (θ,G),s (P D →V θ ,P G ) as the student model; 10) Set the maximum number of iterations for knowledge distillation training, and initialize the current number of iterations to 1; 11) Calculate and measure the difference between the predicted values of the student model and the teacher model using the mean square error; where, L θs,θt and L Gs,Gt respectively represent the differences between the phase angle prediction values and the power generation prediction values of the teacher model and the student model, and the subscripts s and t represent the student model and the teacher model respectively; V θs,i and P Gs,j are the predictions of the student model for the phase angle of the i-th node and the power generation of the j-th generator; V θt,i and P Gt,j are the predicted values of the teacher model for the i-th phase angle and the power generation of the j-th generator unit; 12) Calculate the loss function value of the student model, and the calculation formula is expressed as, Among them, L θs and L Gs respectively represent the loss function values of the phase angle and the generated power; 13) Compare the accuracy of the student model and the teacher model according to the following formula through the loss function value of the student model and the loss function value of the teacher model. If the student model is more accurate, the weight λ of the teacher model is zero. Otherwise, the weight λ of the teacher model linearly decreases as the student model is fine-tuned; Among them, λ θ and λ G are the weights of the teacher model; e and e max are the current iteration number and the maximum iteration number of fine-tuning respectively; λ is the weight of the teacher model under the dynamic annealing mechanism, which increases linearly with the iteration number; L θ,t represents the loss function of the voltage phase angle prediction model f θ,t ; L Gt is the loss function of the power generation prediction model f G,t . 14) Calculate the weighted sum of the teacher loss function and the student loss function and use it for parameter update, calculated by formula (21); L = λ θ L θs,θt +(1 - λ θ )L θs +λ G L Gs,Gt +(1 - λ G )L Gs (21) Among them, L is the comprehensive loss function used for knowledge distillation, which is composed of L θs,θt , L θs , L Gs,Gt and L Gs combined; 15) Repeat steps 11) to 14) until the number of iterations reaches the upper limit, indicating that the small sample learning of the optimal power flow is completed.
2. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that, The historical data described in step 1) includes: the load demand levels of each node recorded in several states in the historical operation of the power system, the actual active power output values of each generator, and the phase angles of the voltage phasors of each node.
3. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that, The simulation data described in step 1) includes: the load demand levels of each node, the actual active power output values of each generator, and the phase angles of the voltage phasors of each node solved by the power flow equation.
4. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that 5. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that, In step 2), if there is only load data in the training samples, it is unlabeled data; if there are both load data, power generation data, and phase angle data in the training samples, it is labeled data; the labeled data is divided into phase angle labels and power generation labels according to the phase angle variable and the power generation variable. L H (X, Z l ) = ||X - Z l ||² 2 (8) Among them, L H is the loss function in the pre-training stage; ||X - Z l ||2 is the Euclidean norm of the residual vector formed by the input feature and the reconstructed feature.
6. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that, The loss function used in step 4) is the mean square error function, calculated using the following formula: Among them, P D represents the variable composed of the load demands of each node, V θ represents the vector composed of the voltage phase angles of each node, f θ,t is the voltage phase angle prediction model used to predict the phase angle V θ ; P G represents the power generation vector of the generator, f G,t is the power generation prediction model used to predict the power generation P G ; f θ,t (P D →V θ ) indicates that the knowledge learned by the voltage phase angle prediction model f θ,t is the correlation relationship from the node load to the node phase angle, f G,t (P D →P G ) then indicates that the knowledge learned by the power generation prediction model f G,t is the correlation relationship from the node load to the power generation.
7. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that, In step 6), for the voltage phase angle prediction model f θ,t The loss function involved in the fine-tuning training is an improved dynamic scaling loss function, which is expressed as follows; where n θ is the number of samples of the phase angle label; a i is the attention weight related to the phase angle label data of the i-th node; and p i is the attention weight related to the phase angle prediction value of the i-th node; Ω N is the set of nodes of the power system; V θt,i is the predicted value of the phase angle of the i-th node of the power system; V θ,i represents the phase angle value of the i-th node of the power system in the label data; L θ,t represents the loss function of the voltage phase angle prediction model f θ,t ; Attention weight a i The mathematical relationship with the tag data is as follows: Among them, P[V θ,i is the proportion of the category corresponding to the numerical value of V θ,i , which is determined by dividing the value range of all V θ,i into 20 intervals and counting the number of phase angle labels falling into the same interval; The attention weight p of the i-th node i Determined based on the power operation result of the predicted value p i = V θt,i r (1 - V θt,i ) r + 0.5(11) The input vector and prediction vector of the voltage phase angle prediction model and the power generation prediction model in step 5) are expressed by the following formula:
8. A small-sample learning method for optimal power flow prediction in a power system according to claim 1, characterized in that, In step 7), for the power generation prediction model f G,t The loss function involved in the fine-tuning training is an improved dynamic scaling loss function, which is expressed as follows: Among them, n G is the sample quantity of the power generation quantity label; a j is the attention weight of the j-th generator related to the power generation quantity label; and p j is the attention weight of the j-th generator related to the power generation quantity prediction value; Ω G is the set of power generation equipment in the power system; P Gt,j is the predicted value of the power generation quantity of the j-th generator in the power system by the power generation quantity prediction model f G,t ; P G,j represents the power generation quantity label of the j-th generator in the power system; L Gt is the loss function of the power generation quantity prediction model f G,t . a j The mathematical relationship with the power generation label is that Among them, P[P G,j is the proportion of the category where the numerical value of P G,j lies. It is determined by dividing the value range of all P G,j into 20 intervals and counting the number of power generation labels falling into the same interval; Attention weight p of the j-th generator j Determined based on the power operation result of the predicted value p j = P Gt,j r (1 - P Gt,j ) r + 0.5(14) where r is set to 1. where r is set to 1.
Citation Information
Patent Citations
Power system probabilistic-optimal power flow calculation method based on stacked denoising autoencoder
CN109599872A
Demand response complementary electricity price system and method for high-component flexible load
CN112150190A