Power system state estimation method fusing transfer learning and generative adversarial network

By integrating transfer learning and generating adversarial network methods, the problem of insufficient sample data under the new topology of the power system is solved, the accuracy and generalization ability of the state estimation model are improved, and the generalization of the data-driven model is achieved.

CN120235469APending Publication Date: 2025-07-01FUJIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313624.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When the topology of the power system changes, the sample data is insufficient under the new topology, resulting in the problems of overfitting and degradation of estimation accuracy in traditional data-driven methods.

Method used

Using the method of fusion transfer learning and generative adversarial network (GAN), the DA-CNN-BiLSTM neural network is trained on the original topology to generate measurement data that simulates the new topology, and the CWGAN-div model is used for data augmentation to expand the data set of the new topology. Then, transfer learning is performed, partial layer structure is frozen and non-frozen parts are fine-tuned to obtain the state estimation model of the new topology.

Benefits of technology

The accuracy and generalization ability of the state estimation model under the new topology is improved, and the problem that small sample data sets cannot effectively characterize measurement information is solved, achieving higher generalization of data-driven models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235469A_ABST
    Figure CN120235469A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power system state estimation method fusing transfer learning and a generative adversarial network. The method comprises the following steps: adding measurement noise to power flow data of an original topology to obtain measurement data of the original topology, and preprocessing the measurement data; building a DA-CNN-BiLSTM neural network, and training by using the preprocessed original topology measurement data set to obtain a source domain model; adding measurement noise to the power flow data of the new topology to obtain a small sample new topology measurement data set simulating actual measurement; performing data enhancement on the new topology measurement data set by using a CWGAN-div model to obtain an expanded new topology measurement data set; carrying out transfer learning on the number of layers of a frozen part by using the source domain model and the expanded new topology measurement data set, and carrying out fine tuning training on an unfrozen part to obtain a target domain model; and during online estimation, new topology real-time section measurement information is input into the target domain model to predict and obtain a current system state quantity. According to the invention, the problem of insufficient estimation precision of the original topology estimation model during topology change is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system management, and particularly to a power system state estimation method integrating transfer learning and generative adversarial network. Background Art

[0002] When the topology of the power system changes and the measured data of the new topology is insufficient, the accuracy of the power system state estimation model trained with the original topology will decrease significantly when applied to the new topology.

[0003] As a key link to ensure the safety of user power consumption, with the large-scale access of distributed energy and the increasing proportion of new loads such as electric vehicles in the distribution network, the uncertainty problem of "source-load" becomes increasingly serious. The feeder switches of the distribution network operate more frequently, resulting in changes in the network structure of the distribution network. It is difficult to include all topological data information in the historical database. When the topology changes, there is little known data under the new topology, and traditional data-driven methods relying on historical data for training will cause overfitting. Therefore, it is necessary to invent a method that can make up for the difficulty of state estimation caused by the small sample data of the new topology. Summary of the Invention

[0004] The purpose of the present invention is to provide a power system state estimation method integrating transfer learning and generative adversarial network.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A power system state estimation method integrating transfer learning and generative adversarial network, which includes the following steps:

[0007] Step 1: Add measurement noise to the power flow data of the original topology to obtain the measured data of the original topology, and perform preprocessing including normalization, dimension conversion, and data tiling on the measured data set of the original topology;

[0008] Step 2: Build a DA-CNN-BiLSTM neural network, and use the preprocessed measured data set of the original topology to train the DA-CNN-BiLSTM neural network to obtain a source domain model;

[0009] Step 3: Add measurement noise to the power flow data of the new topology to obtain a small sample measured data set of the new topology that simulates actual measurements;

[0010] Step 4: Use the CWGAN-div model to perform data augmentation on the measured data set of the new topology to obtain an expanded measured data set of the new topology;

[0011] Step 5: Based on the source domain model and the expanded measured data set of the new topology, perform transfer learning, freeze some layers and perform fine-tuning training on the unfrozen part to obtain a new topology state estimation model, that is, a target domain model;

[0012] Step 6, during online estimation, input the real-time sectional measurement information of the new topology into the target domain model to obtain the current system state variables. Compare with the true values and calculate the performance indicators, and then output the results.

[0013] Further, step 1 specifically includes the following steps:

[0014] Step 1-1, perform power flow calculation on the original topology structure from the historical operating load to obtain multi-sectional historical power flow data as the true value of the original topology power flow;

[0015] Step 1-2, add Gaussian white noise to the multi-sectional power flow data to generate multi-sectional measurement information, simulating the original topology measurement data set; and perform preprocessing such as normalization, dimension conversion, and data tiling on the original topology measurement data set.

[0016] Further, the DA-CNN-BiLSTM neural network in step 2 includes two parts. One part is the SE-CNN formed by the SE-Net module with channel attention mechanism and CNN; the other part is the FA-BiLSTM formed by the feature attention (FA) module and BiLSTM; SE-CNN is used to extract spatial features; FA-BiLSTM is used to bidirectionally explore temporal features. The dual attention mechanism strengthens the correlation between the input features and target features of the CNN-BiLSTM network, improving the model performance.

[0017] Further, the SE-CNN in step 2 includes a two-dimensional grouped convolution operation module and a squeeze-and-excitation module, and specifically performs the following steps:

[0018] Step 2-1-1, the two-dimensional grouped convolution operation module divides the input data X into g groups, and can generate output feature maps g times that of the conventional convolution with the same number of parameters, and stack them to obtain the unweighted feature map U;

[0019] Step 2-1-2, the squeeze-and-excitation module aggregates the information contained in each channel through the squeeze operation to form a multi-dimensional statistic containing the importance information of each channel;

[0020] Specifically, the squeeze in the present invention is implemented by global average pooling and global max pooling, and the data obtained by pooling them separately are added. The pooling process is to compress the unweighted feature map U along the spatial dimension H×W to obtain a C-dimensional statistic z∈R C , where C represents the number of groups of convolution kernels, each group of convolution kernels forms a 1D feature map, and the c-th element is:

[0021]

[0022] where u c (i, j) is the element in the i-th row and j-th column of the input data.

[0023] Step 2-1-3: A parameterized gating mechanism is constructed through two fully connected layers. The first fully connected layer reduces the input C-dimensional statistic z to C / r dimensions (followed by a ReLU activation function), and the second fully connected layer raises its dimension to C dimensions (followed by a sigmoid activation function); that is:

[0024] s = F ex (z, W) = sigmoid(f(z, W)) = sigmoid(W2Relu(W1z)) (7)

[0025] where W1 and W2 are the parameters of the two fully connected layers respectively, r is the scaling parameter, and s is the weight representing the importance of each channel in the feature map U;

[0026] Step 2-1-4: Multiply the obtained weight s and the unweighted feature map U channel by channel to form the final output The specific expression is as follows:

[0027]

[0028] where c represents a random number, represents the c-th channel weighted feature map; U c represents the c-th unweighted feature map output by the transformation part; s c represents the c-th weight output by the excitation part; C represents the number of convolution kernel groups, represents the C-th channel weighted feature map, which is also the last output channel weighted feature map.

[0029] Furthermore, in Step 2, the FA-BiLSTM includes a feature attention module FA and a BiLSTM network; the specific implementation is as follows:

[0030] Step 2-2-1: Use two-dimensional conventional convolution operation (followed by a sigmoid activation function) to reduce the dimension to a one-dimensional attention weight map containing feature information:

[0031] e = sigmoid(W e X1 + b e ) (10)

[0032] where e is the combination of attention weight coefficients corresponding to a single feature on the input feature map; W e is the parameter of the convolution layer; b e is the set bias vector.

[0033] Step 2-2-2: Multiply the input feature map X by the obtained weight map e to get the weighted feature map U1;

[0034] U1 = eX (11).

[0035] Specifically, aggregate the information contained in the feature map through two-dimensional conventional convolution to form a one-dimensional statistic that includes the importance information of each feature on the feature map;

[0036] Step 2-2-3: Normalize the features of the obtained weighted feature map through the Softmax function to reduce the influence of the data dimension on the output result; prevent overfitting. The normalization formula for a single feature x is:

[0037]

[0038] Step 2-2-4: Pass the normalized result x' through the activation function swish to add a non-linear factor to solve the defect of insufficient expression ability of the linear model, and obtain a new weighted feature map U2.

[0039] Step 2-2-5: Perform grouped convolution on the new weighted feature map U2, set the vaild calculation method for pointwise convolution, further extract relevant features, and perform normalization processing (followed by the swish activation function) to obtain the final output feature map

[0040] Step 2-2-6: Flatten the data through the flatten function to obtain a feature sequence with allocated dynamic weights, and input it into the BiLSTM model; when processing long sequence data, the forward and backward LSTM layers of the BiLSTM process each time step, and then linearly combine the outputs of all time steps to form the final output.

[0041] Furthermore, Step 3 specifically includes the following steps:

[0042] Step 3-1: Perform power flow calculation on the new topological structure to obtain multi-section historical power flow data as the true value of the new topological power flow;

[0043] Step 3-2: Add Gaussian white noise to the true value of the new topological power flow to generate multi-section measurement information, and obtain a small-sample new topological measurement data set that simulates actual measurements; and perform normalization processing, dimension conversion, and data tiling preprocessing on the small-sample new topological measurement data set.

[0044] Further, in step 4, the Wasserstein divergence is introduced on the basis of GAN to replace the JS divergence. At the same time, the conditional variable y is incorporated into the conditional generative adversarial network to generate the CWGAN-div model that conforms to the required data type. The conditional information and random noise are input into the generator of the CWGAN-div model, enabling the generator to generate new samples under the guidance of the conditions. Then, the generated samples, the target data measured by the small-sample new topology, and the conditional information are jointly input into the discriminator for determination.

[0045] The objective function of the CWGAN-div model is as follows:

[0046]

[0047] where: E[·] is the expectation function; x is the input data; z is the random noise; D[·] is the discriminator function; G[z] is the generator function; P r is the data distribution of x; P z is the data distribution of z, and P z is the normal distribution of N(0, 1); y is the conditional information, that is, y is the measurement data that conforms to the true measurement distribution; is the sampling of the linear combination of the real data and the generated data; p u is the data distribution of.

[0048] Further, in step 5, during transfer learning, the structure before the fully connected layer of the source domain model is frozen to save parameters, and the fully connected layer is used for fine-tuning the target task to obtain the target domain model.

[0049] The present invention adopts the above technical solutions, a power system state estimation method that combines transfer learning and conditional Wasserstein generative adversarial network (CWGAN-div), combines the advantages of both to improve the estimation performance of the model under topological transformation, and enables the data-driven model to have higher generalization. The CWGAN-div model is used to perform data augmentation on the small-sample measurement data of the new topology after topological transformation, adds conditional information to make the generated data close to the distribution of the measurement data, and the two are mixed to obtain a new topology data set after data augmentation for transfer learning training, improving the model training effect, and solving the problem that the small-sample data set may not be able to effectively represent the measurement information of the data set. Through transfer learning, the effective information in the measurement data of the original topology can be extracted and applied to the model training of the new topology, thereby improving the accuracy of the new topology state estimation model; when fine-tuning training is performed by transfer learning, the dual attention convolutional bidirectional long short-term memory neural network (DA-CNN-BiLSTM) model is used as the basic model. This model strengthens the spatio-temporal feature screening ability of CNN-BiLSTM through the dual attention mechanism, screens important features according to the obtained importance measure, and dynamically mines the correlation between the measured quantity and the state quantity, and has better feature extraction performance than a single model. The present invention solves the problem of insufficient estimation accuracy of the original topology estimation model when the topology changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The following further describes the present invention in detail with reference to the drawings and specific embodiments;

[0051] Figure 1 It is a schematic flowchart of the power system state estimation method that combines transfer learning and generative adversarial network of the present invention;

[0052] Figure 2 It is a schematic diagram of the CWGAN-div model structure of the present invention;

[0053] Figure 3 It is a schematic diagram of the principle architecture of the DA-CNN-BiLSTM transfer learning of the present invention;

[0054] Figure 4 It is a schematic diagram of the SE-CNN structure of the present invention.

[0055] Figure 5 It is a schematic diagram of the FA-BiLSTM structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application.

[0057] As Figures 1 to 5As shown in one of them, the present invention discloses a power system state estimation method integrating transfer learning and generative adversarial network, and trains an estimation model for a new topology under topological transformation. First, a CWGAN-div model is constructed, and the new topology data is input into the model to direct generate qualified augmented data for data augmentation. Secondly, a topology is selected as the reference topology (source domain), and the historical data of the source domain is used for offline training; based on the source domain model, the new topology samples with data augmentation are fine-tuned using transfer learning to obtain a new topology state estimation model. Finally, when applying online, the real-time measurements of the new topology collected are input into the estimation model to obtain the state quantities at the current moment.

[0058] The power system state estimation method integrating transfer learning and generative adversarial network disclosed by the present invention includes the following steps:

[0059] Step 1, measurement noise is added to the power flow data of the original topology to obtain the original topology measurement data, and preprocessing such as normalization processing, dimension conversion, and data tiling is performed on the original topology measurement data set;

[0060] Step 2, a DA-CNN-BiLSTM neural network is built, and the DA-CNN-BiLSTM neural network is trained using the preprocessed original topology measurement data set to obtain a source domain model;

[0061] Step 3, measurement noise is added to the power flow data of the new topology to obtain a small sample new topology measurement data set simulating actual measurements;

[0062] Step 4, the CWGAN-div model is used to perform data augmentation on the new topology measurement data set to obtain an augmented new topology measurement data set;

[0063] Step 5, based on the source domain model and the augmented new topology measurement data set, transfer learning is performed, some layers are frozen and the unfrozen parts are fine-tuned to obtain a new topology state estimation model, that is, a target domain model;

[0064] Step 6, during online estimation, the real-time cross-section measurement information of the new topology is input into the target domain model to obtain the current system state quantities, and the results are output by comparing with the true values and calculating performance indicators.

[0065] Further, step 1 specifically includes the following steps:

[0066] Step 1-1, perform power flow calculation on the original topology structure from historical operating loads to obtain multi-section historical power flow data as the true value of the original topology power flow;

[0067] Step 1-2, Gaussian white noise is added to the multi-section tidal current data to generate multi-section measurement information, simulating the original topological measurement data set; and preprocessing such as normalization, dimension conversion, and data tiling is performed on the original topological measurement data set.

[0068] Furthermore, the DA-CNN-BiLSTM neural network in Step 2 includes two parts. One part is the SE-CNN formed by the SE-Net module with channel attention mechanism and CNN; the other part is the FA-BiLSTM formed by the feature attention (FA) module and BiLSTM; the SE-CNN is used to extract spatial features; the FA-BiLSTM is used to bidirectionally explore temporal features. The correlation between the input features and the target features of the CNN-BiLSTM network is strengthened through the dual attention mechanism, improving the model performance.

[0069] Furthermore, the SE-CNN in Step 2 includes a two-dimensional grouped convolution operation module and a squeeze-and-excitation module, and specifically performs the following steps:

[0070] Step 2-1-1, the two-dimensional grouped convolution operation module divides the input data X into g groups, and can generate output feature maps g times that of the conventional convolution with the same number of parameters, and stacks them to obtain the unweighted feature map U;

[0071] Step 2-1-2, the squeeze-and-excitation module aggregates the information contained in each channel through the squeeze operation to form a multi-dimensional statistic containing the importance information of each channel;

[0072] Specifically, the squeeze (Squeeze) of the present invention is implemented by global average pooling and global maximum pooling, and the data obtained by pooling them separately are added. The pooling process is to compress the unweighted feature map U along the spatial dimension H×W to obtain a C-dimensional statistic z∈R C , where C represents the number of groups of convolution kernels, and each group of convolution kernels forms a 1D feature map, and the c-th element is:

[0073]

[0074] where, u c (i,j) is the element in the i-th row and j-th column of the input data.

[0075] Step 2-1-3, a parameterized gating mechanism is formed by two fully connected layers. The first fully connected layer reduces the input C-dimensional statistic z to C / r dimensions (followed by a ReLU activation function), and the second fully connected layer raises it to C dimensions (followed by a sigmoid activation function); that is:

[0076] s = F ex(z, W) = sigmoid(f(z, W)) = sigmoid(W2Relu(W1z)) (7)

[0077] Among them, W1 and W2 are the parameters of two fully connected layers respectively, r is the scale parameter, and s is the weight representing the importance of each channel in the feature map U;

[0078] Step 2-1-4: Multiply the obtained weight s and the unweighted feature map U channel by channel to form the final output Specifically, the expression is as follows:

[0079]

[0080] Among them, c represents a random number, represents the weighted feature map of the c-th channel; U c represents the unweighted feature map of the c-th output of the transformation part; s c represents the weight of the c-th output of the excitation part; C represents the number of convolution kernel groups, represents the weighted feature map of the C-th channel, which is also the weighted feature map of the last output channel.

[0081] Furthermore, in step 2, the FA-BiLSTM includes a feature attention module FA and a BiLSTM network; the specific implementation is as follows:

[0082] Step 2-2-1: Use two-dimensional conventional convolution operation (followed by sigmoid activation function) to reduce the dimension to a one-dimensional attention weight map containing feature information:

[0083] e = sigmoid(W e X1 + b e ) (10)

[0084] Among them, e is the combination of attention weight coefficients corresponding to a single feature on the input feature map; W e is the parameter of the convolution layer; b e is the set bias vector.

[0085] Step 2-2-2: Multiply the input feature map X and the obtained weight map e to obtain the weighted feature map U1;

[0086] U1 = eX (11)

[0087] Specifically, the information contained in the feature map is aggregated through two-dimensional conventional convolution to form a one-dimensional statistic containing the importance information of each feature on the feature map;

[0088] Step 2-2-3: Normalize the features of the obtained weighted feature map through the Softmax function to reduce the impact of data dimensions on the output results; prevent overfitting. The normalization formula for a single feature x is:

[0089]

[0090] Step 2-2-4: Pass the normalized result x′ through the activation function swish to introduce non-linear factors to address the deficiency of the linear model's expressive ability, obtaining a new weighted feature map U2.

[0091] Step 2-2-5: Perform grouped convolution on the new weighted feature map U2, set the vaild calculation method for pointwise convolution, further extract relevant features, and perform normalization (followed by the swish activation function) to obtain the final output feature map

[0092] Step 2-2-6: Flatten the data through the flatten function to obtain a feature sequence with assigned dynamic weights, and input it into the BiLSTM model; when processing long sequence data, the forward and backward LSTM layers of the BiLSTM process each time step, and then linearly combine the outputs of all time steps to form the final output.

[0093] Furthermore, Step 3 specifically includes the following steps:

[0094] Step 3-1: Perform power flow calculation on the new topological structure to obtain multi-section historical power flow data as the true value of the new topological power flow;

[0095] Step 3-2: Add Gaussian white noise to the true value of the new topological power flow to generate multi-section measurement information, obtaining a small-sample new topological measurement data set that simulates actual measurements; and perform normalization processing, dimension conversion, and data tiling preprocessing on the small-sample new topological measurement data set.

[0096] Furthermore, in Step 4, the Wasserstein divergence is introduced to replace the JS divergence on the basis of the GAN, and at the same time, the conditional variable y is combined with the conditional generative adversarial network to generate the CWGAN-div model that conforms to the required data type; input the conditional information and random noise into the generator of the CWGAN-div model, so that the generator generates new samples under the condition of guidance; then input the generated samples, the target data of the small-sample new topological measurement, and the conditional information into the discriminator for determination.

[0097] The objective function of the CWGAN-div model is:

[0098]

[0099] where: E[·] is the expectation function; x is the input data; z is the random noise; D[·] is the discriminator function; G[z] is the generator function; P r is the data distribution of x; P z is the data distribution of z, P z is the normal distribution of N(0,1); y is the conditional information, that is, y is the measurement data conforming to the true measurement distribution; is the sampling of the linear combination of the real data and the generated data; p u is the data distribution of.

[0100] Furthermore, in the transfer learning in step 5, the structure before the fully connected layer of the source domain model is frozen to save parameters, and the fully connected layer is used for fine-tuning of the target task to obtain the target domain model.

[0101] The specific principle of the present invention is described in detail as follows:

[0102] The power system state estimation method of the present invention is composed of the CWGAN-div model and the transfer learning model. The basic model DA-CNN-BiLSTM used in transfer learning is mainly divided into two parts. One part is SE-CNN formed by the SE-Net module and CNN; the other part is FA-BiLSTM formed by the FA module and BiLSTM.

[0103] CWGAN-div model: GAN is an unsupervised deep learning model based on statistics and game theory. Through continuous game between the generator G and the discriminator D, the potential feature connection between data is obtained, and "pseudo data" conforming to the corresponding distribution law is generated. The GAN objective function is:

[0104]

[0105] In formula (1): E[·] is the expectation function; x is the input data; z is the random noise; D[·] is the discriminator function; G[z] is the generator function; P r is the data distribution of x; P z is the data distribution of z. During the iteration process, G hopes that the objective function is as large as possible, and D hopes that the objective function is as small as possible. The two reach the optimum in the confrontation process.

[0106] When the conventional GAN uses the JS (Jensen-Shannon) divergence for training, problems such as model collapse and gradient disappearance will occur. Therefore, WGAN-div solves the above problems by introducing the Wasserstein divergence instead of the JS divergence, making the updates of G and D more stable. The WGAN-div objective function is:

[0107]

[0108] In formula (2): is the sampling of the linear combination of real data and generated data; p u is the data distribution; Experimental verification shows that the model has the best effect when k = 2 and p = 6.

[0109] The conventional GAN model cannot control the type of generated data. Therefore, in the present invention, a conditional variable y is added in combination with a conditional generative adversarial network to generate data of the required type to form a CWGAN-div model. The objective function of the model is:[[]]

[0110]

[0111] In formula (3): P z is the normal distribution of N(0,1); y is the conditional information, and in the present invention, y is the measurement data conforming to the real measurement distribution. In the present invention, the conditional information and random noise are first input into the generator, so that the generator generates new samples under the guidance of the conditions; then the generated samples, the target data of the new topology measurement of the small samples, and the conditional information are jointly input into the discriminator for determination. The CWGAN-div framework is as Figure 2 shown.

[0112] Transfer learning model: There are two basic concepts in transfer learning, namely domain D and task T. Among them, domain D is divided into the source domain and D S the target domain D T two parts, and task T is divided into the source domain task and T S the target domain task T T two parts. Different domains D have different feature spaces X and marginal probability distributions P(X), and the expression is:[[]]

[0113]

[0114] Different tasks T have different label spaces Y and conditional probability distributions P(X|Y) corresponding to the labels. Since it is difficult to give the specific form of the conditional probability distribution P(X|Y), the mapping function f is usually used to replace the conditional probability distribution P(X|Y), and the mapping function f learns from the feature-label pairs. The expression is:[[]]

[0115]

[0116] Therefore, the definition of transfer learning is: Given the source domain D S and the source task T S , the target domain D T and the target task T T , in D S ≠D T or T S≠T T In the case of, use D S and T S model to improve the target task function f T effect.

[0117] As Figure 3 shown, based on the transfer learning graph of the dual attention convolutional bidirectional long short-term memory neural network (DA-CNN-BiLSTM) model, the structure before the fully connected layer of the source domain model is frozen to save parameters, and the fully connected layer is used for fine-tuning of the target task to obtain the target domain model.

[0118] The DA-CNN-BiLSTM model structure is mainly composed of SE-CNN and FA-BiLSTM. The SENet module with channel attention mechanism is added to the CNN network to form SE-CNN to extract spatial features; the FA module is added to the BiLSTM network to form FA-BiLSTM to explore temporal features bidirectionally. The dual attention mechanism strengthens the correlation between the input features and target features of the CNN-BiLSTM network, improving the model performance. The SE-CNN structure is as Figure 4 shown, divided into four parts:

[0119] 1) Transformation: As Figure 4 shown, the left part is the two-dimensional grouped convolution operation process. The input data X is divided into g groups, and the operation with the same number of parameters can generate g times the output feature map of the conventional convolution, and the unweighted feature map U is obtained by stacking.

[0120] 2) Squeeze: Through the squeeze operation, the information contained in each channel is aggregated in the spatial dimension to form a multi-dimensional statistic containing the importance information of each channel. The squeeze in the present invention is realized by global average pooling and global max pooling, and the data obtained by pooling them separately are added. The pooling process is to compress the unweighted feature map U along its spatial dimension H×W to obtain the multi-dimensional statistic z∈R C , where the c-th element is:

[0121]

[0122] In the formula, u c (i,j) is the element in the i-th row and j-th column of the input data.

[0123] 3) Excitation: A parameterized gating mechanism is formed by two fully connected layers. The first fully connected layer reduces the input C-dimensional statistic z to C / r dimensions (followed by the ReLU activation function), and the second fully connected layer raises it to C dimensions (followed by the sigmoid activation function), that is:

[0124] s = F ex (z, W) = sigmoid(f(z, W)) = sigmoid(W2Relu(W1z)) (7)

[0125] In Equation (7), W1 and W2 are the parameters of two fully connected layers respectively, r is the scale parameter, and s is the weight representing the importance of each channel in the feature map U.

[0126] 4) Scaling: Multiply the obtained weight s by the unweighted output U channel by channel to form the final output

[0127] The FA-BiLSTM module structure is as Figure 5 shown, and it contains two parts, namely:

[0128] 1) Scaling: Aggregate the information contained in the feature map through two-dimensional conventional convolution to form a one-dimensional statistic containing the information of the importance of each feature on the feature map. During the process, the input C1'-dimensional feature map X1 is reduced to a one-dimensional attention weight map containing feature information by two-dimensional conventional convolution operation (followed by sigmoid activation function):

[0129] e = sigmoid(W e X1 + b e ) (10)

[0130] In Equation (10), e is the combination of attention weight coefficients corresponding to a single feature on the input feature map; W e is the parameter of the convolutional layer; b e is the set bias vector.

[0131] Then, multiply the input feature map X by the obtained weight map e to get the weighted feature map U1:

[0132] U1 = eX (11)

[0133] 2) Transformation: Normalize the features of the obtained weighted feature map through the Softmax function to reduce the influence of the data dimension on the output result and prevent overfitting. The normalization formula for a single feature x is:

[0134]

[0135] Pass the normalization result x' through the activation function swish to add non-linear factors to solve the defect of insufficient expression ability of the linear model, and obtain a new weighted feature map U2.

[0136] Then, perform grouped convolution on U2, set the valid calculation method for pointwise convolution, further extract relevant features, and perform normalization (followed by the swish activation function) to obtain the final output feature map

[0137] Finally, Perform data flattening through the flatten function to obtain a feature sequence with allocated dynamic weights, and input it into the BiLSTM model. When processing long sequence data, the forward and backward LSTM layers of BiLSTM need to process each time step, and then linearly combine the outputs of all time steps to form the final output. However, since the subsequent information processing of BiLSTM depends on the processing results of previous signals, it may lead to insufficient attention to some important features or being restricted by the early processing results. By combining with the FA module to form FA-BiLSTM, it can dynamically allocate attention weights to individual features on the input feature channels of BiLSTM, dynamically focus on the key information in the sequence, rather than treating all information equally. This helps to solve the problem of information dilution in long sequences, enabling the model to better handle the temporal dependence in the sequence, thereby improving the concurrent performance.

[0138] Effect description: To verify the effectiveness of the method of the present invention, the proposed method is experimentally tested in IEEE33 and IEEE118 node distribution systems.

[0139] The source domain is set as follows: Using the household load data of London, UK as the historical database, with a 5-minute sample section, 3000 groups of actual loads are generated. The training set and test set are divided in a ratio of 7:3 for training. The power flow results obtained by reducing the multi-section actual loads to node loads for power flow calculation are used as the true values. Verification is carried out by adding different Gaussian noises to the power flow true values to simulate actual measurements and bad data.

[0140] The target domain is set as follows: Assume that there are 50 groups of historical data in the training set after topological changes and 300 groups in the test set. The dataset after data augmentation is fine-tuned based on the source domain model to obtain a new topological state estimation model.

[0141] Based on the original standard model, the present invention changes the topological connection relationship of the network to construct a new topological structure as the target domain topology. The construction method of the new topology and its relationship with the original topology are shown in Table 1.

[0142] Table 1 New topological structure

[0143] Transformed topology Relationship with the connection of the original topology foundation 33-node topology 1 Connect 9 - 15, 12 - 22, disconnect 14 - 15, 11 - 12 33-node topology 2 Connect 9 - 15, 25 - 29, disconnect 14 - 15, 28 - 29 118-node topology 1 Connect 17 - 27, 110 - 118, disconnect 16 - 17, 109 - 110 118-node topology 2 Connect 8 - 24, 73 - 91, disconnect 23 - 24, 72 - 73

[0144] Performance test of the CWGAN-div model: To intuitively reflect the performance of the CWGAN-div model, the present invention quantitatively tests the performance of the CWGAN-div model in generating data using the 2-norm error, calculates the 2-norm errors of the generated data and the target measurement data with respect to the real data, and its calculation formula is:

[0145]

[0146] In the formula, T is the number of time sections, N is the number of measurements, is the measurement data or the generated data, and x is the real data. Taking the 2-norm error of the node voltage amplitude (per unit value) in different systems and topologies as an example, as shown in Table 2.

[0147] Table 2 2-Norm Errors of Different Topologies

[0148]

[0149] Estimation accuracy test: The accuracy test of the present invention adds noise to 50 groups of multi-section power flow data to simulate actual measurements, and the parameter settings are as follows: Gaussian noise with a standard deviation of 0.02 and a mean of 0 is added to construct power measurements; errors with standard deviations of 0.005 and 0.002 and a mean of 0 are added to construct voltage amplitude and phase angle measurements.

[0150] The method of the present invention is compared with WLS, DA-CNN-BiLSTM (abbreviated as DCBiLSTM) directly trained with small samples, and DCBiLSTM-TL with small sample direct transfer learning. The mean absolute error (MAE), maximum absolute error (MaxAE), and root mean square error (RMSE) are used to compare the differences between different methods.

[0151]

[0152]

[0153] In the formula, is the estimated value, x is the real value, i is the node number, t is the time section, and T is the total number of time sections.

[0154] The overall error indicators of each method in the IEEE33 topology 1 test set are shown in Table 2.

[0155] Table 3 Indexes of State Estimation Results for 33-Node Topology 1

[0156]

[0157] As can be seen from Table 3, compared with DCBiLSTM-TL, DCBiLSTM, and WLS, the average absolute error of the voltage amplitude of the method of the present invention decreased by 52.94%, 88.51%, and 70% respectively; the average absolute error of the voltage phase angle decreased by 27.56%, 85.77%, and 73.8% respectively. Moreover, its maximum absolute error and root mean square error are also smaller than those of the other three methods.

[0158] The overall error indicators of each method in the IEEE33 topology 2 test set are shown in Table 4.

[0159] Table 4 Indexes of state estimation results for 33-node topology 2

[0160]

[0161] As can be seen from Table 4, compared with DCBiLSTM-TL, DCBiLSTM, and WLS, the average absolute error of the voltage amplitude of the method of the present invention decreased by 36.95%, 90.06%, and 74.56% respectively; the average absolute error of the voltage phase angle decreased by 28.52%, 87.24%, and 74.75% respectively. Moreover, its maximum absolute error and root mean square error are also smaller than those of the other three methods.

[0162] The overall error indicators of each method in the IEEE118 topology 1 test set are shown in Table 5.

[0163] Table 5 Indexes of state estimation results for 118-node topology 1

[0164]

[0165]

[0166] As can be seen from Table 5, compared with DCBiLSTM-TL, DCBiLSTM, and WLS, the average absolute error of the voltage amplitude of the method of the present invention decreased by 30.3%, 90.41%, and 68.91% respectively; the average absolute error of the voltage phase angle decreased by 27.35%, 87.6%, and 68.65% respectively. Moreover, its maximum absolute error and root mean square error are also smaller than those of the other three methods. At the same time, the maximum absolute error of the amplitude of DCBiLSTM-TL based on transfer learning with small samples is 10.96% higher than that of the traditional WLS.

[0167] The overall error indicators of each method in the IEEE118 topology 2 test set are shown in Table 6.

[0168] Table 6 Indexes of state estimation results for 118-node topology 2

[0169]

[0170] As can be seen from Table 6, compared with DCBiLSTM-TL, DCBiLSTM, and WLS, the average absolute error of the voltage amplitude of the method of the present invention decreased by 45.65%, 89.91%, and 68.75% respectively; the average absolute error of the voltage phase angle decreased by 36.19%, 87.71%, and 65.04% respectively. Moreover, its maximum absolute error and root mean square error are also smaller than those of the other three methods.

[0171] Through the above simulations, it is proved that when the measurement data of the power system is insufficient during the transformation of the topology, the estimation accuracy of DCBiLSTM is the worst and cannot meet the requirements; the accuracy of DCBiLSTM-TL is better than that of the traditional WLS algorithm in most cases, but its maximum error is relatively large in the 118-node topology 1, indicating that when the measurement data is less, relying solely on transfer learning may produce relatively large deviations at certain moments; the three indicators of the method of the present invention are the lowest in the simulation, having higher estimation accuracy than other methods, and being able to adapt to different topological changes with a certain degree of generalization.

[0172] Robustness test: Bad data is an inevitable situation in the operation of the power system. In the present invention, a part of the measurements are randomly selected in the test set and mixed Gaussian errors are added to simulate the measurements containing bad data, and the robustness of the method of the present invention is tested. The bad data simulation method is to add 20% mixed Gaussian error to the power measurement, 5% mixed Gaussian error to the amplitude measurement, and 2% Gaussian error to the phase angle measurement. Table 7 shows the average absolute error of each method on the node topology after adding bad data to the test set.

[0173] Table 7 Robustness test of different methods

[0174]

[0175] As can be seen from the above table, the average absolute error of the method of the present invention is the lowest in the above four cases. Taking the 118-node topology 1 as an example, the average absolute error of the amplitude decreased by 35.48%, 91.45%, and 80.86% compared with DCBiLSTM-TL, DCBiLSTM, and WLS respectively, and the average absolute error of the phase angle decreased by 45.82%, 86.71%, and 52.61% respectively. The results show that the method of the present invention has stronger robustness than DCBiLSTM-TL, DCBiLSTM, and WLS in the presence of bad data, and can better meet the safe operation requirements of the power system.

[0176] Computational efficiency test: As can be seen from Table 8, the computational efficiency of the traditional WLS is significantly affected by the system scale, while the method of the present invention is less affected. When the two algorithms adopt the same measurement configuration, the WLS estimation time increases from 0.0152 s to 0.0549 s after the system nodes change, and the method of the present invention increases from 0.0356 s to 0.0424 s. In IEEE118, the computational efficiency of the method of the present invention is 12.5% higher than that of WLS. It can be concluded that as the system scale increases, the computational efficiency of WLS will gradually lag behind the method of the present invention. Therefore, the method of the present invention better meets the real-time requirements of large-scale system state estimation.

[0177] Table 8 Single-section calculation time of different methods

[0178]

[0179] The state estimation driven by data of the present invention can fully exploit the effective information in historical data. By training historical data to obtain the mapping relationship between measurement quantities and state quantities for state estimation without complex physical modeling, it has gradually become a research hotspot. In the power system state estimation method that combines transfer learning and CWGAN-div, the CWGAN-div model performs data augmentation on the new topology dataset, amplifying some less effective sample information, preventing it from being ignored due to too few samples during transfer learning, and effectively enhancing the dataset's ability to represent information; transfer learning applies the effective information of the original topology dataset with sufficient data samples to the new topology model with few samples. The combination of the two enables the data-driven model to have good generalization performance in the case of topological changes and insufficient measurement data for the new topology. Finally, detailed and objective experimental analysis through simulation shows the effectiveness and feasibility of the proposed method, with certain application prospects.

[0180] The present invention adopts the above technical solutions, and provides a power system state estimation method that combines transfer learning and conditional Wasserstein generative adversarial network (CWGAN-div). By integrating the advantages of both, the estimation performance of the model under topological transformation is improved, enabling the data-driven model to have higher generalization ability. The CWGAN-div model is used to perform data augmentation on the small-sample measurement data of the new topology after topological transformation. By adding conditional information, the generated data is made to be close to the distribution of the measurement data. The two are mixed to obtain a new topology data set after data augmentation for transfer learning training, improving the model training effect and solving the problem that the small-sample data set may not be able to effectively represent the measurement information of the data set. Through transfer learning, the effective information in the measurement data of the original topology can be extracted and applied to the model training of the new topology, thereby improving the accuracy of the new topology state estimation model. When performing fine-tuning training with transfer learning, the dual-attention convolutional bidirectional long short-term memory neural network (DA-CNN-BiLSTM) model is used as the basic model. This model strengthens the spatio-temporal feature screening ability of CNN-BiLSTM through the dual-attention mechanism, screens important features according to the obtained importance measure, and dynamically mines the correlation between the measured quantity and the state quantity, having better feature extraction performance than a single model. The present invention solves the problem of insufficient estimation accuracy of the original topology estimation model during topological changes.

[0181] Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Without conflict, the embodiments and features in the present application can be combined with each other. Usually, the components of the embodiments of the present application described and illustrated in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

Claims

1. A power system state estimation method integrating transfer learning and generative adversarial network, characterized by: It includes the following steps: Step 1, adding measurement noise to the power flow data of the original topology to obtain the original topology measurement data, and performing normalization, dimension conversion and data tiling preprocessing on the original topology measurement data set; Step 2: Build a DA-CNN-BiLSTM neural network and use the preprocessed original topology measurement data set to train the DA-CNN-BiLSTM neural network to obtain a source domain model; Step 3, adding measurement noise to the power flow data of the new topology to obtain a small sample new topology measurement data set that simulates actual measurement; Step 4, using the CWGAN-div model to perform data enhancement on the new topology measurement dataset to obtain an expanded new topology measurement dataset; Step 5: Perform transfer learning based on the source domain model and the expanded new topology measurement data set, freeze some layers and perform fine-tuning training on the unfrozen parts to obtain a new topology state estimation model, i.e., the target domain model. Step 6: During online estimation, the real-time cross-sectional measurement information of the new topology is input into the target domain model to obtain the current system state quantity, which is then compared with the true value and the performance index is calculated to output the result.

2. The power system state estimation method integrating transfer learning and generative adversarial network according to claim 1 is characterized in that: Step 1 specifically includes the following steps: Step 1-1, calculate the flow of the original topology structure from the historical operating load, and obtain the multi-section historical flow data as the true value of the original topology flow; Step 1-2, Gaussian white noise is added to the multi-section flow data to generate multi-section measurement information, simulating the original topological measurement data set; and the original topological measurement data set is normalized, dimensionally converted and pre-processed with data tiling.

3. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 1 is characterized in that: The DA-CNN-BiLSTM neural network in step 2 consists of two parts, one of which is SE-CNN formed by the SE-Net module of the channel attention mechanism and CNN; the other is FA-BiLSTM formed by the feature attention module FA and BiLSTM; SE-CNN is used to extract spatial features; FA-BiLSTM is used to bidirectionally explore temporal features.

4. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 3 is characterized in that: In step 2, SE-CNN includes a two-dimensional grouped convolution operation module and a squeeze-excitation module, and specifically performs the following steps: Step 2-1-1, the two-dimensional grouped convolution operation module divides the input data X into g groups, and performs operations with the same parameter amount to generate an output feature map that is g times the output of the conventional convolution, and stacks them to obtain an unweighted feature map U; Step 2-1-2, the squeeze-excitation module aggregates the information contained in each channel in the spatial dimension through a squeeze operation to form a multidimensional statistic z containing the importance information of each channel; Step 2-1-3, a parameterized gating mechanism is constructed through two fully connected layers. The first fully connected layer reduces the input C-dimensional statistic z to C / r dimensions (followed by the ReLU activation function), and the second fully connected layer increases its dimension to C dimensions (followed by the sigmoid activation function); that is: s=F ex (z,W)=sigmoid(f(z,W))=sigmoid(W2Relu(W1z)) (7) Among them, W1 and W2 are the parameters of the two fully connected layers, r is the scale parameter, and s is the weight representing the importance of each channel in the feature map U; Step 2-1-4, multiply the obtained weight s by the unweighted feature map U by channel to form the final output The specific expression is as follows: Where c represents a random number. represents the weighted feature map of the cth channel; U c represents the cth unweighted feature map output by the transformation part; s c represents the cth weight of the output of the excitation part; C represents the number of convolution kernel groups, Represents the Cth channel weighted feature map, which is also the last output channel weighted feature map.

5. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 4 is characterized in that: In step 2-1-2, the squeezing operation is implemented by global average pooling and global maximum pooling, and the data after the two poolings are added; the pooling process is to compress the unweighted feature map U along the spatial dimension H×W to obtain the C-dimensional statistic z∈R C , C represents the number of convolution kernel groups, each group of convolution kernels forms a 1-dimensional feature map, where the cth element is: Among them, u c (i,j) is the element in the i-th row and j-th column of the input data.

6. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 3 is characterized in that: In step 2, FA-BiLSTM includes a feature attention module FA and a BiLSTM network; the specific steps are as follows: Step 2-2-1, use two-dimensional convolution operation to reduce the dimension into a one-dimensional attention weight map e containing feature information: e=sigmoid(W e X1+b e ) (10) Among them, e is the attention weight coefficient combination corresponding to a single feature on the input feature map; W e is the convolution layer parameter; b e is the bias vector set; Step 2-2-2, multiply the input feature map X by the obtained weight map e to obtain the weighted feature map U1; U1=eX (11); Step 2-2-3, normalize the features of the obtained weighted feature map U1 through the Softmax function to reduce the impact of the data dimension on the output result; prevent overfitting, and the normalization formula for a single feature x is: Step 2-2-4, the normalized result x′ is passed through the activation function swish, adding nonlinear factors to solve the defect of insufficient expression ability of the linear model, and obtaining a new weighted feature map U2; Step 2-2-5, perform group convolution on the new weighted feature map U2, set the vaild calculation method to perform point-by-point convolution to extract relevant features, and perform normalization to obtain the final output feature map Step 2-2-6, The data is flattened by the flatten function to obtain a feature sequence with dynamic weights assigned to it, which is then input into the BiLSTM model. When BiLSTM processes long sequence data, the forward and backward LSTM layers process each time step, and then linearly combine the outputs of all time steps to form the final output.

7. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3-1, perform flow calculation on the new topology structure to obtain multi-section historical flow data as the true value of the new topology flow; Step 3-2, add Gaussian white noise to the true value of the new topology flow to generate multi-section measurement information, and obtain a small sample new topology measurement data set that simulates actual measurement; and perform normalization, dimension conversion and data tiling preprocessing on the small sample new topology measurement data set.

8. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 1, characterized in that: In step 4, Wasserstein divergence is introduced on the basis of GAN to replace JS divergence, and the conditional variable y is added in combination with the conditional generative adversarial network to generate a CWGAN-div model that meets the required data type; the conditional information and random noise are input into the generator of the CWGAN-div model, so that the generator generates new samples under the guidance of the conditions; then the generated samples, the target data of the new topology measurement of the small samples and the conditional information are input into the discriminator for judgment.

9. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 8 is characterized in that: The objective function of the CWGAN-div model is: Where: E[·] is the expected function; x is the input data; z is random noise; D[·] is the discriminator function; G[z] is the generator function; P r is the data distribution of x; P z is the data distribution of z, P z is the normal distribution of N(0,1); y is the conditional information, that is, y is the measurement data that conforms to the actual measurement distribution; is the sampling of the linear combination of real data and generated data; p u for data distribution.

10. The method for power system state estimation integrating transfer learning and generative adversarial network according to claim 1, characterized in that: In step 5, during transfer learning, the structure before the fully connected layer of the source domain model is frozen to save the parameters, and the fully connected layer is used for fine-tuning the target task to obtain the target domain model.