Dynamic power distribution network deduction method based on graph neural network perception and privacy protection type Transform model

By introducing graph neural networks and differential privacy technologies into the Transformer model, the problems of insufficient interpretability, neglect of graph structure and privacy leakage in distribution network data processing are solved, and more accurate distribution network status deduction and privacy protection are achieved.

CN120087218APending Publication Date: 2025-06-03HOHAI UNIV
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510230157.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing Transformer model has problems such as insufficient interpretability, ignoring graph structure and privacy leakage when processing distribution network data.

Method used

The perception mechanism and differential privacy technology based on graph neural network (GNN) are used to extract the spatial characteristics of the dynamic topological structure of the distribution network, and differential privacy optimization technology is used during the training process to ensure the privacy protection of the model.

Benefits of technology

It improves the interpretability of the Transformer model for the state deduction of the distribution network, makes full use of the topological structure information of the distribution network, and effectively protects the user's privacy data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087218A_ABST
    Figure CN120087218A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic power distribution network deduction method based on graph neural network perception and a privacy protection type Transform model. The method comprises the steps of collecting source load storage historical data in a power distribution network and performing standardization processing; performing time sequence segmentation; establishing a dynamic topological structure model of the power distribution network; extracting spatial features of the dynamic topological structure of the power distribution network by using a graph neural network; the spatial features are converted into time sequence data, the time sequence data are combined with time sequence data obtained after time sequence segmentation, and a source network load storage time sequence data set is established; differential privacy processing is carried out; training a Transform model, in the training process, using a differential privacy optimization technology, adding noise to the gradient when the gradient is updated each time, and finally obtaining a trained DP-Transform model; and inputting a specific scene, and deducing the state of the power distribution network by using a DP-Transform model. According to the method, the interpretability of deducing the state of the power distribution network by using the Transform is improved, so that the Transform model can fully receive the graph structure input of the network topology structure of the power distribution network, and privacy protection on user information in the deducing process is considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power grids, relates to power system distribution networks and artificial intelligence technologies, and particularly relates to a dynamic distribution network deduction method based on a graph neural network perception and privacy protection type Transformer model. Background Art

[0002] With the large-scale application of renewable energy and the gradual development of the power system towards cleaner and lower-carbon directions, the operation and management of distribution networks are facing unprecedented challenges. In traditional centralized power systems, the power supply system can relatively stably meet the load demand. However, with the access of distributed energy sources (such as photovoltaic, wind power, etc.) and energy storage systems, the distribution network is gradually transformed from a centralized single system into a distributed and multi-source interconnected complex system, and the deduction of its future operating state is becoming increasingly important for the safe, reliable and efficient operation of the power system. Traditional distribution network deduction methods mainly rely on physical models, such as power flow calculation, equipment state estimation, etc. These methods can provide accurate theoretical basis for the operation deduction of distribution networks. However, with the wide application of new technologies such as the energy Internet, renewable energy generation, and large-scale energy storage systems, the limitations of traditional physical models are becoming increasingly obvious. In order to adapt to the dynamically changing distribution network environment, researchers have begun to apply data-driven AI models to distribution network deduction.

[0003] In recent years, Transformer has gradually become one of the important research directions in the field of distribution network deduction due to its superior performance in time series data modeling and long dependence relationship capture. Compared with traditional recurrent neural networks (such as RNN, LSTM, etc.), Transformer has higher parallel computing efficiency and stronger feature expression ability, especially showing great advantages when dealing with complex and multi-time scale dynamic data.

[0004] However, the existing Transformer still has the following problems when applied:

[0005] 1. Although Transformer performs well in processing time series data and multi-variable features, its "black box" characteristic leads to insufficient interpretability; although its attention mechanism can weigh the importance of different time steps in the input sequence, its internal decision-making process is often difficult to intuitively explain, and it is difficult to clearly identify which input features or moments have a key impact on the deduction result.

[0006] 2. The distribution network has a complex topological structure, and the spatial correlation and mutual dependence between nodes are relatively strong. Traditional machine learning methods usually cannot handle this kind of structured data well and input it into the Transformer model, resulting in the Transformer model ignoring the graph structure of the input data and only relying on serialized input.

[0007] 3. When dealing with the distribution network deduction task, the Transformer model usually requires a large amount of time-series data for training and deduction, and this time-series data often contains sensitive information of users, such as load curves, electricity consumption habits, etc., which may lead to privacy leakage problems. Summary of the Invention

[0008] Object of the Invention: In order to overcome the deficiencies in the prior art, a dynamic distribution network deduction method based on a graph neural network perception and privacy-protected Transformer model is provided, which improves the interpretability of using the Transformer for the distribution network state deduction, enables the Transformer model to fully receive the graph structure input of the distribution network topology structure, and takes into account the privacy protection of user information during the deduction process.

[0009] Technical Solution: To achieve the above object, the present invention provides a dynamic distribution network deduction method based on a graph neural network perception and privacy-protected Transformer model, including the following steps:

[0010] S1: Collect the historical data of the power sources, loads, and energy storage in the distribution network and perform standardization processing;

[0011] S2: Perform time-series segmentation on the historical data after the standardization processing in step S1;

[0012] S3: Establish a dynamic topology structure model of the distribution network;

[0013] S4: Use a graph neural network to extract the spatial features of the dynamic topology structure of the distribution network;

[0014] S5: Convert the spatial features extracted in step S4 into time-series data, and combine it with the time-series data after time-series segmentation in step S2 to establish a time-series data set of the power sources, network, loads, and energy storage;

[0015] S6: Perform differential privacy processing on the time-series data set of the power sources, network, loads, and energy storage;

[0016] S7: Use the time-series data set of the power sources, network, loads, and energy storage after differential privacy processing to train the constructed Transformer model. During the training process, use differential privacy optimization technology to add noise to the gradient at each gradient update, and finally obtain the trained DP-Transformer model;

[0017] S8: Input a specific scenario and use the DP-Transformer model to perform distribution network state deduction.

[0018] Furthermore, the historical data of the power sources, loads, and energy storage in the distribution network collected in step S1 includes the output power P of each generator in the power grid gen(t, k), the load demand P of each node in the power grid load (t, n), the charge and discharge state E of each energy storage device s (t, k);

[0019] The Z-score normalization method is used to standardize all data so that the mean of each feature is zero and the standard deviation is one; the formula is as follows:

[0020]

[0021] Among them, X(t) represents the original data (generator output, load demand, energy storage state), μ X is the mean of this feature, σ X is the standard deviation of this feature; X N (t) is the standardized data after the original data is normalized.

[0022] Furthermore, in order to enable the Transformer model to perform time series prediction based on historical data in step S2, it is necessary to perform time series segmentation on the historical data standardized in step S1. The purpose of time series segmentation is to divide the historical data according to a fixed time window so that the input data at each time step matches the corresponding output data, and then it is used for training and deduction.

[0023] The time series segmentation specifically includes:

[0024] Set the time window size T, and each window contains T consecutive time steps; in this way, the historical data is divided into multiple subsequences, and each subsequence will contain an input window and an output prediction value; specifically, each input window is a data sequence of length T, which is used to predict the future power grid state;

[0025] Based on the standardized historical data set {X 标 (1), X 标 (2)..., X 标 (T), X 标 (T + 1)..., X 标 (N)} given in step S1, each time window contains T time steps, the input data is the power grid state in the past T steps, and the output data is the future power grid state; the specific formula is as follows:

[0026] X input (t) = {X 标 (t - T + 1), X 标 (t - T + 2),..., X 标 (t)} (4)

[0027] X output (t) = {X 标(t + 1)} (5)

[0028] To generate sufficient training samples, a sliding window technique is used. Each time a time step is slid, new data points enter the input window and the corresponding output targets are updated. In this way, the historical data will be divided into multiple time series samples for model training.

[0029] Furthermore, the establishment of the distribution network dynamic topology structure model in step S3 includes:

[0030] The distribution network topology structure is represented by an adjacency matrix in graph theory. In the graph structure, nodes represent various components in the power grid (such as substations, loads, generators, energy storage devices, etc.), and edges represent the power transmission lines between nodes; the topology structure of the distribution network is represented by an adjacency matrix, where the element A of the matrix ij indicates whether there is a connection (power transmission line) between nodes;

[0031] The definition formula of the adjacency matrix is as follows:

[0032]

[0033] In an actual distribution network, the topology structure of the distribution network will change dynamically with the operating state of the power grid. When a fault occurs in the power grid, a switching operation is performed, or the load is adjusted, the topology structure may change. By defining a dynamic adjacency matrix of the power grid topology to simulate the dynamic changes of the power grid topology, the specific formula is as follows:

[0034]

[0035] Among them, A(t - 1) represents the topology matrix of the previous time step, and A new represents the new topology matrix at the current time step t. When a fault or operation adjustment occurs in the power grid, the topology structure will change.

[0036] Furthermore, the process of using a graph neural network to extract the spatial features of the distribution network dynamic topology structure A(t) in step S4 includes:

[0037] The goal of the graph neural network is to update the information of each node through the adjacency matrix, so that the representation of each node can contain the information of its adjacent nodes; through multiple iterations, the representation of the node will aggregate the information from multiple neighbor nodes, thereby extracting the spatial features between nodes; the graph convolution operation of the graph neural network is the core of the GNN, and its basic formula is as follows:

[0038]

[0039] ReLU(x) = max(0, x) (10)

[0040] where, H (l) is the node representation matrix of the l-th layer, representing the feature vector of each node; W (l) is the weight matrix of the l-th layer, representing the parameters in the graph convolution operation; is the normalized adjacency matrix, representing the connection relationship between nodes, D is the degree matrix of nodes, representing the number of connections of each node. In this way, it can ensure that the information transmission of nodes with different degrees is balanced; ReLU is a commonly used activation function, which directly returns the value when the input is positive, and when the input is negative, it "truncates" it to 0; this non-linear transformation enables ReLU to introduce non-linear characteristics in the neural network, thereby helping the network to fit complex patterns and improving the training speed and performance of the model;

[0041] In order to enable the representation of each node to contain more levels of adjacent node information, the spatial feature extraction of the distribution network network structure is realized by means of multi-layer stacking. The graph convolution operation of each layer updates the representation of the nodes once, and finally forms a deep node representation. The specific formula is as follows:

[0042]

[0043] where, H i (L) is the spatial feature of node i at the L-th layer, representing the state of the node, considering the dependence relationship between it and its adjacent nodes; H l (L) is the spatial feature of line l at the L-th layer, representing the state of the line; H i (0) is the initial feature of node i, which is a feature vector of all 1s or all 0s, representing the initial node connection situation; H l (0) is the initial feature of line l, indicating whether the line is disconnected; W (l) is the learning weight in each layer of graph convolution operation; L is the number of stacked layers. Through multi-layer graph convolution, the features of nodes gradually aggregate the information of neighbor nodes from more levels.

[0044] Furthermore, the establishment of the source-network-load-storage time-series data set in step S5 specifically includes:

[0045] When constructing the time-series data set, the spatial features (such as node voltage, line power flow, etc.) extracted by the GNN are normalized by Z-score, and then combined with the time-series data after time-series segmentation to form a complete time-series data set as follows:

[0046]

[0047] Among them, is the spatial feature (such as node voltage, power flow, etc.) extracted by node i at time step t through the GNN; is the power flow feature of line l at time step t; X N,load (t), X N,gen (t), X N,ES (t) is the time series data of source, load and storage obtained in step S2 after being standardized in step S1;

[0048] Combine the data of all time steps to form a complete time series data set for subsequent training and deduction of the Transformer model:

[0049]

[0050] Furthermore, the differential privacy processing in step S6 specifically includes:

[0051] In the preprocessing stage of the time series data set, noise is added to the data. The magnitude of the noise is controlled by the privacy budget, and the privacy budget determines the strength of privacy protection; Gaussian noise is used for noise addition. In this way, it can be ensured that even if the data is leaked, the attacker cannot precisely know what the original data is because the introduction of noise will perturb the original data; the specific formula is:

[0052]

[0053] In the formula, σ 1 represents the amount of noise; ρ represents the privacy budget in the privacy protection process. The smaller ρ is, in order to achieve stronger privacy protection, the noise needs to be increased, that is, σ is increased 1 ; Δs is the sensitivity, which represents the maximum output difference caused by one processing to adjacent data sets; ln(1.25 / δ) is a term related to δ, and the purpose is to allow a lower privacy "failure" probability when δ is small, and a larger σ 1 is required to protect privacy; X date (t) is the original time series data set obtained in step S5, is the time series data set after adding noise, and N(0, σ 1 2 ) represents Gaussian noise with a mean of 0 and a variance of σ 1 2 .

[0054] Furthermore, the process of training the constructed Transformer model using the source-grid-load-storage time series data set after differential privacy processing in step S7 specifically includes:

[0055] From the time series data set processed by differential privacy Among them, training samples with a batch size of n are selected Among them represents the sample with an input quantity of i, and y(i) represents the target output (node voltage and network power flow at future moments);

[0056] When the sample is input into the Transformer, it first passes through the embedding layer, and the dimension of the embedding layer is d mod , and the mapping after passing through the embedding layer is expressed as:

[0057]

[0058] Among them, W e and λ e are learnable parameters, with values between 0 and 1, used to map the input features to the dimensions required by the model; E t (i) is the representation of the original input data in the high-dimensional space after passing through the embedding layer, with a shape of It provides the input feature representation for the subsequent calculations of the model;

[0059] After embedding the data, it will go through two sub-stages of forward propagation in the Transformer model: the multi-head self-attention mechanism and the forward network; adding positional encoding to E t (i) obtained from formula (17), Z t (i) can be obtained, and the specific formula is as follows:

[0060]

[0061] Z t (i) = E t (i) + PE t (i)(20)

[0062] In the formula, pos is the position index, for example (t = 1, 2, 3,... T); i is the index of the positional encoding, i = 1, 2,... d mod -1; PE t (i) is the positional encoding calculated according to the position t, with a shape of Z t (i) is the embedding vector with positional information, and the shape is also

[0063] In the multi-head self-attention, finally, a three-dimensional tensor with a shape of (B, S, d mod ) is required for calculation, and all Z t (i) are combined into the large tensor Z, and the specific formula is as follows:

[0064]

[0065] Z = (Z 1, Z 2 ,...Z S )(22)

[0066] In the formula, the first dimension B is the batch size, storing B samples; the second dimension S is the sequence length, and each sample contains S time steps internally; the third dimension d mod is the embedding dimension; Z i is a two-dimensional tensor merged from Z t (i) for S time steps; Z is the final large tensor used for subsequent calculations;

[0067] In multi-head attention, Z integrates the information of all samples and all time steps. Injecting Z into the self-attention mechanism of the Transformer, the specific formula is as follows:

[0068] Q = ZW Q , K = ZW K , V = ZW V (23)

[0069]

[0070] In the formula, Q (query) is used to calculate the matching degree with other positions; K (key) is used to be matched by the query; V (value) is used to weighted summarize the actual information; W Q , W K , W V is a learnable projection matrix used to extract Q, K, V from Z. These weight matrices are used as the parameters of the model. In the backpropagation process of step S8, calculate the gradient of the loss function and use the gradient descent algorithm to update the values of these weight matrices; Attention(Q, K, V) is single-head self-attention; QK T is to calculate the dot product of Q and K to calculate the similarity between each position and other positions in the sequence; is the scaling factor, taking values between 0 and 1, used to maintain numerical stability and appropriate gradients; softmax(x) is to normalize the similarity to obtain the attention weights; x i is the similarity score of a certain position; S is the sequence length, that is, all positions where attention is allocated; Perform exponential operation on the similarity; is the sum of the exponents of all scores, used to ensure that the sum of all attention weights is 1;

[0071] The Transformer generally sets h heads, and each head has independent {ZW i Q , ZW iK ,ZW i V}, after projecting the input Z, single-head attention is performed respectively, and then the outputs of h heads are concatenated on the last dimension, and a learnable matrix is used for linear transformation to obtain the final multi-head attention result. The specific formula is as follows:

[0072] head i = Attention(ZW i Q ,ZW i K ,ZW i V )(26)

[0073] M(Z) = Con(head 1 ,...,head h )W O (27)

[0074] Among them, head is the single-head self-attention. i ranges from 1 to h, capturing different subspace dimensions and different types of dependencies; Con(head 1 ,...,head h ) is to stack the output vectors of all heads; W O is a learnable matrix to ensure that the dimension of the output result is consistent with E t (i); M(Z) is the final multi-head self-attention result;

[0075] The above is the work for a single layer. For the multi-head self-attention result M l (Z) of the l-th layer and the input X l of this layer, after performing a residual connection and then adding layer normalization, Y l is obtained; Y l is used as the input of the forward network, and two-layer linear transformation and activation are performed independently for each position, and then a residual connection is performed again and layer normalization is added to obtain the final output X l+1 of this layer. The specific formula is as follows:

[0076] U l = ReLU(0, Y l W 1 + b 1 )W 2 + b 2 (28)

[0077]

[0078] X l+1 = Layer(Y l + U l ).(30)

[0079] In the formula, the weight matrix W 1 ,W 2 and the bias term b 1 ,b 2 are all learnable parameters obtained through back-propagation learning, with initial values ​​of 1 and 0, updated by the gradient descent method during training, and the goal is to minimize the loss function; Layer(x) is the layer normalization function; x is the input activation (such as the node voltage and load at each position); μ and σ are the mean and standard deviation of the input respectively; γ is a learnable scaling factor, with an initial value of 1, and updated through back-propagation learning; ω is a learnable bias term, with an initial value of 0, and updated through back-propagation learning;

[0080] According to this step, it is continuously iterated. Under the condition that the total number of layers is L, the output X of the top layer can be obtained. L+1 , the final output of the model is obtained through specific projection, and the deduction results of the power grid state at each time step / location are obtained.

[0081]

[0082] Where W out is the projection matrix, which is randomly initialized at the beginning of training and updated by back-propagation learning through formula (38). Its function is to transform the model dimension into the target dimension; b out is the bias vector, with a value of 0-1, which is used to improve flexibility and fitting ability;

[0083] All the learnable parameters in the above training process together constitute θ, as follows:

[0084] θ={W e ,b e ,{W Q ,W K ,W V ,W O},{W 1 ,b 1 ,W 2 ,b 2},W out ,b out ,…} (32)

[0085] Furthermore, in step S7, differential privacy optimization technology is used to add noise to the gradient each time the gradient is updated, specifically including:

[0086] Through the differential privacy SGD (Private SGD) method, noise is added to the gradient each time the gradient is calculated to ensure that the gradient of each data point does not leak too much information. The specific formula is as follows:

[0087]

[0088] ||g i || 2 ≤C (34)

[0089]

[0090]

[0091] where g i represents the gradient of parameter θ, C is a clipping threshold, and if the norm of a certain g i is greater than C, it is scaled to a norm exactly equal to C; Equation (33) is similar to Equation (15) and uses the privacy budget to limit the noise added during the process; Equations (34)-(36) are the gradient clipping process, aiming to control the influence of the gradient of each sample on the model update within a controllable range, so that differential privacy can obtain a clear sensitivity upper bound when adding noise later; the clipped gradient is denoted as Within a batch, these gradients are averaged, representing the average gradient within the batch; Equation (37) adds noise to the gradient, and N(0,σ 2 2 C 2 ) represents a multi-dimensional Gaussian noise vector with the same dimension as , being the finally obtained differentially private gradient.

[0092] Finally, the differentially private gradient is used to update the model parameters; if the learning rate is k (taking values from 0 to 1), the update formula for one iteration is:

[0093]

[0094] In the formula, ξ(θ) is the loss function of the model deduction under this θ. The smaller the loss function, the more accurate the deduction; and y i are the accurate value and the true value of the future state of the power grid respectively; the optimization process of the loss function is to take the derivative of ξ(θ) with respect to θ and iteratively update θ until the model approaches the optimal fit on the training data.

[0095] Furthermore, the specific process of using the DP-Transformer model for power distribution network state deduction in step S8 includes:

[0096] The DP-Transformer model and its parameters have been trained and finalized in the previous steps. Now, only the current moment or scenario information needs to be input, and the node voltages, line power flows, etc. of the distribution network at future moments are output.

[0097] The "scenario" used in the present invention refers to the known quantities of the power grid at a single moment, such as the current load, generator output, energy storage state, network topology, etc., which are the assumed situations for the model to deduce the future node voltages and line power flows.

[0098] Each feature in the scenario is combined into a vector R(t), where t represents the current moment, and the dimension of R(t) is d input , integrating the required electrical quantities (but only for one moment), and the specific formula is as follows:

[0099] R n,i (t) = [P 1oad,i (t), P gen,i (t), E storage,i (t), V n,i (t)] T (40)

[0100] R l,j (t) = [P flow,j (t)] T (41)

[0101] R(t) = [R n,i (t); R l,j (t)] (42)

[0102]

[0103] In the formula, the feature vector R n,i (t) of each node i contains features such as load, generator output, energy storage state, and node voltage; the feature vector R l,j (t) of each line j contains the line power flow feature; the feature vectors of all nodes and lines are concatenated by rows, and finally the overall input matrix R(t) is obtained, where the number of nodes is N n , and the number of lines is N l ;

[0104] R(t) is used as the initial input for model inference. In the deduction stage, the trained DP-Transformer performs forward propagation (constructed embedding layer, multi-head attention mechanism, forward network), and finally obtains the power grid state at the next moment (or multiple steps). According to the dimension of the output layer, the model can simultaneously output N v node voltages and N f line power flows, and the specific formula is as follows:

[0105]

[0106] In the formula, t 0 is the initial moment of the deduction; Δt is a fixed time interval, such as 15 minutes or 1 hour, depending on the label setting during training; is the deduced result; is the voltage of each node obtained by the deduction, is the power flow of each line obtained by the deduction.

[0107] The deduced future node voltages of the power grid can be used in power grid dispatching. If it is predicted that the voltages of some nodes exceed the limit, the dispatcher can carry out dispatching or reactive power compensation in advance; the deduced future power flows can be used in power flow analysis. If the line power flow approaches the safety limit, shunt or switch dispatching can be taken in advance.

[0108] Advantageous effects: Compared with the prior art, the present invention has the following advantages:

[0109] 1. The present invention makes full use of the complex topological structure information of the distribution network, captures the spatial relationship and mutual dependence between nodes through the graph neural network (GNN), so that the model can more accurately reflect the dynamic interaction between various components in the distribution network. This method breaks through the limitation of only using time series data in traditional distribution network deduction, combines topological structure features, and effectively improves the prediction ability of the model for the future operation state of the distribution network.

[0110] 2. Aiming at the problem of privacy leakage in distribution network deduction, the present invention introduces differential privacy technology to ensure the privacy protection of users' load data and electricity consumption behavior during model training and deduction. By adding noise to the time series data or encrypting the output results, it is ensured that users' sensitive information will not be leaked. BRIEF DESCRIPTION OF THE DRAWINGS

[0111] Figure 1 is the flowchart of the method of the present invention;

[0112] Figure 2 is the structural block diagram of the DP-Transformer model. DETAILED DESCRIPTION OF THE INVENTION

[0113] The present invention will be further clarified below in conjunction with the drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.

[0114] Such as Figure 1As shown in the figure, the present invention provides a dynamic distribution network deduction method based on a graph neural network perception and privacy protection type Transformer model, which includes the following steps:

[0115] S1: Collect the historical data of power sources, loads, and energy storage in the distribution network and perform standardization processing;

[0116] The historical data of power sources, loads, and energy storage collected in the distribution network includes the output power P gen (t,k) of each generator in the power grid, the load demand P load (t,n) of each node in the power grid, and the charge and discharge state E s (t,k) of each energy storage device;

[0117] Adopt the Z-score normalization method to perform standardization processing on all data, so that the mean of each feature is zero and the standard deviation is one; the formula is as follows:

[0118]

[0119] Among them, X(t) represents the original data (generator output, load demand, energy storage state), μ X is the mean of this feature, σ X is the standard deviation of this feature; X N (t) is the standard data after the original data is normalized.

[0120] S2: Perform time series segmentation on the historical data after the standardization processing in step S1;

[0121] In order to enable the Transformer model to perform time series prediction based on historical data, it is necessary to perform time series segmentation on the historical data standardized in step S1. The purpose of time series segmentation is to divide the historical data according to a fixed time window, so that the input data of each time step matches the corresponding output data, and then be used for training and deduction.

[0122] The time series segmentation specifically includes:

[0123] Set the time window size T, and each window contains T consecutive time steps; in this way, the historical data is divided into multiple subsequences, and each subsequence will contain an input window and an output prediction value; specifically, each input window is a data sequence with a length of T, which is used to predict the future power grid state;

[0124] Based on the standardized historical data set {X 标 (1), X 标 (2)..., X 标 (T), X 标 (T + 1)..., X 标(N)}, where each time window contains T time steps. The input data is the grid state in the past T steps, and the output data is the future grid state. The specific formula is as follows:

[0125] X input (t) = {X 标 (t - T + 1), X 标 (t - T + 2),..., X 标 (t)} (4)

[0126] X output (t) = {X 标 (t + 1)} (5)

[0127] To generate sufficient training samples, the sliding window technique is used. Each time a time step is slid, new data points enter the input window and update the corresponding output targets. In this way, the historical data will be divided into multiple time series samples for model training.

[0128] S3: Establish a dynamic topological structure model of the distribution network;

[0129] The topological structure of the distribution network is represented by an adjacency matrix in graph theory. In the graph structure, nodes represent various components in the power grid (such as substations, loads, generators, energy storage devices, etc.), and edges represent the power transmission lines between nodes; the topological structure of the distribution network is represented by an adjacency matrix, where the element A ij of the matrix indicates whether there is a connection (power transmission line) between nodes;

[0130] The definition formula of the adjacency matrix is as follows:

[0131]

[0132] In an actual distribution network, the topological structure of the distribution network will change dynamically with the operating state of the power grid. When a fault occurs in the power grid, a switch operation is performed, or a load adjustment is made, the topological structure may change. The dynamic change of the grid topology is simulated by defining a dynamic adjacency matrix of the grid topology. The specific formula is as follows:

[0133]

[0134] where A(t - 1) represents the topological matrix at the previous time step, and A new represents the new topological matrix at the current time step t. When a fault occurs in the power grid or an operation adjustment is made, the topological structure will change.

[0135] S4: Use a graph neural network to extract the spatial features of the dynamic topological structure A(t) of the distribution network;

[0136] The goal of a graph neural network is to update the information of each node through an adjacency matrix so that the representation of each node can contain the information of its adjacent nodes; through multiple iterations, the representation of the node will aggregate the information from multiple neighbor nodes, thereby extracting the spatial features between nodes; the graph convolution operation of the graph neural network is the core of the GNN, and its basic formula is as follows:

[0137]

[0138] ReLU(x) = max(0, x)(10)

[0139] In the formula, H (l) is the node representation matrix of the l-th layer, representing the feature vector of each node; W (l) is the weight matrix of the l-th layer, representing the parameters in the graph convolution operation; is the normalized adjacency matrix, representing the connection relationship between nodes, D is the degree matrix of nodes, representing the number of connections of each node. In this way, it can be ensured that the information transmission of nodes with different degrees is balanced; ReLU is a commonly used activation function, which directly returns the value when the input is positive, and when the input is negative, it "truncates" it to 0; this non-linear transformation enables ReLU to introduce non-linear characteristics in the neural network, thereby helping the network to fit complex patterns and improving the training speed and performance of the model;

[0140] In order to enable the representation of each node to contain more levels of adjacent node information, the spatial feature extraction of the distribution network network structure is realized by means of multi-layer stacking. The graph convolution operation of each layer updates the representation of the node once, and finally forms a deep node representation. The specific formula is as follows:

[0141]

[0142] In the formula, H i (L) is the spatial feature of node i at the L-th layer, representing the state of the node, considering the dependence relationship between it and its adjacent nodes; H l (L) is the spatial feature of line l at the L-th layer, representing the state of the line; H i (0) is the initial feature of node i, which is a feature vector of all 1s or all 0s, representing the initial node connection situation; H l (0) is the initial feature of line l, indicating whether the line is disconnected; W (l) is the learning weight in the graph convolution operation of each layer; L is the number of stacked layers. Through multi-layer graph convolution, the features of the nodes gradually aggregate the information from more levels of neighbor nodes.

[0143] S5: Convert the spatial features extracted in step S4 into time-series data, and combine it with the time-series data after time-series segmentation in step S2 to establish a source network-load-storage time-series dataset;

[0144] When constructing the time-series dataset, perform Z-score normalization on the spatial features (such as node voltage, line power flow, etc.) extracted by the GNN, and then combine them with the time-series data after time-series segmentation to form a complete time-series dataset as follows:

[0145]

[0146] Among them, is the spatial feature (such as node voltage, power flow, etc.) extracted by the GNN for node i at time step t; is the power flow feature of line l at time step t; X N,load (t), X N,gen (t), X N,ES (t) is the source-load-storage time-series data obtained in step S2 after being standardized in step S1;

[0147] Combine the data of all time steps to form a complete time-series dataset for subsequent training and deduction by the Transformer model:

[0148]

[0149] S6: Perform differential privacy processing on the source network-load-storage time-series dataset;

[0150] In the preprocessing stage of the time-series dataset, add noise to the data. The magnitude of the noise is controlled by the privacy budget, and the privacy budget determines the strength of privacy protection; use Gaussian noise to add noise. In this way, it can be ensured that even if the data is leaked, the attacker cannot accurately know what the original data is because the introduction of noise will perturb the original data; the specific formula is:

[0151]

[0152] In the formula, σ 1 represents the amount of noise; ρ represents the privacy budget in the privacy protection process. The smaller ρ is, the more noise needs to be added to achieve stronger privacy protection, that is, increase σ 1 ; Δs is the sensitivity, indicating the maximum output difference caused by a single processing to adjacent datasets; ln(1.25 / δ) is a term related to δ, aiming to allow a lower privacy "failure" probability when δ is small, and a larger σ 1 is required to protect privacy; X date (t) is the original time-series dataset obtained in step S5, is the time series data set after adding noise, and N(0,σ 1 2 ) represents Gaussian noise with a mean of 0 and a variance of σ 1 2 .

[0153] S7: Train the constructed Transformer model using the time series data set of source network load and storage after differential privacy processing. During the training process, use differential privacy optimization technology to add noise to the gradient at each gradient update, and finally obtain the trained DP-Transformer model;

[0154] Refer to Figure 2 , and the process of training the constructed Transformer model using the time series data set of source network load and storage after differential privacy processing specifically includes:

[0155] Select training samples with a batch size of n from the time series data set processed by differential privacy , where represents the sample with the input quantity of i , and y(i) represents the target output (node voltage and network power flow at future moments); When inputting the sample into the Transformer, first pass it through the embedding layer, and the dimension of the embedding layer is d

[0156] , and the mapping after passing through the embedding layer is expressed as: mod , where W

[0157]

[0158] and λ e and e are learnable parameters, and their values are between 0 and 1, used to map the input features to the dimensions required by the model; E t (i) is the representation of the original input data in the high-dimensional space after passing through the embedding layer, and its shape is , which provides the input feature representation for the subsequent calculations of the model;

[0159] After embedding the data, it will go through two sub-stages of forward propagation in the Transformer model: the multi-head self-attention mechanism and the forward network; add position encoding to E t (i) obtained from formula (17) to obtain Z t (i), and the specific formula is as follows:

[0160]

[0161] Z t (i) = E t(i) + PE t (i)(20)

[0162] Where pos is the position index, e.g., (t = 1, 2, 3, … T); i is the index of the position encoding, i = 1, 2, … d mod -1; PE t (i) is the position encoding calculated according to position t, with the shape of Z t (i) is the embedding vector with position information, and the shape is also

[0163] In the multi-head self-attention, finally, a three-dimensional tensor with the shape of (B, S, d mod ) is needed for calculation. All Z t (i) are merged into the large tensor Z, and the specific formula is as follows:

[0164]

[0165] Z = (Z 1, Z 2 ,... Z S ) (22)

[0166] Where the first dimension B is the batch size, storing B samples; the second dimension S is the sequence length, and each sample contains S time steps inside; the third dimension d mod is the embedding dimension; Z i is the two-dimensional tensor merged by Z t (i) with S time steps; Z is the finally obtained large tensor, which is used for subsequent calculations;

[0167] In the multi-head attention, Z integrates the information of all samples and all time steps. Injecting Z into the self-attention mechanism of the Transformer, the specific formula is as follows:

[0168] Q = ZW Q , K = ZW K , V = ZW V (23)

[0169]

[0170] Where Q (query) is used to calculate the matching degree with other positions; K (key) is used to be matched by the query; V (value) is used to weighted summarize the actual information; W Q , W K , W Vis a learnable projection matrix used to extract Q, K, and V from Z. These weight matrices are parameters of the model. During the backpropagation process in step S8, the gradient of the loss function is calculated, and the values of these weight matrices are updated using the gradient descent algorithm; Attention(Q, K, V) is single-head self-attention; QK T is to perform a dot product on Q and K to calculate the similarity between each position and other positions in the sequence; is a scaling factor, with a value between 0 and 1, used to maintain numerical stability and an appropriate gradient; softmax(x) normalizes the similarity to obtain attention weights; x i is the similarity score at a certain position; S is the sequence length, that is, all positions for attention allocation; performs an exponential operation on the similarity; is the sum of the exponentials of all scores, used to ensure that the sum of all attention weights is 1;

[0171] Transformer generally sets h heads, and each head has independent {ZW i Q ,ZW i K ,ZW i V}. After projecting the input Z, single-head attention is performed respectively, and then the outputs of the h heads are concatenated on the last dimension, and a linear transformation is performed with a learnable matrix to obtain the final multi-head attention result. The specific formula is as follows:

[0172] head i =Attention(ZW i Q ,ZW i K ,ZW i V )(26)

[0173] M(Z)=Con(head 1 ,...,head h )W O (27)

[0174] Among them, head is single-head self-attention, i ranges from 1 to h, capturing different subspace dimensions and different types of dependencies; Con(head 1 ,...,head h ) is to stack the output vectors of all heads; W O is a learnable matrix to ensure that the dimension of the output result is consistent with E t (i); M(Z) is the final multi-head self-attention result;

[0175] The above is the work for a single layer. For the multi-head self-attention result M of the l-th layer l (Z) and the input X of this layer l , after performing a residual connection and then applying layer normalization, we obtain Y l ; Use Y l as the input to the feed-forward network. Independently perform two-layer linear transformation and activation for each position, and then perform a residual connection again and apply layer normalization to obtain the final output X of this layer l+1 , and the specific formula is as follows:

[0176] U l = ReLU(0, Y l W 1 + b 1 )W 2 + b 2 (28)

[0177]

[0178] X l+1 = Layer(Y l + U l ).(30)

[0179] In the formula, the weight matrices W 1 , W 2 and the bias terms b 1 , b 2 are all learnable parameters obtained through backpropagation learning. The initial values are 1 and 0, and they are updated by the gradient descent method during training. The goal is to minimize the loss function; Layer(x) is the layer normalization function; x is the input activation (such as the node voltage, load, etc. at each position); μ and σ are the mean and standard deviation of the input respectively; γ is a learnable scaling factor with an initial value of 1 and is updated through backpropagation learning; ω is a learnable bias term with an initial value of 0 and is updated through backpropagation learning;

[0180] Continuously iterate according to this step. Under the condition that the total number of layers is L, the output X of the top layer can be obtained L+1 , and the final output of the model is obtained through a specific projection, and the deduction results of the power grid state at each time step / position

[0181]

[0182] In the formula, W out is the projection matrix, which is randomly initialized at the beginning of training and is updated through backpropagation learning according to formula (38). Its role is to transform the model dimension to the target dimension; b out is the bias vector, with a value range of 0 - 1, used to improve flexibility and fitting ability;

[0183] All learnable parameters in the above training process together constitute θ, as follows:

[0184] θ = {W e , b e , {W Q , W K , W V , W O , {W 1 , b 1 , W 2 , b 2}, W out , b out , …} (32)

[0185] Using differential privacy optimization technology, noise is added to the gradient during each gradient update, specifically including:

[0186] Through the differential privacy SGD (Private SGD) method, noise is added to the gradient each time the gradient is calculated to ensure that the gradient of each data point does not leak too much information. Its specific formula is as follows:

[0187]

[0188] ||g i || 2 ≤ C (34)

[0189]

[0190] where g i represents the gradient of the parameter θ, and C is a clipping threshold. If the norm of a certain g i is greater than C, then it is scaled to a norm exactly equal to C; Formula (33) is similar to Formula (15) and uses the privacy budget to limit the noise added during the process; Formulas (34)-(36) are the gradient clipping process, aiming to control the impact of the gradient of each sample on the model update within a controllable range, so that differential privacy can obtain a clear sensitivity upper bound when adding noise later; The gradient after clipping is denoted as Within a batch, these gradients are averaged, represents the average gradient within the batch; Formula (37) adds noise to the gradient, and N(0, σ 2 2 C 2 ) represents a multi-dimensional Gaussian noise vector with the same dimension as , is the finally obtained differential privacy gradient.

[0191] Finally, use the differential privacy gradient to update the model parameters; if the learning rate is κ (taking values from 0 to 1), the update formula for one iteration is:

[0192]

[0193] where ξ(θ) is the loss function deduced by the model under this θ. The smaller the loss function, the more accurate the deduction; and y i are the accurate value and the true value of the future state of the power grid respectively; the optimization process of the loss function is to take the derivative of ξ(θ) with respect to θ and iteratively update θ until the model approaches the optimal fit on the training data.

[0194] S8: Input a specific scenario and use the DP-Transformer model to deduce the state of the distribution network;

[0195] The DP-Transformer model and its parameters have been trained and finalized in the previous steps. Now, only the current moment or scenario information needs to be input, and the node voltage, line power flow, etc. states of the distribution network at the future moment are output.

[0196] The "scenario" used in the present invention refers to the known quantities of the power grid at a single moment, such as the current load, generator output, energy storage state, network topology, etc., which are the assumed situations for the model to deduce the future node voltage and line power flow.

[0197] Combine each feature in the scenario into a vector R(t), where t represents the current moment, and the dimension of R(t) is d input , integrating the required electrical quantities (but only for one moment), and the specific formula is as follows:

[0198] R n,i (t) = [P 1oad,i (t), P gen,i (t), E storage,i (t), V n,i (t)] T (40)

[0199] R l,j (t) = [P flow,j (t)] T (41)

[0200] R(t) = [R n,i (t); R l,j (t)] (42)

[0201]

[0202] where, the feature vector R n,i(t) contains the characteristics of load, generator output, energy storage status and node voltage; the characteristic vector R of each line j l,j (t) contains the line flow characteristics; the characteristic vectors of all nodes and lines are concatenated row by row, and finally the overall input matrix R(t) is obtained, with N nodes. n , the number of lines is N l ;

[0203] R(t) is used as the initial input of the model reasoning. In the deduction stage, the trained DP-Transformer performs forward propagation (constructed embedding layer, multi-head attention mechanism, forward network), and finally obtains the power grid state at the next moment (or multiple steps). According to the output layer dimension, the model can simultaneously output N v The node voltage and N f The specific formula for the line flow is as follows:

[0204]

[0205] Where, t 0 is the initial time of the simulation; Δt is a fixed time interval, such as 15 minutes or 1 hour, depending on the label setting during training; For the deduction result; To deduce the voltage of each node, The tidal currents of each route are deduced.

[0206] The deduced future node voltages of the power grid can be used for power grid dispatching. If it is predicted that the voltage of some nodes exceeds the limit, the dispatchers can carry out dispatching or reactive power compensation in advance. The deduced future trends can be used for trend analysis. If the line trends approach the safety limit, diversion or switch dispatching can be carried out in advance.

[0207] Based on the above scheme, the present invention has the following advantages compared with the prior art:

[0208] 1. The present invention makes full use of the complex topological structure information of the distribution network, and captures the spatial relationship and interdependence between nodes through the graph neural network (GNN), so that the model can more accurately reflect the dynamic interaction between various components in the distribution network. This method breaks through the limitation of using only time series data in traditional distribution network deduction, and combines the topological structure characteristics to effectively improve the model's ability to predict the future operating status of the distribution network.

[0209] 2. To address the privacy leakage problem in distribution network simulation, this paper introduces differential privacy technology to ensure the privacy protection of users' load data and electricity consumption behavior during model training and simulation. By adding noise to time series data or encrypting the output results, it is ensured that users' sensitive information will not be leaked.

[0210] To verify the actual effect of the method of the present invention, in this embodiment, the above two performance data are presented by comparison through simulation experiments, which are specifically as follows:

[0211] To verify the topological modeling advantage based on the graph neural network (GNN), an IEEE 33-node system is used to construct a simulation scenario. The experiment shows that: compared with the deduction method of the traditional Transformer (only using time series data), after the GNN of the present invention integrates the topological structure, the mean square error (MSE) of the node voltage deduction at future moments of the distribution network drops from 0.032 to 0.018 (a decrease of 43.7%), and the correlation coefficient (R 2 ) of the key line power flow deduction increases from 0.82 to about 0.9, proving that the introduction of topological space features effectively improves the dynamic deduction ability.

[0212] The privacy protection effect is tested on the CEC-500 user load data set: when the differential privacy parameters are δ = 0.5 and ρ = 0.01 (strong privacy protection), the mean absolute error (MAE) of the load prediction of the traditional federated learning scheme increases by 26.3% due to noise interference. However, through the noise injection strategy optimization and gradient clipping mechanism of the present invention, the MAE only increases by 9%, which has a smaller impact on the deduction accuracy, and keeps the leakage risk of users in the power consumption mode reduced by more than 90%, verifying the effectiveness of privacy protection and maintaining the usability of the deduction results.

Claims

1. A dynamic distribution network deduction method based on graph neural network perception and privacy-preserving Transformer model, characterized in that: The steps include: S1: Collect historical data of sources, loads and storage in the distribution network and perform standardized processing; S2: Perform time series segmentation on the historical data after the standardization process in step S1; S3: Establish a dynamic topology model of the distribution network; S4: Extracting spatial features of dynamic topological structure of distribution network using graph neural network; S5: convert the spatial features extracted in step S4 into time series data, and combine it with the time series data after time series segmentation in step S2 to establish a source-grid-load-storage time series data set; S6: Perform differential privacy processing on the source-grid-load-storage time series data set; S7: Use the source-grid-load-storage time series data set processed with differential privacy to train the constructed Transformer model. During the training process, use differential privacy optimization technology to add noise to the gradient each time the gradient is updated, and finally obtain the trained DP-Transformer model; S8: Input a specific scenario and use the DP-Transformer model to simulate the distribution network state.

2. According to claim 1, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The source-load-storage historical data in the distribution network collected in step S1 includes the output power P of each generator in the power grid. gen (t, k), load demand P of each node in the power grid load (t,n), the charge and discharge status E of each energy storage device s (t,k); The Z-score normalization method is used to standardize all data so that the mean of each feature is zero and the standard deviation is one; the formula is as follows: Among them, X(t) represents the original data, μ X is the mean of the feature, σ X is the standard deviation of the feature; X N (t) is the standard data after the original data is normalized.

3. According to claim 1, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The timing segmentation in step S2 specifically includes: Set the time window size T, each window contains T consecutive time steps; in this way, the historical data is divided into multiple subsequences, each of which will contain an input window and an output prediction value; specifically, each input window is a data sequence of length T, which is used to predict the future power grid state; Based on the standardized historical data set {X 标 (1),X 标 (2)...,X 标 (T),X 标 (T+1)...,X 标 (N)}, each time window contains T time steps, the input data is the power grid status of the past T steps, and the output data is the future power grid status; the specific formula is as follows: X input (t)={X 标 (t-T+1),X 标 (t-T+2),...,X 标 (t)} (4) X output (t)={X 标 (t+1)} (5) Each time a time step is slid, a new data point enters the input window and the corresponding output target is updated.

4. According to claim 1, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The establishment of the dynamic topology model of the distribution network in step S3 includes: The topological structure of the distribution network is represented by the adjacency matrix in graph theory. In the graph structure, nodes represent the components in the power grid, and edges represent the power transmission lines between nodes. The topological structure of the distribution network is represented by the adjacency matrix, where the elements A of the matrix are ij Indicates whether there is a connection between nodes; The adjacency matrix definition formula is as follows: The dynamic changes of the power grid topology are simulated by defining the dynamic adjacency matrix of the power grid topology. The specific formula is as follows: Among them, A(t-1) represents the topological matrix of the previous time step, A new Represents the new topology matrix at the current time step t. When a fault occurs in the power grid or the operation is adjusted, the topology structure will change.

5. According to claim 4, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The process of using the graph neural network to extract the spatial features of the dynamic topological structure A(t) of the distribution network in step S4 includes: The goal of the graph neural network is to update the information of each node through the adjacency matrix so that the representation of each node can contain the information of its adjacent nodes; through multiple iterations, the representation of the node will aggregate the information from multiple neighboring nodes, thereby extracting the spatial features between nodes; the graph convolution operation of the graph neural network is the core of GNN, and its basic formula is as follows: ReLU(x)=max(0,x)(10) In the formula, H (l) is the node representation matrix of the lth layer, representing the feature vector of each node; W (l) is the weight matrix of the lth layer, representing the parameters in the graph convolution operation; is the normalized adjacency matrix, which indicates the connection relationship between nodes. D is the degree matrix of the node, which indicates the number of connections of each node. ReLU is an activation function, which directly returns the value when the input is positive, and "truncates" it to 0 when the input is negative. In order to allow the representation of each node to contain more levels of adjacent node information, the spatial feature extraction of the distribution network structure is realized by stacking multiple layers. The graph convolution operation of each layer updates the node representation once, and finally forms a deep node representation. The specific formula is as follows: In the formula, H i (L) is the spatial feature of node i at layer L, representing the state of the node; H l (L) is the spatial feature of line l at layer L, representing the state of the line; H i (0) is the initial feature of node i, which is a feature vector of all 1 or all 0, indicating the initial node connection status; H l (0) is the initial characteristic of line l, indicating whether the line is disconnected; W (l) is the learning weight in each layer of graph convolution operation; L is the number of stacked layers. Through multi-layer graph convolution, the features of the nodes gradually aggregate the information of neighboring nodes from more levels.

6. According to claim 1, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The establishment of the source-grid-load-storage time series data set in step S5 specifically includes: When constructing a time series data set, the spatial features extracted by GNN are normalized by Z-score and then combined with the time series data after time series segmentation to form a complete time series data set as follows: in, Is a node i Spatial features extracted by GNN at time step t; is the power flow characteristic of line l at time step t; X N,load (t),X N,gen (t),X N,ES (t) is the source-load-storage time series data obtained in step S2 after being normalized in step S1; Combine the data from all time steps to form a complete time series dataset:

7. According to claim 1, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The differential privacy processing in step S6 specifically includes: In the stage of preprocessing the time series data set, noise is added to the data. The size of the noise is controlled by the privacy budget, which determines the strength of privacy protection. Gaussian noise is used to add noise. The specific formula is: Where σ1 represents the amount of noise; ρ represents the privacy budget in the privacy protection process; Δs is the sensitivity, which indicates the maximum output difference caused by one processing to adjacent data sets; ln(1.25 / δ) is a term related to δ; X date (t) is the original time series data set obtained in step S5, is the time series data set after adding noise, N(0,σ1 2 ) means that the mean is 0 and the variance is σ1 2 Gaussian noise.

8. According to claim 7, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: The process of training the constructed Transformer model using the source-grid-load-storage time series data set after differential privacy processing in step S7 specifically includes: From the differentially private time series dataset In the example, we select a training sample with a batch size of n. in Indicates that the input quantity is i Sample, y(i) represents the target output; When the sample is input into the Transformer, First pass through the embedding layer, the dimension of the embedding layer is d mod , the mapping after the embedding layer is expressed as: Among them, W e and λ e It is a learnable parameter with a value between 0 and 1, which is used to map the input features to the dimensions required by the model; E t (i) is the representation of the original input data in high-dimensional space after the embedding layer, and its shape is After the data is embedded, it will go through two sub-stages of forward propagation in the Transformer model: the multi-head self-attention mechanism and the forward network; E obtained by formula (17) t (i) Add position coding to obtain Z t (i), the specific formula is as follows: Z t (i)=E t (i)+PE t (i) (20) Where pos is the position index; i is the index of the position code, i = 1, 2, ... d mod -1;PE t (i) is based on the position t The calculated positional encoding has the shape Z t (i) is an embedding vector with position information, and its shape is In multi-head self-attention, we finally need a shape of (B, S, d mod ) to calculate the three-dimensional tensor, and put all Z t (i) Merge into a large tensor Z. The specific formula is as follows: From=(From 1, Z2,…Z S ) (22) In the formula, the first dimension B is the batch size, storing B samples; the second dimension S is the sequence length, each sample contains S time steps; the third dimension d mod is the embedding dimension; Z i is the Z of S time steps t (i) The two-dimensional tensor is merged into Z; Z is the final large tensor, which is used for subsequent calculations; In multi-head attention, Z integrates the information of all samples and all time steps, and injects Z into the self-attention mechanism of Transformer. The specific formula is as follows: Q=ZW Q ,K=ZW K ,V=ZW V (23) In the formula, Q is used to calculate the matching degree with other positions; K is used to match the query; V is used to weightedly summarize the actual information; W Q , W K , W V is a learnable projection matrix used to extract Q, K, and V from Z. These weight matrices are used as parameters of the model. In the back propagation process of the subsequent step S8, the gradient of the loss function is calculated and the values ​​of these weight matrices are updated using the gradient descent algorithm. Attention (Q, K, V) is a single-head self-attention. QK T It is to do a dot product of Q and K and calculate the similarity of each position with other positions in the sequence; is a scaling factor, which takes a value between 0 and 1 to maintain numerical stability and appropriate gradients; softmax(x) normalizes the similarity to obtain the attention weight; x i is the similarity score of a certain position; S is the sequence length, i.e., all positions where attention is allocated; Perform an exponential operation on the similarity; is the exponential sum of all scores, used to ensure that the sum of all attention weights is 1; Transformer sets h heads, each with independent {ZW i Q ,ZW i K ,ZW i V }, after projecting the input Z, perform single-head attention separately, then concatenate the outputs of the h heads in the last dimension, and use a learnable matrix for linear transformation to obtain the final multi-head attention result. The specific formula is as follows: head i =Attention(ZW i Q ,ZW i K ,ZW i V ) (26) M(Z)=Con(head1,...,head h )W O (27) Among them, head is a single-head self-attention, i ranges from 1 to h, capturing different subspace dimensions and different types of dependencies; Con(head1,...,head h ) is the stacking of the output vectors of all heads; W O is a learnable matrix, ensuring that the dimension of the output result is the same as E t (i) Consistent; M(Z) is the final multi-head self-attention result; The above is the work for a single layer. For the multi-head self-attention result M of the lth layer l (Z) and the input X of this layer l , do residual connection and normalize the additional layer to get Y l ; Y l It is used as the input of the forward network, and two layers of linear transformation and activation are performed independently at each position, and then residual connection is performed again and an additional layer is normalized to obtain the final output X of this layer. l+1 , the specific formula is as follows: <h2 style=";text-align:left;direction:ltr">U<h2 style=";text-align:left;direction:ltr"> l <h2 style=";text-align:left;direction:ltr"> =ReLU(0,Y<h2 style=";text-align:left;direction:ltr"> l <h2 style=";text-align:left;direction:ltr"> W1+b1)W2+b2 (28) X l+1 =Layer(Y l +U l ). (30) Wherein, weight matrices W1, W2 and bias terms b1, b2 are learnable parameters obtained through back-propagation learning, with initial values ​​of 1 and 0, and are updated by the gradient descent method during training, with the goal of minimizing the loss function; Layer(x) is the layer normalization function; x is the input activation; μ and σ are the mean and standard deviation of the input, respectively; γ is a learnable scaling factor, with an initial value of 1, and is updated through back-propagation learning; ω is a learnable bias term, with an initial value of 0, and is updated through back-propagation learning; According to this step, it is continuously iterated. Under the condition that the total number of layers is L, the output X of the top layer can be obtained. L+1 , the final output of the model is obtained through specific projection, and the deduction results of the power grid state at each time step / location are obtained. Where W out is the projection matrix, which is randomly initialized at the beginning of training and updated by back-propagation learning through formula (38). Its function is to transform the model dimension into the target dimension; b out is the bias vector, with a value of 0-1, which is used to improve flexibility and fitting ability; All the learnable parameters in the above training process together constitute θ, as follows: θ={W e ,b e ,{W Q ,W K ,W V ,W O },{W1,b1,W2,b2},W out ,b out ,…} (32)。 9. According to claim 8, a dynamic distribution network deduction method based on graph neural network perception and privacy protection Transformer model is characterized in that: In step S7, differential privacy optimization technology is used to add noise to the gradient each time the gradient is updated, specifically including: Through the differential privacy SGD method, noise is added to the gradient each time the gradient is calculated to ensure that the gradient of each data point does not leak too much information. The specific formula is as follows: ||g i ||2≤C (34) Among them, g i represents the gradient of the parameter θ, C is a clipping threshold, if a g i If the norm of is greater than C, it is scaled to a norm of exactly C; Formula (33) uses the privacy budget to limit the noise added in the process; Formulas (34)-(36) are the gradient clipping process, the purpose of which is to control the impact of the gradient of each sample on the model update within a controllable range, so that differential privacy can obtain a clear sensitivity upper limit when noise is added later; the gradient after clipping is recorded as Within a batch, these gradients are averaged. represents the average gradient within the batch; Formula (37) adds noise to the gradient, N(0,σ2 2 C 2 ) represents a dimension and The same multidimensional Gaussian noise vector, is the final differential privacy gradient; Finally, using differentially private gradients To update the model parameters; if the learning rate is κ, the update formula for one iteration is: Where ξ(θ) is the loss function of the model deduction under θ; and i are the accurate value and true value of the future state of the power grid respectively; the optimization process of the loss function is to derive ξ(θ) with respect to θ and iteratively update θ until the model approaches the optimal fit on the training data.

10. A dynamic distribution network deduction method based on graph neural network perception and privacy-preserving Transformer model according to claim 9, characterized in that: The specific process of using the DP-Transformer model to perform distribution network state deduction in step S8 includes: Combine each feature in the scene into a vector R(t), where t represents the current moment and the dimension of R(t) is d input , integrate the required electrical quantity, the specific formula is as follows: R n,i (t)=[P 1oad,i (t),P gen,i (t),E storage,i (t),V n,i (t)] T (40) R l,j (t)=[P flow,j (t)] T (41) R(t)=[R n,i (t);R l,j (t)] (42) In the formula, the feature vector R of each node i is n,i (t) contains the characteristics of load, generator output, energy storage status and node voltage; the characteristic vector R of each line j l,j (t) contains the line flow characteristics; the characteristic vectors of all nodes and lines are concatenated row by row, and finally the overall input matrix R(t) is obtained, with N nodes. n , the number of lines is N l ; R(t) is used as the initial input of the model reasoning. In the deduction stage, the trained DP-Transformer performs forward propagation and finally obtains the power grid state at the next moment. According to the output layer dimension, the model can simultaneously output N v The node voltage and N f The specific formula for the line flow is as follows: In the formula, t0 is the initial time of the simulation; Δt is a fixed time interval; For the deduction result; To deduce the voltage of each node, The tidal currents of each route are deduced.

Citation Information

Cited By

  • Intelligent optimization method for inverting new material structure parameters

    CN120356589A

  • An intelligent optimization method for inverting structural parameters of new materials

    CN120356589B

  • Distribution network line facility defect detection method based on line magnetic field variable characteristics

    CN120522512A

  • Mineral resource reserve intelligent management method and system based on big data

    CN120724481A

  • Intelligent management method and system for mineral resources reserves based on big data

    CN120724481B