Active power distribution network fault positioning and identification method and system based on space-time diagram network

Through the fault diagnosis model based on the spatio-temporal graph network, combined with dynamic data augmentation and transfer learning strategies, the accuracy and adaptability of fault location and identification in the active distribution network are solved, and the synchronous judgment of fault location and type in complex scenarios is realized, and the accuracy and adaptability of fault diagnosis are improved.

CN120354254AActive Publication Date: 2025-07-22SHANDONG UNIV

Patent Information

Application Number
CN202510854703.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The fault location and identification methods in the active distribution network in the prior art have problems such as poor training data quality, poor adaptability to scene transformation, and low fault diagnosis accuracy. It is especially difficult to achieve synchronous judgment of fault location and type in complex scenarios.

Method used

A fault diagnosis model based on spatiotemporal graph network is adopted, and a fault diagnosis model is built for fault location and identification through dynamic data augmentation method and transfer learning strategy, combined with multi-scale attention mechanism and dynamic nuclearized maximum mean difference.

Benefits of technology

It improves the adaptability and reliability of fault location and identification, improves the accuracy of fault diagnosis, shortens fault location and identification time, reduces power outage losses, and enhances the model's adaptability to complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354254A_ABST
    Figure CN120354254A_ABST
Patent Text Reader

Abstract

The invention discloses an active power distribution network fault positioning and identification method and system based on a space-time diagram network, and relates to the technical field of active power distribution network fault positioning and identification, and the method comprises the steps: obtaining data of each node of an active power distribution network after a fault occurs, building structured diagram data, marking fault nodes and types, and constructing a sample data set; pre-training the fault diagnosis model based on the space-time diagram network by taking the sample data set as source domain data; in each iteration process, a multi-scale adversarial disturbance addition method based on gradient optimization is adopted to generate adversarial sample data, through a data category dynamic balance mechanism and a confidence evaluation mechanism, the data are screened and added to a data set, and the updated data set is utilized to train a model; introducing target domain data, and finely tuning the pre-training model by adopting a transfer learning strategy based on dynamic coring maximum mean difference; and inputting actually acquired node data into the model to realize accurate positioning of fault nodes and accurate identification of fault types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of active distribution network fault location and identification, and particularly relates to a method and system for active distribution network fault location and identification based on a spatio-temporal graph network. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, the development requirements of high-reliability power grids and smart power grids have promoted the widespread popularity of real-time intelligent measurement devices such as phasor measurement units (PMUs), providing new opportunities for the application of machine learning methods with data-driven as the core in actual distribution systems. However, the widespread access of distributed power sources and new types of loads has significantly changed the network structure, power flow direction, and load characteristics of the distribution network, resulting in more complex operating characteristics and fault mechanisms. These problems have reduced the applicability of traditional fault location and identification methods and deteriorated the identification ability, leading to an increased risk of misjudgment and missed judgment of faults, seriously affecting the safe and stable operation of the active distribution network and the traceability, investigation, and maintenance after faults.

[0004] To solve the problem of fault diagnosis in active distribution networks, artificial intelligence technology has been widely applied in this field due to its feature extraction ability and powerful non-linear fitting ability. Currently, the mainstream artificial intelligence diagnosis methods include expert systems, artificial neural networks, Bayesian networks, fuzzy set theory, Petri nets, and various optimization algorithms. Considering the complex and changeable topology structure and operating conditions of active distribution networks and the large-scale data, the processing and analysis of fault data largely depend on the selection of methods. However, most of the current fault diagnosis and location methods combined with artificial intelligence technology only use a single model or algorithm, often failing to fully explore the spatio-temporal correlation characteristics of distribution network fault data during the data processing process and lacking the dynamic processing ability for changing data, resulting in low accuracy of fault location and identification and poor adaptability to scene changes. In addition, the existing methods also have the limitation of relying too much on the quality of training data, and using additional models such as traditional generative adversarial networks and variational autoencoders to enhance the data will increase the consumption of training resources. Summary of the Invention

[0005] To solve the deficiencies of the above-mentioned existing technologies, the present invention provides a method and system for fault location and identification of active distribution networks based on spatio-temporal graph networks. By designing a dynamic data augmentation method, introducing a transfer learning strategy, and building a fault diagnosis model based on spatio-temporal graph networks to perform fault location and identification on active distribution networks, it realizes the synchronous judgment of fault locations and fault types in complex scenarios, improves the adaptability and reliability of fault location and identification, and solves the problems of poor training data quality, poor adaptability to scenario changes, and low fault diagnosis accuracy in existing methods.

[0006] In the first aspect, the present invention provides a method for fault location and identification of active distribution networks based on spatio-temporal graph networks.

[0007] A method for fault location and identification of active distribution networks based on spatio-temporal graph networks includes: Obtain the data of each node of the target-domain active distribution network, and input the node data into the fault diagnosis model to locate the fault node and identify the fault type; the training process of the fault diagnosis model is as follows: Obtain the data of each node of the active distribution network after a fault occurs, build a structured graph data, and label the fault nodes and types to construct a sample data set; Use the sample data set as the source-domain data to pre-train the fault diagnosis model based on spatio-temporal graph networks; in each iteration process, use the multi-scale adversarial perturbation addition method based on gradient optimization to generate adversarial sample data, and through the data category dynamic balance mechanism and the confidence evaluation mechanism, screen the data and add it to the data set, and use the updated data set to train the model; Introduce the target-domain data, and use the transfer learning strategy based on dynamic kernel maximum mean discrepancy to fine-tune the pre-trained model.

[0008] In a further technical solution, the node data includes the amplitudes and phase angles of the three-phase voltages at the current node within a T-time step window.

[0009] In a further technical solution, using the sample data set as the training data set to pre-train the fault diagnosis model based on spatio-temporal graph networks includes: In each iteration process, generate multi-scale adversarial perturbations generated at multiple step lengths and add them to the sample data in the training data set to generate adversarial sample data, and accumulate the gradients of the perturbations added this time to optimize the generation of perturbations in the next iteration; Based on the sample data and the generated adversarial sample data, introduce the data category dynamic balance mechanism, monitor the distribution of all samples of each fault type, judge whether the category balance judgment condition is met, and according to the category balance judgment result, combined with the confidence evaluation mechanism, update the training data set; Continue to train the fault diagnosis model using the updated training dataset, optimize the model parameters to meet the optimization goal of minimizing the classification error for the current iteration, and then perform the next round of iteration until the maximum number of iterations is reached to complete the model training.

[0010] A further technical solution is to update the training dataset according to the category balance judgment result in combination with the confidence evaluation mechanism, including: If the balance judgment condition is not met, adversarial samples with a confidence higher than the set confidence threshold are selected for augmentation according to the confidence of the generated adversarial samples to update the training dataset; Conversely, if the balance judgment condition is met, it is determined that the current training dataset is a data-balanced training dataset, and the training dataset is not updated.

[0011] A further technical solution is that the fault diagnosis model is built using a spatio-temporal graph network based on a multi-scale attention mechanism, including a time module, a space module, and an identification module; Among them, the sample data is input into the spatio-temporal graph network. The time series data of each node in the sample data first extracts multi-time scale features through the time module, and through the multi-scale spatio-temporal attention mechanism, the time features of each node are obtained; The time features of all nodes are input into the space module to extract multi-scale space features, and the multi-scale space features are fused through a graph attention network introducing a multi-hop mechanism and a gating mechanism to obtain spatio-temporal features; The spatio-temporal features are input into the identification module. After feature transformation, they are converted into a category probability distribution through a softmax layer, and the fault node location and fault category recognition results are output.

[0012] A further technical solution is to introduce target domain data and use a transfer learning strategy based on dynamic kernel maximum mean discrepancy to fine-tune the pre-trained model, including: Freeze the time module and space module of the pre-trained model, and use the frozen pre-trained model to extract the features of the target domain data; Calculate the dynamic kernel maximum mean discrepancy loss between the target domain data features and the source domain data features to achieve feature alignment; Construct a total loss function based on the dynamic kernel maximum mean discrepancy loss and the fault location and recognition task loss, and fine-tune the frozen pre-trained model in combination with the target domain features. Continuously optimize the parameters of the identification module and the kernel weights through the iterative process of forward propagation and backward propagation until the maximum number of iteration rounds is reached to complete the fine-tuning of the pre-trained model.

[0013] In the second aspect, the present invention provides an active distribution network fault location and recognition system based on a spatio-temporal graph network.

[0014] An active distribution network fault location and identification system based on a spatio-temporal graph network, comprising: A data acquisition module, configured to acquire the data of each node of the active distribution network after a fault occurs, construct structured graph data, label the fault node and its type, and construct a sample data set; A model pre-training module, configured to pre-train a fault diagnosis model based on a spatio-temporal graph network using the sample data set as source domain data; in each iteration process, generate adversarial sample data by using a multi-scale adversarial perturbation addition method based on gradient optimization, screen the data through a data category dynamic balancing mechanism and a confidence evaluation mechanism, add the data to the data set, and train the model using the updated data set; A model fine-tuning module, configured to introduce target domain data and fine-tune the pre-trained model by using a transfer learning strategy based on dynamic kernel maximum mean discrepancy; A fault location and identification module, configured to acquire the data of each node of the target domain active distribution network, input the node data into the fault diagnosis model, and locate the fault node and identify the fault type.

[0015] In a third aspect, the present invention further provides an electronic device, comprising: a memory, configured to store executable instructions; a processor, configured to implement the above-mentioned active distribution network fault location and identification method based on a spatio-temporal graph network when executing the executable instructions stored in the memory.

[0016] In a fourth aspect, the present invention further provides a computer-readable storage medium, storing executable instructions, configured to cause a processor to implement the above-mentioned active distribution network fault location and identification method based on a spatio-temporal graph network when executing the executable instructions.

[0017] In a fifth aspect, the present invention further provides a computer program product, which includes executable instructions stored in a computer-readable storage medium; wherein, when a processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the above-mentioned active distribution network fault location and identification method based on a spatio-temporal graph network is implemented.

[0018] The above one or more technical solutions have the following beneficial effects: 1. The present invention proposes a method and system for fault location and identification of active distribution networks based on spatio-temporal graph networks. By designing a dynamic data augmentation method, introducing a transfer learning strategy, and building a fault diagnosis model based on spatio-temporal graph networks to perform fault location and identification on active distribution networks, it realizes the synchronous judgment of fault locations and fault types in complex scenarios. Among them, the dynamic data augmentation method is used to effectively process the original data to solve the problem of poor quality of training data. The spatio-temporal graph network based on the multi-scale attention mechanism is used to locate and identify known faults, improving the accuracy of fault diagnosis. The transfer learning strategy is used to cope with the changes in the operation scenarios of active distribution networks and improve the adaptability to scenario changes. Through the above fault location and identification method with strong adaptability and high reliability, the positions and types of faults such as short circuits and line breaks in the distribution network can be effectively identified, and the problems of low accuracy and long training time for fault location and identification in active distribution networks can be targeted solved, which helps to improve the intelligent perception level of active distribution networks, shorten the fault location and identification time of short circuits, line breaks, etc., reduce power outage losses, and help the rapid recovery of the normal operation state of active distribution networks.

[0019] 2. The present invention proposes a dynamic data augmentation method that combines adversarial sample generation and dynamic balancing mechanism. By embedding the dynamic data augmentation method into the iterative training process of the fault location and identification model, adversarial samples are continuously generated, the proportion of each type of sample is automatically identified, and the number of minority class samples is expanded, providing a high-quality data basis for the fault location and identification tasks. Moreover, by dynamically generating training samples with adversarial properties and class balance, the consumption of computing resources can be significantly reduced, and the adaptability of the model to out-of-distribution samples can be improved.

[0020] 3. The present invention proposes a spatio-temporal graph network model based on the multi-scale attention mechanism. This spatio-temporal graph network model is divided into a time module, a space module, and an identification module. The multi-scale time and space attention mechanisms are respectively introduced into the time and space modules to extract the multi-hop spatial correlation features between different nodes in the distribution network and the temporal dependence features at different time scales, so as to flexibly capture and dynamically fuse the spatial topological information and temporal fluctuation features at different scales in the operation data of the active distribution network, and input them into the identification module for fault location and identification, improving the accuracy of the model for locating and identifying complex faults.

[0021] 4. The present invention proposes an unsupervised transfer learning strategy for the target domain based on dynamic kernel maximum mean discrepancy. By using the multi-kernel maximum mean discrepancy method to align the distribution differences between the source domain and target domain data, the adaptability of the model to unknown operation scenarios is effectively improved, the learning ability of the model to dynamic distribution changes is enhanced, and the fault states under complex and variable operation conditions can be better identified.

[0022] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Brief Description of the Drawings

[0023] The attached drawings of the specification that form a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0024] Figure 1 It is the overall flowchart of the active distribution network fault location and identification method based on the spatio-temporal graph network described in the embodiments of the present invention; Figure 2 It is the flowchart of dynamic data enhancement and model training in the embodiments of the present invention; Figure 3 It is the architecture diagram of the fault diagnosis model based on the spatio-temporal graph network in the embodiments of the present invention; Figure 4 It is the flowchart of transfer learning based on DK-MMD in the embodiments of the present invention. Detailed Embodiments

[0025] It should be noted that the following detailed description is exemplary only for describing the specific embodiments, aiming to provide a further explanation of the present invention, and is not intended to limit the exemplary embodiments according to the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0026] Embodiment 1 This embodiment provides an active distribution network fault location and identification method based on a spatio-temporal graph network. This method is based on a multi-task learning framework for fault location and identification, effectively processes the original data using a dynamic data enhancement method, locates and identifies known faults through a spatio-temporal graph network based on a multi-scale attention mechanism, and uses a transfer learning strategy to cope with the changes in the operating scenarios of the active distribution network, which can specifically solve the problems of low accuracy and long training time for active distribution network fault location and identification. Specifically, as Figure 1 shown, this method specifically includes the following steps: Step S1, obtain the data of each node of the active distribution network after a fault occurs, build a structured graph data, label the fault nodes and types, and construct a sample data set.

[0027] Specifically, for an active distribution network, in this embodiment, distribution network measurement devices such as micro synchronous phasor measurement units, advanced metering infrastructure, and smart meters are used to collect voltage data of each node in the distribution network after different types of faults occur, and non-Euclidean structured graph data is constructed.

[0028] Among them, for the generation of graph data, it is as follows: In graph theory, a graph consists of nodes and edges. Regarding the buses of the active distribution network as the nodes of the graph and the connecting lines as the edges of the graph, the real-time measurement data of the active distribution network can be associated with the topological structure through graph data. Then, the graph data can be expressed as: (1) Wherein, is the node set, n is the number of nodes, represents the n th node, is the edge set, m is the number of edges. If there is a straight line between nodes and , then ([[]] i , j ) and ([[]] j , i ) are both included in , thus forming a bidirectional graph topology. Due to the bidirectionality of power flow in the active distribution network, the constructed graph data should be an undirected graph.

[0029] Furthermore, each node is associated with time series voltage phase measurement. Then, the node feature data is the amplitudes T and phase angles of the three-phase voltages (A, B, C) within the time-step window. Its mathematical expression is: (2) Wherein, represents the feature vector of the t th data point of the i th node.

[0030] The overall input feature tensor of the structured graph constructed above is . Assuming that the voltage phases of the key nodes are observable, for the unobservable nodes, their features are filled with zeros to maintain the consistency of the structure.

[0031] After that, the obtained data is labeled with the fault nodes and fault types, and the labeled data is used as sample data to construct a sample data set.

[0032] ​​Step S2: Use the sample data set as the source domain data to pre-train the fault diagnosis model based on the spatio-temporal graph network; in each iteration process, use the multi-scale adversarial perturbation addition method based on gradient optimization to generate adversarial sample data, and through the data category dynamic balance mechanism and the confidence evaluation mechanism, screen the data and add it to the data set, and use the updated data set to train the model.

[0033] Specifically, this embodiment proposes a dynamic data augmentation method, directly embed this dynamic data augmentation method into the training process of the model, continuously generate adversarial samples and dynamically augment the minority class samples, screen and retain the samples with higher confidence, and gradually achieve the balance of the data distribution during the training process.

[0034] Furthermore, the above-mentioned proposed dynamic data augmentation and the model training process are as Figure 2 shown. Use the sample data set obtained in the above step S1 as the training data set. For the sample data in this data set, first generate adversarial samples by adding multi-scale perturbations, and aggregate the gradients of all scales in the way of cumulative gradients to facilitate the update of the perturbations in the next iteration. Then, based on the original samples and the generated adversarial samples, calculate the number of samples of each category, and use the dynamic balance mechanism to judge whether the samples are evenly distributed. If not, screen the minority class samples with high confidence for supplementation and add them to the training data set. If they are already balanced, use the balanced data set to continue training the model. The above training process continues until the maximum number of iterations is reached, until the training of the fault diagnosis model is completed. The above process is specifically as follows: Step S2.1: In each iteration process, generate multi-scale adversarial perturbations generated by multi-step lengths and add them to the sample data in the training data set to generate adversarial sample data, and accumulate the gradients of the perturbations added this time to optimize the generation of perturbations in the next iteration.

[0035] In this embodiment, the gradient-based adversarial perturbation iteratively enhances the node features to generate adversarial samples that can be generalized beyond the initial input. In adversarial training, the goal of sample generation can be summarized as a min-max optimization problem (Min-Max Optimization). The goal of this min optimization is to minimize the classification error under the maximum perturbation by adjusting the model parameters , and this max optimization is to maximize the current classification error by adjusting the adversarial perturbation . The mathematical formula of the min-max optimization problem is as follows: (3) where represents the classification loss of the model, represents the norm limit of the perturbation, which is used to control the generation range of the adversarial samples, Represents a parameterized classifier, Represents the predicted classification result, Represents the actual classification result.

[0036] That is, the generation of adversarial samples adopted in this embodiment is based on the feature perturbation addition method optimized by the gradient direction. Its core idea is to mine the vulnerable areas of the model by maximizing the classification loss of the model. However, considering the addition of single-scale perturbations in traditional gradient optimization methods, it is difficult to cover multiple feature distributions outside the original area. For this reason, on the basis of traditional PGD (Projected Gradient Descent) perturbations, this embodiment introduces multi-scale adversarial perturbations generated by multi-step lengths for adversarial training. By introducing multi-step lengths, diverse adversarial samples are generated to maximize the current classification error. The calculation formula for this perturbation is: (4) Where, Represents the th perturbation scale; Represents the step length of the multi-scale perturbation, simulating abnormal fluctuations of different degrees; Is a projection operation used to limit the perturbation within the constraint range; Represents the gradient of the perturbation.

[0037] Furthermore, during the generation of the above multi-scale perturbations, each perturbation will introduce a new set of gradient updates, which will significantly increase the computational consumption of the iterative process. To reduce the computational cost, this embodiment introduces cumulative gradient optimization to aggregate the gradients of all scales in one backpropagation for the update of the model in the next iteration. The formula is: (5) Where, Represents the total number of steps of the multi-scale perturbation; Represents the learning rate; Is the gradient with respect to the model parameters.

[0038] Through the above cumulative gradient, the update of the model in each iterative optimization can comprehensively consider the gradient directions of all perturbation scales, avoiding the local optimum caused by single-scale perturbations and the additional consumption brought by repeated calculations.

[0039] Step S2.2: Based on the sample data and the generated adversarial sample data, introduce a data category dynamic balance mechanism, monitor the distribution of all samples of each fault type, determine whether the category balance judgment condition is satisfied, and update the training data set according to the category balance judgment result in combination with the confidence evaluation mechanism.

[0040] Specifically, since the imbalance of data categories not only affects the robustness of model training but also further worsens the distribution bias of adversarial samples, this embodiment proposes a category dynamic balancing mechanism. By monitoring the distribution of samples of each category during each round of iteration, the minority class dataset is gradually expanded until the balance of the category distribution is achieved. Among them, the balance judgment condition is: (6) Among them, represents the number of samples of a category, represents the tolerance deviation of the category distribution.

[0041] Furthermore, if the above balance judgment condition is not met, then based on the confidence evaluation mechanism, according to the confidence of the generated adversarial samples, the adversarial samples with a confidence higher than the set confidence threshold are selected for expansion and added to the dataset to update the training dataset; conversely, if the above balance judgment condition is met, then the current training dataset is judged as a training dataset with balanced data.

[0042] In this embodiment, during the sample expansion process, combined with the confidence evaluation mechanism, only high-confidence samples are retained for expansion, thereby reducing the negative impact of pseudo-samples. Among them, the calculation formula for confidence is: (7) In this embodiment, only when and the sample belongs to the minority class, it is added to the balanced dataset. In the above formula, represents the th adversarial sample, represents the set confidence threshold for screening high-confidence samples.

[0043] Step S2.3: Use the updated training dataset to continue training the fault diagnosis model, optimize the model parameters to meet the optimization goal of minimizing the classification error for the current iteration, and then perform the next round of iteration, that is, loop and iterate the above steps S2.1~S2.3, re-iterate the generation process and the model training process until the maximum number of iterations is reached, and complete the training of the fault diagnosis model.

[0044] As an implementation method, during the above iterative training process, the number of iterations of model training may not be equal to the number of iterations of adversarial sample generation. For example, when the model is iteratively trained for 200 rounds, it may be that after 30 rounds of iteration, the initially constructed sample dataset reaches category balance after data augmentation and has been updated to a balanced dataset. At this time, the balanced dataset can be directly used to train the model for the subsequent iterative rounds.

[0045] In this embodiment, the data augmentation process is embedded in the training process of the spatio-temporal graph network model without relying on an additional adversarial model. During the entire training process, while generating multi-scale adversarial perturbations, the parameters of the spatio-temporal graph network model are updated, enabling parallel computing of parameter and perturbation updates, which can effectively reduce the training cost.

[0046] Furthermore, during the above training process, using the sample data in the training sample dataset as the source domain data, the fault diagnosis model based on the spatio-temporal graph network is pre-trained.

[0047] This embodiment formulates the fault diagnosis of the active distribution network as a multi-task learning problem, which includes two main tasks: 1) Fault type recognition: Classify the fault type among multiple predefined categories. This task belongs to the graph-level multi-classification problem, and the model needs to capture the overall features and patterns in the graph to achieve the discrimination of the fault type.

[0048] 2) Fault location: Predict the exact node where the fault occurs. This task is a node-level classification problem, requiring the model to have the ability to independently judge the state of each node, so as to achieve the accurate location of the fault node.

[0049] Based on the above multi-tasks, combined with the updated training sample dataset , where is the k th sample graph, including the type recognition label and the node-level location label, design the spatio-temporal graph network model with a shared spatio-temporal backbone and two task-specific heads and , and its overall goal is to jointly minimize the weighted sum of the losses of the two tasks, which is: (8) where, is the cross-entropy loss for fault type recognition, is the node-level classification loss for fault location; the weights and control the relative importance of the two tasks and are adjusted according to experience.

[0050] Furthermore, the fault diagnosis model based on the spatio-temporal graph network built in this embodiment is divided into three layers: a time module, a space module, and an identification module. By introducing a multi-scale time and space attention mechanism, it deeply excavates the node time series features and network structure features at different scales, and improves the feature extraction ability of the model. Specifically, the architecture of the spatio-temporal graph network model based on the multi-scale attention mechanism is as Figure 3 shown, including: (1) Space module For graph data for each node , its neighbor node set is denoted as . The basic graph attention network updates the node representation by weighted aggregation of the central node and its neighbor nodes through the attention mechanism. Then the basic update rule of the node output feature is: (9) where is the new feature that fuses the neighborhood information for each node, is the concatenation operator, K represents the number of heads of the multi-head attention mechanism, is the non-linear activation function, is the learnable linear transformation matrix; represents the attention weight between node and node , reflecting the importance of node to node . Its specific calculation formula is: (10) (11) where is the unnormalized attention weight, which becomes after being normalized by the Softmax function; LeakyReLU is an activation function, usually used to introduce non-linearity and prevent gradient vanishing; is a learnable attention vector.

[0051] To mine the broader spatial relationship information of the network, in this embodiment, a multi-hop mechanism and a gating mechanism are introduced on the basis of the basic graph attention layer for multi-scale spatial feature fusion. For neighbor nodes with different hop numbers , the update rule of the output feature of each central node changes to: (12) where H is the hop number of the neighbor node, representing the aggregation range of different scale spatial information; is the gating coefficient, dynamically generated through the gating mechanism, used to balance the contributions of different scale spatial features in node update. Its calculation formula is: (13) where is the dynamically adjusted gating weight matrix; is an activation function that restricts the gating coefficient to between.

[0052] (2) Time module Since most graph neural networks such as graph convolutional networks and graph attention networks focus on mining the spatial characteristics of data and do not fully consider the temporal dependence of node features, this embodiment introduces a multi-scale temporal attention mechanism to enhance the network's ability to extract temporal features.

[0053] Specifically, the time series features of each node are initially extracted through temporal convolution. To quickly extract temporal features, the convolution operation is performed through a one-dimensional convolutional layer. Assuming M time scales are defined, and the size of each time scale window is respectively. Further, for the m th time scale, the temporal attention mechanism is defined as: (14) where represents the attention output at the m th time scale, , and are the query, key, and value matrices generated on the time step window respectively, which are the three basic elements of the self-attention mechanism; is the scaling factor of the attention head, which is used to prevent the numerical value from being too large.

[0054] The multi-scale temporal attention mechanism also introduces a gating mechanism, and its calculation formula is: (15) where, is the gating coefficient at the m th time scale, which reflects the importance of the m th scale in the fusion process.

[0055] (3) Identification module As the output layer of the spatio-temporal graph network, the identification module is mainly used to realize the identification of the fault type and the determination of the fault location of the distribution network nodes. Based on the multi-scale spatio-temporal feature vectors and extracted by the previous two modules, a joint representation vector is formed through feature concatenation operation and input into a multi-layer perceptron network to complete the fault location and identification tasks.

[0056] First, the feature vectors are compressed and structured through max-pooling and flattening operations to obtain an input vector with a fixed dimension . Subsequently, the input undergoes feature transformation through several fully connected layers with non - linear activation functions. The output of its layer is expressed as: (16) where, and are the weight matrix and bias term of the layer respectively, represents the activation function, .

[0057] The final output is transformed into a class probability distribution through the softmax layer, which is defined as follows: (17) where, represents the prediction result of fault type identification or fault node location; for the fault type identification task, represents the probability of each fault type; for the node - level location task, represents the probability that each node is a fault point.

[0058] In this embodiment, the sample data is input into the spatio - temporal graph network. The time - series data of each node in the sample data first extracts multi - time - scale features through the time module, and through the multi - scale spatio - temporal attention mechanism, the time features of each node are obtained; the time features of all nodes are then input into the space module to extract multi - scale space features, and through the graph attention network introducing the multi - hop mechanism and the gating mechanism for multi - scale space feature fusion, spatio - temporal features are obtained; the spatio - temporal features are finally input into the identification module, and after feature transformation, they are converted into a class probability distribution through the softmax layer, and the fault node location and fault class identification results are output.

[0059] Furthermore, to improve the adaptability of the fault diagnosis model, this embodiment also introduces a transfer learning strategy, which mainly includes two processes: using source - domain data for model pre - training and using the target domain for fine - tuning. First, in the pre - training stage (i.e., the process of step S2 above), dynamic data augmentation is performed on the source - domain data, which is sequentially input into the time module, space module, and identification module (MLP network), and the total loss of fault location and identification is calculated and the model parameters are optimized through backpropagation.

[0060] Secondly, in the fine - tuning stage, this embodiment freezes the time module and the space module, extracts the features of the target - domain data, realizes feature alignment by calculating the DK - MMD loss between the source - domain and target - domain features, and jointly constructs the total loss function with the loss of the fault location and identification task. Through the iterative process of forward propagation and backpropagation, the parameters of the identification module and the kernel weights are continuously optimized. After reaching the maximum number of iteration rounds, the fault location and identification results of the target domain are output, that is: Step S3: Introduce the target domain data and use a transfer learning strategy based on dynamic kernelized maximum mean discrepancy to fine-tune the pre-trained model.

[0061] As Figure 4 shown, first, freeze the temporal module and spatial module of the pre-trained model, and use the frozen pre-trained model to extract the features of the target domain data; second, calculate the dynamic kernelized maximum mean discrepancy loss between the target domain data features and the source domain data features to achieve feature alignment.

[0062] Among them, the design of the dynamic kernelized maximum mean discrepancy loss includes: Considering that transfer learning aims to improve the performance of the model in a new domain by transferring the knowledge of the source domain to the target domain. However, when there are differences in the distributions of the source domain and target domain data, direct transfer will lead to a decline in model performance. Therefore, in this embodiment, based on the distribution alignment method based on maximum mean discrepancy (MMD), by improving multi-kernel maximum mean discrepancy (MK-MMD), a dynamic kernelized maximum mean discrepancy method (DK-MMD) is proposed to enhance the flexible alignment ability for different distribution differences.

[0063] MMD is a non-parametric statistic for measuring the difference between two distributions. Its core idea is to map the samples of the source domain distribution and the target domain distribution to the Reproducing Kernel Hilbert Space (RKHS), and measure the distance between the distributions by calculating the kernel mean discrepancy of the source domain and target domain samples. The difference between distributions is usually described by a statistic. If and respectively represent the mean embeddings of the distributions and in the RKHS, then MMD is defined as: (18) (19) Among them, is the kernel function, is the RKHS formed by the kernel function.

[0064] The logic of the above MMD calculation formula comes from the mean embedding theory of distributions, that is, any distribution can be embedded into a high-dimensional space through a kernel function, so that its characteristic mean can directly reflect the distribution characteristics. Among them, the kernel mean embedding does not depend on the explicit construction of the mapping function , but is directly determined by the kernel function , so even if there may be different forms of mapping functions in the same RKHS, this does not affect the result of the distribution mean embedding. This embedding theoretically satisfies consistency and uniqueness, and thus is suitable for distribution alignment.

[0065] By expanding , we can obtain: (20) Further using the properties of the kernel function, the mean embedding is expressed as an expectation: (21) Therefore, the formula of MMD can be written as: (22) Since the distributions and are usually unknown, the calculation of MMD needs to be based on the empirical estimation of finite samples. Let the numbers of samples in the source domain and the target domain be and respectively, the source domain samples be , and the target domain samples be . Taking the sample mean as an unbiased estimate of the distribution expectation, the empirical estimation formula of MMD can be obtained as: (23) The above transformation from the theoretical formula to the empirical formula is based on the assumption of independent and identically distributed samples. When the sample size is sufficiently large, the empirical formula is an asymptotic approximation of the theoretical formula.

[0066] However, MMD only uses a single kernel function and cannot comprehensively capture the differences between complex distributions. MK-MMD improves this deficiency. MK-MMD introduces multiple kernel functions and their weights , and constructs a total kernel through multiple weighted kernels to capture more complex distribution differences. Based on the theoretical basis of MMD, the calculation formula of MK-MMD is: (24) where is the empirical estimation formula of MMD using the -th kernel function.

[0067] Considering that although MK-MMD can overcome the deficiencies of single-kernel MMD to a certain extent, its kernel weights are usually fixed and cannot dynamically adapt to the requirements of the task. For this reason, this embodiment proposes Dynamic Kernelized Maximum Mean Discrepancy (DK-MMD). Its core idea is to dynamically learn the kernel weights by introducing a parametric model (such as MLP) so that it can adapt to the feature changes of the source domain and the target domain. The calculation formula of DK-MMD is: (25) Among them, is the parameterized kernel weight, which is dynamically adjusted by learnable parameters (such as MLP weights); The optimization of is achieved by minimizing the total loss function.

[0068] Under the transfer learning framework, the total loss function of the model consists of the task loss and the DK-MMD loss: (26) Among them, is the total loss of the fault location and identification task, and the cross-entropy loss is adopted in this embodiment; is the weight coefficient of the distribution alignment loss.

[0069] Compared with the MK-MMD method with fixed kernel weights, the proposed DK-MMD can dynamically adjust the kernel weights through a lightweight parameterized model, which can not only more flexibly adapt to the distribution changes between the source domain and the target domain, but also more intuitively reflect the contribution of different kernel functions to the distribution alignment, improving the interpretability of the model.

[0070] Finally, based on the above dynamic kernelized maximum mean discrepancy loss and the fault location and identification task loss, a total loss function is constructed, and the frozen pre-trained model is fine-tuned in combination with the target domain features. The parameters and kernel weights of the identification module are continuously optimized through the iterative process of forward propagation and backward propagation until the maximum number of iteration rounds is reached, and the fine-tuning of the pre-trained model is completed.

[0071] Step S4, obtain the data of each node of the active distribution network in the target domain, and input the node data into the fault diagnosis model to locate the fault node and identify the fault type. Specifically, input the actually obtained node data of the active distribution network in the target domain into the above-mentioned trained fault diagnosis model, and output the fault node and its fault type.

[0072] Embodiment 2 This embodiment provides an active distribution network fault location and identification system based on a spatio-temporal graph network, including: A data acquisition module, which is used to acquire the data of each node of the active distribution network after a fault occurs, build structured graph data, label the fault nodes and types, and construct a sample data set; A model pre-training module, which is used to pre-train a fault diagnosis model based on a spatio-temporal graph network with the sample data set as the source domain data; in each iteration process, the multi-scale adversarial perturbation addition method based on gradient optimization is used to generate adversarial sample data, and through the data category dynamic balance mechanism and the confidence evaluation mechanism, the data is screened and added to the data set, and the model is trained with the updated data set; A model fine-tuning module for introducing target domain data and fine-tuning a pre-trained model by adopting a transfer learning strategy based on dynamic kernel maximum mean discrepancy. A fault location and identification module for obtaining data of each node of the target domain active distribution network, inputting the node data into a fault diagnosis model to locate the fault node and identify the fault type.

[0073] Embodiment III This embodiment provides an electronic device, including: a memory for storing executable instructions; a processor for implementing the above method provided in this embodiment when executing the executable instructions stored in the memory.

[0074] Embodiment IV This embodiment also provides a computer-readable storage medium storing executable instructions, which when executed by a processor, will cause the processor to execute the above method provided in this embodiment.

[0075] Embodiment V This embodiment provides a computer program product, which includes executable instructions, and the executable instructions are computer instructions; the executable instructions are stored in a computer-readable storage medium. When a processor of an electronic device reads the executable instructions from the computer-readable storage medium and the processor executes the executable instructions, the electronic device is caused to execute the above method provided in this embodiment.

[0076] The steps involved in Embodiments II to V above correspond to those in Method Embodiment I, and the specific implementation manners can be referred to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.

[0077] Those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0078] The above are only the preferred embodiments of the present invention. Although the specific implementation manners of the present invention have been described in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solution of the present invention are still within the protection scope of the present invention.

Claims

1. A fault location and identification method for active distribution networks based on spatio-temporal graph networks, characterized in that, Including: Obtain the data of each node in the target-domain active distribution network, input the node data into the fault diagnosis model to locate the fault node and identify the fault type. Among them, the training process of the fault diagnosis model is as follows: Obtain the data of each node in the active distribution network after a fault occurs, construct a structured graph data, and label the fault node and its type to construct a sample data set. Use the sample data set as the source-domain data to pre-train the fault diagnosis model based on the spatio-temporal graph network. In each iteration process, use the multi-scale adversarial perturbation addition method based on gradient optimization to generate adversarial sample data. Through the data category dynamic balance mechanism and the confidence evaluation mechanism, screen the data and add it to the data set, and use the updated data set to train the model. Introduce the target-domain data and use the transfer learning strategy based on dynamic kernel maximum mean discrepancy to fine-tune the pre-trained model.

2. The method for fault location and identification of an active distribution network based on a spatio-temporal graph network according to claim 1, wherein, The node data includes the amplitude and phase angle of the three-phase voltage at the current node within a T-time-step window.

3. The method for fault location and identification of an active distribution network based on a spatio-temporal graph network according to claim 1, wherein, Using the sample data set as the training data set to pre-train the fault diagnosis model based on the spatio-temporal graph network includes: In each iteration process, generate multi-scale adversarial perturbations generated by multiple step lengths and add them to the sample data in the training data set to generate adversarial sample data, and accumulate the gradient of this added perturbation to optimize the generation of perturbations in the next iteration. Based on the sample data and the generated adversarial sample data, introduce the data category dynamic balance mechanism, monitor the distribution of all samples of each fault type, judge whether it meets the category balance judgment condition, and according to the category balance judgment result, combined with the confidence evaluation mechanism, update the training data set. Use the updated training data set to continue training the fault diagnosis model, optimize the model parameters to meet the optimization goal of minimizing the classification error in the current iteration, and then perform the next round of iteration until the maximum iteration number is reached to complete the model training.

4. The method for active distribution network fault location and identification based on a spatio-temporal graph network according to claim 3, characterized in that, According to the category balance judgment result, combined with the confidence evaluation mechanism, update the training data set, including: If the balance judgment condition is not met, select the adversarial samples with a confidence higher than the set confidence threshold for expansion according to the confidence of the generated adversarial samples, and update the training data set. On the contrary, if the balance judgment condition is met, judge that the current training data set is a data-balanced training data set and do not update the training data set.

5. The method for fault location and identification of an active distribution network based on a spatio-temporal graph network according to claim 1, characterized in that, The fault diagnosis model is built using a spatio-temporal graph network based on a multi-scale attention mechanism, including a time module, a space module, and an identification module. Among them, input the sample data into the spatio-temporal graph network. The time series data of each node in the sample data first extracts multi-time-scale features through the time module, and through the multi-scale spatio-temporal attention mechanism, obtains the time features of each node. Input the time features of all nodes into the space module to extract multi-scale space features, and perform multi-scale space feature fusion through the graph attention network introducing the multi-hop mechanism and the gating mechanism to obtain spatio-temporal features. Input the spatio-temporal features into the identification module. After feature transformation, convert them into a category probability distribution through the softmax layer, and output the fault node location and fault category recognition results.

6. The active distribution network fault location and identification method based on the spatio-temporal graph network according to claim 1, characterized in that, Introduce target domain data, and adopt a transfer learning strategy based on dynamic kernel maximum mean discrepancy to fine-tune the pre-trained model, including: Freeze the temporal module and spatial module of the pre-trained model, and use the frozen pre-trained model to extract the features of the target domain data; Calculate the dynamic kernel maximum mean discrepancy loss between the features of the target domain data and the features of the source domain data to achieve feature alignment; Construct a total loss function based on the dynamic kernel maximum mean discrepancy loss and the fault location and identification task loss, and fine-tune the frozen pre-trained model by combining the target domain features. Continuously optimize the parameters and kernel weights of the identification module through the iterative process of forward propagation and backward propagation until the maximum number of iterations is reached, and complete the fine-tuning of the pre-trained model.

7. An active distribution network fault location and identification system based on a spatio-temporal graph network, characterized in that, Including: A data acquisition module for acquiring the data of each node of the active distribution network after a fault occurs, building a structured graph data, annotating the fault nodes and types, and constructing a sample data set; A model pre-training module for pre-training a fault diagnosis model based on a spatio-temporal graph network using the sample data set as the source domain data; In each iteration process, use the multi-scale adversarial perturbation addition method based on gradient optimization to generate adversarial sample data, screen the data through the data category dynamic balance mechanism and the confidence evaluation mechanism, add the data to the data set, and use the updated data set to train the model; A model fine-tuning module for introducing target domain data and fine-tuning the pre-trained model by adopting a transfer learning strategy based on dynamic kernel maximum mean discrepancy; A fault location and identification module for acquiring the data of each node of the target domain active distribution network, inputting the node data into the fault diagnosis model, and locating the fault node and identifying the fault type.

8. An electronic device, characterized in that, Including: A memory for storing executable instructions; A processor, when executing the executable instructions stored in the memory, implements the method for active distribution network fault location and identification based on a spatio-temporal graph network according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Stored with executable instructions for causing the processor to implement the method for active distribution network fault location and identification based on a spatio-temporal graph network according to any one of claims 1-6 when executing the executable instructions.

10. A computer program product, characterized in that, The computer program product includes executable instructions, and the executable instructions are stored in a computer-readable storage medium; When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the method for active distribution network fault location and identification based on a spatio-temporal graph network according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Intelligent fault prediction method and system for filling packaging machine

    CN113779724A

  • Rolling bearing cross-working-condition fault detection method based on migration convolutional neural network

    CN114152442A

  • Fault diagnosis and prediction method and system for distribution transformer

    CN117092421A

  • Photovoltaic power prediction method based on multi-scale space-time diagram attention convolutional network

    CN117154704A

  • Active power distribution network abnormal state sensing method and system based on data enhancement

    CN118395363A

Cited By

  • Power distribution network line fault positioning method, device and equipment

    CN120669059B

  • Transfer learning-based cross-regional electric meter anomaly detection method and system, and electric energy meter

    CN120822157A

  • A cross-region electric meter anomaly detection method and system based on transfer learning and an electric energy meter

    CN120822157B

  • Event-driven architecture-oriented spatio-temporal feature extraction method and system for multi-type fault power failure events

    CN120995281A

  • Event-driven architecture-oriented spatiotemporal feature extraction method and system for multi-type fault power outage events

    CN120995281B