Master-Slave Cooperative Fault Detection Method, Device and Medium Based on Mamba Network

The Mamba network addresses inefficiencies in power grid fault diagnosis by integrating time-series and graph convolutional networks with prior knowledge, enhancing accuracy and reducing complexity for rapid fault detection.

CN119885083BActive Publication Date: 2025-07-15HEFEI UNIV OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510346687.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-15
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing power grid fault diagnosis methods have limitations when processing complex data, especially inefficient in timing characteristics and graph structure data processing, which is difficult to meet the millisecond response requirements, and the error detection rate is high under newly connected distributed power supplies and severe weather conditions, which cannot meet the rapid diagnosis and repair requirements of modern power systems.

Method used

The main-coupled collaborative fault detection method based on Mamba network is adopted, and the adjacency matrix and edge feature matrix of the graph are constructed, combined with the timing feature extraction network LSTM, graph convolution neural network, prior knowledge and feature fusion module are combined, and the joint loss function is used for training, which optimizes information dissemination and calculation efficiency, and improves the accuracy and efficiency of the model.

Benefits of technology

It realizes efficient processing of timing characteristics and graph structures in power grid fault diagnosis, reduces calculation complexity, improves model performance and computing efficiency, and can perform fault detection in real time or near real time, reduces delay, adapts to new access to distributed power supplies and reduces false detection rates in bad weather.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119885083B_ABST
    Figure CN119885083B_ABST
Patent Text Reader

Abstract

The present invention discloses a main-distribution collaborative fault detection method, device and medium based on the Mamba network, including: S1. Measuring the main-distribution network fault data and constructing a training set; S2. Representing the topological structure of the main-distribution network through the adjacency matrix of the graph, constructing the adjacency matrix and edge feature matrix of the graph, and defining the weights of the edges to obtain the edge feature matrix E and the feature matrix X as prior knowledge; S3. Constructing a Mamba network model based on the prior knowledge and measurement data, where the Mamba network model includes a time series feature extraction network LSTM, a graph convolutional neural network, a prior knowledge and feature fusion module, a Softmax layer and a Mamba network; S4. Using the training set and adopting a joint loss function to train and test the Mamba network model. The present invention can optimize information propagation, has a low computational complexity, effectively improves the unbalanced data processing ability in the main-distribution collaborative power grid, and improves the model training accuracy and effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network fault detection, and specifically relates to a main-distribution collaborative fault detection method, device and storage medium based on the Mamba network. Background Art

[0002] With the continuous growth of power demand and the increasing complexity of the power system structure, the stable operation and fault diagnosis of the power grid have become increasingly important. The main-distribution collaborative power grid (abbreviated as the main-distribution collaborative system), as the core part of the modern power system, is mainly composed of the main power grid (high-voltage transmission grid) and the distribution network (low-voltage distribution network). The main task of the main-distribution collaborative system is to ensure the reliable transmission, dispatching and distribution of electric power, and to achieve rapid diagnosis and repair in case of faults.

[0003] In the traditional power grid management system, the existing fault diagnosis methods mainly rely on sensors to collect power grid data, and then process and analyze them through traditional algorithms. For example, the gradient boosting decision tree (GBDT) algorithm is used for distribution network fault detection, but it may have certain limitations when dealing with complex data. In addition, deep learning methods, such as models based on convolutional neural networks (CNNs) and recurrent neural networks (RNNs), are also often applied to power grid fault diagnosis, but they often have some deficiencies when dealing with time series features and graph structure data.

[0004] Therefore, the prior art also uses a fault diagnosis method based on graph neural networks (GNNs), which can effectively process the topological structure information of the power grid, capture the complex relationships between nodes, and thus improve the accuracy and efficiency of fault diagnosis. For example, a power grid fault location method based on a deep graph convolutional neural network uses graph convolution operations to process the topological information of the power grid and realizes the accurate location of the fault position.

[0005] In addition, a hybrid model that combines a time series feature extraction network (such as a long short-term memory network, LSTM) and a graph convolutional neural network can also be used to simultaneously process the time series features and topological structure information of power grid data. This method can capture the operating state of the power grid more comprehensively and improve the accuracy of fault diagnosis. For example, the existing patent "Cloud-edge collaborative fault diagnosis method" with the patent publication number CN119312141A uses a federated learning framework to implement model iteration, but there are problems such as high occupancy rate of edge-side computing resources (3.2GB of memory is consumed per station per day on average) and model update delay (the average lag time from the occurrence of the fault is 8 minutes), which are difficult to meet the relay protection requirements of millisecond-level response; another example is the existing patent "High-resistance fault detection device" with the patent publication number CN118937891A, which improves the detection efficiency through the fundamental wave phasor accumulation algorithm, but its computational complexity is still quadratic with the number of distribution network nodes, and the single detection time exceeds 2 seconds in an urban distribution network with more than 5000 nodes.

[0006] Alternatively, the prior art also combines spatio-temporal joint modeling techniques. For example, the existing patent "Irregular Graph Structure Prediction Model" with the patent publication number CN119044671A constructs a two-dimensional feature space through time-series data and topological structures. However, the model training depends on complete device historical data and has insufficient adaptability to newly connected distributed power sources. In a distribution network with a photovoltaic penetration rate of more than 30%, the prediction error exceeds 18%. Another example is the existing patent "Mobile Robot Detection System" with the patent publication number CN118818218A, which can improve the multi-point fault response speed. However, its image recognition algorithm has a false detection rate as high as 42% in rainy and foggy weather, and the single-machine inspection coverage radius is less than 2 kilometers.

[0007] Therefore, this application specifically proposes a main-distribution collaborative fault detection method based on the Mamba network to solve the above technical problems. Summary of the Invention

[0008] The main object of the present invention is to provide a main-distribution collaborative fault detection method based on the Mamba network to solve the technical problems proposed in the background art.

[0009] The present invention adopts the following technical solutions to solve the above technical problems:

[0010] A main-distribution collaborative fault detection method based on the Mamba network, comprising:

[0011] S1. Measure and collect main-distribution network fault data and construct a training set ;

[0012] S2. Represent the topological structure of the main-distribution network through the adjacency matrix of the graph, construct the adjacency matrix and edge feature matrix of the graph, and define the weights of the edges to obtain the edge feature matrix E and the feature matrix X as prior knowledge;

[0013] S3. Based on the prior knowledge and measurement data, construct a Mamba network model for main-distribution collaborative fault detection. The Mamba network model includes a time-series feature extraction network LSTM, a graph convolutional neural network, a prior knowledge and feature fusion module, a Softmax layer, and a Mamba network;

[0014] S4. Use the training set and adopt a joint loss function to train and test the Mamba network model.

[0015] Preferably, the construction process of the training set in the step S1 includes:

[0016] S11. Collect the three-phase voltage, three-phase current, and phase angle data of the faulty equipment in the main-distribution network and construct a fault data classification set, denoted as ;

[0017] represents a three-phase voltage data set, and , represents the three-phase voltage data of the th fault data;

[0018] represents a three-phase current data set, and , represents the three-phase current data set of the th fault data;

[0019] Among them, , represents the total number of faults;

[0020] S12. Construct a classification set of fault data The label information set is denoted as , represents the label value of the th fault data, and , is the number of fault types;

[0021] S13. Randomly shuffle the fault data set with labels as the training set .

[0022] Preferably, the specific operation process of step S2 includes:

[0023] S21. Construct the adjacency matrix A and edge feature matrix E of the graph. The adjacency matrix A is used to describe the connection relationship between the main and distribution network nodes. The phase angle and power of the main and distribution network are used as the edge features of the graph. Among them:

[0024] The adjacency matrix A is a matrix of , is the number of nodes, and the element in the matrix represents whether node and node are connected and the connection strength between them. There is:

[0025]

[0026] The edge feature matrix E is a matrix with a dimension of , is the number of edges, is the number of features of each edge. The element in the matrix represents the edge between node and node . The feature of edge is represented as: , where is the phase angle of this side , is the power of this side ;

[0027] S22. Collect the actual phase angle and power data of the main and distribution network fault equipment, and construct a fault data classification set, denoted as , represents the three-phase phase angle data set, and , represents the th three-phase phase angle data of the th fault data, represents the three-phase power data set, and , represents the three-phase power data of the

[0028] S23. Construct the feature matrix of each node in the adjacency matrix A, which is used to represent the electrical parameters of the node. The node feature matrix has a dimension of , where is the number of nodes, is the feature dimension of each node.

[0029] Preferably, the time series feature extraction network LSTM in the S3 step includes a forgetting gate, an input gate, a memory update unit, and an output gate, where:

[0030] The forgetting gate is used to input the data at the th sampling moment of the th fault data , and then extract the fault characteristic information of the th fault data at the th time step in the time series feature extraction unit ;

[0031] The input gate is used to calculate and obtain the input fault information of the th fault data at the th time step in the memory unit, the fault modulation information , and the fault information to be updated;

[0032] The memory update unit is used to calculate and obtain the memory information of the th fault data at the th time step in the memory unit ,

[0033] The specific calculation formula is as follows:

[0034]

[0035] Among them, represents the Hadamard product;

[0036] The output gate is used to calculate and obtain the th fault data at the time step of the memory unit , and combines the current memory unit state to obtain the final hidden state output.

[0037] Preferably, in the S3 step, the graph convolutional neural network is used to aggregate the features of a node and its neighbor nodes, and at the same time, the edge features participate in the propagation of node information through feature weighting, and finally the softmax function is used to convert the feature output of each node into a class label, where:

[0038] The specific expression formula for aggregating features is:

[0039]

[0040] Among them, is the node feature matrix of the th layer, and there is an initial node feature ; is the learnable weight matrix of the th layer; is the adjacency matrix after normalization, and is the degree matrix of the node, which is used to represent the connection strength of the node; sigmoid is a non-linear activation function;

[0041] The specific expression formula for feature weighting is:

[0042]

[0043] Among them, represents the feature of the node in the th layer, represents the feature of the node in the th layer, is an element in the matrix , represents the set of neighbor nodes of the node , is the node and the node between the power features, is the node The phase angle feature between and the node is the weight matrix of the current graph convolutional layer. In the graph convolution of each layer, the feature vector of the node is continuously updated through the graph convolution operation. is a set of trainable hyperparameters.

[0044] Preferably, in the step S3, the prior knowledge and feature fusion module is used to dynamically fuse the features extracted by LSTM and GCN through the Gaussian kernel function and correlation calculation, and retain the specified prior knowledge. The specific implementation steps are as follows:

[0045] L1. Perform dimensional transformation on the temporal feature extracted by LSTM at the moment, so that it is feature-aligned with the prior knowledge feature extracted by the graph convolutional network at the moment. Then, calculate the similarity between the dimension-transformed temporal feature and the prior knowledge feature through a learnable Gaussian kernel function;

[0046] L2. After calculating the similarity between the features, set a dynamic threshold to screen the prior knowledge features, and there is: ;

[0047] Among them, represents the similarity between the temporal feature extracted by LSTM and the prior knowledge feature extracted by the graph convolutional network (GCN).

[0048] L3. According to the retained prior knowledge features, use the splicing method to fuse the temporal feature and the prior knowledge feature . Concatenate the two feature vectors along the feature dimension to obtain a higher-dimensional fused feature vector . Then, input the fused feature vector into the Softmax layer for output and final prediction, and there is:

[0049]

[0050]

[0051] Among them, is the weight matrix of the Softmax layer;

[0052] L4. Through the Mamba network structure, input the input at the current moment and the state Perform selective update of the feature dynamic adjustment state and obtain the updated state , there are:

[0053]

[0054] Among them, is the selective update function, is the parameter set of the Mamba network model.

[0055] Preferably, the joint loss function in the step S4 has the following specific expression formula:

[0056]

[0057] Among them, is the cross-entropy loss function, is the Focal loss function, is defined as the improved NENum function of the sample and the category , there are:

[0058] The cross-entropy loss function has the following calculation formula:

[0059]

[0060] Among them, is the number of categories, is the th sample's true label on the category , is the probability that the model predicts the th sample belongs to the category ;

[0061] The Focal loss function has the following calculation formula:

[0062]

[0063] Among them, is the weight of the category , is the exponent of the weighting factor, used to control the weight of difficult-to-classify samples.

[0064] Preferably, the calculation and acquisition process of the improved NENum function includes:

[0065] L1. Calculate and obtain the scatter degree of the category , and the calculation formula is:

[0066]

[0067] Among them, is the covariance matrix of the class and is used to represent the distribution of the sample data of the class ; is the direction vector in the feature space and is used to indicate the classification boundary direction; is the transpose of the vector matrix of ;

[0068] L2. Combining the temporal and regional differences of power grid data, the weights for obtaining samples are improved, and there is a definition:

[0069]

[0070] Among them, is the stationarity of the sample in the time series, is the regional difference factor of the class and is used to reflect the data differences in different regions of the power grid; , , are hyperparameters and are used to control the influences of the dispersion, stationarity, and regional factors respectively; is the weighted function of the sample and the class ; is the number of samples of the class and is used to adjust the influence of the stationarity hyperparameter on the final weight to compensate for the potential bias caused by the difference in the number of samples;

[0071] L3. According to the prediction probability , the weighted function for obtaining is adjusted, and there is:

[0072]

[0073] Among them, is the hyperparameter for adjusting the weighted sensitivity, is the preset threshold value and is used to determine whether to regard the sample as noise, is the prediction probability, is the sample 's prior probability;

[0074] L4. The improved NENum function is obtained, and there is:

[0075]

[0076] Among them, is the hyperparameter of the stationarity loss.

[0077] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.

[0078] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which, when executed by the processor, causes the processor to execute the steps of the above method.

[0079] As can be seen from the above technical solutions, the present invention provides a main and distribution collaborative fault detection method based on the Mamba network. Compared with the prior art, the present invention has the following advantages:

[0080] 1. By setting a prior knowledge and feature fusion module, the present invention calculates the similarity between the time series features and the prior knowledge features using the Gaussian kernel function and performs dimensional alignment adjustment, which can conveniently determine the threshold to screen and retain the prior knowledge after correlation degree calculation, avoid the interference of noise, and make the prior knowledge in the fusion process targeted and accurate.

[0081] 2. Compared with the traditional recurrent neural network and self-attention model, by adopting the Mamba network to construct the network model, the present invention dynamically adjusts the state of the input features for selective update, which can optimize information propagation, does not need to rely on the complete device historical data, and has a low computational complexity, which can greatly improve the performance and computational efficiency of the model.

[0082] 3. The present invention can efficiently process time series features and graph structures in the process of distribution network fault diagnosis, and is trained by adopting a joint loss function in the process of model training and testing. Based on the traditional NENum loss function, considering the characteristics of the main and distribution collaborative power grid, the consideration of the stationarity of the time series is added, which effectively improves the ability to process unbalanced data in the main and distribution collaborative power grid and improves the training accuracy and effect of the model.

[0083] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The accompanying drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0085] Figure 1Schematic diagram of the overall system process of the present invention;

[0086] Figure 2 Schematic diagram of the LSTM structure of the present invention;

[0087] Figure 3 Schematic diagram of the data processing process of the prior knowledge and feature fusion module of the present invention. Specific implementation manners

[0088] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0089] In the embodiment, please refer in detail to Figures 1 to 3 .

[0090] As Figure 1 shown. The main and distribution coordinated fault detection method based on the Mamba network proposed in the embodiment of the present invention is carried out according to the following steps:

[0091] Step 1. Construct the training set of the network ;

[0092] Step 1.1. Collect the three-phase voltage, three-phase current and phase angle data of the faulty equipment, and construct a fault data classification set, denoted as , where:

[0093] represents the three-phase voltage data set, , represents the th three-phase voltage data of the kth fault data, and , represents the a-phase voltage data of the kth fault data at the tth sampling moment, represents the b-phase voltage data of the kth fault data at the tth sampling moment, represents the c-phase voltage data of the kth fault data at the tth sampling moment;

[0094] represents the three-phase current data set, , represents the th three-phase current data set of the kth fault data, and , represents the a-phase current data of the kth fault data at the tth sampling moment, represents the b-phase current data at the t-th sampling moment of the k-th fault data, and represents the c-phase current data at the t-th sampling moment of the k-th fault data; 1 ≤ k ≤ K, where K represents the total number of faults; 1 ≤ t ≤ T, where T represents the total sampling time;

[0095] Step 1.2: Construct a set for classifying fault data and the set of label information is denoted as , represents the label value of the -th fault data, and , is the number of fault types;

[0096] Step 1.3: Randomly shuffle the fault data set with labels and use it as the training set ; and , represents the k-th fault data, and , represents the three-phase voltage and current data at the t-th sampling moment of the k-th fault data;

[0097] Step 2: Construct the prior knowledge of the network;

[0098] Step 2.1: Construct the adjacency matrix and edge feature matrix of the graph. The topological structure of the main and distribution networks is represented by the adjacency matrix of the graph. The adjacency matrix A is used to describe the connection relationship between nodes (generators, loads, substations). The features of the edges (phase angles and powers) will be used as the edge features of the graph, denoted by the matrix E. The adjacency matrix A is an N×N matrix, where N is the number of nodes, representing the connection relationship between nodes. The element represents whether nodes and are connected and the strength of their connection. If there is an edge between nodes and , , otherwise, .

[0099] Step 2.2: Define the weights of the edges,

[0100] Collect the phase angle and power data of the faulty equipment, and construct a set for classifying fault data, denoted as , where:

[0101] represents the three-phase phase angle data set, and , represents the three-phase phase angle data of the -th fault data, and , represents the phase angle data of phase A at the t-th sampling moment of the k-th fault data, represents the phase angle data of phase B at the t-th sampling moment of the k-th fault data, represents the phase angle data of phase C at the t-th sampling moment of the k-th fault data;

[0102] represents the three-phase power data set, and , represents the -th three-phase power data of the fault data, and , represents the power data of phase A at the t-th sampling moment of the k-th fault data, represents the power data of phase B at the t-th sampling moment of the k-th fault data, represents the power data of phase C at the t-th sampling moment of the k-th fault data;

[0103] Step 2.3, construct the edge feature matrix E and the feature matrix X

[0104] Store the features (phase angle and power) of the edges in the edge feature matrix E. The dimension of the edge feature matrix E is M×G, where M is the number of edges and G is the number of features per edge. Each row represents the features of an edge, including the elements in the phase angle and power matrices of the edge represents the node and the node of the edge. At this time, the features of the edge are represented as: , where is the phase angle of this edge , is the power of this edge .

[0105] Finally, the edge feature matrix E will include the phase angle and power features of all edges. The feature matrix of each node is used to represent the electrical parameters of the node. The dimension of the node feature matrix is N×F, where N is the number of nodes and F is the feature dimension of each node.

[0106] Step 3, construct a Mamba network driven by prior knowledge and measurement data, including: a time series feature extraction network LSTM, a graph convolutional neural network, a prior knowledge and feature fusion module, a Mamba network, and a Softmax layer;

[0107] Step 3.1, construct the time series feature extraction network LSTM as shown in Figure 2

[0108] ​The LSTM network consists of four parts, namely, the forget gate, the input gate, the memory update unit, and the output gate. Among them, the forget gate takes the data at the t-th sampling moment of the k-th fault data as the input to the forget gate. The forget gate performs feature extraction operations to obtain the fault feature information of the k-th fault data at the t-th time step in the time series feature extraction unit : :

[0109]

[0110] In the formula, represents the hidden state of the fault information of the k-th fault data at the -th time step in the forget unit; when , let ; is the activation function; and respectively represent the data at the t-th sampling moment of the k-th fault data and the hidden state of the fault information of the k-th fault data at the t-th time step in the forget unit ; is the weight matrix of the forget unit, represents the bias vector of the forget unit of the k-th fault data ; The forget gate uses the following formula to obtain the fault retention information of the k-th fault data at the t-th time step in the memory unit :

[0111]

[0112] In the formula, × represents the outer product operation, is the memory information of the -th fault data at the -th time step in the memory unit;

[0113] Step 3.1.1, Input gate

[0114] The input gate uses the following formula to obtain the input fault information of the k-th fault data at the t-th time step in the memory unit , the fault modulation information and the fault information to be updated ;

[0115]

[0116]

[0117]

[0118] Wherein, is the th fault data fault information to be updated at the t-th time step of the memory unit, represents the Hadamard product (element-wise multiplication);

[0119] and respectively represent the data at the t-th sampling moment of the th fault data and the th fault data fault information hidden state at the time step of the memory unit, represents the input bias vector of the th fault data

[0120] and respectively represent the data at the t-th sampling moment of the th fault data and the th fault data fault information hidden state at the time step of the memory unit, represents the bias vector of the th fault data

[0121] ⊙ represents the element-wise multiplication operation; is the activation function;

[0122] Step 3.1.2. The memory update unit obtains the memory information of the th fault data at the t-th time step of the memory unit by the following formula :

[0123]

[0124] Wherein, is the memory unit state of the

[0125] th fault data at the t-th time step. At time is the memory unit state of the previous time step, is the output of the forget gate, indicating how much of the memory information of the previous time step is forgotten, is the current candidate memory information, representing the modulation information of the current input data;

[0126] Update the memory cell state through the weighted sum of the forget gate and the candidate memory.

[0127] Step 3.1.3, the output gate obtains the k-th fault data using the following formula The signal at the t-th time step of the memory cell :

[0128]

[0129] where, is the output gate value of the k-th fault data at the t-th time step, and is the weight matrix of the output gate, is the input data of the k-th fault data at the t-th time step, is the hidden state (fault information) of the previous time step, is the bias term of the output gate.

[0130] Through the output gate and the current memory cell state, the final hidden state (i.e., the output) is:

[0131]

[0132] where, is the hidden state output of the k-th fault data at the t-th time step, is the current memory cell state, containing the information updated through the input gate and the forget gate, is the activation function, used for non-linear transformation of the memory cell state, is the output information controlled by the output gate.

[0133] Step 3.2, construct a graph convolutional neural network

[0134] Step 3.2.1, construct a graph convolutional layer

[0135] Construct a corresponding graph convolutional neural network for processing according to the topological structure processed above, and transform it into prior knowledge to improve the accuracy of network anomaly detection. The main operation of the graph convolutional neural network is to aggregate the features of a node and the features of its neighbor nodes, which can be expressed by the formula:

[0136]

[0137] where, is the node feature matrix of the layer, and there is an initial node feature ; is the The learnable weight matrix of the layer; is the adjacency matrix The normalized matrix, and is the degree matrix of the nodes, used to represent the connection strength of the nodes; is the non-linear activation function, and here the ReLU non-linear function is used.

[0138] Step 3.2.2, Fusion and Propagation of Edge Features

[0139] During the graph convolution process, in order to update the features of the nodes, not only the node features need to be aggregated, but also the edge features (power and phase angle) need to be processed. The main reason is that these edge features greatly affect the way of information propagation. In graph convolution, edge features usually participate in the propagation of node information in a weighted manner. It can be expressed by the formula:

[0140]

[0141] where, represents the feature of the -th layer node , represents the feature of the -th layer node , is an element in the matrix , represents the set of neighbor nodes of node , is the power feature between node and node , is the phase angle feature between node and node , is the weight matrix of the current graph convolution layer, is a set of trainable hyperparameters.

[0142] In order to capture more detailed global information and local information, here a multi-layer graph convolutional neural network is adopted. In the graph convolution of each layer, the feature vector of the nodes is continuously updated through the graph convolution operation. Specifically, the graph convolution of each layer aggregates the features of the neighbor nodes and uses the aggregated information to update the features of the current node. In this way, as the network deepens, the graph convolution of each layer can not only capture the local information of the neighboring nodes, but also gradually integrate the global information of the distant nodes through the multi-layer structure, so as to obtain a richer node representation. Finally, the softmax function is used to convert the feature output of each node into a class label.

[0143] Step 3.3, Prior Knowledge and Feature Fusion Module

[0144] In traditional neural network methods, the fusion of prior knowledge and measurement data is usually carried out through fixed weight allocation. However, fixed weights often fail to fully exploit the most valuable prior knowledge, resulting in interference in the use of the model. Especially when dealing with large and complex systems, the degree of influence of fault locations on different nodes also varies.

[0145] To overcome this problem, this prior knowledge and feature fusion module proposes a dynamic fusion operation method based on Gaussian function and correlation calculation, referring to Figure 3 , which can more precisely fuse the features extracted by LSTM and GCN and retain the prior knowledge beneficial to the final prediction.

[0146] Step 3.3.1, Learnable correlation calculation

[0147] In this part, the similarity between the temporal features extracted by LSTM and the prior knowledge features extracted by the graph convolutional network (GCN) is calculated by a learnable Gaussian function. Since the dimensions of these two types of features may be different, the temporal features must be dimensionally transformed first to align with the GCN features. The calculation formula is as follows:

[0148]

[0149] where is the temporal feature extracted by the LSTM network at the t-th moment, is the prior knowledge feature extracted by the GCN network at the t-th moment. is the Euclidean distance between the temporal feature and the prior knowledge feature. is a trainable hyperparameter that controls the width of the Gaussian function and determines the sensitivity of the similarity. According to the calculated similarity value, the similarity degree between the temporal feature and the prior knowledge is judged. The higher the similarity, the more similar the two are.

[0150] Step 3.3.2, Retention of prior knowledge

[0151] After calculating the similarity between the features, it is then necessary to decide whether to retain the prior knowledge extracted by GCN according to the similarity. To ensure that only the most relevant prior knowledge is retained, a dynamic threshold is set. When the similarity is greater than this threshold, the corresponding prior knowledge is retained; when the similarity is less than or equal to the threshold, the prior knowledge at that moment is discarded. This retention mechanism makes the prior knowledge in the fusion process targeted and precise. The retention condition formula is:

[0152]

[0153] Among them, is the prior knowledge extracted by the GCN at the t-th moment. If the similarity is greater than the threshold , the current prior knowledge is retained. If the similarity is less than or equal to the threshold , the prior knowledge is set to 0, that is, the prior knowledge at this moment is discarded. Here, the threshold is set to 0.6, and this threshold can effectively retain the prior knowledge strongly correlated with the temporal features, thus avoiding the interference of noise.

[0154] Step 3.3.3. Measurement feature and prior knowledge fusion

[0155] After completing the retention of prior knowledge, the next step is the fusion of features. According to whether the prior knowledge is retained, the temporal features extracted by the LSTM and the prior knowledge features extracted by the GCN are fused. If the prior knowledge at the current moment is retained, the splicing method is used for fusion. The two feature vectors are spliced along the feature dimension to obtain a higher-dimensional feature vector:

[0156]

[0157] Among them, is the temporal feature extracted by the LSTM network. is the prior knowledge feature extracted by the GCN network. The fusion method used here is splicing. The LSTM feature and the GCN feature are spliced along the feature dimension to obtain a higher-dimensional feature vector. The key to this step is that due to the screening and retention of prior knowledge after the correlation degree calculation, the spliced fusion feature vector will contain more valuable temporal information and prior knowledge information, thus improving the accuracy of the model. This fusion feature vector will be further input to the output layer for final prediction, which is expressed by the formula as follows:

[0158]

[0159] Among them, is the weight matrix of the output layer, which determines how the fusion feature is mapped to the final prediction result. Through this fusion method, the model can effectively utilize the complementarity of temporal features and prior knowledge, thus improving the accuracy of prediction.

[0160] Step 3.4. Construct the Mamba network

[0161] Mamba is a linear time series modeling method based on a selective state space model. Compared with traditional recurrent neural networks and self-attention models, Mamba can efficiently process long sequences with linear time complexity and improve the model's performance and computational efficiency by introducing a selective mechanism to optimize information propagation. In Mamba, the core idea of the selective state space model is to dynamically adjust the state update method according to the input features. Given the input at the current time step and the state at the previous time step, Mamba decides which information needs to be passed and which information can be ignored by selectively updating the state. Specifically, the formula for state update is:

[0162]

[0163] where is a selective update function and are the model's parameters. Through this selective update mechanism, the model can efficiently capture important features in the input sequence and suppress irrelevant information, thereby improving the ability to model long-term dependencies. Mamba optimizes memory usage and computational efficiency when running on modern hardware (such as GPUs) by introducing hardware-friendly algorithms. By avoiding complex attention mechanisms, Mamba can provide higher throughput during training and inference and reduce computational resource consumption. Mamba has made significant progress in multiple fields, especially in tasks such as language modeling, audio processing, and genomics, showing superior performance. In the language modeling task, Mamba not only outperforms Transformer models of comparable scale but also shows higher efficiency and lower computational resource consumption during inference. Mamba also provides an efficient and powerful way to process long sequence data and shows leading performance in multiple tasks. In the main and distribution coordinated power grid, there is a large amount of data, and the Mamba network can perform fine-tuning computationally through hardware optimization algorithms to enable the network to operate efficiently. Therefore, Mamba can detect faults in the operating state of the main and distribution coordination in real-time or near real-time, reducing latency.

[0164] Step 4: Model training and testing

[0165] In the main and distribution coordinated power grid, data is affected by various factors, such as load fluctuations, voltage changes, and equipment failures. To effectively handle the imbalance problem in power grid data, the following aspects are considered: First, power grid data is time-series, with volatility and trends, so the model must pay attention to its time dependence; Second, there are differences in data collection frequency and quality in different regions, resulting in the influence of regional differences; Finally, the failure modes of power grid equipment occur at different times and locations, leading to class imbalance. Therefore, when training the model, a joint loss function is used for training, which is composed of the cross-entropy loss function, the Focal loss function, and the improved NENum function. Among them, the cross-entropy loss function is the most commonly used algorithm in binary classification and multi-classification. This method mainly measures the performance of the model by calculating the difference between the predicted probability and the true label, and can be expressed by the formula:

[0166]

[0167] where: is the number of classes, is the th sample's true label in class , is the probability that the model predicts the th sample belongs to class .

[0168] Focal Loss is a commonly used loss function when dealing with class imbalance problems, especially in object detection. It assigns a larger weight to difficult-to-classify samples, thus avoiding the model from over-focusing on easy-to-classify samples. The Focal loss introduces a weighting factor on the basis of the cross-entropy loss to reduce the attention to easy-to-classify samples. It can be expressed by the formula:

[0169]

[0170] where, is the probability that the model predicts the th sample belongs to class , is the th sample's true label in class , is the weight of class , usually used to alleviate the class imbalance problem, is the exponent of the weighting factor, used to control the weight of difficult-to-classify samples.

[0171] Regarding the class imbalance problem in the power grid, the data scatter degree is used to measure the data scatter degree of each class. The scatter degree reflects the distribution of data in the feature space and can be measured by the covariance matrix of class data. For class data, its scatter degree is defined, and it can be expressed by the formula as:

[0172]

[0173] where, is the covariance matrix of class and is used to represent the distribution of the sample data of class , is the direction vector in the feature space and is used to indicate the classification boundary direction, is the transpose of the vector matrix of

[0174] To process the imbalanced data, each sample needs to be weighted. Combining the temporal and regional differences of power grid data, an improved sample weight calculation formula is proposed. Specifically, the weight of sample is defined as:

[0175]

[0176] where, is the scatter degree of class , is the stationarity of sample in the time series and is used to measure whether the sample has abnormal fluctuations, is the regional difference factor of class and is used to reflect the data differences in different regions of the power grid, is the weighted function adjusted according to the prediction probability and is defined as:

[0177]

[0178] where, is the hyperparameter for adjusting the weighted sensitivity, is the threshold value and is used to determine whether to regard the sample as noise, , , are hyperparameters that respectively control the influences of the scatter degree, stationarity, and regional factor.

[0179] According to the characteristics of the main and distribution coordinated power grid, the consideration of the time series stationarity is added. The loss function is:

[0180]

[0181] Among them, is the weighted weight of the sample , is the hyperparameter of the stationarity loss, used to control the influence of the stationarity loss, is the sample of the stationarity loss in the time series, used to measure whether the sample is an abnormal fluctuation.

[0182] Finally, the combined loss of the above three loss functions is used to train the model, which is expressed by the formula:

[0183]

[0184] By combining factors such as temporality, regional differences, and dispersion, this function can effectively improve the ability to process unbalanced data in the main and distribution collaborative power grid and perform outstandingly in the fault detection task of the main and distribution collaborative power grid.

[0185] In a specific embodiment, data collection and experiments are carried out on different data. In the main and distribution collaborative power grid, for the five categories of the power grid fault classification task, including single-phase grounding, three-phase grounding, phase-to-phase short circuit, high-resistance grounding, and normal data, experiments are carried out through different power grid fault classification tasks, and the following table shows the data results:

[0186]

[0187] Among them, the classification accuracy represents the correct recognition rate of the model for each category. The classification accuracy of "normal data" is relatively high, reaching 99%, while some fault categories such as "phase-to-phase short circuit" are more difficult, and the classification accuracy is 95%.

[0188] In order to comprehensively evaluate the performance of the model for different categories, the comparative experiment also introduces the F1 score, which is the harmonic mean of precision and recall. For example, the F1 score of the "normal data" category is 0.97, showing its good comprehensive performance, while the F1 score of the "phase-to-phase short circuit" category is 0.93, indicating that the model is also relatively effective in identifying fault categories.

[0189] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.

[0190] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0191] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, it causes the computer to execute any one of the main distribution collaborative fault detection methods based on the Mamba network in the above embodiments.

[0192] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. The explanations, examples and beneficial effects of the relevant content can refer to the corresponding parts in the above method.

[0193] The embodiments of the present application further provide an electronic device, including a processor, a communication interface, a memory and a communication bus. Among them, the processor, the communication interface and the memory complete communication with each other through the communication bus.

[0194] The memory is used to store a computer program.

[0195] The processor is used to implement the above-mentioned main distribution collaborative fault detection method based on the Mamba network when executing the program stored in the memory.

[0196] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0197] The communication interface is used for communication between the above electronic device and other devices.

[0198] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0199] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0200] It should also be noted that the electronic device further includes a terminal device, which can also be referred to as a terminal, user equipment, mobile station, mobile terminal, etc. The terminal device can be a mobile phone, smart TV, wearable device, tablet computer, computer with wireless transceiver function, virtual reality terminal device, augmented reality terminal device, wireless terminal in industrial control, wireless terminal in driverless, wireless terminal in remote surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0201] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive), etc.

[0202] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

[0203] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0204] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes Scenario A, or Scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality of" means more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

Claims

1. A master-slave collaborative fault detection method based on the Mamba network, characterized in that Including: S1. Measure and collect the main and distribution network fault data and construct a training set ; S2. Represent the topological structure of the main and distribution networks through the adjacency matrix of the graph, construct the adjacency matrix and edge feature matrix of the graph, and define the weights of the edges to obtain the edge feature matrix E and the feature matrix X as prior knowledge; S3. Build a Mamba network model for main-distribution collaborative fault detection based on prior knowledge and measurement data. The Mamba network model includes a time series feature extraction network LSTM, a graph convolutional neural network, a prior knowledge and feature fusion module, a Softmax layer, and a Mamba network; S4. Use the training set And use the combined loss function to train and test the Mamba network model. The combined loss function The specific expression formula is as follows: Among them, is the cross-entropy loss function, is the Focal loss function, is defined as the sample and the category of the improved NENum function; The calculation acquisition process of the improved NENum function includes: L1. Calculate and obtain the category of the scatter degree, and the calculation formula is as follows: Among them, is the covariance matrix of the class and is used to represent the distribution of the sample data of the class . is the direction vector in the feature space and is used to indicate the classification boundary direction. is the transpose of the vector matrix of L2. Improve the sample acquisition by combining the temporal and regional differences in grid data of weights There is a definition as follows: Among them, is the stationarity of the sample in the time series, is the regional difference factor of the category used to reflect the data differences in different regions of the power grid, and and are hyperparameters used to control the influence of the dispersion, stationarity, and regional factor respectively, is the sample and the category weighting function, is the number of samples of the category used to adjust the influence of the stationarity hyperparameter on the final weight to compensate for the potential bias caused by the difference in the number of samples; L3. According to the predicted probability Adjust the acquisition weighting function , we have: Among them, is a hyperparameter for adjusting the weighted sensitivity, is a preset threshold value used to determine whether a sample is regarded as noise, is the predicted probability, is the sample 's average predicted probability; L4. Obtain the improved NENum function, and there is: Among them, is the hyperparameter of the stationarity loss.

2. The main and distribution coordination fault detection method based on the Mamba network according to claim 1, wherein The training set in step S1 is constructed as follows: S11. Collect the three-phase voltage, three-phase current, and phase angle data of the main and distribution network fault devices, and construct a fault data classification set, denoted as ; represents a three-phase voltage data set, and , represents the three-phase voltage data of the th fault data; represents a three-phase current data set, and , represents the three-phase current data set of the th fault data; Among them, , represents the total number of faults; S12. Construct a classified set of fault data The set of label information is denoted as , indicating the label value of the th fault data, and , being the number of fault types; S13. Take the fault data set with labels Randomly shuffle the order and use it as the training set .

3. The main and distribution collaborative fault detection method based on the Mamba network according to claim 1, wherein The specific operation process of the S2 step includes: S21. Construct the adjacency matrix A and edge feature matrix E of the graph. The adjacency matrix A is used to describe the connection relationship between the nodes of the main and distribution networks. The phase angle of the main and distribution networks and the power of the main and distribution networks are used as the edge features of the graph, where: The adjacency matrix A is a matrix, where represents the number of nodes, and the elements in the matrix indicate whether node and node are connected and the connection strength between them, and there is: The edge feature matrix E is a matrix with a dimension of , being the number of edges, being the number of features of each edge. The element in the matrix represents the edge between node and node . The feature of edge is represented as: , where is the phase angle of this edge , and is the power of this edge . S22. Collect the actual phase angles and power data of the main and distribution network fault devices, and construct a fault data classification set, denoted as , represents the three-phase phase angle data set, and , represents the three-phase phase angle data of the th fault data,[ represents the three-phase power data set, and , represents the three-phase power data of the th fault data; S23. Construct the feature matrix of each node in the adjacency matrix A , which is used to represent the electrical parameters of the node. The node feature matrix has a dimension of , where is the number of nodes, is the feature dimension of each node.

4. The main and distribution collaborative fault detection method based on the Mamba network according to claim 1, wherein In the S3 step, the time series feature extraction network LSTM includes a forget gate, an input gate, a memory update unit, and an output gate, where: Forgotten gate, used for inputting the fault data of Article After the data at the sampling moment is extracted, the fault data of Article fault characteristic information at the time step of the timing feature extraction unit ; Input gate, used to calculate and obtain the th fault data at the time step of the memory unit for input fault information , fault modulation information and fault information to be updated ; A memory update unit for calculating and obtaining the th fault data at the time step memory information of the memory unit , The specific calculation formula is: Among them, represents the Hadamard product; Output gate, used to calculate and obtain the th fault data at the time step of the memory unit signal, and combine the current state of the memory unit to obtain the final hidden state output.

5. The main and distribution coordination fault detection method based on the Mamba network according to claim 1, characterized in that, In the S3 step, the graph convolutional neural network is used to aggregate the features of a node and its neighbor nodes, and at the same time, the edge features participate in the propagation of node information through feature weighting, and finally, the softmax function is used to convert the feature output of each node into a class label, where: The specific expression formula for aggregating features is: Among them, is the node feature matrix of the -th layer, and there is an initial node feature ; is the learnable weight matrix of the -th layer; is the adjacency matrix after normalization, and is the degree matrix of nodes, which is used to represent the connection strength of nodes; sigmoid is a non-linear activation function; The specific expression formula for feature weighting is: Among them, represents the characteristics of the layer nodes, represents the characteristics of the layer nodes, is an element in the matrix and represents the set of neighbor nodes of node is the power characteristic between node and node ; is the phase angle characteristic between node and node ; is the weight matrix of the current graph convolutional layer, and in each layer of graph convolution, the feature vector of the node is continuously updated through the graph convolution operation. is a set of trainable hyperparameters.

6. The master-slave collaborative fault detection method based on the Mamba network according to claim 1, wherein In the S3 step, the prior knowledge and feature fusion module is used to perform dynamic fusion of the features extracted by LSTM and GCN through the Gaussian kernel function and correlation calculation, and retain the specified prior knowledge. The specific implementation steps include: L1. At the time, extract the temporal features and perform dimensional transformation to align them with the prior knowledge features extracted by the graph convolutional network at the time. Then, calculate the similarity between the dimensional-transformed temporal features and the prior knowledge features through a learnable Gaussian kernel function; After calculating the similarity between features, set a dynamic threshold , and filter the prior knowledge features, we have: ; Among them, represents the similarity between the temporal features extracted by LSTM and the prior knowledge features extracted by the graph convolutional network (GCN); similarity; L3. According to the retained prior knowledge features, the splicing method is used to fuse the temporal features and the prior knowledge features , the two feature vectors are spliced along the feature dimension to obtain a higher-dimensional fused feature vector , and the fused feature vector is input to the output of the Softmax layer for final prediction, and we have: Among them, is the weight matrix of the Softmax layer; L4. Through the Mamba network structure, input the input at the current moment and the state at the previous moment to perform selective updates of the feature dynamic adjustment state and obtain the updated state , there is: Among them, is the selective update function, is the parameter set of the Mamba network model.

7. A computer-readable storage medium, characterized in that, There is a computer program stored. When the computer program is executed by a processor, the processor executes the steps of the method according to any one of claims 1 to 6.

8. A computer device, characterized in that, Including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Power distribution network fault detection device

    CN118818218A

  • Power distribution network high-resistance fault detection method and device and computer equipment

    CN118937891A

  • Fault detection method and device and electronic equipment

    CN119044671A

  • Cloud-side cooperative fault diagnosis method and system for intelligent substation equipment

    CN119312141A

  • Power distribution network fault line selection method considering topological structure change

    CN119438800A