Power distribution network fault diagnosis method and device based on large model, terminal and medium
By constructing a dynamic graph attention network model and a cross-modal attention mechanism to integrate multimodal data characteristics, the problem of low accuracy in distribution network fault diagnosis is solved, and efficient fault diagnosis is achieved in dynamic topology and complex scenarios.
Patent Information
- Application Number
- CN202510764101.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing distribution network fault diagnosis technology relies on static models and insufficient fusion of multimodal data, resulting in low fault diagnosis accuracy and prone to misjudgment, especially in scenarios such as frequent access to distributed new energy and dynamic reconstruction of network topology.
By constructing a dynamic graph attention network model, the multimodal characteristics of topology, electrical quantity and environmental data are integrated using the cross-modal attention mechanism to reflect the topology dynamic changes of the distribution network in real time, and enhance the robustness of fault diagnosis.
It improves the accuracy of fault diagnosis of distribution networks, reduces the risk of misjudgment caused by topological mismatch and multimodal data distribution differences, and improves the diagnostic reliability in complex scenarios.
Smart Images

Figure CN120337157A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of distribution networks, and in particular, to a distribution network fault diagnosis method, device, terminal, and medium based on a large model. Background Art
[0002] With the large-scale access of high-proportion renewable energy and new types of electrical loads, the operation of modern distribution networks presents strong uncertainty and multi-time-space coupling characteristics, posing higher requirements for the real-time performance and accuracy of fault diagnosis.
[0003] Currently, most of the distribution network fault diagnosis technologies based on large models are based on fixed topologies and artificial rule reasoning. Although this method is effective in simple scenarios, it relies on static models and preset thresholds, and it is difficult to adapt to complex changes such as the frequent access of distributed new energy and the dynamic reconstruction of network topologies. It is easy to cause deviation in fault location due to topological mismatch, resulting in missed and misjudged fault diagnosis and location, and there is a technical problem of low accuracy. Summary of the Invention
[0004] This application provides a distribution network fault diagnosis method, device, terminal, and medium based on a large model, which is used to solve the technical problem of low accuracy of existing distribution network fault diagnosis technologies.
[0005] To solve the above technical problem, the first aspect of this application provides a distribution network fault diagnosis method based on a large model, including:
[0006] Obtain distribution network operation data, where the distribution network operation data includes: topology data, electrical quantity data, and environmental data;
[0007] According to the topology data and electrical quantity data in the distribution network operation data, construct a dynamic graph attention network model, where the dynamic graph attention network model is used to reflect the real-time topological dynamic change relationship of the distribution network;
[0008] Based on a preset multi-modal large model, perform feature fusion processing on different modal distribution network operation data through a cross-modal attention mechanism to obtain a modal fusion feature, and based on the modal fusion feature, obtain a fault diagnosis result of the distribution network.
[0009] Preferably, the constructing a dynamic graph attention network model according to the distribution network operation data includes:
[0010] According to the distribution network operation data, construct a node feature matrix, where the node feature matrix is used to reflect the node features of each distribution network node at different time steps;
[0011] Construct the dynamic adjacency matrix according to the node feature matrix, where the dynamic adjacency matrix is used to reflect the electrical coupling strength between different distribution network nodes at different times;
[0012] Construct a dynamic graph attention network model according to the dynamic adjacency matrix.
[0013] Preferably, it further includes:
[0014] Based on the selected target node, calculate the attention coefficient between the target node and its neighbor nodes at the current time step through a preset spatial attention coefficient calculation formula;
[0015] Through an activation function, perform feature aggregation on the node features of each neighbor node and the attention coefficient to obtain the updated node features of the target node at the next time step;
[0016] Summarize the node features of the target node up to the current time step, and perform time-dependent feature extraction through an LSTM network to obtain the time feature sequence of the target node;
[0017] Obtain the latest node features of the target node according to the weighted sum of the updated node features and the time series features.
[0018] Preferably, the feature fusion processing of the distribution network operation data of different modalities through the cross-modal attention mechanism to obtain the modality fusion features includes:
[0019] According to the distribution network operation data, perform feature encoding processing through multiple preset modality encoders to obtain an electrical quantity modality matrix, a topology modality matrix, and an environment modality matrix respectively;
[0020] According to the electrical quantity modality matrix, the topology modality matrix, and the environment modality matrix, taking the electrical quantity modality matrix as the modality alignment benchmark, obtain QKV features through a preset QKV mapping calculation formula;
[0021] According to the QKV features, combine the attention weight calculation formula to obtain the attention weight, and then based on the sum of the attention weight and the electrical quantity modality matrix, obtain the modality fusion features.
[0022] Preferably, the specific QKV mapping calculation formula is:
[0023]
[0024] In the formula, is the electrical quantity modality matrix, is the combined matrix of the topology modality matrix and the environment modality matrix, , , Weight matrices for Q feature, K feature, and V feature respectively.
[0025] Preferably, it further includes:
[0026] Based on a preset multi-modal data sample of the distribution network, combined with a loss function based on structural causal constraints, the multi-modal large model is trained to determine the convergence degree of the multi-modal large model based on the output of the loss function. When the convergence degree reaches a preset convergence threshold or the number of training iterations reaches a preset number threshold, the multi-modal large model is output.
[0027] Preferably, after obtaining the distribution network operation data, it further includes:
[0028] Preprocess the distribution network operation data, where the preprocessing includes: data cleaning, normalization, and denoising.
[0029] Meanwhile, a second aspect of the present application provides a distribution network fault diagnosis device based on a large model, including:
[0030] An operation data acquisition unit for acquiring distribution network operation data, where the distribution network operation data includes: topological data, electrical quantity data, and environmental data;
[0031] A dynamic graph network construction unit for constructing a dynamic graph attention network model according to the topological data and electrical quantity data in the distribution network operation data, where the dynamic graph attention network model is used to reflect the real-time topological dynamic change relationship of the distribution network;
[0032] A distribution network fault diagnosis unit for performing feature fusion processing on different modalities of distribution network operation data through a cross-modal attention mechanism based on a preset multi-modal large model to obtain a modality fusion feature, and obtaining a fault diagnosis result of the distribution network based on the modality fusion feature.
[0033] A third aspect of the present application provides a distribution network fault diagnosis terminal based on a large model, including: a memory and a processor;
[0034] The memory is used to store program code, and the program code is used to implement a distribution network fault diagnosis method provided in the first aspect of the present application;
[0035] The processor is used to read and execute the program code.
[0036] A fourth aspect of the present application provides a computer-readable storage medium, in which program code is stored, and the program code is used to be read and executed by a processor to implement a distribution network fault diagnosis method provided in the first aspect of the present application.
[0037] As can be seen from the above technical solutions, the present application has the following advantages:
[0038] The solution provided by the present application is based on the operating data such as the topology, electrical quantities, and operating environment of the distribution network obtained. By constructing a dynamic graph attention network model, it reflects the real-time topological dynamic changes of the distribution network, and uses a cross-modal attention mechanism to fuse multi-modal data features, eliminates the interference of the distribution differences of multi-modal data on feature extraction, and enhances the robustness of the model for fault diagnosis operations under complex fault modes. Thus, it solves the problems of false alarms and missed detections in fault diagnosis caused by insufficient static models and multi-modal data fusion in traditional methods, and improves the accuracy of distribution network fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of an embodiment of a distribution network fault diagnosis method based on a large model provided by the present application.
[0041] Figure 2 It is a schematic structural diagram of an embodiment of a distribution network fault diagnosis device based on a large model provided by the present application.
[0042] Figure 3 It is a schematic structural diagram of an embodiment of a distribution network fault diagnosis terminal based on a large model provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The embodiments of the present application provide a distribution network fault diagnosis method, device, terminal, and medium based on a large model, which are used to solve the technical problem of low accuracy in existing distribution network fault diagnosis technologies.
[0044] In order to make the invention purpose, features, and advantages of the present application more obvious and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the embodiments described below are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0045] First, a detailed description of an embodiment of a distribution network fault diagnosis method based on a large model provided by the present application is as follows:
[0046] Please refer to Figure 1 , a distribution network fault diagnosis method based on a large model provided in this embodiment includes:
[0047] Step 101, obtain distribution network operation data;
[0048] Among them, the distribution network operation data includes: topological data, electrical quantity data, and environmental data;
[0049] Step 102, construct a dynamic graph attention network model according to the topological data and electrical quantity data in the distribution network operation data;
[0050] Among them, the dynamic graph attention network model is used to reflect the real-time topological dynamic change relationship of the distribution network;
[0051] Step 103, based on a preset multimodal large model, perform feature fusion processing on distribution network operation data of different modalities through a cross-modal attention mechanism to obtain a modality fusion feature, and based on the modality fusion feature, obtain a fault diagnosis result of the distribution network.
[0052] It should be noted that the solution of this application proposes to obtain distribution network operation data including topological data, electrical quantity data, and environmental data, construct a dynamic graph attention network model to reflect real-time topological dynamic changes, use a multimodal large model to fuse different modality data through a cross-modal attention mechanism to obtain a modality fusion feature, and perform fault diagnosis based on this feature. Among them, the dynamic graph attention network model refers to a graph structure that describes the electrical coupling strength between nodes changing over time through a dynamic adjacency matrix, and specifically can be implemented by adjusting the connection weights between nodes with real-time updated attention coefficients, and is used to capture the impact of topological dynamic reconstruction on the fault propagation path. The cross-modal attention mechanism refers to determining the association weight by mapping and calculating the query, key-value features of different modality data, and specifically can use a multi-head attention layer to align the potential associations of the electrical quantity, topology, and environment modalities, and is used to eliminate the distribution differences of multi-source data. The modality fusion feature refers to a multimodal joint representation after weighted aggregation, and specifically can use a residual connection to retain the original feature information, and is used to enhance the representation ability of fault features.
[0053] Specifically, the distribution network operation data is input into the dynamic graph attention network after preprocessing. This network constructs a dynamic adjacency matrix according to features such as node voltage and current, aggregates neighbor node information through a spatial attention coefficient, and extracts the evolution law of node states by combining time series analysis. The multimodal large model extracts deep features of each modality through an independent encoder, generates cross-modal attention weights based on the electrical quantity modality, and performs weighted fusion on the topology and environment features. The fused features are input into a classifier to output the fault type and location information.
[0054] Compared with the prior art, this solution can adjust the node connection relationship in real time through a dynamic graph structure to adapt to the dynamic scenario of distribution network reconfiguration. It can accurately capture the changes in the node coupling relationship under dynamic topology, eliminate the interference of multi-modal data distribution differences on feature extraction, enhance the diagnostic robustness under complex fault modes, effectively reduce the misjudgment risk caused by single data dimension and topological mismatch. At the same time, semantic associations between data are established through a cross-modal attention mechanism to improve the effectiveness of feature fusion.
[0055] Based on the above basic embodiment, further, for the dynamic graph attention network model in step 102, the specific construction steps are as follows:
[0056] Construct a node feature matrix according to the distribution network operation data;
[0057] Construct a dynamic adjacency matrix according to the node feature matrix;
[0058] Construct a dynamic graph attention network model according to the dynamic adjacency matrix.
[0059] It should be noted that in this embodiment, it is further proposed to construct a node feature matrix according to the distribution network operation data, construct a dynamic adjacency matrix according to the node feature matrix, and construct a dynamic graph attention network model according to the dynamic adjacency matrix. The node feature matrix is used to reflect the node feature data of each distribution network node at different time steps, and the dynamic adjacency matrix is used to reflect the electrical coupling strength between different distribution network nodes at different times.
[0060] Among them, the node feature matrix refers to a multi-dimensional data set formed by electrical parameters such as voltage, current, power and time step information. Specifically, it can be realized by periodically sampling and feature recombination of real-time collected electrical quantity data using a sliding time window, and is used to characterize the dynamic behavior pattern of distribution network nodes in a continuous time series. Its expression can be:
[0061]
[0062] Among them, is the number of distribution network nodes; is the feature dimension, such as voltage, current, power, etc.; is the time step.
[0063] The dynamic adjacency matrix refers to a time-varying weight matrix generated based on the electrical quantity change rate and topological connection relationship between nodes. Specifically, it can be realized by calculating the difference in electrical coupling degree between adjacent nodes using the dynamic time warping algorithm, and is used to capture the real-time fluctuations in the association strength between nodes during the dynamic reconfiguration of the network topology. The matrix expression and the calculation formula examples of each element in the matrix are as follows:
[0064]
[0065] Among them, It can change with time and operating status, and reflect the real-time electrical coupling strength between nodes.
[0066] Dynamic adjacency matrix calculation formula:
[0067]
[0068] In the formula, represents node , The active power between them. The higher the power flow of a node, the higher the correlation degree.
[0069] Specifically, after obtaining the operation data of the distribution network, first slice the electrical quantity parameters such as voltage, current, and power of each node according to a preset time step, form a three-dimensional tensor structure including the number of nodes, feature dimensions, and time steps as the node feature matrix. Subsequently, according to the change trend of the electrical quantities of adjacent nodes in the node feature matrix, calculate the electrical coupling strength of different time segments, and generate a dynamic adjacency matrix in combination with the topological connection relationship, where the weights of the matrix elements are dynamically adjusted according to the electrical quantity correlation degree. Finally, input the dynamic adjacency matrix and the node feature matrix into the graph attention network framework, and realize the dynamic aggregation and update of node features through the graph convolution layer and the attention mechanism, and construct a dynamic graph attention network model that can adapt to real-time topological changes.
[0070] Compared with the existing technology, this solution can update the electrical coupling strength between nodes in real time through the dynamic adjacency matrix, enabling the model to adaptively capture the topological dynamic change characteristics and avoiding misjudgment of faults caused by topological mismatch.
[0071] It can be understood that in order to further improve the node feature representation ability of this solution and reduce the misjudgment risk caused by network structure changes, the update logic example of the node feature matrix provided in this embodiment is as follows:
[0072] Based on the selected target node, calculate the attention coefficient between the target node and its neighbor nodes at the current time step through a preset spatial attention coefficient calculation formula;
[0073] Through the activation function, perform feature aggregation on the node features of each neighbor node and the attention coefficient to obtain the updated node feature of the target node at the next time step;
[0074] Summarize the node features of the target node up to the current time step, and perform time-dependent feature extraction through the LSTM network to obtain the time feature sequence of the target node;
[0075] According to the weighted sum of the updated node feature and the time series feature, obtain the latest node feature of the target node.
[0076] It should be noted that further based on the selected target node, through a preset spatial attention coefficient calculation formula, the attention coefficient between the target node and its neighbor nodes at the current time step is calculated; through an activation function, the node features of each neighbor node are aggregated with the attention coefficient to obtain the updated node features of the target node at the next time step; the node features of the target node up to the current time step are summarized, and the time-dependent features are extracted through an LSTM network to obtain the time feature sequence of the target node; according to the weighted sum of the updated node features and the time series features, the latest node features of the target node are obtained.
[0077] Among them, the spatial attention coefficient refers to dynamically calculating the degree of association between the target node and its neighbor nodes. Specifically, for nodes i and j, the attention coefficient at time t can be calculated by calculating the cosine similarity or dot product similarity after linear transformation, which is used to capture the coupling strength between nodes under topological dynamic changes. Among them, the spatial attention coefficient mentioned in this embodiment The calculation formula is:
[0078]
[0079] In the formula, represents the hidden state of node i at time t, represents the hidden state of node j at time t, is a learnable parameter matrix, is a vector concatenation operation. LeakyReLU (Leaky Rectified Linear Unit) is an improved activation function, which is mainly used to solve the defect that the traditional ReLU (Rectified Linear Unit) has a zero gradient in the negative value interval.
[0080] The activation function refers to a function that introduces non-linear transformation, such as the ReLU or Sigmoid function, which can be specifically used for non-linear mapping of the aggregated features to enhance the model's expression ability. Feature aggregation refers to weighted superposition of the features of neighbor nodes according to the attention coefficient, which can be specifically implemented by weighted summation or concatenation, and is used to fuse the spatial association information between multiple nodes.
[0081] The LSTM network refers to a long short-term memory recurrent neural network, which can be specifically implemented through hidden state transmission and time gating mechanisms, and is used to extract the dependence relationship of node features in the time dimension. The time series feature refers to the feature change trend of the target node at historical time steps, which can be specifically encoded by intercepting time series segments through a sliding window, and is used to reflect the dynamic evolution law of the node state.
[0082] Specifically, for the selected target node, first, based on the node feature matrix at the current time step, the spatial attention coefficient calculation formula is used to quantify the electrical coupling strength between this node and its adjacent nodes. For example, the current mutation or voltage fluctuation of the neighbor nodes will dynamically affect the feature update of the target node through the attention coefficient. Subsequently, the features of the neighbor nodes are multiplied by their corresponding attention coefficients and then accumulated, and a non-linear transformation is performed through the activation function to generate the preliminary updated features of the target node at the next moment, and its expression is as follows:
[0083]
[0084] represents the hidden state of node i at time t+1, is the neighbor set of node i, defined by defined, is the learnable parameter matrix, is the activation function.
[0085] Meanwhile, the feature sequences of the target node at the current and historical time steps are input into the LSTM network, and its memory unit is used to capture the fault propagation pattern in time series, and the expression is as follows:
[0086]
[0087] is the cell state sequence, is the hidden state sequence.
[0088] Finally, the updated features after spatial aggregation and the time features extracted by LSTM are weighted and fused. For example, higher weights are given to the time features to reflect the causal time series of fault propagation, so as to generate the latest multi-dimensional feature representation of the target node, and the expression is as follows:
[0089]
[0090] is the learnable parameter used to balance the importance of spatio-temporal features.
[0091] This solution uses a dynamic spatial attention mechanism to adjust the correlation weights between nodes in real time. At the same time, it combines the LSTM network to mine deep time series patterns, enabling the node features to fuse the spatial dynamic correlation and time causal evolution information simultaneously, and dynamically adjusting the model parameters according to the fault severity to enhance the attention to key areas:
[0092]
[0093] Among them, , represents the fault confidence of node i, is the adjustment coefficient, is the correlation weight before adjustment, is the correlation weight after adjustment.
[0094] Through the above technical solution, the representation ability of node features can be improved, enabling the fault diagnosis model to more accurately capture the abnormal propagation path under dynamic topology, reducing the misjudgment risk caused by network structure changes, and enhancing the ability to extract fault features of multiple time scales, thereby significantly improving the reliability of diagnosis results.
[0095] Furthermore, this embodiment also proposes to perform feature fusion processing on the operation data of the distribution network in different modalities through a cross-modal attention mechanism to obtain modal fusion features, including: according to the operation data of the distribution network, performing feature encoding processing through multiple preset modal encoders to respectively obtain an electrical quantity modal matrix, a topological modal matrix, and an environmental modal matrix; according to the electrical quantity modal matrix, the topological modal matrix, and the environmental modal matrix, using the electrical quantity modal matrix as the modal alignment benchmark, through a preset QKV mapping calculation formula, obtaining QKV features; according to the QKV features, combining with the attention weight calculation formula, obtaining attention weights, and then based on the sum of the attention weights and the electrical quantity modal matrix, obtaining modal fusion features.
[0096] Among them, the modal encoder refers to a neural network module designed for different data modalities, which can specifically be implemented using a convolutional neural network or a fully connected layer, and is used to map heterogeneous data such as original electrical quantities, topological structures, and environmental parameters to a unified feature space. The modal encoders involved in this embodiment mainly include an electrical quantity encoder, which inputs (time-series electrical quantities) and outputs an electrical quantity modal matrix .
[0097]
[0098] A topological encoder that inputs the dynamic graph (adjacency matrix ) generated by DGAT (Dynamic Graph Attention Network) and outputs a topological modal matrix .
[0099]
[0100] Among them, is the geographical data of each node of the distribution network.
[0101] An environmental encoder that inputs environmental image features and meteorological features and outputs an environmental modal matrix .
[0102]
[0103] The QKV (Query-Key-Value) mapping calculation formula refers to a mathematical expression that takes the electrical quantity modal matrix as a reference and generates query, key, and value feature vectors through linear transformation. Specifically, a trainable weight matrix can be selected to project the input data to achieve cross-modal feature alignment. The attention weight calculation formula refers to calculating the similarity between the query feature and the key feature to generate the correlation strength coefficient between different modalities. Specifically, the dot product attention mechanism can be selected to dynamically allocate the contribution of each modality to fault diagnosis.
[0104] The QKV mapping calculation formula mentioned in this embodiment, its expression is specifically:
[0105]
[0106] In the formula, is the electrical quantity modal matrix, is the combined matrix of the topological modal matrix and the environmental modal matrix, , , are the weight matrices of the Q feature, K feature, and V feature respectively.
[0107] Specifically, the electrical quantity modal matrix contains the time series features of real-time measurement data such as voltage and current. The topological modal matrix reflects the dynamic connection relationship between network nodes, and the environmental modal matrix integrates external influencing factors such as temperature and humidity. After performing QKV mapping with the electrical quantity modal as the reference, by calculating the matching degree between the key features corresponding to the topological and environmental modalities and the reference query features, the attention weights reflecting cross-modal correlation are generated. These weights are used to adjust the fusion ratio of different modal features. For example, in thunderstorm weather, the humidity feature of the environmental modality may be given a higher weight. Finally, the weighted multi-modal features are superimposed with the reference electrical quantity features to form a fusion feature vector containing cross-modal association information, providing more complete input information for the downstream fault classifier. Through the introduction of the cross-modal attention mechanism, this scheme establishes a dynamic association model between modalities in the feature space, enabling external interferences such as meteorological condition changes and topological reconstruction events to be quantitatively evaluated and incorporated into the diagnostic decision, effectively solving the feature conflict problem caused by static fusion rules. Through the above technical solutions, this embodiment realizes the efficient cooperation of multi-modal data at the feature level and overcomes the problem of false fault judgment caused by the single data dimension in traditional methods. For example, in the scenario of frequent switching of distributed power sources, by dynamically adjusting the association weights between the topological modality and the electrical quantity modality, it is possible to accurately distinguish the normal fluctuations caused by topological changes from the real fault signals, reduce the false alarm probability caused by network reconstruction, and improve the diagnostic robustness under complex working conditions.
[0108] Furthermore, the solution of this embodiment may further include the following steps:
[0109] Based on a preset multimodal data sample of the distribution network, combined with a loss function based on structural causal constraints, the multimodal large model is trained to determine the convergence degree of the multimodal large model based on the output of the loss function. When the convergence degree reaches the preset convergence threshold or the number of training iterations reaches the preset number threshold, the multimodal large model is output.
[0110] Among them, the loss function with structural causal constraints refers to introducing prior knowledge of the power grid fault propagation logic to constrain the causal rationality of the model output. Specifically, it can be implemented by combining causal inference theory and a regularization loss term. For example, a causal graph structure constraint term is added to the loss function to limit the direction of model parameter update. The multimodal data sample refers to a distribution network operation data set containing multi-dimensional information such as electrical quantities, topology, and environment. Specifically, a sample can be constructed by combining historical fault records, real-time monitoring data, and meteorological environment data. For example, voltage, current, power data, and temperature, humidity, and wind speed data are aligned according to the time stamp. Among them, the convergence threshold refers to the error threshold for determining whether to terminate the iteration during the model training process. Specifically, the fault diagnosis accuracy or the change rate of the loss value on the validation set can be used as the judgment basis. For example, when the validation set accuracy fluctuates less than 1% for three consecutive iterations, it is considered convergent.
[0111] Specifically, in the training process, a batch of data is first extracted from the multimodal data sample and input into the multimodal large model. After the model outputs the fault diagnosis result, the difference between the prediction result and the true label is calculated in combination with the loss function with structural causal constraints. The loss function consists of a data fitting term and a causal constraint term. The former measures the prediction error, and the latter constructs a causal graph according to the power grid fault propagation path to constrain the distribution of model feature weights. For example, the causal constraint term can adopt a graph regularization method based on the adjacency matrix to ensure that the fault association strength between nodes conforms to the actual physical law. During the training process, the output of the loss function is used for backpropagation to update the model parameters. When the performance index of the model on the validation set reaches the preset threshold or the number of iterations reaches the upper limit, the training terminates and the optimized multimodal large model is output.
[0112] In this solution, by introducing structural causal constraints, prior knowledge such as the power grid topology connection relationship and fault propagation path is encoded into the loss function, so that the model training process not only focuses on the prediction accuracy, but also forces the learning of feature associations that conform to the actual causal logic, thereby improving the rationality and robustness of the diagnosis result.
[0113] Furthermore, after step 101 and before step 102 in this embodiment, it may further include:
[0114] Preprocess the data, and the preprocessing includes data cleaning, normalization, and denoising.
[0115] Among them, data cleaning refers to correcting or removing outliers or missing values in the operation data of the distribution network. Specifically, interpolation methods can be used to fill in missing data segments, or abnormal fluctuation data can be identified and filtered based on the statistical threshold of a sliding window. This step can eliminate the interference of sensor acquisition errors on the subsequent model input.
[0116] Among them, normalization refers to converting electrical measurement data with different dimensions to a unified numerical range. Specifically, the Z-score normalization method can be used to map voltage and current features to a distribution space with a mean of 0 and a variance of 1. This step eliminates the magnitude differences between multi-modal data to improve the model convergence efficiency. Among them, denoising refers to suppressing high-frequency interference signals in the operation data of the distribution network. Specifically, wavelet transform can be used to perform multi-scale decomposition on the current waveform and then remove high-frequency noise components. This step helps to retain the effective frequency band information of fault characteristics.
[0117] Specifically, during the acquisition process of distribution network operation data, it is vulnerable to factors such as sensor drift and communication interference, resulting in data with noise, dimensional differences, and incompleteness. By cleaning the original data, missing values caused by communication interruptions can be repaired, and abnormal mutation points caused by equipment failures can be removed; through normalization, electrical measurement data with different dimensions such as voltage and power are converted into standardized feature vectors, facilitating the subsequent collaborative training of the dynamic graph attention network and the multi-modal large model; through wavelet denoising, time-frequency decomposition of the current waveform is performed to filter out high-frequency noise interference and retain the low-frequency components of fault characteristics, thereby improving the model's ability to identify weak fault signals.
[0118] Through the above technical solutions, this application can effectively suppress the problem of model overfitting caused by training data noise or missingness, enhance the generalization ability of the multi-modal large model in complex dynamic scenarios, make the fault diagnosis results conform to the data statistical law and meet the physical operation constraints of the power grid, and significantly reduce the probability of false alarms and missed detections in the new energy high-penetration environment.
[0119] Furthermore, knowledge distillation and lightweight deployment can also be adopted for the multi-modal large model to accelerate fault diagnosis response and improve the operation stability of the distribution network. The loss function is set as:
[0120]
[0121] In the formula, are the hidden layer feature representations of the teacher model and the student model respectively, λ is the weight parameter, and the KL divergence is used to measure the output probability distribution P of the teacher model and the student model.
[0122] Through the above distillation design, the student model can achieve an order-of-magnitude efficiency improvement with minimal accuracy loss, providing key technical support for this method.
[0123] The above is an embodiment of a distribution network fault diagnosis method based on a large model provided by this application. The following is a detailed description of an embodiment of a distribution network fault diagnosis device based on a large model provided by this application.
[0124] Please refer to Figure 2 , a distribution network fault diagnosis device based on a large model provided in this embodiment includes:
[0125] An operation data acquisition unit 201, configured to acquire distribution network operation data, where the distribution network operation data includes: topology data, electrical quantity data, and environmental data;
[0126] A dynamic graph network construction unit 202, configured to construct a dynamic graph attention network model according to the topology data and electrical quantity data in the distribution network operation data, where the dynamic graph attention network model is used to reflect the real-time topology dynamic change relationship of the distribution network;
[0127] A distribution network fault diagnosis unit 203, configured to perform feature fusion processing on different modalities of distribution network operation data through a cross-modal attention mechanism based on a preset multi-modal large model to obtain a modality fusion feature, and obtain a fault diagnosis result of the distribution network based on the modality fusion feature.
[0128] As Figure 3 shown, this embodiment also provides a distribution network fault diagnosis terminal based on a large model, including: a memory 33 and a processor 31, where the memory 33 and the processor 31 can be connected through a communication bus 34;
[0129] The memory 33 is used to store program codes, and the program codes are used to implement a distribution network fault diagnosis method as provided in the above embodiment;
[0130] The processor 31 is used to read and execute the program codes.
[0131] Among them, the memory refers to a hardware device with a data storage function, and can be specifically implemented by a solid-state drive or a flash chip, and is used to store program codes including dynamic graph attention network construction logic, multi-modal feature fusion algorithms, and fault diagnosis rules. The processor refers to a computing unit with a data processing function, and can be specifically implemented by a multi-core central processing chip or a graphics processing chip, and is used to parse the program codes and perform real-time topology dynamic modeling, cross-modal data alignment, and fault reasoning operations. The program codes refer to software modules including computer instruction sequences, and can be specifically written and implemented in Python or C++ languages, and are used to convert distribution network operation data into dynamic graph structure features, and realize the collaborative analysis of electrical quantity, topology, and environmental data through a cross-modal attention mechanism.
[0132] Specifically, when the terminal is running, after the program code pre - set in the memory is loaded by the processor, it first obtains real - time topology data, electrical measurement information, and meteorological environment monitoring data from the external data acquisition system, and performs data cleaning and normalization operations. Subsequently, the processor constructs a dynamic graph attention network model based on the topological connection relationship and the electrical quantity time - series characteristics, captures the dynamic coupling relationship between nodes through the spatial attention mechanism, and extracts the evolution law of the time - dimension characteristics in combination with the LSTM network. In the multi - modal fusion stage, the processor inputs the encoded features of different modalities into the cross - modal attention layer, aligns the topological and environmental features based on the electrical quantity data, calculates the correlation weights between modalities through QKV mapping, and generates a fused global feature vector. Finally, the processor performs fault type classification and location calculation based on the fused features, and outputs the diagnostic results to the human - machine interaction interface or the control execution system.
[0133] The present application also provides a computer - readable storage medium, in which program code is stored. The program code is used to be read and executed by the processor to implement a distribution network fault diagnosis method based on a large model as provided in the above - mentioned embodiment.
[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above - described terminal, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0135] In several embodiments provided by the present application, it should be understood that the disclosed terminal, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the shown or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other forms.
[0136] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0137] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0138] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0140] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0141] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A method for diagnosing faults in a distribution network based on a large model, characterized in that, Including: Obtain the operation data of the distribution network, where the operation data of the distribution network includes: topology data, electrical quantity data, and environmental data; Construct a dynamic graph attention network model according to the topology data and electrical quantity data in the operation data of the distribution network, where the dynamic graph attention network model is used to reflect the real-time topological dynamic change relationship of the distribution network; Based on a preset multi-modal large model, perform feature fusion processing on the operation data of the distribution network in different modalities through a cross-modal attention mechanism to obtain a modality fusion feature, so as to obtain a fault diagnosis result of the distribution network based on the modality fusion feature.
2. The method for diagnosing faults in a distribution network based on a large model according to claim 1, characterized in that The constructing a dynamic graph attention network model according to the operation data of the distribution network includes: Construct a node feature matrix according to the operation data of the distribution network, where the node feature matrix is used to reflect the node features of each distribution network node at different time steps; Construct the dynamic adjacency matrix according to the node feature matrix, where the dynamic adjacency matrix is used to reflect the electrical coupling strength between different distribution network nodes at different times; Construct a dynamic graph attention network model according to the dynamic adjacency matrix.
3. The method for diagnosing faults in a distribution network based on a large model according to claim 2, wherein, It also includes: Based on the selected target node, calculate the attention coefficient between the target node and its neighbor nodes at the current time step through a preset spatial attention coefficient calculation formula; Through an activation function, perform feature aggregation on the node features of each neighbor node and the attention coefficient to obtain the updated node feature of the target node at the next time step; Summarize the node features of the target node up to the current time step, and perform time-dependent feature extraction through an LSTM network to obtain the time feature sequence of the target node; Obtain the latest node feature of the target node according to the weighted sum of the updated node feature and the time series feature.
4. A method for diagnosing faults in a distribution network based on a large model according to claim 1, characterized in that, The performing feature fusion processing on the operation data of the distribution network in different modalities through a cross-modal attention mechanism to obtain a modality fusion feature includes: According to the operation data of the distribution network, perform feature encoding processing through multiple preset modality encoders to obtain an electrical quantity modality matrix, a topology modality matrix, and an environmental modality matrix respectively; According to the electrical quantity modality matrix, the topology modality matrix, and the environmental modality matrix, taking the electrical quantity modality matrix as the modality alignment benchmark, obtain QKV features through a preset QKV mapping calculation formula; According to the QKV features, combine the attention weight calculation formula to obtain attention weights, and then obtain the modality fusion feature based on the sum of the attention weights and the electrical quantity modality matrix.
5. The method for diagnosing faults in a distribution network based on a large model according to claim 4, characterized in that, The specific QKV mapping calculation formula is: wherein, is the electrical quantity modal matrix, is the combined matrix of the topological modal matrix and the environmental modal matrix, , , are the weight matrices of the Q feature, the K feature, and the V feature, respectively.
6. The method for diagnosing faults in a distribution network based on a large model according to claim 4, wherein, It also includes: Based on a preset distribution network multi-modal data sample, combined with a loss function based on structural causal constraints, train the multi-modal large model, so as to determine the convergence degree of the multi-modal large model based on the output of the loss function. When the convergence degree reaches a preset convergence threshold or the number of training iterations reaches a preset number threshold, output the multi-modal large model.
7. A method for diagnosing faults in a distribution network based on a large model according to claim 1, characterized in that, After obtaining the operation data of the distribution network, it also includes: Preprocess the operation data of the distribution network, where the preprocessing includes: data cleaning, normalization, and denoising.
8. A distribution network fault diagnosis device based on a large model, characterized in that, Including: An operation data acquisition unit for acquiring operation data of a distribution network, where the operation data of the distribution network includes: topology data, electrical quantity data, and environmental data; A dynamic graph network construction unit for constructing a dynamic graph attention network model according to the topology data and electrical quantity data in the operation data of the distribution network, where the dynamic graph attention network model is used to reflect the real-time topology dynamic change relationship of the distribution network; A distribution network fault diagnosis unit for performing feature fusion processing on operation data of different modalities of the distribution network through a cross-modal attention mechanism based on a preset multi-modal large model to obtain a modality fusion feature, so as to obtain a fault diagnosis result of the distribution network based on the modality fusion feature.
9. A distribution network fault diagnosis terminal based on a large model, characterized in that, Comprising: A memory and a processor; The memory is used to store program code, and the program code is used to implement a distribution network fault diagnosis method according to any one of claims 1 to 7; The processor is used to read and execute the program code.
10. A computer-readable storage medium, characterized in that Program code is stored in the computer-readable storage medium, and the program code is used to be read and executed by the processor to implement a distribution network fault diagnosis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Topology identification and fault diagnosis analysis method, device and equipment for active power distribution network
CN117609824A
Power distribution network fault positioning method based on FSGATv2 and topological similarity
CN119619734A
Social network user viewpoint classification method and device, computing equipment and storage medium
CN120011563A
Topology-inspired neural network autoencoding for electronic system fault detection
US20190286506A1
Fault locating method and apparatus applied to power distribution network, device, and medium
WO2024187506A1
Cited By
Power distribution network fault intelligent diagnosis device and method based on AI large model
CN121352035A
Industrial production fault intelligent diagnosis method and system based on multi-modal data
CN121365206A
Power distribution network fault diagnosis method and system based on dynamic graph attention network and cross-modal fusion
CN121412750A
Power distribution network fault diagnosis method and system based on dynamic graph attention network and cross-modal fusion
CN121412750B
Distribution transformer analysis method and device based on intelligent fusion terminal, equipment and medium
CN121432026A