Distribution network fault identification method, device and electronic equipment
By constructing a feature alignment method between the multi-head attention mechanism model and the large language model, the problem of insufficient fusion of dual-modal data is solved, the accuracy and efficiency of distribution fault detection is improved, and training time is reduced.
Patent Information
- Application Number
- CN202411272344.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-11
AI Technical Summary
In the detection of distribution faults in the prior art, dual-mode semantic extraction is insufficient research, and it is difficult to effectively integrate the deep semantic correlation and knowledge transfer between text and ID table data, resulting in low detection accuracy and efficiency.
By building a basic prediction model and large language model for faults of the distribution system based on multi-head attention mechanism, fine-grained feature alignment of ID table features and text features is achieved, information is reconstructed using mask features, and combined with the advantages of dual-modal features, fine-tuning of downstream fault detection tasks is carried out.
It improves the accuracy and efficiency of power distribution fault detection, reduces the number of model parameters and training time, and enhances the universality and scalability of the framework.
Smart Images

Figure CN119125764B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of distribution system fault detection, and particularly relates to a method for identifying faults in a distribution network, a device for identifying faults in a distribution network, an electronic device, and a corresponding storage medium. Background Technique
[0002] With the rapid development of information technology, various digital distribution fault detection technologies are emerging continuously, bringing new opportunities for improving the reliability and safety of the distribution system. Among them, intelligent fault diagnosis technologies based on big data analysis and artificial intelligence have attracted particular attention. Traditional fault diagnosis methods usually rely on expert experience and regular inspections, and it is difficult to detect hidden fault problems in a timely manner. The emergence of big data and artificial intelligence has injected new impetus into fault diagnosis. These emerging technologies can effectively solve the limitations of traditional methods: a large amount of device operation data can provide a rich information basis for fault diagnosis, and by using methods such as machine learning to deeply analyze these data, hidden fault patterns and abnormal rules can be discovered. Artificial intelligence algorithms such as deep learning have excellent pattern recognition and automatic diagnosis capabilities, and can realize comprehensive monitoring of device status and intelligent fault diagnosis. These technologies have strong self-adaptability and self-learning capabilities, and can continuously optimize and improve the fault diagnosis model as the usage of the device changes. With the continuous development and maturity of emerging technologies such as big data and artificial intelligence, they will surely play an increasingly important role in the intelligent management of distribution equipment. In the future, intelligent fault diagnosis systems based on these cutting-edge technologies will be widely applied in the power industry, making great contributions to improving the safety and reliability of the distribution system.
[0003] On the other hand, with the development of deep learning technology, large language models (LLMs) have been a major breakthrough in the field of machine learning in recent years. These models learn a vast amount of semantics and knowledge by training on a massive amount of text data and have achieved remarkable results in many natural language processing tasks. For example, with the rise of BERT and GPT, a new paradigm for processing text data has begun to attract the attention of researchers. These models can understand and extract deep semantic relationships in text data through pre-training on large-scale corpora. Applying these powerful large language models to the fault detection of power distribution equipment can leverage their advantages in text understanding and knowledge expression. Specifically, text data such as the operation logs and maintenance records of power distribution equipment can be input into the large language model for analysis to identify abnormal patterns and fault symptoms. Large language models are good at extracting hidden patterns and rules from a large amount of historical data and can quickly and accurately detect abnormal situations in equipment operation. Compared with traditional rule-based or statistical method-based fault diagnosis, large language models can more intelligently discover complex fault patterns and significantly improve the accuracy and efficiency of fault detection. In addition, large language models can also be combined with other machine learning technologies such as deep learning to further enhance the fault diagnosis ability. For example, features extracted by the large language model can be used to train a high-performance fault classification model. Generally speaking, as a powerful AI technology, large language models are deeply integrating into the intelligent operation and maintenance of power distribution equipment and making important contributions to improving the reliability and safety of power distribution systems.
[0004] Although the applications of LLMs are becoming increasingly widespread in various tasks and fields, the research on dual-modal semantic extraction in the application of power distribution fault detection lags behind relatively. Dual-modal semantic extraction involves extracting and fusing information from data in two modalities (such as text and ID tables), which poses higher requirements for the overall understanding ability and flexibility of the model. Current research mainly focuses on single-modal or limited modal fusion, and the exploration of the deep semantic associations and knowledge transfer between dual-modal data is still insufficient. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide a power distribution network fault identification method, device, and electronic device. This application aims to perform fine-grained feature-level alignment between an attention mechanism model based on ID table features and a large language model based on text features to improve the accuracy of downstream power distribution fault detection tasks. Through joint pre-training, it is allowed that the masked features of one modality can borrow information from the other modality when reconstructing information, thereby aligning features between different modalities. In addition, these two models are jointly fine-tuned on downstream power distribution fault detection tasks to integrate their advantages in processing different modality features, so as to solve at least some of the problems in the background technology.
[0006] To achieve the above object, a method for identifying distribution network faults is provided in this application. The method includes:
[0007] Determine the real-time operation data of the distribution system for the modal features to be detected. The modal features include ID table features and text features; respectively construct a basic distribution system fault prediction model based on the multi-head attention mechanism for processing the ID table features and a large language model for distribution system fault detection for processing the text features; construct dual-modal features according to the ID table features and text features to achieve fine-grained feature alignment between the basic distribution system fault prediction model and the large language model for distribution system fault detection, and align the feature representations of different modalities at the instance-level coarse-grained level; apply the aligned large language model for distribution system fault detection and the basic distribution system fault prediction model to downstream tasks, and fine-tune them for distribution fault detection.
[0008] Optionally, the ID table features are constructed through the following steps: preset a two-dimensional ID table with a dimension of R N×F where N represents the number of timestamps and F represents the total number of distribution fault features; each row of the two-dimensional ID table represents the instance operation data n at a timestamp, and each column represents a distribution fault feature f; according to the data features of the real-time operation data of the distribution system, encode the row and column values of the two-dimensional ID table to generate the ID table features corresponding to the real-time operation data of the distribution system.
[0009] Optionally, the text features are constructed through the following steps: according to the data features of the real-time operation data of the distribution system, adopt the hard prompt template method, and generate the corresponding text features after text splicing according to the text name of the distribution fault feature and the text feature value corresponding to the distribution fault feature in each instance input.
[0010] Optionally, a basic power distribution system fault prediction model based on a multi-head attention mechanism for processing the ID table features is constructed, including: determining the model category and model structure of the basic power distribution system fault prediction model before training, where the model category is a multi-classification model, and the model structure includes an embedding layer, a feature interaction layer, and a prediction layer; the embedding layer is configured to: convert the high-dimensional sparse ID table features into a low-dimensional dense embedding matrix; the power distribution fault feature f is converted into an embedding vector through a look-up table method by its corresponding representation matrix; the feature interaction layer is configured to: take the embedding matrix as the input and output a dense representation vector for each power distribution terminal instance running data n; adopt a high-low order cross model with a two-tower structure to design complex interaction operations for each feature to obtain low-order or high-order feature cross terms; the prediction layer is configured to: after concatenating the representation vectors output by the high-low order cross, input them into a neural network or a linear regression module, and then connect the Sigmoid function to adjust the range to obtain the classification probability of the final fault type; after training the basic power distribution system fault prediction model before training with training samples in an end-to-end manner using binary cross-entropy loss, the basic power distribution system fault prediction model is obtained.
[0011] Optionally, a large language model for power distribution system fault detection for processing the text features is constructed, including: converting the text features into tokens form based on a large language model pre-trained for power distribution fault detection; inputting the obtained tokens into an encoder network based on Transformer, and obtaining a context-consistent representation after processing; performing power distribution fault classification by adding randomly initialized attention heads on the context-consistent representation to obtain the fault classification label corresponding to the text features.
[0012] Optionally, construct dual-modal features based on the ID table features and text features to achieve fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection, and align the feature representations of different modalities at the instance-level coarse-grained level, including: extracting a certain proportion of feature domains from the ID table features and text features respectively for masking operations to obtain the interference forms of the ID table features and text features, namely masked ID table features and masked text features; inputting the ID table features and masked ID table features into the basic power distribution system fault prediction model for encoding, and inputting the text features and masked text features into the large language model for power distribution system fault detection for encoding to obtain the corresponding outputs; performing joint reconstruction based on the output after encoding the ID table features and the output after encoding the masked text features, and performing joint reconstruction based on the output after encoding the text features and the output after encoding the masked ID table features; establishing coarse-grained hierarchical feature alignment between the dual-modalities according to the instance-level comparative learning method; combining the optimization objectives in the joint reconstruction and coarse-grained hierarchical feature alignment processes, and training based on the overall loss function in the dual-modal alignment pre-training stage of the basic power distribution fault prediction model and the large language model for power distribution system fault detection.
[0013] Optionally, extract a certain proportion of feature domains from the ID table features and text features respectively for masking operations to obtain the interference forms of the ID table features and text features, including: extracting a certain proportion of feature domains from the ID table features, and replacing the corresponding features with an additional masking symbol, which is shared by all feature domains; replacing the features of the ID table's feature vector based on the encoding method with a specific symbol to obtain the masked ID table features, and the masked ID table features include a zero vector with a dimension of v f ; and extracting a certain proportion of feature domains from the text features, replacing the features with specific characters, and obtaining the masked text features by adopting the hard prompt template method.
[0014] Optionally, input the ID table features and masked ID table features into the basic power distribution system fault prediction model for encoding, and input the text features and masked text features into the large language model for power distribution system fault detection for encoding to obtain the corresponding outputs, including: inputting the ID table features and masked ID table features into the basic power distribution system fault prediction model, and after encoding through the embedding layer and feature interaction layer in the basic power distribution system fault prediction model, obtaining the corresponding outputs; and inputting the text features and masked text features into the large language model for power distribution system fault detection, and obtaining the corresponding outputs through the set of hidden layer states of the last layer in the large language model for power distribution system fault detection.
[0015] Optionally, joint reconstruction is performed based on the output encoded by text features and the output encoded by mask ID table features, including: inputting the output encoded by text features and the output encoded by mask ID table features into a self-attention mechanism module in the form of an encoded vector tuple to obtain an aggregated output; for each mask ID table feature, passing it through an independent multi-layer perceptron network and a Softmax function to obtain a distribution function on the candidate features of the aggregated output; calculating the reconstructed features of the ID table features using the cross-entropy loss function of the distribution function and the ID table features for all mask features; calculating the reconstruction loss based on the reconstructed features using a noise contrast estimation function; joint reconstruction is performed based on the output encoded by ID table features and the output encoded by mask text features, including: concatenating the output encoded by ID table features and the output encoded by mask text features in the form of an encoded vector tuple to obtain a concatenated vector; inputting the concatenated vector into the prediction layer in the basic power distribution system fault prediction model to obtain a reconstructed feature distribution function; using the cross-entropy loss function of the feature distribution function and text features as the reconstruction loss.
[0016] Optionally, a coarse-grained hierarchical feature alignment between the two modalities is established according to the instance-level comparison learning method, including: using two independent linear layers and normalization layers to project the text input corresponding tokens vector output and the ID table feature output into a first multi-dimensional vector and a second multi-dimensional vector; calculating the comparison loss at the instance level of the first multi-dimensional vector and the second multi-dimensional vector using the InfoNCE function; obtaining the final loss based on the comparison loss; obtaining the overall loss of the basic power distribution system fault prediction model and the large language model for power distribution system fault detection based on the reconstruction loss and the final loss.
[0017] Optionally, the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model are applied to downstream tasks and fine-tuned for power distribution fault detection, including: in the downstream power distribution system fault detection task, jointly fine-tuning the basic power distribution system fault detection model and the large language model for power distribution system fault using supervised click signals; setting randomly restarted linear output layers in the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model respectively, and outputting the estimated fault classification probabilities respectively; obtaining the final fault classification probability through weighted calculation based on the fault classification probabilities of the two models respectively; comparing the estimated results corresponding to the final fault classification probability with the true labels respectively, and using the cross-entropy loss function to measure the gap to fine-tune and optimize the model.
[0018] In some embodiments of the present application, a distribution network fault identification device is further provided. The device includes: a feature confirmation module, configured to determine the real-time operation data of the power distribution system for the modal features to be detected, where the modal features include ID table features and text features; a model construction module, configured to respectively construct a basic power distribution system fault prediction model based on a multi-head attention mechanism for processing the ID table features and a large language model for power distribution system fault detection for processing the text features; a feature alignment module, configured to construct dual-modal features according to the ID table features and the text features to achieve fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection, and align the feature representations of different modalities at the instance-level coarse-grained level; and an application deployment module, configured to apply the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model to downstream tasks, and fine-tune them for power distribution fault detection.
[0019] In the present application, an electronic device is further provided, including: at least one processor; a memory connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the at least one processor implements the foregoing distribution network fault identification method by executing the instructions stored in the memory.
[0020] In the present application, a machine-readable storage medium is further provided. Instructions are stored on the machine-readable storage medium, and when the instructions are executed by a processor, the processor is configured to execute and implement the foregoing distribution network fault identification method.
[0021] In the present application, a computer program product is further provided, including a computer program, and the computer program implements the foregoing distribution network fault identification method when executed by a processor.
[0022] The above technical solutions have the following beneficial effects:
[0023] (1) Appropriate models can be selected according to different scenarios, which helps to enhance the generality and scalability of the framework.
[0024] (2) The parameterized models can be plug-and-play. Through the joint reconstruction at the feature level between dual-modalities, the model pre-training task of fine-grained feature-level interaction and alignment is achieved, and the parameterized models can be well inserted into downstream power distribution fault detection tasks for fine-tuning.
[0025] (3) The number of parameters and training time of the LLM and the multi-head attention mechanism model are reduced. Because the LLM and the multi-head attention mechanism model will provide more dual-modal information during joint pre-training, which is conducive to the rapid convergence of training. In downstream power distribution fault detection tasks, the model fine-tuning process will greatly reduce the training time compared with the original supervised learning.
[0026] Other features and advantages of the embodiments of the present application will be described in detail in the following detailed implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the following detailed implementation, they are used to explain the embodiments of the present application, but do not constitute a limitation to the embodiments of the present application. In the drawings:
[0028] Figure 1 Schematically shows a schematic diagram of the steps of the distribution network fault identification method according to an embodiment of the present application;
[0029] Figure 2 Schematically shows a technical block diagram of the distribution network fault identification method according to an embodiment of the present application;
[0030] Figure 3 Schematically shows a fine-tuning block diagram of the downstream distribution fault detection task according to an embodiment of the present application;
[0031] Figure 4 Schematically shows a structural diagram of the distribution network fault identification device according to an embodiment of the present application;
[0032] Figure 5 Schematically shows an internal structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The following details the specific implementation manners of the embodiments of the present application with reference to the drawings. It should be understood that the specific implementation manners described herein are only for explaining and illustrating the embodiments of the present application, and are not used to limit the embodiments of the present application.
[0034] Figure 1 Schematically shows a schematic diagram of the steps of the distribution network fault identification method according to an embodiment of the present application. As Figure 1 shown, a distribution network fault identification method includes:
[0035] S01. Determine the real-time operation data of the power distribution system for the modal features to be detected, where the modal features include ID table features and text features;
[0036] S02. Respectively construct a basic prediction model for power distribution system faults based on the multi-head attention mechanism for processing the ID table features and a large language model for power distribution system fault detection for processing the text features;
[0037] S03. Construct dual-modal features based on the ID table features and text features to achieve fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection, and align the feature representations of different modalities at the instance-level coarse-grained hierarchy;
[0038] S04. Apply the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model to downstream tasks, and fine-tune them for power distribution fault detection.
[0039] Through the above implementation methods, first, various modal features of the ID table features and text features are selected for detection, reducing or avoiding the detection errors of a single modality. Furthermore, fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection is provided, allowing the masked features of one modality to borrow information from another modality when reconstructing information, thereby aligning features between different modalities and improving the accuracy of power distribution fault detection. Finally, jointly fine-tune the above two models on the downstream power distribution fault detection task to integrate their advantages in processing different modal features.
[0040] In some implementation manners of the present application, the ID table features are constructed through the following steps: preset a two-dimensional ID table with a dimension of R N×F The real-time operation data features of the power distribution system are presented in the form of a two-dimensional ID table, and the dimension of the two-dimensional ID table is set to R N×F , where N represents the number of timestamps and F represents the total number of power distribution fault features; each row of the two-dimensional ID table represents the instance operation data n at a timestamp, and each column represents a power distribution fault feature f;
[0041] According to the data features of the real-time operation data of the power distribution system, encode the row and column values of the two-dimensional ID table to generate the ID table features corresponding to the real-time operation data of the power distribution system. Here, the encoding is preferably one-hot encoding. This encoding process includes:
[0042]
[0043]
[0044]
[0045] where x v represents the marker indication of the feature value text v in the feature domain, x f is the one-hot representation vector of the power distribution fault feature f, is the ID table feature vector, and its dimension is R V , V is the total number of power distribution fault features, and v f is the number of features within the power distribution fault feature f, then
[0046] In some embodiments of the present application, the text feature is constructed through the following steps: According to the data characteristics of the real-time operation data of the power distribution system, in the way of a hard prompt template, and based on the text name of the power distribution fault feature and the text feature value corresponding to the power distribution fault feature in each instance input, the corresponding text feature is generated after text splicing operation. According to the real-time operation data characteristics of the power distribution system, the corresponding text feature is generated in the way of a hard prompt template, that is:
[0047]
[0048]
[0049] where k f represents the text name of the power distribution fault feature f, w f represents the text feature value corresponding to the power distribution fault feature f in each instance input, and the operator represents the splicing operation between texts.
[0050] In some embodiments of the present application, a basic power distribution system fault prediction model based on a multi-head attention mechanism for processing the ID table feature is constructed, including:
[0051] Determine the model category and model structure of the basic power distribution system fault prediction model before training. The model category is a multi-classification model, that is: the power distribution fault detection problem is modeled as a multi-classification problem, that is, the power distribution fault prediction problem. Its basic form is to perform multi-classification using data of multiple features, where y n ∈ {0, 1, 2} is the true label of the fault type of the power distribution terminal instance n. When there is no fault in the power distribution terminal of instance n, then y n = 0. When the power distribution terminal of instance n has an overheating fault, then y n = 1. When the power distribution terminal of instance n has a voltage imbalance fault, then y n = 2. The basic power distribution system fault prediction model classifies by estimating the probability of each fault type and according to the network structure of the basic power distribution system fault prediction model, which is abstracted into a three-layer network, namely an embedding layer, a feature interaction layer, and a prediction layer.
[0052] The embedding layer is configured to: convert the high-dimensional sparse ID table feature into a low-dimensional dense embedding matrix; in the power distribution fault feature f, the corresponding representation matrix is converted into an embedding vector by looking up a table. In the embedding layer, the high-dimensional sparse ID table feature vector is converted into a low-dimensional dense embedding matrix E n , and in the power distribution fault feature f, the corresponding representation matrix is converted into an embedding vector by looking up a table, which is expressed as:
[0053]
[0054] Among them, represents the one-hot vector of the eigenvalue in the power distribution fault feature f, is the feature matrix of the power distribution fault feature f, D is the embedding dimension, and the complete embedding matrix is represented as E = [e1; e2; …; e f ; …; e F ∈ R F×D .
[0055] The feature interaction layer is configured to: take the embedding matrix as the input, and output a dense representation vector for each power distribution terminal instance running data n; adopt a high-order and low-order cross model with a two-tower structure to design complex interaction operations for each feature to obtain low-order or high-order feature cross terms; in the feature interaction layer, take the embedding matrix E as the input, and output a dense representation vector v n for each power distribution terminal instance n; adopt a high-order and low-order cross model with a two-tower structure to design complex interaction operations for each feature to obtain low-order or high-order feature cross terms.
[0056] The prediction layer is configured to: after concatenating the representation vectors output by the high-order and low-order cross, input them into a neural network or a linear regression module, and then connect the Sigmoid function to adjust the range to obtain the classification probability of the final fault type; in the prediction layer, after concatenating the representation vectors output by the high-order and low-order cross, input them into a neural network or a linear regression module, and then connect the Sigmoid function to adjust the range to obtain the final user click probability
[0057] After training the basic power distribution system fault prediction model before training in an end-to-end manner using binary cross-entropy loss with training samples, the basic power distribution system fault prediction model is obtained. This binary cross-entropy loss is expressed as:
[0058]
[0059] In some embodiments of the present application, a large language model for power distribution system fault detection for processing the text features is constructed, including: converting the text features into the form of tokens based on a large language model pre-trained for power distribution fault detection; inputting the obtained tokens into an encoder network based on Transformer, and after processing, obtaining a context-consistent representation; performing power distribution fault classification by adding randomly initialized attention heads on the context-consistent representation to obtain the fault classification label corresponding to the text features. Specifically, select a large language model pre-trained for power distribution fault detection as the text features for power distribution fault detection Feature processing model; based on a pre-trained large language model, convert text features into the form of tokens; in natural language processing, tokens can be words, phrases or characters in a sentence. These tokens are the basic units for the language model to process text, and they form the basis for the model to understand and generate text. Then, use a Transformer-based encoder network to process the tokens obtained in the previous step and output a context-consistent representation w n ; perform power distribution fault classification by adding randomly initialized attention heads to w n to obtain a fault classification label
[0060] In some embodiments of the present application, construct bimodal features based on ID table features and text features to achieve fine-grained feature alignment between the power distribution system fault basic prediction model and the power distribution system fault detection large language model, and align the feature representations of different modalities at the instance-level coarse-grained level, including:
[0061] Extract a certain proportion of feature domains from the ID table features and text features respectively for masking operations to obtain the interference forms of the ID table features and text features, that is, masked ID table features and masked text features; that is: obtain the interference forms of the ID table features and text features through data masking, and denote them as and To output the masked ID features uniformly extract a certain proportion of feature domains from the original for masking operations; to output text features uniformly extract a certain proportion of feature domains from the original for masking operations.
[0062] Input the ID table features and the masked ID table features into the power distribution system fault basic prediction model for encoding, input the text features and the masked text features into the power distribution system fault detection large language model for encoding, and obtain the corresponding outputs; according to the output v n after encoding the ID table features and the output after encoding the masked text features for joint reconstruction, and according to the output w n after encoding the text features and the output after encoding the masked ID table features for joint reconstruction;
[0063] Establish coarse-grained hierarchical feature alignment between bimodals according to the instance-level comparative learning method; combine the optimization objectives in the joint reconstruction and coarse-grained hierarchical feature alignment processes, and train based on the overall loss function in the bimodal alignment pre-training stage of the power distribution fault prediction basic model and the power distribution system fault detection large language model.
[0064] Further, in some embodiments, a specific implementation of the masking operation is provided. A certain proportion of feature domains are extracted from the ID table features, and the corresponding features are replaced with an additional masking symbol, which is shared by all feature domains; the feature vectors of the ID table based on the encoding method are replaced with specific symbols to obtain the masked ID table features, and the masked ID table features include a zero vector with a dimension of v f Specifically, in the generation process of the masked ID features an additional masking symbol [MASK] is used to replace the corresponding features, which is shared by all feature domains, and the set of masked feature domain indices is denoted as the one-hot encoded ID feature vector masks the feature with the symbol 0, so is expressed as follows:
[0065]
[0066] wherein, is a zero vector with a dimension of v f Specifically, in the generation process of the masked text features, a certain proportion of feature domains are extracted from the text features, the features are replaced with specific characters, and the masked text features are obtained by using the hard prompt template method. Specifically, in the generation process of the text features
[0067] is expressed as follows: During the generation process, is expressed as follows:
[0068]
[0069] wherein, if the text features mask the features with the symbol <unknown>, then the set of masked tokens indices is denoted as
[0070] In some embodiments of the present application, the ID table features and the masked ID table features are input into the basic prediction model of the power distribution system fault for encoding, and the text features and the masked text features are input into the large language model for power distribution system fault detection for encoding to obtain the corresponding outputs, including:
[0071] The ID table features and the masked ID table features are input into the basic prediction model of the power distribution system fault, and after being encoded by the embedding layer and the feature interaction layer in the basic prediction model of the power distribution system fault, the corresponding outputs are obtained; exemplarily, the feature vector represented by the ID table features is used as the input, and the data is encoded by using the basic power distribution fault prediction model, which is expressed as:
[0072]
[0073] wherein, f CTR(·) represents the basic power distribution fault prediction model that only includes the embedding layer and the feature interaction layer. represents the output representation vector after high and low order cross, D ID represents the output dimension of the feature interaction layer.
[0074] Moreover, the text feature and the masked text feature are input into the large language model for power distribution system fault detection, and the corresponding output is obtained through the set of hidden layer states of the last layer of the large language model for power distribution system fault detection. Exemplarily, taking the feature vector of the text modality as the input and using the large language model for data encoding, it is expressed as:
[0075]
[0076] where f LLM (·) represents the pre-trained large language model, w n =[w n,l ∈R L×Dtext represents the set of hidden layer states of the last layer of the large language model, L is the number of tokens of the text feature D text represents the hidden layer output dimension of the large language model, w n,1 represents the [CLS] tokens vector of the compressed representation of the entire text input.
[0077] In some embodiments of the present application, joint reconstruction is performed according to the output after text feature encoding and the output after masked ID table feature encoding, including: inputting the output after text feature encoding and the output after masked ID table feature encoding into the self-attention mechanism module in the form of an encoded vector tuple to obtain an aggregated output; taking the encoded masked ID feature output and the text feature output w n , in the form of an encoded vector tuple as the input, designing a self-attention mechanism module and using it to aggregate and w n , and its output is expressed as:
[0078]
[0079] where K = V = w n , represents the trainable attention parameter matrix, represents the scaling factor, and the multi-head attention mechanism is integrated by h different self-attention mechanism modules, and its aggregated output process is:
[0080] u n = MultiHead(Q, K, V) = [head1||head2||…||headh W O
[0081] head i = Attention(QW i Q ,KW i K ,VW i V )
[0082] where the symbol || represents the concatenation operation between vectors, and W i Q , W i K , W i V and W O represent the parameter matrices of Q, K, V, and the aggregation representation mapping, respectively.
[0083] For each masked ID table feature, after passing through an independent multi-layer perceptron network and the Softmax function, a distribution function on the candidate features of the aggregated output is obtained; for each masked ID feature design an independent multi-layer perceptron network g f (·), and immediately follow it with a Softmax function to calculate the distribution p n,f ∈R V , expressed as:
[0084] z n,f = g f (u n ), z n,f ∈R V
[0085]
[0086] where V represents the size of the entire feature space.
[0087] For all masked features, the cross-entropy loss function of the distribution function and the ID table features is used to calculate the reconstructed features of the ID table features; the noise contrast estimation function is used to calculate the reconstruction loss based on the reconstructed features. Then, for all masked features, the cross-entropy loss function is used to calculate the ID feature reconstruction loss, expressed as:
[0088]
[0089] where represents the number of masked features, and loss BCE (·) represents the cross-entropy loss function;
[0090] The reconstruction loss is calculated using the noise contrast estimation function, expressed as:
[0091]
[0092] where δ(·) represents the Sigmoid function, t represents the positive feature exponent, and q represents the sampled noise feature exponent.
[0093] Joint reconstruction is performed based on the output after encoding the ID table features and the output after encoding the masked text features, including: concatenating the output after encoding the ID table features and the output after encoding the masked text features in the form of an encoded vector tuple to obtain a concatenated vector; specifically, the encoded ID table feature output v n and the text feature output are used as inputs in the form of an encoded vector tuple for each masked tokens feature and their concatenation operation is:
[0094]
[0095] The concatenated vector is input into the prediction layer to obtain the reconstructed feature distribution, expressed as:
[0096]
[0097] where M represents the vocabulary size, and g LLM (·) represents a two-layer perceptron module at the decoder end of the large language model;
[0098] Finally, the large language model is pre-trained using the cross-entropy loss function, expressed as:
[0099]
[0100] where is the original l-th token.
[0101] Coarse-grained hierarchical feature alignment between the two modalities is established according to the instance-level comparison learning method, including: using two independent linear layers and normalization layers to project the output of the tokens vector corresponding to the text input and the output of the ID table features into a first multi-dimensional vector and a second multi-dimensional vector; the [CLS] tokens vector represents the text input w n,1 , and using two independent linear layers and normalization layers to project the [CLS] tokens output and the ID feature output v n into d-dimensional vectors, denoted as the first multi-dimensional vector and the second multi-dimensional vector
[0102] The InfoNCE function is used to calculate the comparison loss between the first multi-dimensional vector and the second multi-dimensional vector at the instance level, which is expressed as:
[0103]
[0104]
[0105] where B represents the size of the training batch, τ represents the temperature hyperparameter, and sim(·) represents the similarity function.
[0106] The final loss is obtained based on the comparison loss, including: summing and averaging the two instance-level loss functions to obtain the final loss, which is expressed as:
[0107]
[0108] Based on the reconstruction loss and the final loss, the overall loss of the basic power distribution system fault prediction model and the large language model for power distribution system fault detection is obtained, and this overall loss is expressed as:
[0109]
[0110] Figure 2 Schematically shows the technical block diagram of the power distribution network fault identification method according to the embodiments of the present application. As Figure 2 shown, ID table features and text features are respectively extracted from the real-time operation data, namely: timestamp: 2024-07-01:20:00, terminal ID: G002,..., voltage / V: 240, fault type: overheat fault, and respectively pass through the embedding layer and the feature interaction layer in the pre-trained LLM and the basic power distribution system fault prediction model to obtain the original tokens and original features, as well as the InfoNCE instance-level contrast loss.
[0111] In some embodiments of the present application, the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model are applied to downstream tasks and fine-tuned for power distribution fault detection, including:
[0112] In the downstream power distribution system fault detection task, use the supervised click signal to jointly fine-tune the basic power distribution system fault detection model and the large language model for power distribution system fault;
[0113] Random restart linear output layers are respectively set in the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model, and respectively output the estimated fault classification probabilities, which are expressed as:
[0114]
[0115]
[0116] Among them, linear(·) represents the linear output layer of the corresponding model.
[0117] The final fault classification probability is obtained through weighted calculation based on the fault classification probabilities of the two models respectively, and is expressed as:
[0118]
[0119] Among them, λ∈[0,1] is an adjustable parameter, which is used to balance the weights of the losses between the distribution fault detection basic model and the large language model.
[0120] Compare the estimation results corresponding to the final fault classification probability with the true labels respectively, and use the cross-entropy loss function to measure the gap to fine-tune and optimize the model. This gap is expressed as:
[0121]
[0122] Figure 3 Schematically shows the fine-tuning block diagram of the downstream distribution fault detection task according to the embodiment of the present application. As Figure 3 shown, extract the ID table features and text features from the real-time operation data, namely: timestamp: 2024-07-01:20:00, terminal ID: G002,... voltage / V: 240, fault type: overheat fault, and respectively obtain and through the pre-trained LLM and the basic prediction model of the distribution system fault respectively, and thus obtain the final fault classification probability In this embodiment, in the downstream distribution fault detection task, the model fine-tuning process will greatly reduce the training time compared with the original supervised learning.
[0123] Through the above embodiments, through the joint reconstruction at the feature level between the two modalities, the model pre-training task of fine-grained feature-level interaction and alignment is realized, reducing the number of parameters and the training time. At the same time, the appropriate model can be selected according to different scenarios, which helps to enhance the generality and scalability of the framework.
[0124] Based on the same inventive concept, the present application also provides a distribution network fault identification device. Figure 4 Schematically shows the structural schematic diagram of the distribution network fault identification device according to the embodiment of the present application. As Figure 4As shown in the figure, the device includes: a feature confirmation module for determining the real-time operation data of the power distribution system for the detected modal features, where the modal features include ID table features and text features; a model construction module for respectively constructing a basic power distribution system fault prediction model based on the multi-head attention mechanism for processing the ID table features and a large language model for power distribution system fault detection for processing the text features; a feature alignment module for constructing dual-modal features according to the ID table features and text features to achieve fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection, and aligning the feature representations of different modalities at the instance-level coarse-grained level; and an application deployment module for applying the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model to downstream tasks, and fine-tuning for power distribution fault detection.
[0125] In some alternative embodiments of the present application, the ID table features are constructed through the following steps: preset a two-dimensional ID table with dimensions of R N×F , where N represents the number of timestamps and F represents the total number of power distribution fault features; each row of the two-dimensional ID table represents the instance operation data n at a timestamp, and each column represents a power distribution fault feature f; according to the data features of the real-time operation data of the power distribution system, encode the row and column values of the two-dimensional ID table to generate the ID table features corresponding to the real-time operation data of the power distribution system.
[0126] In some alternative embodiments of the present application, the text features are constructed through the following steps: according to the data features of the real-time operation data of the power distribution system, adopt the hard prompt template method, and generate the corresponding text features after text splicing according to the text name of the power distribution fault feature and the text feature value corresponding to the power distribution fault feature in each instance input.
[0127] In some alternative embodiments of the present application, a basic power distribution system fault prediction model based on a multi-head attention mechanism for processing the ID table features is constructed, including: determining the model category and model structure of the basic power distribution system fault prediction model before training, where the model category is a multi-classification model, and the model structure includes an embedding layer, a feature interaction layer, and a prediction layer; the embedding layer is configured to: convert the high-dimensional sparse ID table features into a low-dimensional dense embedding matrix; the power distribution fault feature f is converted into an embedding vector through a look-up table method by its corresponding characterization matrix; the feature interaction layer is configured to: take the embedding matrix as the input and output a dense characterization vector for each power distribution terminal instance running data n; adopt a high-low order cross model with a two-tower structure to design complex interaction operations for each feature to obtain low-order or high-order feature cross terms; the prediction layer is configured to: after splicing the representation vectors output by the high-low order cross, input them into a neural network or a linear regression module, and then adjust the range through a Sigmoid function to obtain the classification probability of the final fault type; after training the basic power distribution system fault prediction model before training with training samples in an end-to-end manner using binary cross-entropy loss, the basic power distribution system fault prediction model is obtained.
[0128] In some alternative embodiments of the present application, a large language model for power distribution system fault detection for processing the text features is constructed, including: converting the text features into tokens form based on a large language model pre-trained for power distribution fault detection; inputting the obtained tokens into an encoder network based on Transformer, and after processing, obtaining a context-consistent representation; performing power distribution fault classification by adding randomly initialized attention heads on the context-consistent representation to obtain the fault classification label corresponding to the text features.
[0129] In some alternative embodiments of the present application, a bimodal feature is constructed based on the ID table feature and the text feature to achieve fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection, and the feature representations of different modalities are aligned at the instance-level coarse-grained level, including: extracting a certain proportion of feature domains from the ID table feature and the text feature respectively for masking operations to obtain the perturbed forms of the ID table feature and the text feature, namely the masked ID table feature and the masked text feature; inputting the ID table feature and the masked ID table feature into the basic power distribution system fault prediction model for encoding, and inputting the text feature and the masked text feature into the large language model for power distribution system fault detection for encoding to obtain the corresponding outputs; performing joint reconstruction based on the output after encoding the ID table feature and the output after encoding the masked text feature, and performing joint reconstruction based on the output after encoding the text feature and the output after encoding the masked ID table feature; establishing coarse-grained level feature alignment between the two modalities according to the instance-level comparative learning method; and training based on the overall loss function in the bimodal alignment pre-training stage of the basic power distribution fault prediction model and the large language model for power distribution system fault detection by combining the optimization objectives in the joint reconstruction and the coarse-grained level feature alignment processes.
[0130] In some alternative embodiments of the present application, a certain proportion of feature domains are extracted from the ID table feature and the text feature respectively for masking operations to obtain the perturbed form of the ID table feature and the perturbed form of the text feature, including: extracting a certain proportion of feature domains from the ID table feature and replacing the corresponding features with an additional masking symbol, which is shared by all feature domains; replacing the features of the ID table based on the encoding method with specific symbols to obtain the masked ID table feature, and the masked ID table feature includes a zero vector with a dimension of v f ; and extracting a certain proportion of feature domains from the text feature and replacing the features with specific characters to obtain the masked text feature in the form of a hard prompt template.
[0131] In some alternative embodiments of the present application, the ID table feature and the masked ID table feature are input into the basic power distribution system fault prediction model for encoding, and the text feature and the masked text feature are input into the large language model for power distribution system fault detection for encoding to obtain the corresponding outputs, including: inputting the ID table feature and the masked ID table feature into the basic power distribution system fault prediction model, and obtaining the corresponding output after encoding through the embedding layer and the feature interaction layer in the basic power distribution system fault prediction model; and inputting the text feature and the masked text feature into the large language model for power distribution system fault detection, and obtaining the corresponding output through the set of hidden layer states of the last layer of the large language model for power distribution system fault detection.
[0132] In some alternative embodiments of the present application, joint reconstruction is performed based on the output encoded by the text features and the output encoded by the masked ID table features, including: inputting the output encoded by the text features and the output encoded by the masked ID table features into a self-attention mechanism module in the form of an encoded vector tuple to obtain an aggregated output; for each masked ID table feature, after passing through an independent multi-layer perceptron network and a Softmax function, obtaining a distribution function on the candidate features of the aggregated output; calculating the reconstructed features of the ID table features using the cross-entropy loss function of the distribution function and the ID table features for all masked features; calculating the reconstruction loss based on the reconstructed features using a noise contrastive estimation function.
[0133] Joint reconstruction is performed based on the output encoded by the ID table features and the output encoded by the masked text features, including: concatenating the output encoded by the ID table features and the output encoded by the masked text features in the form of an encoded vector tuple to obtain a concatenated vector; inputting the concatenated vector into the prediction layer in the basic power distribution system fault prediction model to obtain a reconstructed feature distribution function; using the cross-entropy loss function of the feature distribution function and the text features as the reconstruction loss.
[0134] In some alternative embodiments of the present application, coarse-grained hierarchical feature alignment between two modalities is established according to the instance-level comparative learning method, including: projecting the text input corresponding tokens vector output and the ID table feature output into a first multi-dimensional vector and a second multi-dimensional vector using two independent linear layers and normalization layers; calculating the comparative loss of the first multi-dimensional vector and the second multi-dimensional vector at the instance level using the InfoNCE function; obtaining the final loss according to the comparative loss; obtaining the overall loss of the basic power distribution system fault prediction model and the large language model for power distribution system fault detection based on the reconstruction loss and the final loss.
[0135] In some alternative embodiments of the present application, applying the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model to downstream tasks and fine-tuning them for power distribution fault detection, including: in the downstream power distribution system fault detection task, jointly fine-tuning the basic power distribution system fault detection model and the large language model for power distribution system fault using supervised click signals; setting randomly restarted linear output layers in the aligned large language model for power distribution system fault detection and the basic power distribution system fault prediction model respectively, and outputting the estimated fault classification probabilities respectively; obtaining the final fault classification probability through weighted calculation based on the fault classification probabilities of the two models respectively; comparing the estimated results corresponding to the final fault classification probability with the true labels respectively, and using the cross-entropy loss function to measure the gap to fine-tune and optimize the model.
[0136] The specific limitations of each functional module in the above-mentioned distribution network fault identification device can be referred to the limitations of the distribution network fault identification method in the foregoing text, which will not be elaborated here. Each module in the above system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the electronic device in the form of hardware, or stored in the memory in the electronic device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules. It also has the advantages of performing fine-grained feature-level alignment between the attention mechanism model based on the ID table feature and the large language model based on the text feature to improve the accuracy of the downstream power distribution fault detection task, and at the same time reducing the training duration of the model.
[0137] In some embodiments of the present application, an electronic device is further provided, including: at least one processor; a memory connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the at least one processor executes the foregoing distribution network fault identification method. Its internal structure diagram can be as Figure 5 shown. Figure 5 Schematically shows the internal structure diagram of the electronic device according to the embodiment of the present application. The electronic device includes a processor A01, a network interface A02, a memory (not shown in the figure), and a database (not shown in the figure) connected through a system bus. Among them, the processor A01 of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes an internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02, and a database (not shown in the figure). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A04. The network interface A02 of the electronic device is used to communicate with an external terminal through a network connection. The computer program B02, when executed by the processor A01, implements a distribution network fault identification method.
[0138] Those skilled in the art can understand that Figure 5 the structure shown in
[0139] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0140] In an implementation provided by the present application, a computer program product is provided, including a computer program which, when executed by a processor, implements the foregoing method for identifying faults in a distribution network.
[0141] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0142] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks.
[0145] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0146] The memory may include non-permanent memory in the form of computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0147] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0148] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0149] The above are only examples of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for identifying faults in a distribution network, characterized in that, The method includes: Determining real-time operation data of the distribution system for the detected modal features, where the modal features include ID table features and text features; Respectively constructing a basic distribution system fault prediction model based on a multi-head attention mechanism for processing the ID table features and a large language model for distribution system fault detection for processing the text features; Constructing dual-modal features based on the ID table features and text features to achieve fine-grained feature alignment between the basic distribution system fault prediction model and the large language model for distribution system fault detection, and aligning the feature representations of different modalities at the instance-level coarse-grained level; Applying the aligned large language model for distribution system fault detection and the basic distribution system fault prediction model to downstream tasks, and fine-tuning for distribution fault detection, including: in the downstream distribution system fault detection task, jointly fine-tuning the basic model for distribution system fault detection and the large language model for distribution system fault using supervised click signals; setting randomly restarted linear output layers in the aligned large language model for distribution system fault detection and the basic distribution system fault prediction model respectively, and respectively outputting estimated fault classification probabilities; obtaining the final fault classification probability through weighted calculation based on the fault classification probabilities of the two models respectively; comparing the estimated results corresponding to the final fault classification probability with the true labels respectively, and using the cross-entropy loss function to measure the gap to fine-tune and optimize the model.
2. The method according to claim 1, wherein The ID table features are constructed through the following steps: The preset dimension is R N×F of the two-dimensional ID table, where N represents the number of timestamps and F represents the total number of power distribution fault characteristics; each row of the two-dimensional ID table represents the instance operation data n at a timestamp, and each column represents a power distribution fault characteristic f; According to the data features of the real-time operation data of the distribution system, encoding to obtain the row and column values of the two-dimensional ID table, and generating the ID table features corresponding to the real-time operation data of the distribution system.
3. The method according to claim 1, characterized in that The text features are constructed through the following steps: According to the data features of the real-time operation data of the distribution system, adopting the hard prompt template method, and generating corresponding text features through text splicing operation according to the text name of the distribution fault feature and the text feature value corresponding to the distribution fault feature in each instance input.
4. The method according to claim 2, wherein Constructing a basic distribution system fault prediction model based on a multi-head attention mechanism for processing the ID table features, including: Determining the model category and model structure of the basic distribution system fault prediction model before training, where the model category is a multi-classification model, and the model structure includes an embedding layer, a feature interaction layer, and a prediction layer; The embedding layer is configured to: convert the high-dimensional sparse ID table features into a low-dimensional dense embedding matrix; in the distribution fault feature f, convert it into an embedding vector through a look-up table from its corresponding representation matrix; The feature interaction layer is configured to: take the embedding matrix as input, and output a dense representation vector for each distribution terminal instance operation data n; adopt a high-low order cross model with a two-tower structure to design complex interaction operations for each feature to obtain low-order or high-order feature cross terms; The prediction layer is configured to: after splicing the representation vectors output by the high-low order cross, input them into a neural network or a linear regression module, and then adjust the range through the Sigmoid function to obtain the classification probability of the final fault type; After training the pre-training distribution system fault prediction basic model with training samples in an end-to-end manner using binary cross-entropy loss, the distribution system fault basic prediction model is obtained.
5. The method according to claim 3, wherein Construct a large language model for distribution system fault detection for processing the text features, including: Convert the text features into the form of tokens based on a large language model pre-trained for distribution fault detection; Input the obtained tokens into an encoder network based on Transformer, and after processing, obtain a context-consistent representation; Perform distribution fault classification by adding randomly initialized attention heads on the context-consistent representation to obtain the fault classification label corresponding to the text features.
6. The method according to claim 1, wherein Construct dual-modal features according to the ID table features and text features to achieve fine-grained feature alignment between the distribution system fault basic prediction model and the distribution system fault detection large language model, and align the feature representations of different modalities at the instance-level coarse-grained level, including: Extract a certain proportion of feature domains from the ID table features and text features respectively for masking operations to obtain the perturbed forms of the ID table features and text features, namely masked ID table features and masked text features; Input the ID table features and masked ID table features into the distribution system fault basic prediction model for encoding, and input the text features and masked text features into the distribution system fault detection large language model for encoding to obtain the corresponding outputs; Perform joint reconstruction based on the output after encoding the ID table features and the output after encoding the masked text features, and perform joint reconstruction based on the output after encoding the text features and the output after encoding the masked ID table features; Establish coarse-grained hierarchical feature alignment between the dual-modalities according to the instance-level comparative learning method; combine the optimization objectives in the joint reconstruction and coarse-grained hierarchical feature alignment processes, and train based on the overall loss function in the dual-modal alignment pre-training stage of the distribution fault prediction basic model and the distribution system fault detection large language model.
7. The method according to claim 6, wherein Extract a certain proportion of feature domains from the ID table features and text features respectively for masking operations to obtain the perturbed form of the ID table features and the perturbed form of the text features, including: Extract a certain proportion of feature fields from the ID table features, and replace the corresponding features with an additional mask symbol, which is shared by all feature fields; replace the features in the feature vector of the ID table based on the encoding method with specific symbols to obtain the masked ID table features, and the masked ID table features include a zero vector with a dimension of v f ; and Extract a certain proportion of feature domains from the text features, replace the features with specific characters, and obtain the masked text features in the form of a hard prompt template.
8. The method according to claim 6, wherein Input the ID table features and masked ID table features into the distribution system fault basic prediction model for encoding, and input the text features and masked text features into the distribution system fault detection large language model for encoding to obtain the corresponding outputs, including: Input the ID table features and masked ID table features into the distribution system fault basic prediction model, and after encoding by the embedding layer and feature interaction layer in the distribution system fault basic prediction model, obtain the corresponding outputs; and Input the text features and masked text features into the distribution system fault detection large language model, and obtain the corresponding outputs through the set of hidden layer states of the last layer of the distribution system fault detection large language model.
9. The method according to claim 6, characterized in that, Perform joint reconstruction based on the output after encoding the text features and the output after encoding the masked ID table features, including: The output after encoding the text features and the output after encoding the mask ID table features are input into the self-attention mechanism module in the form of an encoded vector tuple to obtain an aggregated output; For each mask ID table feature, after passing through an independent multi-layer perceptron network and the Softmax function, a distribution function on the candidate features of the aggregated output is obtained; The cross-entropy loss function of the distribution function and the ID table features is used to calculate the reconstructed features of the ID table features for all mask features; The noise contrastive estimation function is used to calculate the reconstruction loss based on the reconstructed features; Joint reconstruction is performed according to the output after encoding the ID table features and the output after encoding the masked text features, including: The output after encoding the ID table features and the output after encoding the masked text features are concatenated in the form of an encoded vector tuple to obtain a concatenated vector; After the concatenated vector is input into the prediction layer in the basic power distribution system fault prediction model, a reconstructed feature distribution function is obtained; The cross-entropy loss function of the feature distribution function and the text features is used as the reconstruction loss.
10. The method according to claim 6, characterized in that, Coarse-grained hierarchical feature alignment between the two modalities is established according to the instance-level comparison learning method, including: Two independent linear layers and normalization layers are used to project the text input corresponding tokens vector output and the ID table feature output into a first multi-dimensional vector and a second multi-dimensional vector; The InfoNCE function is used to calculate the comparison loss of the first multi-dimensional vector and the second multi-dimensional vector at the instance level; The final loss is obtained according to the comparison loss; Based on the reconstruction loss and the final loss, the overall loss of the basic power distribution system fault prediction model and the large language model for power distribution system fault detection is obtained.
11. A distribution network fault identification device, characterized in that, The device includes: A feature confirmation module for determining the real-time operation data of the power distribution system for the modal features to be detected, where the modal features include ID table features and text features; A model construction module for respectively constructing a basic power distribution system fault prediction model based on the multi-head attention mechanism for processing the ID table features and a large language model for power distribution system fault detection for processing the text features; A feature alignment module for constructing bimodal features according to the ID table features and the text features to achieve fine-grained feature alignment between the basic power distribution system fault prediction model and the large language model for power distribution system fault detection, and aligning the feature representations of different modalities at the instance-level coarse-grained hierarchy; and An application deployment module, which is used to apply the aligned large language model for distribution system fault detection and the basic model for distribution system fault prediction to downstream tasks, and is fine-tuned for distribution fault detection, including: in the downstream distribution system fault detection task, jointly fine-tuning the basic model for distribution system fault detection and the large language model for distribution system fault with supervised click signals; setting randomly restarted linear output layers in the aligned large language model for distribution system fault detection and the basic model for distribution system fault prediction respectively, and outputting estimated fault classification probabilities respectively; obtaining the final fault classification probability through weighted calculation based on the fault classification probabilities of the two models respectively; comparing the estimated results corresponding to the final fault classification probability with the true labels respectively, and using the cross-entropy loss function to measure the gap to fine-tune and optimize the model.
12. An electronic device, characterized in that, Comprising: At least one processor; A memory connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the at least one processor realizes the steps of the distribution network fault identification method according to any one of claims 1 to 10 by executing the instructions stored in the memory.
13. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the distribution network fault identification method according to any one of claims 1 to 10 are realized.
14. A computer program product, comprising a computer program which, when executed by a processor, realizes the distribution network fault identification method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Large model recommendation method based on cross-modal semantic extraction
CN118709131A