Multi-mode transformer fault tracing inference method based on knowledge graph enhancement and related system
By combining the knowledge graph and the Transformer self-attention mechanism model, the problems of large errors and poor prediction results in transformer fault inference are solved, and more accurate fault cause traceability and efficient utilization of multimodal data are achieved.
Patent Information
- Application Number
- CN202510321051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
When the existing large-text graphic models process multimodal data of power equipment, there are problems such as large errors and poor prediction results. Especially in transformer fault inference, image object detection is poor, and error propagation risks are high. The differences between picture and text modalities affect model understanding and fault inference.
Using a multimodal transformer fault tracing reasoning method based on knowledge graph enhancement, the transformer image is divided into image blocks and converted into high-dimensional vectors through the Transformer self-attention mechanism model, and fuses with text description data to extract visual features and knowledge graph information, generate images embedded in high-dimensional vectors, and perform fault cause inference.
It improves the accuracy and robustness of transformer fault inference, reduces errors, enhances the model's inference ability, can make accurate judgments in complex fault modes, and has better generalization ability.
Smart Images

Figure CN120258135A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of transformer fault reasoning, and specifically relates to a multi-modal transformer fault causal reasoning method and a related system based on knowledge graph enhancement. Background Art
[0002] As the power system transforms towards intelligence and digitalization, the multimodal data of power equipment has exploded. These data include text descriptions, equipment images and other forms, providing a rich source of information for equipment status assessment, fault prediction and intelligent operation and maintenance. However, traditional data analysis methods are difficult to fully utilize the potential value of these heterogeneous data and cannot meet the increasingly complex power system management needs.
[0003] In recent years, with the rapid development of artificial intelligence technology, especially the breakthrough of big model technology, new possibilities have been provided for processing and analyzing multimodal data. Big picture and text models have shown strong capabilities in image understanding, text analysis, and cross-modal reasoning, bringing new opportunities for power equipment management. However, existing big picture and text models still have limitations in processing professional domain knowledge, making it difficult to fully utilize the professional knowledge and experience of the power industry.
[0004] As an important way to represent structured knowledge, knowledge graphs can effectively organize and express the professional knowledge, operating rules and relationships of power equipment. Combining knowledge graphs with large graph and text models is expected to break through the current technical bottleneck and achieve in-depth understanding and efficient use of multimodal data of power equipment.
[0005] At present, the research on large-scale graph and text models enhanced by knowledge graphs is still in its infancy. How to effectively inject domain knowledge into large models, how to achieve the organic integration of knowledge graphs and multimodal data, and how to improve the reasoning ability and interpretability of models are all key issues that need to be solved. At the same time, the particularity of power equipment, such as a wide variety of equipment, complex operating environment, and diverse failure modes, also brings unique challenges to the application of large-scale graph and text models enhanced by knowledge graphs.
[0006] There are still a series of problems with existing reasoning methods: The two-stage knowledge graph extraction model of target detection and relationship prediction has high computational cost and high risk of error propagation caused by target detection. Image target detection has poor adaptability to the detected target and easily introduces a large amount of irrelevant background information, which affects the subsequent relationship prediction effect. There is a large difference between image and text modalities, which affects model understanding and subsequent fault reasoning. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of existing methods such as large errors and poor prediction effects, and to provide a multi-modal transformer fault causal reasoning method and related system based on knowledge graph enhancement.
[0008] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a multi-modal transformer fault root cause reasoning method based on knowledge graph enhancement, including the following steps: Collect transformer images and text description data corresponding to the transformer images; Divide the transformer image into several image blocks, and convert the image blocks into high-dimensional image vectors; Randomly select several high-dimensional image vectors, and extract the visual features of the high-dimensional image vectors through the Transformer self-attention mechanism model; Stitch the extracted visual features with the text description data corresponding to the transformer image, and convert the stitched text into a high-dimensional image embedding vector; Randomly select several high-dimensional image embedding vectors, and extract the image knowledge graph information of the high-dimensional image embedding vectors through the Transformer self-attention mechanism model; Obtain the transformer fault cause according to the image knowledge graph information.
[0009] A further improvement of the present invention lies in that the specific method of dividing the transformer image into several image blocks and converting the image blocks into high-dimensional image vectors is as follows: Adopt Label the transformer image to form several labeled images; where is the abscissa of the center point of the rotation box, is the ordinate of the center point of the rotation box, is the width of the rotation box, is the height of the rotation box, is the rotation angle of the rotation box; Preprocess the several labeled images to obtain high-dimensional image vectors that meet the size requirements.
[0010] A further improvement of the present invention lies in that when using to label the transformer image, the width of the rotation box is the long side of the rectangle, and the height of the rotation box is the short side of the rectangle; the rotation angle of the rotation box is the included angle formed by the positive direction of the vertical axis of the image and any long side of the rectangle in the clockwise direction. , the labeling direction of the rotation box is consistent with the top orientation of the transformer. If there is partial occlusion in the transformer image, it is labeled according to the unoccluded part.
[0011] A further improvement of the present invention lies in that the specific method of preprocessing the several labeled images to obtain high-dimensional image vectors that meet the size requirements is as follows: Perform size mapping on a number of labeled images, randomly flip the labeled images after size mapping, splice and fuse all the labeled images to obtain a fused image; Normalize the fused image according to the preset mean and variance to obtain an image high-dimensional vector that meets the size requirements.
[0012] A further improvement of the present invention lies in that a number of image high-dimensional vectors are randomly selected, and the specific method for extracting the visual features of the image high-dimensional vectors through the Transformer self-attention mechanism model is as follows: Randomly select a number of image high-dimensional vectors, and extract the picture features of the selected image high-dimensional vectors; Use the Transformer self-attention mechanism model to flatten the picture features and learn the flattened image features; Calculate the relationship prediction map according to the flattened image features; Generate the visual features of the image high-dimensional vectors according to the relationship prediction map and the picture features of the image high-dimensional vectors.
[0013] A further improvement of the present invention lies in that the extracted visual features are spliced with the text description data corresponding to the transformer image, and the specific method for converting the spliced text into an image embedding high-dimensional vector is as follows: According to the text description data corresponding to the transformer image, select the visual features corresponding to the text description data corresponding to the transformer image; Splice the visual features to a high-dimensional vector of one dimension of the text description data corresponding to the transformer image converted into an image embedding vector as the image embedding high-dimensional vector.
[0014] A further improvement of the present invention lies in that a number of image embedding high-dimensional vectors are randomly selected, and the specific method for extracting the image knowledge graph information of the image embedding high-dimensional vectors through the Transformer self-attention mechanism model is as follows: Randomly select a number of image embedding high-dimensional vectors, and extract the picture features and text features of the selected image high-dimensional vectors; Use the Transformer self-attention mechanism model to calculate the relationship prediction map between the picture features and text features and the image embedding high-dimensional vectors; Generate the image knowledge graph information of the image embedding high-dimensional vectors according to the relationship prediction map between the picture features and text features and the image embedding high-dimensional vectors combined with the picture features and text features of the image high-dimensional vectors.
[0015] In a second aspect, the present invention provides a multi-modal transformer fault root cause inference system based on knowledge graph enhancement, including: A data acquisition module for acquiring transformer images and text description data corresponding to the transformer images; An image division module, configured to divide a transformer image into a number of image blocks and convert the image blocks into high-dimensional image vectors; A visual feature extraction module, configured to randomly select a number of high-dimensional image vectors and extract the visual features of the high-dimensional image vectors through a Transformer self-attention mechanism model; A data splicing module, configured to splice the extracted visual features with the text description data corresponding to the transformer image and convert the spliced text into a high-dimensional image embedding vector; An image knowledge graph information acquisition module, configured to randomly select a number of high-dimensional image embedding vectors and extract the image knowledge graph information of the high-dimensional image embedding vectors through a Transformer self-attention mechanism model; A fault root cause analysis module, configured to obtain the transformer fault cause according to the image knowledge graph information.
[0016] A further improvement of the present invention lies in that the function of the image division module is implemented by the following method: Using to label the transformer image to form a number of labeled images; wherein, is the abscissa of the center point of the rotation box, is the ordinate of the center point of the rotation box, is the width of the rotation box, is the height of the rotation box, is the rotation angle of the rotation box; preprocessing the number of labeled images to obtain high-dimensional image vectors meeting the size requirements.
[0017] A further improvement of the present invention lies in that the function of the visual feature extraction module is implemented by the following method: Randomly select a number of high-dimensional image vectors and extract the picture features of the selected high-dimensional image vectors; Use the Transformer self-attention mechanism model to flatten the picture features and learn the flattened image features; Calculate the relationship prediction graph according to the flattened image features; Generate the visual features of the high-dimensional image vectors according to the relationship prediction graph and the picture features of the high-dimensional image vectors.
[0018] A further improvement of the present invention lies in that the function of the data splicing module is implemented by the following method: According to the text description data corresponding to the transformer image, select the visual features corresponding to the text description data corresponding to the transformer image; Splice the visual features into a high-dimensional vector of one dimension of the text description data corresponding to the transformer image converted into an image embedding vector as the high-dimensional image embedding vector.
[0019] A further improvement of the present invention lies in that the function of the image knowledge graph information acquisition module is realized by the following method: Randomly select several images to embed into high-dimensional vectors, and extract the picture features and text features of the selected high-dimensional image vectors; Use the Transformer self-attention mechanism model to calculate the relationship prediction graph between the picture features and text features and the high-dimensional image vectors; According to the relationship prediction graph between the picture features and text features and the high-dimensional image vectors, and combining the picture features and text features of the high-dimensional image vectors, generate the image knowledge graph information of the high-dimensional image vectors.
[0020] In a third aspect, the present invention provides an electronic device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the multi-modal transformer fault root cause reasoning method based on knowledge graph enhancement are realized.
[0021] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the multi-modal transformer fault root cause reasoning method based on knowledge graph enhancement are realized.
[0022] Compared with the prior art, the present invention has the following beneficial effects: The present invention combines the image information of a transformer and the corresponding text description data (i.e., multi-modal data). By converting image patches into high-dimensional vectors and fusing them with text information, a more comprehensive information expression is achieved. The joint processing of image and text information enhances the inference ability of the model, making the prediction results more accurate. The present invention uses a Transformer self-attention mechanism model, which can effectively extract visual features from the high-dimensional vectors of images and extract image knowledge graph information from the high-dimensional vectors of image embeddings. The self-attention mechanism can capture richer features by focusing on the relationships in different regions, reducing the errors caused by local image features. The present invention combines the extracted visual features with the knowledge graph information to obtain the semantic information of the image through the image embedding vector, which helps to more accurately understand the causes of transformer faults. The knowledge graph can provide an inference chain for the causes of faults, reducing the possible errors and uncertainties in the inference process. The present invention combines image features, text description data, and knowledge graph information, which helps to more accurately trace the causes of faults. The inference process enhanced by the knowledge graph can associate the fault modes and specific fault features of the transformer with known knowledge, reducing misjudgments caused by over-reliance on a single data source. Through the enhancement of multi-modal information and the knowledge graph, the model of the present invention can have better robustness and generalization ability when facing diverse fault modes of different transformers. Even in the case of insufficient training data, the model can still perform inferences with the help of the existing experience in the knowledge graph. The present invention can effectively reduce the prediction errors caused by insufficient or misunderstood single-modal information. The combination of visual features and text information, as well as the introduction of the knowledge graph, ensures that the model can make more accurate judgments when facing complex fault modes, improving the problem of poor prediction effects commonly found in traditional methods. In summary, the present invention improves the accuracy and robustness of transformer fault inference through the fusion of multi-modal information, the Transformer self-attention mechanism, the combination of image embedding vectors, and the knowledge graph, effectively overcoming the problems of large errors and poor prediction effects existing in traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flowchart of the present invention; Figure 2 is a system diagram of the present invention; Figure 3 is a flowchart of Embodiment 3; Figure 4 is a system diagram of Embodiment 4. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To further understand the content of the present invention, the following provides a detailed description of the present invention in combination with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and not for limiting it.
[0025] See Figure 1 , a multi-modal transformer fault root cause reasoning method enhanced by a knowledge graph, including the following steps: S1. Collect transformer images and text description data corresponding to the transformer images.
[0026] S2. Divide the transformer image into several image blocks and convert the image blocks into high-dimensional image vectors.
[0027] S3. Randomly select several high-dimensional image vectors and extract the visual features of the high-dimensional image vectors through the Transformer self-attention mechanism model.
[0028] S4. Concatenate the extracted visual features with the text description data corresponding to the transformer image, and convert the concatenated text into a high-dimensional image embedding vector.
[0029] S5. Randomly select several high-dimensional image embedding vectors and extract the image knowledge graph information of the high-dimensional image embedding vectors through the Transformer self-attention mechanism model.
[0030] S6. Obtain the transformer fault cause according to the image knowledge graph information.
[0031] See Figure 2 , a multi-modal transformer fault root cause reasoning system enhanced by a knowledge graph, including: A data acquisition module for collecting transformer images and text description data corresponding to the transformer images; An image division module for dividing the transformer image into several image blocks and converting the image blocks into high-dimensional image vectors; A visual feature extraction module for randomly selecting several high-dimensional image vectors and extracting the visual features of the high-dimensional image vectors through the Transformer self-attention mechanism model; A data concatenation module for concatenating the extracted visual features with the text description data corresponding to the transformer image and converting the concatenated text into a high-dimensional image embedding vector; An image knowledge graph information acquisition module for randomly selecting several high-dimensional image embedding vectors and extracting the image knowledge graph information of the high-dimensional image embedding vectors through the Transformer self-attention mechanism model; A fault root cause module for obtaining the transformer fault cause according to the image knowledge graph information.
[0032] Example 1: For each input target image in this embodiment, it will first be divided into image patches of 16x16 pixel size, and each image patch will be transformed into a high-dimensional image vector through the "Token Embedding" operation. Specifically, the original image data is sliced into pixel patches of [3, 16, 16], resulting in 768 patches.
[0033] In the masking stage, 90% of the high-dimensional image vectors are randomly masked, and the remaining 10% are used as the input to the encoder. The encoder is a standard Transformer self-attention mechanism model. The main task of this model is to flatten the features of the 10% input images, learn the flattened image features, calculate the relationship prediction map based on the flattened image features, and generate the visual features of the high-dimensional image vectors according to the relationship prediction map and the image features of the high-dimensional image vectors, providing rich context information for the subsequent inference process.
[0034] For each input text description information, the knowledge graph information is concatenated behind it, and the combined text is transformed into a high-dimensional vector with the same dimension as the image embedding vector through the "Word Embedding" operation, serving as the high-dimensional image embedding vector. Then, in the masking stage, 90% of the embedding vectors are randomly masked, and the remaining 10% are used as the input to the encoder. The encoder is a standard Transformer self-attention mechanism model. The main task of this model is to calculate the relationship prediction map between the image features, text features, and the high-dimensional image embedding vector from the 10% input. According to the relationship prediction map between the image features, text features, and the high-dimensional image embedding vector, combined with the image features and text features of the high-dimensional image vectors, the image knowledge graph information of the high-dimensional image embedding vector is generated, providing rich context information for the subsequent inference process.
[0035] The above two Transformer self-attention mechanism models refer to the pre-trained model module of CogVLM, and the two Transformers share a normalization layer processing and a multi-head attention mechanism.
[0036] Embodiment 2: This embodiment mainly includes a lightweight image knowledge graph extraction method based on Transformer; a relationship label adjustment algorithm based on an adaptive smoothing mechanism and a graph-text large model module combined with knowledge graph semantic enhancement. It can be effectively applied to transformer fault diagnosis.
[0037] Collect the image data and text descriptions of the transformer, and perform bounding box annotation on the data without detection target annotation. Considering the resolution requirements of the subsequent model and the size of the detection target while keeping the original image size unchanged, scale the image with the smaller pixel size in one direction to 608 to form several image samples of specified sizes.
[0038] The bounding box annotation uses the five-parameter annotation method, that is, for any bounding box, use for annotation and saving. Among them, is the abscissa of the center point of the bounding box, is the ordinate of the center point of the bounding box, is the width of the bounding box, is the height of the bounding box, is the rotation angle of the bounding box. To ensure the uniqueness of the annotation results and no ambiguity in the training results, considering the periodicity of the angle and the replaceability of the width and height, set the following rules: (1) The width of the bounding box is the long side of the rectangle, and the height of the bounding box is the short side of the rectangle. (2) The rotation angle of the bounding box is the angle formed by the positive direction of the vertical axis of the image and any long side of the rectangle in the clockwise direction, . (3) The annotation direction of the bounding box is consistent with the top orientation of the transformer. (4) For some targets whose features cannot be reflected due to occlusion, still perform complete annotation according to their shape structure.
[0039] Next, perform preprocessing on the annotated images. This stage includes multiple steps: First, perform image size mapping on several annotated images, then perform random flipping operations, and finally perform fusion through data splicing. After preprocessing, the system normalizes the data according to the preset mean and variance. The purpose of this series of operations is to generate standardized input data with a size of . Finally, these processed data are fed into the network for training, and the entire process is carried out based on the previously prepared annotation data.
[0040] The network is divided into a lightweight image knowledge graph extraction method LWEG and a text-image large model module combined with knowledge graph semantic enhancement. The lightweight image knowledge graph extraction method consists of an object detection branch and a relationship prediction branch. The object detection branch goes through a backbone network, a Transformer encoder, a Transformer decoder, and an object detection head, and finally generates a knowledge graph that matches the object relationship network.
[0041] First, enter the object detection branch, where the backbone network is used to extract the features of the target image, and the ResNet50 network structure is adopted. This network converts the image data with an input size of to a size of The feature map. Next, it enters the Transformer encoder, and a 1x1 convolutional layer is used to flatten the feature map obtained by the backbone network into After that, the Transformer decoder converts the object input into the corresponding representation form and learns the features of the candidate targets in the input image.
[0042] While the Transformer decoder is learning the target features, it enters the relation extraction branch, which uses binary cross-entropy to calculate the relation prediction map. Subsequently, the relation prediction map and the candidate target features are effectively aggregated to guide the generation of the relation knowledge graph.
[0043] Next, the preprocessed picture data, text description data, and the knowledge graph information obtained by the lightweight image knowledge graph extraction method LWEG are input into the text-image large model module that combines knowledge graph semantic enhancement. For each target image, it is input into the image input path in the pre-trained large model; for the target text and knowledge graph information, they are first concatenated, and the combined text information is input into the text input path in the pre-trained large model. Finally, the trained large model can learn the feature information that combines knowledge graph semantic enhancement, which is beneficial to the understanding ability of the text-image large model and realizes the fault root cause reasoning of the transformer.
[0044] This method can realize the defect recognition and fault reasoning of multi-modal transformer data. The detection accuracy and AP are both above 80%, which better solves the problem of model reasoning distortion caused by the large difference between the picture and text modalities. This method can play an important role in scenarios such as substation operation and maintenance management, transformer emergency fault handling, and power supply guarantee for large industrial users, and has practical application value.
[0045] Example 3: See Figure 3 , this example consists of a target detection branch and a relation prediction branch. The target detection branch passes through the backbone network, Transformer encoder, Transformer decoder, and target detection head respectively, and finally generates a knowledge graph that matches the target relation network.
[0046] This example preprocesses the original picture, including: image size mapping, random flipping, data splicing and fusion, etc. According to the given mean and variance, the data is further normalized to obtain the input data with the size of and is sent to the Transformer self-attention mechanism model for training.
[0047] The data fed into the Transformer self-attention mechanism model enters the object detection branch. The backbone network is used to extract the features of the target image, which consists of common convolutional neural networks. In this embodiment, the ResNet50 network structure is adopted, which goes through a 7x7 convolutional layer, a 3x3 max-pooling downsampling layer, four residual block groups, an average pooling layer, and an FC fully connected layer respectively. During the entire feature extraction process, the input is of the image data, and the output is of the feature map. Among them, and are respectively set to / 32 and / 32.
[0048] The Transformer encoder is used to enhance the image representation of the multi-head attention mechanism. The feature map obtained by the backbone network is flattened into using a 1x1 convolutional layer.
[0049] The Transformer decoder converts the object input into the corresponding representation form. For each target query, a corresponding single target image is detected. N is usually set large enough to cover all the targets in the image. By alternating self-attention layers and cross-attention layers, the features of the candidate targets in the input image are learned. Among the N object queries, the N×N fully connected attention weights are learned as follows:
[0050] where represents the attention weight of the th head in the th layer. and are the query vector and the key vector respectively.
[0051] The object detection head detects N target candidates from the query vectors of the last layer . For each query vector , a linear layer is used to predict the class label of the target, and a three-layer perceptron with a ReLU activation layer is used to obtain the box coordinates .
[0052] In summary, the loss calculation of the object detection branch can be obtained as:
[0053] where represents the loss function of the class label; represents the loss function of the box coordinates; and respectively represent and parameters of the linear combination
[0054] The relation extraction branch is calculated using binary cross - entropy. To make the predicted graph close to the true triple ℇ, first encode the true triple ℇ as a one - hot true graph padded with zeros. Then use the permutation found in object detection to permute the indices of the predicted graph. Finally, through a relation extraction head, generate the final relation prediction graph . From the permuted graph , the loss of relation extraction is calculated as . To solve the problem that the poor adaptability of detected objects affects the relation prediction effect, a relation label adjustment algorithm based on an adaptive smoothing mechanism is proposed.
[0055] The relation label adjustment algorithm based on the adaptive smoothing mechanism first uses the bipartite matching cost to measure the uncertainty of each candidate object. For candidate object , we define the uncertainty as follows:
[0056] where, represents the matching cost, represents the matching cost when is perfectly matched . α is a non - negative hyperparameter representing the minimum uncertainty. Considering the uncertainty, set the value of to ( . By using the relation labels adjusted by uncertainty, the multi - task learning of object detection and relation extraction can be dynamically adjusted according to the quality of detected objects.
[0057] Combining the above object detection and relation extraction, the overall loss of the lightweight image knowledge graph extraction method is as follows:
[0058] Next, the pre - processed picture data, text description data, and the knowledge graph information obtained by the lightweight image knowledge graph extraction method LWEG are input into the text - image large model module combined with knowledge graph semantic enhancement.
[0059] Example 4: Please refer to Figure 4As shown in the figure, the present invention also provides an electronic device 100 for a knowledge graph enhanced multi-modal transformer fault root cause reasoning method; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0060] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the knowledge graph enhanced multi-modal transformer fault root cause reasoning method described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0061] The at least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0062] The memory 101 in the electronic device 100 stores multiple instructions to implement the knowledge graph enhanced multi-modal transformer fault root cause reasoning method. The processor 102 can execute the multiple instructions to achieve: Collect transformer images and text description data corresponding to the transformer images; Divide the transformer image into several image blocks and convert the image blocks into high-dimensional image vectors; Randomly select several high-dimensional image vectors and extract the visual features of the high-dimensional image vectors through the Transformer self-attention mechanism model; Concatenate the extracted visual features with the text description data corresponding to the transformer image, and convert the concatenated text into a high-dimensional image embedding vector; Randomly select several high-dimensional image embedding vectors and extract the image knowledge graph information of the high-dimensional image embedding vectors through the Transformer self-attention mechanism model; Obtain the transformer fault cause according to the image knowledge graph information.
[0063] Embodiment 5: If the modules / units integrated in the electronic device 100 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory).
[0064] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, system, or computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0066] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0067] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A multi-modal transformer fault root cause inference method enhanced by a knowledge graph, characterized in that It includes the following steps: Collect the transformer image and the text description data corresponding to the transformer image; Divide the transformer image into several image blocks and convert the image blocks into high-dimensional image vectors; Randomly select several high-dimensional image vectors and extract the visual features of the high-dimensional image vectors through the Transformer self-attention mechanism model; Concatenate the extracted visual features with the text description data corresponding to the transformer image, and convert the concatenated text into a high-dimensional image embedding vector; Randomly select several high-dimensional image embedding vectors and extract the image knowledge graph information of the high-dimensional image embedding vectors through the Transformer self-attention mechanism model; Obtain the transformer fault cause according to the image knowledge graph information.
2. The multi-modal transformer fault root cause reasoning method enhanced based on a knowledge graph according to claim 1, wherein, The specific method of dividing the transformer image into several image blocks and converting the image blocks into high-dimensional image vectors is as follows: Adopt to label the transformer image to form a number of labeled images; among them, is the abscissa of the center point of the rotation box, is the ordinate of the center point of the rotation box, is the width of the rotation box, is the height of the rotation box, is the rotation angle of the rotation box; Preprocess several labeled images to obtain high-dimensional image vectors that meet the size requirements.
3. The multi-modal transformer fault root cause inference method based on knowledge graph enhancement according to claim 2, wherein Adopt When annotating the transformer image, the width of the rotated bounding box is the long side of the rectangle, and the height of the rotated bounding box is the short side of the rectangle; the rotation angle of the rotated bounding box is the included angle formed clockwise by the positive direction of the vertical axis of the image and any long side of the rectangle. , the annotation direction of the rotated bounding box is consistent with the top orientation of the transformer. If there is partial occlusion in the transformer image, annotation shall be carried out according to the unoccluded part.
4. The multi-modal transformer fault root cause reasoning method enhanced based on a knowledge graph according to claim 2, wherein The specific method of preprocessing several labeled images to obtain high-dimensional image vectors that meet the size requirements is as follows: Perform size mapping on several labeled images, randomly flip the size-mapped labeled images, and splice and fuse all the labeled images to obtain a fused image; Normalize the fused image according to the preset mean and variance to obtain high-dimensional image vectors that meet the size requirements.
5. The method for multi-modal transformer fault root cause reasoning enhanced based on a knowledge graph according to claim 1, wherein The specific method of randomly selecting several high-dimensional image vectors and extracting the visual features of the high-dimensional image vectors through the Transformer self-attention mechanism model is as follows: Randomly select several high-dimensional image vectors and extract the picture features of the selected high-dimensional image vectors; Use the Transformer self-attention mechanism model to flatten the picture features and learn the flattened image features; Calculate the relationship prediction graph according to the flattened image features; Generate the visual features of the high-dimensional image vectors according to the relationship prediction graph and the picture features of the high-dimensional image vectors.
6. The multimodal transformer fault root cause reasoning method based on knowledge graph enhancement according to claim 1, wherein The specific method of concatenating the extracted visual features with the text description data corresponding to the transformer image and converting the concatenated text into a high-dimensional image embedding vector is as follows: According to the text description data corresponding to the transformer image, select the visual features corresponding to the text description data corresponding to the transformer image; Concatenate the visual features to a high-dimensional vector of one dimension of the text description data corresponding to the transformer image converted into an image embedding vector as the high-dimensional image embedding vector.
7. The method for multi-modal transformer fault root cause reasoning enhanced based on a knowledge graph according to claim 1, wherein The specific method of randomly selecting several high-dimensional image embedding vectors and extracting the image knowledge graph information of the high-dimensional image embedding vectors through the Transformer self-attention mechanism model is as follows: Randomly select several high-dimensional image embedding vectors and extract the picture features and text features of the selected high-dimensional image vectors; Use the Transformer self-attention mechanism model to calculate the relationship prediction graph between the picture features and text features and the high-dimensional image embedding vectors; Generate the image knowledge graph information of the high-dimensional image embedding vectors according to the relationship prediction graph between the picture features and text features and the high-dimensional image embedding vectors combined with the picture features and text features of the high-dimensional image vectors.
8. A multi-modal transformer fault root cause inference system enhanced by a knowledge graph, characterized in that, It includes: A data acquisition module, configured to acquire transformer images and text description data corresponding to the transformer images; An image division module, configured to divide the transformer image into a number of image blocks and convert the image blocks into high-dimensional image vectors; A visual feature extraction module, configured to randomly select a number of high-dimensional image vectors and extract the visual features of the high-dimensional image vectors through a Transformer self-attention mechanism model; A data splicing module, configured to splice the extracted visual features with the text description data corresponding to the transformer image and convert the spliced text into a high-dimensional image embedding vector; An image knowledge graph information acquisition module, configured to randomly select a number of high-dimensional image embedding vectors and extract the image knowledge graph information of the high-dimensional image embedding vectors through a Transformer self-attention mechanism model; A fault root cause analysis module, configured to obtain the transformer fault cause according to the image knowledge graph information.
9. The multi-modal transformer fault root cause inference system enhanced based on a knowledge graph according to claim 8, characterized in that The function of the image division module is implemented by the following method: Adopt to label the transformer image to form a number of labeled images; among them, is the abscissa of the center point of the rotation box, is the ordinate of the center point of the rotation box, is the width of the rotation box, is the height of the rotation box, is the rotation angle of the rotation box; Preprocess a number of labeled images to obtain high-dimensional image vectors that meet the size requirements.
10. The multi-modal transformer fault root cause inference system enhanced based on a knowledge graph according to claim 8, characterized in that The function of the visual feature extraction module is implemented by the following method: Randomly select a number of high-dimensional image vectors and extract the picture features of the selected high-dimensional image vectors; Use a Transformer self-attention mechanism model to flatten the picture features and learn the flattened image features; Calculate a relationship prediction graph according to the flattened image features; Generate the visual features of the high-dimensional image vectors according to the relationship prediction graph and the picture features of the high-dimensional image vectors.
11. The multi-modal transformer fault root cause inference system enhanced based on a knowledge graph according to claim 8, wherein The function of the data splicing module is implemented by the following method: According to the text description data corresponding to the transformer image, select the visual features corresponding to the text description data corresponding to the transformer image; Splice the visual features to a high-dimensional vector of one dimension of the text description data corresponding to the transformer image to be converted into an image embedding vector, as the high-dimensional image embedding vector.
12. The multi-modal transformer fault root cause inference system enhanced based on a knowledge graph according to claim 8, wherein The function of the image knowledge graph information acquisition module is implemented by the following method: Randomly select a number of high-dimensional image embedding vectors and extract the picture features and text features of the selected high-dimensional image embedding vectors; Use a Transformer self-attention mechanism model to calculate the relationship prediction graph between the picture features and text features and the high-dimensional image embedding vectors; Generate the image knowledge graph information of the high-dimensional image embedding vectors according to the relationship prediction graph between the picture features and text features and the high-dimensional image embedding vectors and combine the picture features and text features of the high-dimensional image vectors.
13. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the knowledge graph enhanced multi-modal transformer fault root cause analysis and reasoning method according to any one of claims 1 to 7 are implemented.
14. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the steps of the knowledge graph enhanced multi-modal transformer fault root cause analysis and reasoning method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Forging press fault knowledge graph construction method of self-attention mechanism
CN120494816A
Fault diagnosis method and device of transformer, terminal equipment and storage medium
CN120951228A