Complex equipment fault diagnosis method and system based on multi-modal knowledge graph
By constructing a multimodal knowledge graph and introducing graph neural network coding and small sample learning strategies, the shortcomings of multi-source heterogeneous data fusion, fault mechanism knowledge modeling and inference capabilities, and small sample fault recognition in complex equipment fault diagnosis are solved, and high-precision and interpretable predictive fault diagnosis and dynamic fault warning are achieved.
Patent Information
- Application Number
- CN202510685574.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing complex equipment fault diagnosis methods have shortcomings in multi-source heterogeneous data fusion, fault mechanism knowledge modeling and inference capabilities, and small sample fault identification, making it difficult to achieve high-precision and interpretable predictive fault diagnosis.
Using a method based on multimodal knowledge graph, multimodal features are mapped into entity-relationship-attribute triplets by unified semantic ontology, a health knowledge graph of complex equipment is constructed, and graph neural network coding and small sample learning strategies are introduced to generate semantic enhanced fault prototypes to achieve efficient identification and positioning of new fault types.
It significantly improves the comprehensiveness and accuracy of the extraction of fault characteristics of complex equipment, can quickly adapt to new fault categories under the conditions of very few labeled samples, overcomes the problems of data scarcity and category imbalance, and realizes dynamic perception and trend prediction of the equipment's health status, and warns of potential fault risks in advance.
Smart Images

Figure CN120217264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of equipment fault diagnosis, and in particular to a complex equipment fault diagnosis method and system based on a multi-modal knowledge graph. Background Art
[0002] In order to ensure the safe and stable operation of the power system, how to early detect and accurately warn of potential faults such as partial discharge of transformer windings, abnormal vibration of generators, wear of breaker contacts, or loosening of transmission line conductor bolts has become a research hotspot and difficulty. The traditional regular maintenance and after-fact maintenance modes are difficult to balance the two requirements of high reliability and low operation and maintenance costs due to their fixed cycles and lagged responses. Based on this, predictive fault diagnosis has emerged, realizing the transformation from "passive maintenance" to "active prevention", and significantly improving equipment availability and maintenance efficiency.
[0003] The existing predictive fault diagnosis methods can be mainly divided into three categories: one is diagnosis based on thresholds and rules, which realizes real-time alarm by artificially setting the upper and lower limits of indicators such as vibration, temperature, and current. Although it is simple and efficient, it is difficult to cope with working condition changes and new faults; the second is diagnosis based on physical models, which identifies abnormalities by comparing residuals or parameter estimates through mathematical or simulation models. Although it has mechanistic interpretability, it is easily affected by model deviations and inaccurate parameters; the third is data-driven diagnosis, which uses machine learning or deep learning to train large-scale historical data, can mine complex features but relies on a large number of labeled samples, has insufficient generalization ability for new faults and small-sample scenarios, and multi-modal data such as vibration, images, and texts are often processed separately, unable to fully utilize potential causal chains and cross-modal associations. The above methods are difficult to dynamically describe the causal relationship between multi-variables of equipment in practical applications, are also prone to overfitting or failure when samples are extremely scarce or new faults occur, and single-modal analysis means are also difficult to meet the requirements of panoramic perception of "environment - operation - fault" of power equipment.
[0004] In addition, the current fault diagnosis methods for complex equipment mainly rely on single data source modeling, feature engineering, or deep learning techniques trained based on large-scale samples, and have the following significant defects: (1) Insufficient handling of data heterogeneity, making it difficult to fuse multi-modal information such as sensor data, text records, and operation and maintenance knowledge, resulting in one-sided diagnostic results and limited accuracy; (2) Weak knowledge expression and reasoning capabilities, and existing methods are difficult to effectively model the internal structure, functional relationships, and fault evolution mechanisms of equipment, lacking interpretability and reasoning depth; (3) Insufficient ability to identify small samples and early faults. Especially in practical applications, a small number of noisy samples often lead to a significant decline in model performance, affecting the stability and reliability of the system. Therefore, there is an urgent need to develop a small-sample complex equipment fault diagnosis system based on a multi-modal knowledge graph, which fully fuses multi-source heterogeneous data, constructs an intelligent diagnosis system that combines equipment mechanism knowledge and data-driven methods, and realizes high-precision and interpretable predictive fault diagnosis of complex equipment. Summary of the Invention
[0005] To solve the problems of insufficient fusion of multi-source heterogeneous data, weak fault mechanism knowledge modeling and reasoning capabilities, and difficult identification of small-sample faults in the process of predictive fault diagnosis of complex equipment, the present invention provides a complex equipment fault diagnosis method and system based on a multi-modal knowledge graph. Supported by multi-source heterogeneous data such as vibration signals, operation and maintenance texts, and images, multi-modal features are mapped into entity-relationship-attribute triples through a unified semantic ontology to construct an equipment health knowledge graph; and a graph neural network encoding and small-sample learning strategy are introduced. Under the guidance of the structured semantics of the knowledge graph, a semantically enhanced fault prototype is generated to achieve efficient identification and location of new fault types.
[0006] In the first aspect, a complex equipment fault diagnosis method based on a multi-modal knowledge graph provided by the present invention adopts the following technical solutions: A complex equipment fault diagnosis method based on a multi-modal knowledge graph includes: Obtain equipment operation data; Extract multi-modal features from the obtained equipment operation data respectively; Construct a knowledge graph based on the extracted multi-modal features; Enhance the multi-modal features of the knowledge graph through cross-modal contrast learning; Define a small-sample task to initialize the features of the enhanced knowledge graph; Extract subgraphs and perform Mamba enhancement on the knowledge graph after feature initialization; Use a graph neural network to classify faults based on the knowledge graph.
[0007] Further, the multi-modal feature extraction of the obtained equipment operation data is performed, including the original vibration signal data collected by arranging high-precision sensors, and setting the vibration signal data collected by the sensors as X v , using 1D-CNN to extract local features; using the BERT pre-trained language model to perform entity recognition on operation and maintenance logs and operation records, complete the recognition of fault entities, and output corresponding semantic embedding vectors; and using the singular value decomposition (SVD) method to perform dimensionality reduction processing on the image matrix, and input the dimensionally reduced image into a standard CNN network to obtain image feature representations h i .
[0008] Further, the construction of the knowledge graph based on the extracted multi-modal features includes, for the vibration signal features, constructing a feature-fault template dictionary, and using the template matching method and the spectrum analysis method to identify and classify the extracted signal features; for the text data features, for the semantic embedding vectors and entity recognition results extracted by the BERT model in the operation and maintenance logs and operation information, further introduce natural language generation (NLG) to summarize the sentence-level information and extract the key semantic content related to diagnosis; for the image data features, the high-dimensional image features extracted by CNN need to be semantically annotated in combination with computer vision
[0009] Further, the multi-modal feature enhancement of the knowledge graph through cross-modal contrast learning includes inter-modal contrast learning, performing positive and negative contrasts on sample pairs of different modalities, and completing cross-modal feature alignment. Among them, for the different modal feature dimensions of the three modalities, project them into a unified-dimensional contrast learning space. In inter-modal contrast learning, further strengthen the discriminability within each individual modality, quantify the attribute similarity of each sample within the same modality, and generate positive and negative pairs based on this; perform context mask filling for the text data features; add Gaussian blur to the image features, convolve the image with a two-dimensional Gaussian kernel to obtain a blurred enhanced image
[0010] Further, the multi-modal feature enhancement of the knowledge graph through cross-modal contrast learning also includes constructing positive and negative sample pairs within the modality based on the enhanced views of the samples, optimizing the feature distribution structure within the modality, defining the positive sample pair as the pairing between two views of the same sample, and the negative sample pair as the pairing between the sample view and other sample views. To optimize the aggregation and separation effects of the features within the modality, the intra-modal contrast learning uses an improved InfoNCE loss function to maximize the similarity of the positive sample pairs and minimize the similarity of the negative sample pairs. For modality m, its loss form is expressed as: where represents thej The representation vector of a sample after passing through the projection network under the augmentation operation, represents the j representation vector of the sample without augmentation operation after passing through the projection network. τ is the temperature hyperparameter that controls the smoothness of the contrast distribution.
[0011] Furthermore, the definition of the few-shot task and the feature initialization of the augmented knowledge graph include constructing a support set S as a small number of samples for each type of fault in the current task for the model to learn the current task; a query set Q is used to evaluate the learning effect of the model on this task; the knowledge graph is used for structured screening of fault categories. Among them, graph structure metric indicators, eigenvector centrality, and path length in graph theory are used to evaluate the fault nodes in the graph. For each fault category, a third-order neighbor sampling strategy is executed from the knowledge graph, that is, starting from the current node, traversing the graph outward to the third hop, and selecting entity nodes closely related to this category to form a subgraph structure containing semantic neighbors.
[0012] Furthermore, the subgraph extraction and Mamba augmentation of the knowledge graph after feature initialization include introducing a structure metric to dynamically control the sampling hop number S hop for each knowledge graph entity corresponding to a sample e i and dynamically adjusting the neighborhood sampling hop number S according to the topological importance of the node hop and extracting its local neighborhood subgraph G i : Among them, Betweeness represents the betweenness centrality of the node, Degree represents the degree of the node, the parameters α and β are used to weigh the influence of betweenness centrality and the degree of the node, dist represents the path hop number, V i represents the node set, E i represents the edge set.
[0013] Furthermore, the subgraph extraction and Mamba augmentation of the knowledge graph after feature initialization also include introducing the selective state space model Mamba to construct a multi-level and adaptive feature interaction architecture to achieve dynamic parameter adjustment depending on the input. Among them, the core formula of MambaBlock is: Among them, represents the hidden state vector at time step t, represents the gating function, represents the matrix exponential function, A is a learnable state transition matrix, B is an input projection matrix, and I is the identity matrix. represents the input node features. represents the high-order neighbor node features of node v.
[0014] Furthermore, the use of a graph neural network for fault classification based on a knowledge graph includes, in a small sample scenario, compressing the support sample vectors of the same category into a prototype vector, and comparing the similarity between the query sample and the prototypes of each category through a distance metric method to determine its category. Among them, the query test sample x q The distance between and the prototypes of each category is mapped into a probability distribution to quantify the probability that the sample belongs to each fault category: where k is the fault category, and h q represents the vector representation of the query test sample, and d represents the cosine similarity metric.
[0015] In a second aspect, a complex equipment fault diagnosis system based on a multi-modal knowledge graph includes: A data acquisition module configured to acquire equipment operation data; A feature module configured to perform multi-modal feature extraction on the acquired equipment operation data respectively; A graph module configured to construct a knowledge graph based on the extracted multi-modal features; An enhancement module configured to perform multi-modal feature enhancement on the knowledge graph through cross-modal contrast learning; A small sample module configured to define a small sample task and initialize the features of the enhanced knowledge graph; A classification module configured to perform subgraph extraction and Mamba enhancement on the knowledge graph after feature initialization; use a graph neural network to perform fault classification based on the knowledge graph.
[0016] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the described method for complex equipment fault diagnosis based on a multi-modal knowledge graph.
[0017] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the described method for complex equipment fault diagnosis based on a multi-modal knowledge graph.
[0018] In summary, the present invention has the following beneficial technical effects: Compared with the prior art, a complex equipment predictive fault diagnosis method and system based on a multimodal knowledge graph according to the present invention has the following beneficial effects: By constructing a multimodal knowledge graph and integrating vibration signals, equipment logs, and image features, a unified semantic representation of fault information is achieved, significantly improving the comprehensiveness and accuracy of complex equipment fault feature extraction; Using knowledge graph-guided few-shot task construction and prototype network metric learning, the system can still quickly adapt to new fault categories under the condition of extremely few labeled samples, effectively overcoming the problems of data scarcity and class imbalance; At the same time, combining the causal association relationships contained in the knowledge graph, the system can achieve dynamic perception and trend prediction of the equipment health status, and early warning of potential fault risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 FIG. is a schematic diagram of a complex equipment fault diagnosis method based on a multimodal knowledge graph according to Embodiment 1 of the present invention.
[0020] Figure 2 FIG. is a comparison chart of ACC and F1 of each model in Embodiment 1 of the present invention.
[0021] Figure 3 FIG. is a radar chart of different models in Embodiment 1 of the present invention under six indicators.
[0022] Figure 4 FIG. is a schematic diagram of the comparison result warning of the model robustness in Embodiment 1 of the present invention.
[0023] Figure 5 FIG. is a schematic diagram of the comparison result of the lead distribution in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The present invention will be further described in detail below with reference to the accompanying drawings.
[0025] Embodiment 1 Referring to Figure 1 , a complex equipment fault diagnosis method based on a multimodal knowledge graph according to this embodiment includes: Obtain equipment operation data; Respectively perform multimodal feature extraction on the obtained equipment operation data; Construct a knowledge graph based on the extracted multimodal features; Perform multimodal feature enhancement on the knowledge graph through cross-modal contrast learning; Define a few-shot task and initialize the features of the enhanced knowledge graph; Extract subgraphs and perform Mamba enhancement on the knowledge graph after feature initialization; Use a graph neural network to classify faults based on the knowledge graph.
[0026] Specifically: S1. Multimodal feature extraction module During the operation of complex equipment, various types of operation data are often generated, including structural vibration, thermal imaging images, environmental parameters, operation records, and operation and maintenance logs, etc. These data have different modal characteristics, each reflecting the operation state of the equipment in a certain dimension. The information of a single modality often has limitations and cannot comprehensively characterize the health state of the equipment under actual working conditions. The core goal of the multimodal feature extraction module is to make full use of the data from multiple sensing channels, extract the significant features of each data modality through a deep learning model, and achieve a unified feature representation, providing a solid foundation for subsequent knowledge graph construction and reasoning tasks. It specifically includes the following three parts: 1) Vibration signal feature extraction. During the operation of the equipment, the vibration signal is one of the most commonly used and directly important physical quantities reflecting the mechanical state. The original vibration signal data collected by deploying high-precision sensors is set as X v , and 1D-CNN is used to extract local features: Among them, k is the convolution kernel size, d is the hidden layer dimension, W v is weight parameters. Transformer global dependency modeling is used to capture long-range periodicity:
[0027] 2) Text feature extraction. Operation and maintenance logs and operation records are text information manually recorded by front-line maintenance personnel during equipment maintenance, containing a large amount of semantic descriptions related to the equipment operation state, such as fault phenomena, maintenance measures, affected components, etc. This type of text data usually has characteristics such as being unstructured, having a high redundancy degree, and having non-standard semantic expressions. The operation and maintenance log data collected is set as X t , and the BERT pre-trained language model is used to perform entity recognition on operation and maintenance logs and operation records, complete the recognition of fault entities (such as equipment components, fault phenomena, etc.), and output the corresponding semantic embedding vectors.
[0028] Among them, CLS is a special marker used to mark the entire sequence.
[0029] Infrared thermal imaging image feature extraction. Infrared thermal imaging, as a non-contact monitoring method, can obtain the surface temperature distribution image of the equipment in real time, revealing fault precursor features such as temperature rise and hot spot concentration. The image data collected by the infrared thermal imaging sensor is set as Xi , since the thermal imaging image is essentially a grayscale image or a heat distribution map, directly inputting it into a convolutional neural network is vulnerable to noise and redundant information. Therefore, the singular value decomposition (SVD) method is first used to reduce the dimension of the image matrix: Among them, represents the approximate representation of the image using the matrix after SVD dimensionality reduction. U is the left singular vector matrix, Σ is the singular value matrix, and V is the right singular vector matrix. The dimensionality-reduced image is input into the standard CNN network to obtain the image feature representation h i : Among them, represents the approximate representation of the image using the matrix after SVD dimensionality reduction.
[0030] S2. Knowledge graph construction module, Semantically understand and summarize the deep features of multi-source heterogeneous data such as vibration signals, texts, and images obtained from the first module, map them into entities and relationships, and construct a knowledge graph that can express the equipment structure, operating status, and fault factors.
[0031] 1) For the vibration signal features, after the joint modeling of 1D-CNN and Transformer for the vibration signal, high-dimensional time-series feature representations have been obtained. In order to construct the data structure centered on the "entity-relationship-entity" triple in the knowledge graph, these feature vectors need to be further semantically interpreted.
[0032] Specifically, construct a feature-fault template dictionary, and use the template matching method and spectral analysis method to identify and classify the extracted signal features. For example, judge whether there are amplitude anomalies, harmonic peaks, or energy concentration changes at specific frequencies through spectral features. Match these features with typical fault modes to generate standardized description texts, such as: "There is an imbalance in the main shaft", "Inner ring damage of the bearing is detected", "There may be structural looseness". These fault descriptions will serve as the primary semantic entity "fault mode" nodes in the knowledge graph, and further associate the corresponding "equipment components" and "vibration characteristics" attributes. The formula of the template matching method is as follows: Among them, h v represents the vibration signal feature, p i represents the template feature, l i represents the equipment description label, and || || represents the L2 norm.
[0033] 2) For the characteristics of text data, for the semantic embedding vectors and entity recognition results extracted by the BERT model from operation and maintenance logs and operation information, the natural language generation (NLG) technology is further introduced to summarize the sentence-level information and extract the key semantic content related to diagnosis. For example, after identifying the original text fragments such as "the pump body vibrates violently", "abnormal lubrication", and "the bearing has been replaced", the NLG model (such as T5, GPT, etc.) is used to reconstruct and summarize them to obtain structured semantic sentences: "The pump body vibrates violently, presumably due to bearing wear caused by insufficient lubrication", "Regular maintenance records indicate that there is a trend of repeated failures in this component". These contents can be further abstracted into multiple entity nodes in the knowledge graph (such as "pump body", "bearing wear", "insufficient lubrication"), and are connected through relationship types such as "causal relationship".
[0034] In addition, the attribute fields such as "fault type", "severity", and "fault time" can also be extracted to form entities, enhancing the expressive ability of the graph.
[0035] 3) For the characteristics of image data, the high-dimensional image features extracted by CNN need to be semantically annotated in combination with computer vision technologies (such as segmentation results, defect localization, etc.). For example, the regional distribution of temperature hotspots in an infrared image can generate the following description: "There are local overheating areas on the surface of the motor, and the central temperature exceeds 80°C". Such descriptions are transformed into text through rule templates or language generation mechanisms, and then introduced into the graph as the "image diagnosis result" entity to improve the retrievability and inferability of image information, and describe visual information such as temperature anomalies and structural damages. The formula for identifying abnormal temperature hotspot distribution is as follows: where, h i represents the image features extracted by CNN, T represents the temperature matrix, Z represents the composite representation of temperature and semantics, S represents the probability that the calculation area is an abnormal hotspot, and descriptive information is generated using the rule template in combination with coordinates x, y.
[0036] 4) Based on the multi-modal text features, the basic elements of the knowledge graph are defined, including entities (Entity), attributes (Attribute), and relations (Relation), clarifying various equipment components, failure modes, characteristic indicators, and their mutual relations.
[0037] S3. Co-modal contrast learning feature enhancement module, The purpose of the co-modal contrastive learning feature enhancement module is to fully explore the associations and differences between vibration signal features, text data features, and image information features through cross-modal and intra-modal contrastive learning, map the complementary features of each modality to the same representation space, and inject richer and more discriminative vector representations into the nodes of the knowledge graph.
[0038] 1) Inter-modal contrastive learning: Perform positive and negative contrasts on sample pairs of different modalities to complete cross-modal feature alignment. First, since the feature dimensions of the three modalities are different, they need to be projected and mapped into a contrastive learning space of a unified dimension d: Among them, m represents the modality index, taking values of v, t, and i, which represent the three modalities of vibration, text, and image respectively. represents the pre-coded feature vector of sample k in modality m, w represents the linear transformation matrix, and b represents the bias vector. For the convenience of calculating the cosine similarity, L2 normalization is performed on the projection vectors of each modality: Among them, represents the unnormalized vector representation obtained by mapping sample k in modality m, and m represents the modality index.
[0039] Construct positive and negative sample pairs. A training batch contains N nodes, and each node corresponds to three modalities {v k , t k , i k}. Positive sample pairs include cross-modal combinations under the same node Negative sample pairs include pairs between different nodes in the same modality, or pairs between different nodes across modalities. For example . In the negative sample set, select the pair that is most similar (with the largest dot product) to the positive sample but has a different label to increase the training difficulty. For any pair of modalities (m, n) ∈ {(v, t), (v, u), (t, u)}, define the InfoNCE loss as: Among them, sim represents the cosine similarity, N represents the number of samples, represents the temperature parameter, which controls the sharpness of the similarity distribution.
[0040] Sum up the losses of the three pairs of modalities to obtain the overall cross-modal loss :
[0041] 2) Network pruning, while ensuring that the performance of contrastive learning does not decrease significantly, reduces the number of model parameters and speeds up the inference speed, making the subsequent application on large-scale knowledge graphs more efficient. Without significantly sacrificing the cross-modal contrastive learning effect, compress the parameter scale of the projection head as much as possible and improve the inference efficiency. First, for each weight w calculate the importance score s(w), and obtain the importance score s(w) using a gradient sensitivity-based metric: where is the cross-modal InfoNCE loss. After obtaining all s(w), determine the threshold τ globally according to the target pruning ratio p%, and mark the p% of the weights with the lowest scores as redundant. Specifically, construct a binary mask: and update the weight parameters to the sparse form .
[0042] To avoid a significant drop in performance caused by large-scale pruning at once, it is recommended to adopt a progressive pruning strategy: allocate the total pruning ratio p max to T stages, and calculate the immediate pruning rate at each step according to a linear or cosine schedule.
[0043] After pruning, quantify the pruning effect by calculating the model sparsity.
[0044]
[0045] 3) Inter-modal contrastive learning, further strengthen the discriminability within each individual modality, such as different fault categories within the vibration signal, different fault modes in the text description, and different fault locations in the image. Quantify the attribute similarity of each sample within the same modality, and generate positive and negative pairs based on this. Define the cosine similarity between any two samples j and k under modality m as the attribute similarity: Adopt three data augmentation operations for the three modalities. Add signal translation to the vibration signal feature h v , where Δt represents the time offset of the translation, U(-δ,δ) indicates that the offset Δt is randomly sampled from a uniform distribution in the interval [-δ, δ]. δ is the maximum allowable offset, which controls the perturbation intensity.
[0046] Perform context mask filling on the text data features, that is, randomly mask some features in the input text data features and use the language model to fill them. For the image features, add Gaussian blur to the image h i Convolve with a two-dimensional Gaussian kernel to obtain a blurred and enhanced image: Among them, represents the Gaussian kernel function. The values of the image coordinate points (x, y) determine the weight distribution of the blur. σ controls the standard deviation of the Gaussian kernel and controls the blur intensity.
[0047] Based on the enhanced views of the samples as the comparison basis, construct positive and negative sample pairs within the modality to optimize the feature distribution structure within the modality. Define the positive sample pair as the pairing between two views of the same sample, and the negative sample pair as the pairing between the sample view and the views of other samples. To optimize the aggregation and separation effects of the features within the modality, contrastive learning within the modality uses an improved InfoNCE loss function to maximize the similarity of the positive sample pairs and minimize the similarity of the negative sample pairs. For modality m, its loss form is as follows: Among them, represents the representation vector of the j th sample after passing through the projection network under the enhancement operation, represents the representation vector of the j th sample after passing through the projection network without the enhancement operation. τ is the temperature hyperparameter that controls the smoothness of the contrastive distribution. Finally, the within-modality losses of the three modalities are weighted and combined into the overall loss: Among them, represents the trainable weighting coefficient.
[0048] Jointly optimize the cross-modal and within-modal contrast losses to achieve multi-modal collaboration and class balance: Among them, represents the trainable weighting coefficient.
[0049] Fine-tuning and knowledge graph fusion, inject the enhanced multi-modal features into the knowledge graph nodes. Specifically, set the fault instance nodes in the knowledge graph and connect them with the structure nodes newly added to the graph under the new sample conditions (such as equipment components, fault descriptions), and use the weighted fusion of the multi-modal features as the features of the fault instance nodes.
[0050] S4. Small sample task construction module, In actual production, there may be sufficient data for common fault types, but there are often only a very small number of samples for rare and sudden faults. The preparation for small-sample tasks can be achieved by uniformly or based on structural importance extracting multiple fault categories from the knowledge graph, forcing the model to learn to distinguish each fault category using limited samples, thereby improving the diagnostic ability for rare faults.
[0051] 1) Define a small-sample task T. Small-sample learning is based on the idea of "meta-learning". Instead of directly learning the classification model, it learns how to learn classification tasks. Each training is regarded as a task, which contains a support set S (Support Set) and a query set Q (Query Set). The support set S serves as a small number of samples for each type of fault in the current task, used for the model to "learn" the current task. The query set Q is used to evaluate the learning effect of the model on this task. Let the current task be an N-way K-shot learning task, that is, there are N fault categories, each with K samples in the support set and M samples in the query set. The task form is as follows: where N represents the number of categories of the task, K represents the number of support samples for each category, M represents the number of test query samples for each category, x represents the feature, y represents the fault category.
[0052] 2) Use the constructed multi-modal equipment knowledge graph to structurally screen the fault categories to enhance the representativeness and generalization of the training tasks. Specifically, adopt graph structure measurement indicators in graph theory (such as Degree Centrality, Eigenvector Centrality, Shortest Path Length) to evaluate the fault nodes in the graph: high-centrality nodes indicate being associated with multiple entities, representing typical and common faults; low-centrality or marginal nodes may represent rare or long-tail faults. Select N fault categories distributed in different structural positions C , which cover both the backbone knowledge in the core area of the graph and retain the long-tail information in the marginal area, forming a structurally balanced small-sample task configuration.
[0053]
[0054] 3) To improve the semantic integrity and context expression ability of each type of sample, for each fault category C i, perform a third-order neighbor sampling strategy from the knowledge graph, that is, starting from the current node, traverse the graph outward to the third hop, and select entity nodes closely related to this category (such as influencing components, performance symptoms, historical maintenance records, etc.) to form a subgraph structure containing semantic neighbors. The semantic texts of these neighbor nodes are encoded by the pre-trained language model SBERT (Sentence-BERT) to generate the basic feature vectors of each sample. After concatenating the encoded feature vectors with the feature vectors of the fault instance nodes, the feature vectors of each sample are generated. x i,j , complete the preparation for the few-shot task: S5. Subgraph Extraction and Encoding of Graph Neural Network Enhanced by Mamba Using the constructed knowledge graph, extract the structural information around the fault entity corresponding to each sample and extract high-dimensional features through a graph neural network (GNN) to obtain a sample representation enhanced by "structure + semantics".
[0055] 1) Dynamic Neighborhood Subgraph Extraction, Adaptive Structure Sensing Based on Node Importance. To effectively capture the structural context information around the fault entity, extract its local subgraph from the knowledge graph G i =( V i , E i ), where V i is the node set, E i is the edge set. Considering the differences in the structural roles of different fault nodes in the graph, such as core hub nodes and edge isolated nodes, the required neighborhood depths for perception are different. Therefore, introduce a structural index to dynamically control the sampling hop count S hop . For the knowledge graph entity corresponding to each sample e i , dynamically adjust the neighborhood sampling hop count S hop according to the topological importance of the node, and extract its local neighborhood subgraph G i : Among them, Betweeness represents the betweenness centrality of the node, Degree represents the degree of the node, and the parameters α and β are used to weigh the influence of betweenness centrality and the degree of the node. dist represents the path hop count, V i represents the node set, E i represents the edge set.
[0056] 2) Graph neural network encoding, Mamba-enhanced gated state space propagation. To achieve effective propagation and feature fusion of semantic information in subgraphs, based on the traditional gated recurrent unit (GRU) and residual connections, a selective state space model (Mamba) is introduced to construct a multi-level and adaptive feature interaction architecture, realizing input-dependent dynamic parameter adjustment. Node v At the l +1 layer, the state update formula is: Among them, represents layer normalization, which accelerates training and improves stability. MPNN represents the message passing neural network, represents the set of neighbor nodes of node v, represents the hidden state at the current moment. MambaBlock represents the Mamba architecture model, represents the gating function, which realizes the fine-grained fusion of the current input and the previous hidden state . represents the set of high-order neighbor nodes of node v. As a key link in the graph neural network encoding process, the Mamba model solves the core bottleneck of traditional graph neural networks in long-range dependencies. The core formula of MambaBlock is: Among them, represents the hidden state vector at time step t, represents the gating function, represents the matrix exponential function. A is a learnable state transition matrix, B is an input projection matrix, I is the identity matrix, represents the input node features, represents the high-order neighbor node features of node v.
[0057] 3) Graph-level representation generation, readout operation based on structural aggregation. After multiple rounds of propagation by the L-layer graph neural network, the final embedding of each node is obtained. It is necessary to aggregate the entire subgraph representation into a single vector for input to downstream tasks. For this purpose, the structured graph readout function DiffPool is adopted to aggregate nodes into supernodes or graph overall representations in a differentiable manner. The aggregated vector of all node vectors at the L layer is the sample representation: Among them, DiffPool is the readout function, which has structural adaptability and can maintain the topological characteristics and semantic composition of the subgraph during the aggregation process.
[0058] S6. Prototype network classification, In the small-sample scenario, few-shot classification is achieved by calculating the "class prototype" of the support set and measuring the distance from the query test samples to the prototype. The core idea of the prototype network is to compress the support sample vectors of the same class into a prototype vector, and by means of distance measurement, compare the similarity between the query samples and the prototypes of each class, so as to determine their belonging classes. Compared with traditional classifiers, the prototype network does not rely on large-scale parameter optimization and has natural small-sample adaptability and strong generalization ability.
[0059] 1) Prototype vector calculation In a given few-shot task T, it contains a support set S and a query set Q. The support set S contains N classes, with K support samples for each class. Each support sample is encoded by a graph neural network to obtain a structured subgraph encoding. The subgraph encodings with the same fault type C in the support set k are averaged to obtain the prototype vector of fault C k : where N represents the number of samples with the same fault type.
[0060] 2) Distance measurement and probability estimation For any test sample x q in the query set, first obtain its feature vector representation h q by encoding through a graph neural network. Then, calculate the distance between it and each class prototype c k through cosine similarity: where C represents the prototype vector and h T represents the query test vector.
[0061] 3) Loss function To train the prototype network to accurately divide the boundaries between different classes, the cross-entropy loss function is introduced as a supervision signal to minimize the difference between the predicted distribution and the true label distribution: where Q represents the query set of task T, represents the conditional probability that the model predicts the input x q as the correct label y q .
[0062] If there are multiple test samples, batch processing optimization can be adopted. Continuously sample a batch of tasks of size B from the sample set to form a mini-batch, and the overall loss: 4) Classification probabilities of the test samples For actual diagnostic tasks, it is not only necessary to give a clear predicted category, but more importantly, to provide the probability estimates for each fault category to assist experts in evaluating the confidence level of the diagnostic results and forming a risk assessment mechanism. Specifically, the query test sample x q The distances from the prototypes of each category are mapped into a probability distribution to quantify the probability that the sample belongs to each fault category: where k is the fault category, h q represents the vector representation of the query test sample, and d represents the cosine similarity metric.
[0063] S7. Online fault detection and warning Deploy the offline-trained multi-modal knowledge graph and the few-shot prototype network model in the production environment to achieve real-time diagnosis and hierarchical warning of newly arrived streaming data, and provide decision-making basis for operation and maintenance personnel through the interpretability module.
[0064] 1) Data access and synchronous preprocessing The system synchronously collects the multi-modal raw data stream at time t through multi-source sensing devices deployed at key parts of the equipment, mainly including: mechanical dynamic signals collected by vibration sensors (such as accelerometers); text data such as operation logs and alarm information generated by the industrial control system; structural or thermal imaging images collected by visible light or infrared vision cameras. All data are processed by the preprocessing and feature extraction module at the access end, and then uniformly mapped to a certain target entity node in the knowledge graph space, thus completing the semantic alignment with the offline model.
[0065] 2) Threshold determination and multi-level warning Using the trained few-shot prototype network model, the system performs forward inference on each newly input sample to calculate its probability of belonging to each known fault category C k Set a fault recognition probability threshold δ. If there is a probability of a certain fault classification greater than the threshold δ, it is determined that the probability of a fault is relatively high. Based on the "severity" attribute labels associated with each fault node in the knowledge graph, such as "warning level", "severe level", "emergency level", etc., the warning level is jointly determined according to the probability size output by the model and the attribute information in the knowledge graph.
[0066] 3) Result interpretability Calculate the distance between the vector h of the fault sample q and the prototype vectors c of each category k to find the prototype category with the highest similarity k *: Then, search for the entity node e to which the sample belongs in the knowledge graph q , as well as the fault node e corresponding to the prototype k ∗ . Search in the knowledge graph for several paths {Π q} from e k ∗ to e i . Each path is defined as: π=(e q → r j → v j → ⋯ → e k ∗ ), where r j represents the relationship type of the i-th edge in the path, and v j j q is the intermediate node passed through in the path. To measure the credibility of each path, a path confidence function is introduced to calculate a confidence score for each path: where, represents the embedding vector of the entity node e corresponding to the current fault sample in the knowledge graph, r i k ∗ represents the embedding vector of the relationship type of the i-th edge in the path π, indicating the semantic relationship between entities. represents the embedding vector of the target fault node e k ∗ corresponding to the prototype category, and ||·|| represents the L2 norm. Take the top M most likely inference links in descending order of Confidence(Π) for explanation.
[0067] Experimental verification: To verify the effectiveness of the method proposed in the present invention in predictive fault diagnosis of power equipment, multiple groups of comparative experiments are designed in typical power system operation scenarios in this paper, covering key equipment such as transformers (partial discharge monitoring units), generators (bearing and stator vibration monitoring systems), high-voltage circuit breakers (contact wear detection), and transmission lines (bolt loosening image recognition and infrared temperature measurement), covering multi-modal sensing and recording information such as vibration, infrared images, and inspection texts. A total of about 13,000 complete diagnostic cycles of samples are collected. To be close to actual applications, various real working condition factors are introduced in the experimental settings, including power grid load fluctuations, long-term equipment fatigue, false alarms caused by environmental interference, and sudden loss of key data channels, etc., to comprehensively investigate the adaptability of the model to complex scenarios and the fault recognition accuracy.
[0068] The comparison methods include current mainstream fault identification models, such as Transformer dominated by the attention mechanism, MMAN based on the shared attention mechanism, multi-resolution feature fusion model MF-CNN, KGGRN based on the graph neural network, and small-sample adaptability optimization model Proto-MAML. All models are trained under the same training / validation / test split (6:2:2), unified loss function, and optimization strategy to ensure the comparability and consistency of the experimental results. To comprehensively evaluate the performance of the fault diagnosis method, six evaluation indicators are selected: ACC, F1-Score, early warning lead time, miss rate, false alarm rate, and inference latency. The experimental results are as Figure 2 , Figure 3 and shown in Table 1.
[0069] Table 1 Data comparison of different methods under six indicators Model Name ACC F1 Early Warning Lead Miss Rate False Alarm Rate Inference Delay Transformer 80% 79% 17 min 19% 14% 65s MMAN 87% 85% 28 min 16% 13% 72s MF-CNN 85% 82% 23 min 21% 15% 60s KGGRN 88% 87% 30 min 14% 11% 85s Proto-AML 87% 86% 27 min 15% 13% 78s The method of the present invention 93% 91% 42 min 8% 7% 42s From Figure 2 , Figure 3 and Table 1, it can be seen that traditional methods such as Transformer, MMAN, MF-CNN, KGGRN, and Proto-AML show certain recognition capabilities in complex equipment fault diagnosis, but each has its limitations. Although the Transformer method has strong global modeling ability and can model long-term dependencies to a certain extent, due to its native structure's low sensitivity to local features and lack of structural path perception ability, the false alarm rate is relatively high. Although MMAN and KGGRN perform well in terms of accuracy and F1 value, there is still room for improvement in terms of miss rate and false alarm rate. Although MF-CNN and Proto-AML perform well in terms of inference latency, their performance in terms of accuracy and F1 value is poor.
[0070] In contrast, the method of the present invention effectively integrates multi-source heterogeneous data and enhances the expression ability of fault features through the construction of a multi-modal knowledge graph and graph neural network encoding. At the same time, the introduced small-sample learning strategy enables the system to accurately judge even when the sample data is small, improving the adaptability and generalization ability of the model. In the experiment, this method is comprehensively superior to the comparison methods in six dimensions: accuracy, F1 value, early warning lead time, miss rate, false alarm rate, and inference latency, verifying its practicability and efficiency in predictive maintenance tasks. Especially, the accuracy and F1 value reach 93% and 91% respectively, the early warning lead time reaches 42 minutes, the miss rate and false alarm rate are reduced to 8% and 7% respectively, and the inference latency is only 42 seconds, comprehensively surpassing other comparison methods.
[0071] To verify the robustness of the model, the system sets different degrees of dropout rates for the original data in the experiment, and the results are as Figure 4As shown. According to the comparison results of the model robustness under different Dropout ratios in the figure, the method of the present invention (red solid line) shows significant superiority. Its accuracy (ACC) remains the highest under all Dropout ratios, only slightly decreasing from 93% to 88%, indicating extremely strong robustness. In contrast, the accuracies of other models all decrease when the Dropout ratio increases. Especially for the Transformer, the accuracy drops from 80% to 55%, showing sensitivity to Dropout changes.
[0072] To verify the distribution of the early warning lead time of the model, the system tested different models in the experiment, and the results are as Figure 5 shown. According to the comparison results of the early warning lead time distributions of different models in the figure, the median of the early warning lead time of the method of the present invention is the highest, about 40 minutes, indicating that this method performs best in terms of the early warning lead time and can predict faults earlier. The medians of the early warning lead times of other models are all less than 30 minutes, and the distribution range is large, indicating poor prediction effect and poor stability. Therefore, the method of the present invention not only has an advantage in the early warning lead time, but also provides a more stable and reliable fault warning ability.
[0073] Embodiment 2 This embodiment provides a complex equipment fault diagnosis system based on a multi-modal knowledge graph, including: A data acquisition module, configured to acquire equipment operation data; A feature module, configured to perform multi-modal feature extraction on the acquired equipment operation data respectively; A graph module, configured to construct a knowledge graph based on the extracted multi-modal features; An enhancement module, configured to perform multi-modal feature enhancement on the knowledge graph through cross-modal contrast learning; A small sample module, configured to define a small sample task and perform feature initialization on the enhanced knowledge graph; A classification module, configured to perform subgraph extraction and Mamba enhancement on the knowledge graph after feature initialization; and use a graph neural network to perform fault classification based on the knowledge graph.
[0074] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the described method for complex equipment fault diagnosis based on a multi-modal knowledge graph.
[0075] A terminal device, including a processor and a computer-readable storage medium, where the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor to perform the described method for complex equipment fault diagnosis based on a multi-modal knowledge graph.
[0076] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A complex equipment fault diagnosis method based on a multi-modal knowledge graph, characterized in that, Including: Obtain the operation data of the equipment; Perform multi-modal feature extraction on the obtained operation data of the equipment respectively; Construct a knowledge graph based on the extracted multi-modal features; Enhance the multi-modal features of the knowledge graph through cross-modal contrast learning; Define a few-shot task to initialize the features of the enhanced knowledge graph; Extract subgraphs and perform Mamba enhancement on the knowledge graph after feature initialization; Use a graph neural network to perform fault classification based on the knowledge graph.
2. The complex equipment fault diagnosis method based on a multimodal knowledge graph according to claim 1, wherein The multi-modal feature extraction of the obtained equipment operation data is respectively carried out, including the original vibration signal data collected by arranging high-precision sensors, and setting the vibration signal data collected by the sensors as X v , using 1D-CNN to extract local features; using the BERT pre-trained language model to perform entity recognition on operation and maintenance logs and operation records, complete the recognition of fault entities, and output corresponding semantic embedding vectors; and using the singular value decomposition SVD method to perform dimensionality reduction processing on the image matrix, and inputting the dimensionally reduced image into a standard CNN network to obtain an image feature representation h i .
3. The complex equipment fault diagnosis method based on a multi-modal knowledge graph according to claim 2, wherein The construction of the knowledge graph based on the extracted multi-modal features includes, for the vibration signal features, constructing a feature-fault template dictionary, and using the template matching method and the spectrum analysis method to identify and classify the extracted signal features; for the text data features, for the semantic embedding vectors and entity recognition results extracted by the BERT model in the operation and maintenance logs and operation information, further introduce natural language generation NLG to summarize the sentence-level information and extract the key semantic content related to diagnosis; for the image data features, the high-dimensional image features extracted by the CNN need to be combined with computer vision for semantic annotation.
4. A complex equipment fault diagnosis method based on a multi-modal knowledge graph according to claim 3, characterized in that The multi-modal feature enhancement of the knowledge graph through cross-modal contrast learning includes inter-modal contrast learning, performing positive and negative contrasts on sample pairs of different modalities to complete cross-modal feature alignment. Among them, for the different dimensionalities of the three modal features, project them into a unified-dimensional contrast learning space. In the inter-modal contrast learning, further strengthen the discriminability within each individual modality, quantify the attribute similarity of each sample within the same modality, and generate positive and negative pairs based on this; perform context mask filling for the text data features; add Gaussian blur to the image features, convolve the image with a two-dimensional Gaussian kernel to obtain a blurred enhanced image.
5. The complex equipment fault diagnosis method based on a multimodal knowledge graph according to claim 4, wherein The multi-modal feature enhancement of the knowledge graph through cross-modal contrast learning also includes constructing positive and negative sample pairs within the modality based on the enhanced view of the sample as the contrast basis to optimize the feature distribution structure within the modality. Define the positive sample pair as the pairing between two views of the same sample, and the negative sample pair as the pairing between the sample view and the views of other samples. To optimize the aggregation and separation effects of the features within the modality, the intra-modal contrast learning uses an improved InfoNCE loss function to maximize the similarity of the positive sample pairs and minimize the similarity of the negative sample pairs. For modality m, its loss form is expressed as: Among them, represents the j representation vector of the th sample after passing through the projection network under the enhancement operation, j represents the representation vector of the th sample after passing through the projection network without enhancement operation, and τ is the temperature hyperparameter that controls the smoothness of the contrast distribution.
6. The complex equipment fault diagnosis method based on a multi-modal knowledge graph according to claim 5, characterized in that, The definition of the few-shot task to initialize the features of the enhanced knowledge graph includes constructing a support set S as a small number of samples for each type of fault in the current task for the model to learn the current task; a query set Q is used to evaluate the learning effect of the model on this task; use the knowledge graph to perform structured screening of the fault categories. Among them, use the graph structure measurement indicators, eigenvector centrality, and path length in graph theory to evaluate the fault nodes in the graph. For each fault category, execute a third-order neighbor sampling strategy from the knowledge graph, that is, starting from the current node, traverse the graph outward to the third hop, and select the entity nodes closely related to this category to form a subgraph structure containing semantic neighbors.
7. A complex equipment fault diagnosis method based on a multi-modal knowledge graph according to claim 6, characterized in that The subgraph extraction and Mamba enhancement of the knowledge graph after feature initialization include introducing a structural index to dynamically control the sampling hop number S hop for each knowledge graph entity corresponding to a sample e i dynamically adjusting the neighborhood sampling hop number S according to the topological importance of the nodes hop and extracting its local neighborhood subgraph G i : Among them, Betweeness represents the betweenness centrality of a node, Degree represents the degree of a node, the parameters α and β are used to weigh the influence of betweenness centrality and the degree of a node, dist represents the number of path hops, V i represents the set of nodes, E i represents the set of edges.
8. A complex equipment fault diagnosis method based on a multi-modal knowledge graph according to claim 7, characterized in that, The subgraph extraction and Mamba enhancement of the knowledge graph after feature initialization further include introducing the selective state space model Mamba, constructing a multi-level and adaptive feature interaction architecture, and realizing input-dependent dynamic parameter adjustment. The core formula of the MambaBlock is as follows: Among them, represents the hidden state vector at time step t, represents the gating function, represents the matrix exponential function, A is a learnable state transition matrix, B is an input projection matrix, and I is the identity matrix. represents the input node features, represents the high-order neighbor node features of node v.
9. A complex equipment fault diagnosis method based on a multi-modal knowledge graph according to claim 8, characterized in that, The fault classification using a graph neural network based on a knowledge graph includes, in a small-sample scenario, compressing the support sample vectors of the same category into a prototype vector, and comparing the similarity between the query sample and the prototypes of each category through a distance metric method to determine its belonging category. Among them, the query test sample x q The distance between the query test sample and the prototypes of each category is mapped into a probability distribution to quantify the probability that the sample belongs to each fault category: where k is the fault category, h q represents the vector representation of the query test sample, and d represents the cosine similarity metric.
10. A complex equipment fault diagnosis system based on a multi-modal knowledge graph, characterized in that, Including: A data acquisition module configured to acquire equipment operation data; A feature module configured to perform multi-modal feature extraction on the acquired equipment operation data respectively; A graph module configured to construct a knowledge graph based on the extracted multi-modal features; An enhancement module configured to perform multi-modal feature enhancement on the knowledge graph through cross-modal contrast learning; A few-shot module configured to define a few-shot task and perform feature initialization on the enhanced knowledge graph; A classification module configured to perform subgraph extraction and Mamba enhancement on the knowledge graph after feature initialization; and perform fault classification based on the knowledge graph using a graph neural network.
Citation Information
Patent Citations
Database alarm intelligent diagnosis method and system based on multi-modal knowledge graph fusion and small sample learning
CN118467229A
Power grid health assessment and analysis method based on multiple modes
CN118657404A
Small sample power distribution network anomaly detection method and device based on graph contrast learning
CN119046840A
Power equipment knowledge graph completion method and system based on Mamb-GPT model
CN119128166A
Intelligent equipment fault question-answering method and system based on knowledge graph enhancement
CN119226481A
Cited By
Double-prototype drive intelligent fault diagnosis method based on multi-modal knowledge
CN120408422A
Standard digital modeling and verification method and system based on artificial intelligence
CN120409658A
A standard digital modeling and verification method and system based on artificial intelligence
CN120409658B
Early fault early warning method and system based on dynamic evolution of complex industrial map
CN120580828A
An early fault warning method and system based on dynamic evolution of complex industrial maps
CN120580828B