Intelligent fault diagnosis method and device based on label relation enhancement and semantic fusion

By introducing a method based on label relationship enhancement and semantic fusion in traditional fault diagnosis methods, using DistilBERT, K nearest neighbors, graph attention networks and improved capsule networks, the problems of label relationship modeling and semantic information capture in traditional methods are solved, and the accuracy and efficiency of fault diagnosis are improved.

CN120067969APending Publication Date: 2025-05-30HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510023545.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional fault diagnosis methods rely on rules or expert experience, difficulty in modeling relationships and semantic information between labels, difficulty in associating text and labels, and difficulty in dealing with data imbalance, resulting in diagnostic accuracy and inefficiency.

Method used

An intelligent fault diagnosis method based on label relationship enhancement and semantic fusion is adopted, word vector coding is performed through DistilBERT, label correlation matrix is ​​constructed, and the dependence between tags is captured using K nearest neighbors and graph attention networks, text and tag features are fused in combination with a bidirectional attention mechanism, and multi-level semantic information extraction and tag prediction are used to utilize the improved capsule network.

Benefits of technology

The classification accuracy and accuracy of long-tail labels are improved, the ability to handle label relationship modeling and data imbalance problems is improved, the model captures text and label semantic correlations is enhanced, and the richness of feature representation and the generalization ability of classification models is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067969A_ABST
    Figure CN120067969A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent fault diagnosis method and device based on label relation enhancement and semantic fusion, and the method comprises the steps: collecting fault information of industrial equipment, carrying out the preprocessing, generating a labeled fault data set, and carrying out the binary coding of label data; performing word vector coding on the text by using a pre-training model DistilBERT, and capturing implicit features in the text; constructing a label incidence matrix, screening related neighbors by using K neighbors, aggregating semantic information among labels through a graph attention network, and capturing a dependency relationship among the labels; the correlation between the text and the label is calculated by applying a bidirectional attention mechanism, and the features are fused to enhance the mutual perception between the label and the text; and inputting the fused features into an improved capsule network for classification, extracting multi-level semantic information and predicting corresponding labels. By introducing a label relation enhancement and semantic fusion mechanism, label representation is optimized, the problem of scarcity of long-tail labels is relieved, and the accuracy and efficiency of intelligent fault diagnosis are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing and artificial intelligence, and particularly relates to an intelligent fault diagnosis method and device based on enhanced tag relationship and semantic fusion. Background Art

[0002] In the field of industrial equipment fault diagnosis, with the advancement of industrial automation and intelligence, the monitoring and diagnosis of equipment faults have become the key to ensuring the normal operation of equipment and improving production efficiency. The timely discovery and accurate diagnosis of equipment faults can not only avoid production shutdowns, reduce maintenance costs, but also extend the service life of the equipment. With the development of industrial automation and intelligence, the requirements for fault diagnosis technology are getting higher and higher, especially in preventive maintenance and real-time monitoring.

[0003] However, traditional fault diagnosis methods mainly rely on rule engines, expert systems or machine learning-based algorithms such as support vector machines (SVM), decision trees, etc. These methods analyze sensor data, operation logs and equipment operating states to try to predict and diagnose equipment faults. However, traditional fault diagnosis methods have some limitations, mainly reflected in the following aspects:

[0004] 1. Over-reliance on rules or expert experience: Rule-based systems usually rely on expert experience to construct fault models, and it is difficult to deal with complex or dynamically changing fault patterns. Moreover, the scalability of rule systems is poor, and rules need to be continuously updated and adjusted when facing new types of faults.

[0005] 2. Difficulty in modeling the relationships between tags: Traditional fault diagnosis methods usually regard fault tags as independent categories, ignoring the dependency relationships and semantic information between tags. However, in practice, there are often associations between fault tags. For example, some faults may occur simultaneously or affect each other. Traditional methods are difficult to effectively capture these relationships, limiting the expression ability and diagnostic accuracy of the system.

[0006] 3. Difficulty in learning the association between tags and text: Traditional methods usually process text and tags separately and cannot effectively capture the relationships between them. In multi-label text classification, there are often close connections between text features and tags, but traditional methods are difficult to model these associations, thus limiting the diagnostic accuracy.

[0007] 4. Difficulty in dealing with the long-tail problem: In equipment faults, some fault types may occur frequently, while some rare faults may be ignored. Traditional methods are easily affected by frequent fault patterns, resulting in poor diagnostic effects for rare faults. Summary of the Invention

[0008] Objective of the Invention: To overcome the deficiencies of the above-mentioned prior art, the present invention provides an intelligent fault diagnosis method and device based on tag relationship enhancement and semantic fusion.

[0009] Technical Solution: The present invention discloses an intelligent fault diagnosis method based on tag relationship enhancement and semantic fusion, including the following steps:

[0010] Step 1: Collect the fault information of industrial equipment. After data cleaning and preprocessing, perform tag annotation to generate the text information of the labeled fault data set, and perform binary encoding on the tag data.

[0011] Step 2: Use the pre-trained model DistilBERT to perform word vector encoding on the preprocessed and labeled text information in Step 1 to generate context embeddings to capture the hidden features in the text, and obtain the text embedding D.

[0012] Step 3: Construct a tag association matrix through mutual information. Based on the tag association matrix, use K-nearest neighbors to screen the relevant neighbors of the target tag nodes and use graph attention to aggregate information to capture the dependencies and co-occurrence features between tags, and determine the tag embedding H'.

[0013] Step 4: Calculate the correlation between the text and the tags through a tag-text bidirectional attention mechanism, fuse the features to obtain joint features, and strengthen the mutual perception between the text and the tags.

[0014] Step 5: Input the fused joint features into an improved capsule network classifier to extract multi-level semantic information and predict the corresponding tags; the improved capsule network classifier does not use an additional CNN layer in CapsNet for feature extraction, but directly encapsulates the fused joint features and inputs them into the primary capsule layer of CapsNet.

[0015] Further, the specific method of Step 1 is as follows:

[0016] Step 1.1: Obtain the fault records and operation log information from industrial equipment to comprehensively collect the fault conditions of the equipment.

[0017] Step 1.2: Remove duplicates from the collected data, process missing values and outliers, and perform standardization and normalization on numerical data.

[0018] Step 1.3: Obtain the cleaned data set D1 = {data 1 , data 2 ,..., data a ,..., data len(D1)}, where data aThe a-th data in D1, len(D1) is the number of data in D1, and the variable a ∈ [1, len(D1)];

[0019] Step 1.4: Define the label set and perform label annotation on the dataset D1 to generate the labeled fault dataset D2 = {d 1 , d 2 ,..., d b ,..., d len(D2)}, where d b = {data b , label}, label is the label, and data b is the text content;

[0020] Step 1.5: Perform binary encoding on the label data. Each label corresponds to a binary bit, so that the label of each sample can be represented in the form of a vector.

[0021] Furthermore, the specific method of step 2 is as follows:

[0022] Step 2.1: Define the text sequence x = {w 1 , w 2 ,..., w m}, m represents the number of words in the text sequence, and w i represents the i-th word in the text sequence;

[0023] Step 2.2: Construct the pre-trained model DistilBERT, and input the text sequence into the DistilBERT model to obtain the text embedding D ∈ R m×k , where k represents the dimension of each word embedding vector.

[0024] Furthermore, the specific method of step 3 is as follows:

[0025] Step 3.1: Define the label set as Y = {l 1 , l 2 ,..., l K}, where K represents the total number of labels, that is, the number of labels in the label set Y;

[0026] Step 3.2: Define a zero matrix A, where each element is 0;

[0027] Step 3.3: Represent the matrix A as a graph, where the nodes represent the labels and the edges represent the relationship coefficients between the labels;

[0028] Step 3.4: Use the mutual information of the labels to calculate the relationship between the labels, calculate the approximate matrix, and the mutual information is defined as follows:

[0029]

[0030] Among them, p(i) and p(j) respectively represent the probability distributions of labels i and j in the sentence, and p(j, i) represents the joint probability that labels l i and label l j appear in the sentence at the same time. H(j) represents the entropy of label j, which is used to measure the degree of uncertainty of label j, and H(j|i) represents the conditional entropy of the uncertainty of label j under the condition that label i is known;

[0031] Step 3.5: Normalize the calculated mutual information to obtain the relationship coefficient between the initial labels:

[0032] r ij = normal(Σ s g ij ), s)

[0033] Among them, s represents the number of samples that contain both labels i and j, and normal() is the normalization function;

[0034] Step 3.6: Update the label association matrix A: A = [r ij , 0 < i < K, 0 < j < K, and then update the representation of the graph;

[0035] Step 3.7: Given the label association matrix A, use the K-nearest neighbor algorithm to find the k nearest neighbor labels for the target label node e goal to form a set of neighbor labels: N k = {L 1 , L 2 ,..., L k}, L 1 , L 2 ,..., L k ∈ Y, and the label node L a needs to satisfy: D(e goal , L a ) ≤ D(e goal , L′ a ), 1 ≤ a ≤ k;

[0036] Among them, L a ∈ N k , L′ a ∈ Y and the target label node e goal ∈ Y, D(e goal , L a ) represents the distance between the target label node e goal and the label L a , and the formula is:

[0037]

[0038] Among them, E i and E j ∈ d respectively represent the initial feature vectors of e i and e j , where e i , e j ∈Y, and d represents the dimension of the vector;

[0039] Step 3.8: Aggregate the features of the selected neighbor labels through the graph attention network. First, calculate the relationship coefficient between the central label node l i and the best neighbor node l j : e ij =a([Wh i ||Wh j ), j∈N i , where a() is a mapping function, W represents a learnable weight matrix, || represents the vector concatenation operation, h i and h j represent the feature vectors of the label nodes l i and l j , and N i represents the set of neighbor nodes of the label node i;

[0040] Step 3.9: Normalize the relationship coefficient through the Softmax function to calculate the attention coefficient a ij , and the calculation formula is: where LeakyReLU() is an activation function, e ik represents the relationship coefficient between the node l i and any of its neighbor nodes l k , and e ij represents the relationship coefficient between the node l i and the node l j ;

[0041] Step 3.10: Perform a weighted sum of the neighbor features based on the attention coefficient to obtain a new feature representation for each label. The calculation formula is: Then the updated feature representations of all label nodes are represented as a matrix H′={h′ 1 , h′ 2 ,..., h′ K}, where σ() is a non-linear activation function and W represents a learnable weight matrix.

[0042] Furthermore, the specific method of the said step 4 is:

[0043] Step 4.1: First, calculate the attention weights from the label to the text, and the calculation formula is: where, H′ is the label embedding, D is the text embedding, k is the embedding dimension, and T is the transpose operation;

[0044] Step 4.2: Use the attention weight α to obtain the text feature representation related to the label:

[0045] Step 4.3: Then calculate the attention weights from the text to the label, and the calculation formula is:

[0046] Step 4.4: Use the attention weight β to obtain the label feature representation related to the text:

[0047] Step 4.5: Finally, fuse the text embedding with label information and the label embedding with text information to generate the final joint feature, and the calculation formula is: where, concat() represents the concatenation operation.

[0048] Furthermore, the specific method of step 5 is as follows:

[0049] Step 5.1: First, perform a convolution operation on the fused joint feature Z to generate a series of capsules u i ∈R d , and in this way, generate the i-th channel U i for the initial capsule layer, and the calculation formula is: U i = g(K i *Z + b), where, g() is the squash function, K i is the convolution kernel, and b is the bias;

[0050] Step 5.2: To calculate the total input of the high-level capsule, it is necessary to first calculate the prediction vector by multiplying the lower-level capsule u i with the weight matrix W ij , and the calculation formula is:

[0051] Step 5.3: When first entering the capsule for calculation, since there is no dynamic routing, the weight coefficient q ij is initialized to 0;

[0052] Step 5.4: In each iteration of the dynamic routing, each capsule i sends its output vector to all other capsules j and calculates the coupling coefficient c ij , and the calculation formula is:

[0053] Step 5.5: Obtain the total input s of capsule j by weighted summation of all prediction vectors The formula for calculation is as follows: j

[0054] Step 5.6: Use the squash function to compress the output vector of the capsule to ensure that the length of the vector output v of the capsule j is within the range of [0, 1]. The calculation formula is as follows:

[0055] Step 5.7: After each iteration, update the weight coefficient q ij which is used to determine which capsule outputs will be associated with other capsules. The update formula is:

[0056] Step 5.8: After the output vector is calculated for the first time and the weight coefficient is updated, start the dynamic routing. Repeat steps 5.4 to 5.7 until the set number of iterations is reached;

[0057] Step 5.9: Finally, the output vector is processed by the fully connected layer to calculate the classification probability, and according to the predicted class label, quickly locate the cause of the fault and provide corresponding solutions.

[0058] The present invention also discloses an intelligent fault diagnosis device based on label relationship enhancement and semantic fusion, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the above-mentioned intelligent fault diagnosis method based on label relationship enhancement and semantic fusion is implemented.

[0059] Beneficial effects:

[0060] 1. The present invention improves the classification accuracy and precision of long-tail labels: By calculating the relationship coefficient between labels and introducing the relationship between labels, the representation of tail labels is effectively enhanced. By modeling the label dependence relationship, the information of rich labels is effectively utilized to enhance the representation of scarce labels, improving the accuracy and precision of long-tail labels in multi-label text classification, improving the performance of the classification model under imbalanced data, and alleviating the problem of sample scarcity of long-tail labels.

[0061] 2. The present invention improves the method of label relationship modeling: By introducing a neighbor selector to screen the label relationship, the most relevant labels are screened out from the label relationship matrix, avoiding the processing of all label information, reducing the computational burden, focusing on the most relevant label relationships, integrating the most valuable information, improving the extraction accuracy of label-related information, and thus optimizing the feature representation of multi-label text classification tasks. ​

[0062] 3. The present invention enhances the model's ability to handle data imbalance problems: By introducing a graph attention network and a label relationship matrix to integrate the dependency information between labels, the model can learn useful features from head labels to enhance the expressive ability of tail labels, enabling the model to maintain high classification performance and generalization ability when facing data imbalance.

[0063] 4. The present invention improves the ability to capture the semantic relevance between labels and text: By designing a label-text bidirectional attention mechanism to establish the relationship between labels and text, enabling labels and text to mutually focus on relevant parts of each other, achieving semantic fusion of label-text, effectively capturing bidirectional correlation information between text and labels, and enhancing the model's understanding ability of the semantic relationship between text and labels.

[0064] 5. The present invention enhances the representation ability of complex text features: Using an improved traditional capsule network to classify the fused label and text features, improving the traditional capsule network architecture by avoiding the introduction of additional CNN layers, thereby avoiding information loss caused by additional convolution operations, effectively retaining multi-level semantic information, avoiding the information loss problem of traditional convolutional layers, and improving the richness of feature representation and the generalization ability of the classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a flowchart of an intelligent fault diagnosis method based on label relationship enhancement and semantic fusion;

[0066] Figure 2 It is a model diagram of the present invention;

[0067] Figure 3 It is a flowchart of dynamic routing;

[0068] Figure 4 It is a schematic diagram of dynamic routing. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.

[0070] The present invention discloses an intelligent fault diagnosis method and device based on label relationship enhancement and semantic fusion, specifically including the following steps:

[0071] Step 1: Collect the fault information of the device from industrial equipment. After data cleaning and preprocessing, generate a labeled fault data set and perform binary encoding on the label data.

[0072] Step 1.1: Obtain information such as fault records and operation logs from industrial equipment to comprehensively collect the fault conditions of the equipment.

[0073] Step 1.2: Remove duplicates from the collected data, handle missing values and outliers, and perform standardization and normalization on numerical data.

[0074] Step 1.3: Obtain the cleaned dataset D1 = {data 1 , data 2 ,..., data a ,..., data len(D1)}, where data a is the a-th data in D1, len(D1) is the number of data in D1, and the variable a ∈ [1, len(D1)].

[0075] Step 1.4: Define the label set = {high-temperature fault, vibration anomaly, start-up delay, power anomaly, sensor fault, air leakage}, and perform label annotation on the dataset D1 to generate the labeled fault dataset D2 = {d 1 , d 2 ,..., d b ,..., d len(D2)}, where d b = {data b , label}, label is the label, and data b is the text content; (Example: Text: The bearing temperature is too high, reaching 85°C, and there is also abnormal vibration of the equipment; Label: {high-temperature fault, vibration anomaly}).

[0076] Step 1.5: Perform binary encoding on the label data, with each label corresponding to a binary bit, so that the label of each sample can be represented in the form of a vector. (Example: Text: The bearing temperature is too high, reaching 85°C, and there is also abnormal vibration of the equipment; Label: {high-temperature fault, start-up delay}; Encoded label: [1, 0, 1, 0, 0, 0]).

[0077] Step 2: Use the pre-trained model DistilBERT to perform word vector encoding on the preprocessed text to generate context embeddings to capture the hidden features in the text.

[0078] Step 2.1: Define the text sequence x = {w 1 , w 2 ,..., w m}, where the length of the text sequence is m, and w i represents the i-th word in the text sequence.

[0079] Step 2.2: Construct the pre-trained model DistilBERT, and pass the text sequence into the DistilBERT model to obtain the word representation D = {d 1 , d 2 ,..., d m} ∈ R m×k , where d i represents the embedding vector of the i-th word in the text sequence, and k represents the dimension of each word embedding vector.

[0080] Step 3: Construct a label correlation matrix through mutual information, use K-nearest neighbors to filter relevant neighbors, and aggregate information using graph attention to capture the dependencies and co-occurrence features between labels.

[0081] Step 3.1: Define the label set as Y = {l 1 , l 2 ,..., l K}, where K represents the total number of labels.

[0082] Step 3.2: Define a zero matrix A, where each element is 0.

[0083] Step 3.3: Represent the matrix A as a graph, where the nodes represent labels and the edges represent the relationship coefficients between labels.

[0084] Step 3.4: To effectively capture the interaction information between labels, use the mutual information of labels to calculate the relationship between labels, that is, calculate the approximate matrix (proximity matrix). The mutual information is defined as follows:

[0085]

[0086] where p(i) and p(j) represent the probability distributions of labels i and j in the sentence, p(j, i) represents the joint probability that labels l i and l j appear in the sentence at the same time, H(j) represents the entropy of label j, which is used to measure the uncertainty degree of label j, and H(j|i) represents the conditional entropy of the uncertainty of label j under the condition of knowing label i.

[0087] Step 3.5: Normalize the calculated mutual information to obtain the relationship coefficient between labels:

[0088] r ij = normal(Σ s g ij , s), where s represents the number of samples that contain both labels i and j, and normal() is the normalization function.

[0089] Step 3.6: Update the label association matrix A: A = [a ij where 0 < i < K, 0 < j < K, and then update the graph representation, where K represents the total number of labels, i.e., the number of labels in the label set Y.

[0090] Step 3.7: Given the label association matrix A, use the K-nearest neighbor algorithm to find the k closest neighbor labels for the target label node e goal : N k (e goal ) = {L 1 , L 2 ,..., L k}(L 1 , L 2 ,..., L k ∈ Y). The label node L a (1 ≤ a ≤ k) needs to satisfy: D(e goal , L a ) ≤ D(e goal , L′ a ), where L a ∈ N k , L′ a ∈ Y and the target label node e goal ∈ Y, D(e goal , L a ) represents the distance between the target label node e goal and the label L a , and the formula is:

[0091]

[0092] where E i and E j ∈ d represent the initial feature vectors of e i and e j respectively, e i , e j ∈ Y, and d represents the dimension of the vector.

[0093] Step 3.8: Aggregate the features of the selected neighbor labels through the graph attention network. First, calculate the relationship coefficient between the central label node l i and its best neighbor node l j : e ij = a([Wh i || Wh j ), j ∈ N i , where a() is a mapping function, W represents a learnable weight matrix, || represents the vector concatenation operation, h i and hj Denote the label node l i and l j 's eigenvector, N i denotes the set of neighbor nodes of the label node i.

[0094] Step 3.9: Normalize the relationship coefficient through the Softmax function to calculate the attention coefficient a ij , and the calculation formula is: where LeakyReLU() is the activation function, and e ik denotes the relationship coefficient between node i and any of its neighbor nodes k.

[0095] Step 3.10: Based on the attention coefficient, perform weighted summation on the neighbor features to obtain a new feature representation for each label, and the calculation formula is: Then the updated feature representations of all label nodes can be represented as a matrix H′ = {h′ 1 , h′ 2 ,..., h′ K}, where σ() is the non-linear activation function, and W represents the learnable weight matrix.

[0096] Step 4: Calculate the correlation between the text and the label through the bidirectional attention mechanism, fuse the features to obtain the comprehensive representation, and strengthen the mutual perception between the text and the label:

[0097] Step 4.1: First, calculate the attention weight from the label to the text, and the calculation formula is: where H′ is the label embedding, D is the text embedding, k is the embedding dimension, and T is the transpose operation.

[0098] Step 4.2: Use the attention weight α to obtain the text feature representation related to the label:

[0099] Step 4.3: Then calculate the attention weight from the text to the label, and the calculation formula is:

[0100] Step 4.4: Use the attention weight β to obtain the label feature representation related to the text:

[0101] Step 4.5: Finally, fuse the text embedding with label information and the label embedding with text information to generate the final joint feature, and the calculation formula is: where concat() represents the concatenation operation.

[0102] Step 5: Input the fused features into the improved capsule network classifier to extract multi-level semantic information and predict the corresponding labels:

[0103] Step 5.1: First, perform a convolution operation on the fused feature Z to generate a series of capsules u i ∈R d , in this way, generate the i-th channel U i for the initial capsule layer, and the calculation formula is: U i = g(K i * Z + b), where g() is the squash function, K i is the convolution kernel, and b is the bias.

[0104] Step 5.2: To calculate the total input of the high-level capsule, it is necessary to first calculate the prediction vector by multiplying the lower-level capsule u i with the weight matrix W ij , and the calculation formula is:

[0105] Step 5.3: When first entering the capsule for calculation, since there is no dynamic routing, the weight coefficient q ij is initialized to 0.

[0106] Step 5.4: In each iteration of the dynamic routing, each capsule i sends its output vector to all other capsules j and calculates the coupling coefficient c ij according to the vector received from capsule j, and the calculation formula is:

[0107] Step 5.5: Obtain the total input s of capsule j by weighted summing all the prediction vectors j , and the calculation formula is:

[0108] Step 5.6: Use the squash function to compress the output vector of the capsule to ensure that the length of the vector output v j of the capsule is within the range of [0, 1], and the calculation formula is as follows:

[0109] Step 5.7: After each iteration, update the weight coefficient q ij used to determine which capsule outputs will be associated with other capsules, and the update formula for the coupling coefficient is:

[0110] Step 5.8: After the output vector is first calculated and the weight coefficient is updated, start the dynamic routing, and repeat steps 5.4 to 5.7 until the set number of iterations is reached.

[0111] Step 5.9: Finally, the output vector is processed by the fully connected layer to calculate the classification probability, and based on the predicted class label, the cause of the fault is quickly located and the corresponding solution is provided.

[0112] The intelligent fault diagnosis device based on label relationship enhancement and semantic fusion disclosed by the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, the above-mentioned intelligent fault diagnosis method based on label relationship enhancement and semantic fusion is implemented.

[0113] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly. It should not be used to limit the protection scope of the present invention. Any equivalent transformation or decoration made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. An intelligent fault diagnosis method based on label relationship enhancement and semantic fusion, characterized in that: The steps include: Step 1: Collect the fault information of industrial equipment, perform data cleaning and preprocessing, label it, generate the labeled fault data set text information, and perform binary encoding on the label data; Step 2: Use the pre-trained model DistilBERT to encode the word vector of the text information pre-processed and annotated in step 1, generate context embedding to capture the hidden features in the text, and obtain the text embedding D; Step 3: Construct a label association matrix through mutual information, use K nearest neighbors to filter the relevant neighbors of the target label node based on the label association matrix, and use graph attention to aggregate information to capture the dependency and co-occurrence features between labels and determine the label embedding H′; Step 4: Calculate the correlation between text and label through the label-text bidirectional attention mechanism, fuse the features to obtain joint features, and strengthen the mutual perception between text and label; Step 5: Input the fused joint features into the improved capsule network classifier to extract multi-level semantic information and predict corresponding labels; the improved capsule network classifier does not use an additional CNN layer in CapsNet for feature extraction, but directly encapsulates the fused joint features and inputs them into the main capsule layer of CapsNet.

2. The intelligent fault diagnosis method based on label relationship enhancement and semantic fusion according to claim 1 is characterized in that: The specific method of step 1 is: Step 1.1: Obtain fault records and operation log information from industrial equipment and comprehensively collect equipment fault conditions; Step 1.2: Remove duplicates from the collected data, handle missing values ​​and outliers, and standardize and normalize the numerical data; Step 1.3: Get the cleaned data set D1 = {data1, data2, ..., data a ,...,data len(D1) }, where data a is the ath data in D1, len(D1) is the number of data in D1, and the variable a∈[1,len(D1)]; Step 1.4: Define a label set and label the data set D1 to generate the labeled fault data set D2 = {d1, d2, ..., d b ,...,d len(D2) }, where d b ={data b ,label}, label is the label, data b For text content; Step 1.5: Binary encode the label data, with each label corresponding to one binary bit, so that the label of each sample can be represented in the form of a vector.

3. The intelligent fault diagnosis method based on label relationship enhancement and semantic fusion according to claim 1 is characterized in that: The specific method of step 2 is: Step 2.1: Define the text sequence x = {w1,w2,...,w m }, m represents the number of words in the text sequence, w i Represents the i-th word in the text sequence; Step 2.2: Build a pre-trained model DistilBERT, pass the text sequence into the DistilBERT model, and obtain the text embedding D∈R containing context information m×k , k represents the dimension of each word embedding vector.

4. The intelligent fault diagnosis method based on label relationship enhancement and semantic fusion according to claim 1 is characterized in that: The specific method of step 3 is: Step 3.1: Define the label set as Y = {l1,l2,...,l K }, where K represents the total number of tags, that is, the number of tags in the tag set Y; Step 3.2: Define a zero matrix A, where every element is 0; Step 3.3: Represent the matrix A as a graph, where nodes represent labels and edges represent relationship coefficients between labels; Step 3.4: Use the mutual information of the labels to calculate the relationship between the labels and calculate the approximate matrix. The mutual information is defined as follows: Among them, p(i) and p(j) represent the probability distribution of labels i and j in the sentence respectively, and p(j,i) represents the probability distribution of label l i and label l j The joint probability of appearing in the sentence at the same time, H(j) represents the entropy of label j, which is used to measure the uncertainty of label j, and H(j|i) represents the conditional entropy of the uncertainty of label j under the condition of known label i; Step 3.5: Normalize the calculated mutual information to obtain the relationship coefficient between the initial labels: r ij =normal(Σ s g ij (s) Among them, s represents the number of samples containing both labels i and j, and normal() is the normalization function; Step 3.6: Update the label association matrix A by: A = [r ij ], 0<i<K, 0<j<K and then update the representation of the graph; Step 3.7: Given the label association matrix A, use the K nearest neighbor algorithm to find the target label node e goal Find the closest k neighbor labels to form a neighbor label set: N k ={L1,L2,...,L k },L1,L2,...,L k ∈Y, label node L a Need to meet: D(e goal ,L a )≤D(e goal ,L′ a ), 1≤a≤k; Among them, L a ∈N k , L′ a ∈Y and Target label node e goal ∈Y,D(e goal ,L a ) represents the target label node e goal With label L a The distance between them is: Among them, E i and E j ∈ d Respectively represent e i and e j The initial eigenvector of i 、e j ∈Y, d represents the dimension of the vector; Step 3.8: Use the graph attention network to aggregate the features of the selected neighbor labels. First, calculate the center label node l i With the best neighbor node l j The relationship coefficient between: e ij =a([Wh i ||Wh j ]), j∈N i , where a() is the mapping function, W represents the learnable weight matrix, || represents the vector concatenation operation, and h i and h j Represents the label node l i and l j The characteristic vector of i Represents the set of neighbor nodes of label node i; Step 3.9: Normalize the relationship coefficient through the Softmax function and calculate the attention coefficient a ij , the calculation formula is: Among them, LeakyReLU() is the activation function, e ik Represents node l i and any of its neighbor nodes l k The relationship coefficient between ij Represents node l i With node l j The relationship coefficient between Step 3.10: Based on the attention coefficient, the neighbor features are weighted and summed to obtain the new feature representation of each label. The calculation formula is: Then the updated feature representation of all label nodes is represented by a matrix H′={h′1,h′2,...,h′ K }, where σ() is a nonlinear activation function and W represents a learnable weight matrix.

5. The intelligent fault diagnosis method based on label relationship enhancement and semantic fusion according to claim 1 is characterized in that: The specific method of step 4 is: Step 4.1: First, calculate the attention weight from label to text. The calculation formula is: Where H′ is the label embedding, D is the text embedding, k is the embedding dimension, and T is the transposition operation; Step 4.2: Use the attention weight α to obtain the text feature representation related to the label: Step 4.3: Then calculate the attention weight from text to label, the calculation formula is: Step 4.4: Use the attention weight β to obtain the label feature representation related to the text: Step 4.5: Finally, embed the text with label information and tags with text information embedded Fusion to generate the final joint features, the calculation formula is: Among them, concat() represents the concatenation operation.

6. The intelligent fault diagnosis method based on label relationship enhancement and semantic fusion according to claim 1 is characterized in that: The specific method of step 5 is: Step 5.1: First, perform a convolution operation on the fused joint feature Z to generate a series of capsules u i ∈R d , in this way, the i-th channel U is generated for the initial capsule layer i , the calculation formula is: i =g(K i *Z+b), where g() is the squash function, K i is the convolution kernel, b is the bias; Step 5.2: To calculate the total input of the high-level capsule, you need to first calculate the prediction vector By placing the lower capsule u i With the weight matrix W ij Multiply them together and the calculation formula is: Step 5.3: When entering the capsule for the first time for calculation, the dynamic routing is not performed, so the weight coefficient q ij Initialized to 0; Step 5.4: In each iteration of dynamic routing, each capsule i sends its output vector to all other capsules j and calculates the coupling coefficient c based on the vector received from capsule j. ij , the calculation formula is: Step 5.5: Sum all prediction vectors by weight Get the total input s of capsule j j , the calculation formula is: Step 5.6: Use the squash function to compress the capsule output vector to ensure that the capsule vector output v j The length of is in the range [0,1] and is calculated as follows: Step 5.7: After each iteration, update the weight coefficient q ij , which is used to determine which capsule outputs will be associated with other capsules. The update formula is: Step 5.8: After the output vector is calculated for the first time and the weight coefficient is updated, dynamic routing begins, and steps 5.4 to 5.7 are repeated until the set number of iterations is reached; Step 5.9: Finally, the output vector is processed by the fully connected layer to calculate the classification probability. Based on the predicted category label, the cause of the fault is quickly located and the corresponding solution is provided.

7. An intelligent fault diagnosis device based on label relationship enhancement and semantic fusion, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the intelligent fault diagnosis method based on label relationship enhancement and semantic fusion according to any one of claims 1 to 6 is implemented.