Intelligent equipment fault diagnosis method and device based on association embedding and multi-scale space-time capsule network

Through the intelligent device fault diagnosis method based on the associated embedding and multi-scale spatiotemporal capsule network, the problem of handling complex spatiotemporal characteristics and modal interaction information in the prior art is solved, and more accurate and efficient fault diagnosis is achieved.

CN120180263APending Publication Date: 2025-06-20HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243570.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing equipment fault diagnosis methods are difficult to deal with complex spatial and temporal features and modal interaction information, and there are problems such as data imbalance, insufficient multi-scale feature capture and neglected label dependencies.

Method used

The intelligent device fault diagnosis method based on the association embedding and multi-scale spatiotemporal capsule network is adopted. The tag association matrix is ​​constructed through the association embedding, the data samples are expanded using the generative method, and the failure characteristics are extracted through the capsule network.

Benefits of technology

The model's robustness and complex feature aggregation capability of noise interference are improved, and more accurate identification and efficient classification of equipment failures are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180263A_ABST
    Figure CN120180263A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent equipment fault diagnosis method and device based on association embedding and a multi-scale space-time capsule network, and the method comprises the steps: firstly carrying out the cleaning and arrangement of multi-modal data generated in the operation process of equipment, and carrying out the generation type sample expansion and label association modeling of a data sample; a global dependency relationship of multi-modal data on time steps and spatial positions is captured through a self-attention mechanism, the spatial-temporal feature interaction modeling capability is improved, and a stronger expression capability is provided for downstream tasks; according to the method, multi-scale local and global fault features are extracted by adopting expansion convolution, a dynamic routing mechanism of a capsule network is optimized by introducing a Bhatpacaryya coefficient, the expression ability and robustness of a model to a complex mode are enhanced, and finally the fault type or state of equipment is identified. Compared with the prior art, the method can improve the accuracy of fault diagnosis and the robustness to complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of multi-label classification and fault diagnosis, and particularly relates to an intelligent device fault diagnosis method and device based on correlation embedding and multi-scale spatio-temporal capsule network. Background Art

[0002] With the popularization of industrial equipment and intelligent hardware, the multi-modal data (such as vibration signals, temperature signals, operation logs, etc.) generated during their operation provides rich information for fault diagnosis. However, due to diverse data sources, fast-changing dynamic signals, and complex fault modes, traditional single-modal or rule-based diagnosis methods are difficult to meet the requirements of efficient and accurate diagnosis.

[0003] Most of the existing device diagnosis methods rely on empirical rules or simple statistical features. However, these methods cannot handle complex spatio-temporal features and interaction information between modalities. In addition, there are also some attempts to use machine learning for fault diagnosis, but most of these methods have the following problems: insufficient recognition ability for rare faults due to data imbalance, lack of comprehensive capture of multi-scale features, and ignoring the complex dependency relationships between labels.

[0004] Based on the above situation, an intelligent device fault diagnosis method based on correlation embedding and multi-scale spatio-temporal capsule network is proposed. First, a label correlation matrix is constructed through correlation embedding, and a generative method is used to expand data samples to enhance the model's recognition ability for various fault types; a multi-scale spatio-temporal capsule network is introduced, and dilated convolution is used to extract fault features at different scales, while the understanding of time series signals is strengthened through a spatio-temporal sequence encoding module; finally, the capsule network is used for the classification of fault modes, so as to achieve more accurate and intelligent fault diagnosis of power equipment. Summary of the Invention

[0005] Object of the Invention: Aiming at the above problems, the present invention provides an intelligent device fault diagnosis method and device based on correlation embedding and multi-scale spatio-temporal capsule network, which combines the Bhattacharyya coefficient in the capsule network to optimize the routing process of the capsule network, improve the robustness of the model to noise interference and the complex feature aggregation ability; model the spatio-temporal feature interaction relationship through a multi-layer spatio-temporal capsule network, and can achieve more accurate recognition and efficient classification of equipment faults.

[0006] Technical Solution: The present invention discloses an intelligent device fault diagnosis method based on correlation embedding and multi-scale spatio-temporal capsule network, including the following steps:

[0007] Step 1: Clean and preprocess the signal data collected by the device sensor and the text fault information. The specific method is as follows:

[0008] Step 1.1: Define the text data set C = {c1, c2,..., c N}, where c i represents the i-th device log or alarm record, and the sensor signal data set D = {d1, d2,..., d N}, where T is the number of time steps, d d is the sensor feature dimension, and N represents the total number;

[0009] Step 1.2: For each text c i , after cleaning the text c i ′, use the embedding model to transform the cleaned text into an embedding vector;

[0010] Step 1.3: For each sensor signal d i , use the short-time Fourier transform (STFT) to extract the spectral feature f i ;

[0011] Step 1.4: Use a sliding window to generate feature blocks for the text features and the sensor signal data features respectively, where represents the text feature block, represents the sensor signal data feature block, M C , M D respectively represent the divided time steps;

[0012] Step 1.5: Align the text feature E i and the sensor signal data feature F i according to the time step, and splice E i and F i to generate the multimodal feature X i = [E i , F i ;

[0013] Step 1.6: Splice the features for each time step, and finally output the multimodal feature X = {X1, X2,... X M}, where M is the number of time windows.

[0014] Step 2: Use the conditional generator to generate diverse data based on existing samples and perform associated modeling on the labels. The specific method is as follows:

[0015] Step 2.1: Perform generative sample augmentation on the text data features and the sensor signal data features;

[0016] Step 2.2: Perform sample augmentation on the text data features. Use the conditional generative adversarial network, input some real samples E real ∈E and random noise z to generate the augmented text feature E′ = Gtext (z, cond(E real , Y)), where cond(E real , Y) is the conditional vector extracted from the real text features and the corresponding label Y;

[0017] Step 2.3: Augment the sensor signal features. Use a conditional generative adversarial network. Input a partial real sample F real ∈F and random noise z to generate augmented sensor signal features F′ = G sensor (z, cond(F real , Y)), where cond(F real , Y) is the conditional vector extracted from the real sensor signal features and the corresponding label Y;

[0018] Step 2.4: Concatenate the two augmented samples to form the augmented multimodal feature X′ = [E′, F′];

[0019] Step 2.5: Define the label set Y = {y1, y2,..., y l}, where y i represents the i-th label and l represents the total number of labels;

[0020] Step 2.6: Construct a label correlation matrix by statistically analyzing the collinear relationship between labels. The label correlation matrix where count(y i , y j ) represents the number of times the labels y i and y j appear simultaneously, and count(y i ) and count(y j ) represent the number of times y i and y j appear respectively;

[0021] Step 2.7: Optimize the multimodal feature matrix X″ = X′·R using the label correlation matrix R.

[0022] Step 3: Capture the global feature correlation of multimodal data in time steps and spatial positions through the self-attention mechanism to enhance the comprehensive extraction ability of spatial and temporal features; the specific method is as follows:

[0023] Step 3.1: Model the temporal features in the multimodal feature matrix to capture the dependencies between different time steps;

[0024] Step 3.2: First, perform positional encoding on the optimized multimodal feature matrix to generate a positional encoding matrix pn, and concatenate the two feature matrices to obtain the fused feature matrix X pe ;

[0025] Step 3.3: The self-attention mechanism allows the model to consider the information of all time steps at each time step. For the fused feature matrix X pe perform feature mapping to generate query, key, and value vectors, Q = X pe ·W Q , K = X pe ·W K , V = X pe ·W V ; where W Q , W K , W V are all weight matrices, which are used to calculate the query (Q), key (K), and value (V) respectively;

[0026] Step 3.4: Calculate self-attention: where Q·K T is the dot product of the query and the key, which is used to calculate the attention weights. The softmax operation ensures that the weights are normalized, and the final output of the attention-weighted features is obtained by taking the dot product with the value V;

[0027] Step 3.5: Perform temporal modeling on the features of each time step through the self-attention mechanism to obtain the spatio-temporal sequence encoding X att = Attention(Q, K, V);

[0028] Step 3.6: Output the temporal encoding X att , the multi-modal features processed by the self-attention mechanism can capture the dependencies of the time series.

[0029] Step 4: Use dilated convolution to extract local and global features of equipment faults at multiple scales, and at the same time combine the Bhattacharyya coefficient to optimize the dynamic routing algorithm of the capsule network; the specific method is as follows:

[0030] Step 4.1: Perform dilated convolution on the temporal encoding X att to extract features at different scales;

[0031] Step 4.2: The core idea of dilated convolution is to change the receptive field size by adjusting the dilation rate r of the convolution kernel to extract features at different scales. The formula is as follows:

[0032]

[0033] where I represents the elements of the input feature matrix X att for operation with the convolution kernel, H represents the convolution kernel, r represents the dilation rate, i and j represent the indices of the input feature I, and m and n represent the indices of the convolution kernel H;

[0034] Step 4.3: Set different dilation rates r and extract features at different scales;

[0035] Step 4.4: Perform multi-scale feature fusion on the features at different scales obtained after dilated convolution, X m = Concat(DilatedConv(X att ))), where DilatedConv(X att ) represents the operation of extracting features at different scales under dilated convolution;

[0036] Step 4.5: Convert the fused multi-scale feature X m into the initial capsule vector u h|g = W hg ·X m , where W hg represents the transformation matrix used to convert the input feature X m into the capsule prediction vector, and u h|g represents the prediction vector from the h-th layer capsule to the g-th layer capsule;

[0037] Step 4.6: Calculate the Bhattacharyya coefficient between capsules. The Bhattacharyya coefficient measures the similarity between the lower-layer capsules and the higher-layer capsules. The formula is as follows: where v g represents the initial higher-layer capsule vector, u h|g represents the prediction vector of the lower-layer capsule, and BC(v g , u h|g ) represents the similarity between the two capsules;

[0038] Step 4.7: Based on the Bhattacharyya coefficient, update the routing weight b hg between capsules and optimize the connection between capsules. The calculation formula is as follows: b hg (a + 1) = b hg (a) - ln(BC(v g , u h|g ))), where b hg (a) represents the routing weight at the a-th iteration, and ln(BC(v g , u h|g ) represents the weight adjustment value calculated according to the similarity;

[0039] Step 4.8: Convert the routing weight into a routing coefficient through the softmax function, which is used to represent the contribution ratio of each lower-layer capsule to the higher-layer. The calculation formula is as follows: Routing coefficient e hg = softmax(b hg );

[0040] Step 4.9: Weighted sum of the underlying prediction vectors according to the routing coefficients, and the calculation formula is as follows: where s g represents the total input vector of the capsules in the g-th layer. The total input vector s of the high-level capsules g is transformed through the non-linear activation function squash to obtain the output of the high-level capsules;

[0041] Step 4.10: After multiple rounds of iteration, the output of the optimized high-level capsules is finally obtained.

[0042] Step 5: Identify the fault type or status of the device by recursively modeling the interaction relationships of multi-modal features through the multi-layer of the spatio-temporal capsule network; the specific method is as follows:

[0043] Step 5.1: Reconstruct the outputs of all high-level capsules into a matrix P;

[0044] Step 5.2: Perform multi-dimensional feature transformation on the matrix P, P H = P · T H and P V = P · T V where T H and T V represent the correlation tensors in the horizontal and hammer directions respectively, and P H and P V are the feature representations obtained in the horizontal and vertical directions respectively;

[0045] Step 5.3: Combine the features in the horizontal and vertical directions to generate the final feature representation where δ represents the ReLU activation function;

[0046] Step 5.4: Input the final feature P Z into the classifier network, and through the softmax function, obtain the final diagnosis result.

[0047] The present invention also discloses an intelligent device fault diagnosis device based on correlation embedding and multi-scale spatio-temporal capsule network, including a memory, a processor, and a computer program stored on the memory and executable on the processor. It is characterized in that when the computer program is loaded into the processor, the above-mentioned intelligent device fault diagnosis method based on correlation embedding and multi-scale spatio-temporal capsule network is realized.

[0048] Beneficial effects:

[0049] Based on the existing equipment fault information, the present invention uses the fault information to expand diverse samples through a conditional generator, alleviating the problem of data imbalance, and modeling the dependencies between labels through high-order tensors. The self-attention mechanism is used to capture the global dependencies between time steps and positions, improving the feature expression ability; dilated convolutions are adopted to extract local and global features, capturing multi-scale fault features through different dilation rates, enhancing the recognition ability for complex patterns; in the capsule network, the routing process of the capsule network is optimized by combining the Bhattacharyya coefficient, improving the robustness of the model to noise interference and the complex feature aggregation ability; by modeling the spatio-temporal feature interaction relationship through a multi-layer spatio-temporal capsule network, the present invention can achieve more accurate identification and efficient classification of equipment faults. Description of the Drawings

[0050] Figure 1 It is the overall flowchart of the present invention;

[0051] Figure 2 It is the schematic diagram of the capsule network structure;

[0052] Figure 3 It is the flowchart of optimizing the dynamic routing;

[0053] Figure 4 It is the flowchart of fault diagnosis. Detailed Embodiment

[0054] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.

[0055] The present invention discloses an intelligent device fault diagnosis method and device based on associated embedding and multi-scale spatio-temporal capsule network. The method includes the following steps:

[0056] Step 1: Clean and preprocess the signal data collected by the device sensor and the text fault information. The specific method is as follows:

[0057] Step 1.1: Define the text data set C = {c1, c2,..., c N}, where c i represents the i-th device log or alarm record, and the sensor signal data set D = {d1, d2,..., d N}, where T is the number of time steps, d d is the sensor feature dimension, and N represents the total number.

[0058] Step 1.2: For each text c i The text c after cleaningi ', use the embedding model to convert the cleaned text into embedding vectors.

[0059] Step 1.3: For each sensor signal d i Use the short-time Fourier transform (STFT) to extract the spectral feature f i .

[0060] Step 1.4: Use a sliding window to generate feature blocks for the text features and the sensor signal data features respectively, where represents the text feature block, represents the sensor signal data feature block, M C , M D respectively represent the divided time steps.

[0061] Step 1.5: Align the text feature E i and the sensor signal data feature F i , and splice E i and F i to generate the multimodal feature X i = [E i , F i .

[0062] Step 1.6: Splice the features for each time step, and finally output the multimodal feature X = {X1, X2,... X M}], where M is the number of time windows.

[0063] Step 2: Use the conditional generator to generate diverse data based on existing samples and perform associated modeling on the labels. The specific method is as follows:

[0064] Step 2.1: Perform generative sample augmentation on the text data features and the sensor signal data features:

[0065] Step 2.2: Perform sample augmentation on the text data features. Use the conditional generative adversarial network. Input a part of the real samples E real ∈ E and the random noise z to generate the augmented text feature E' = G text (z, cond(E real , Y)), where cond(E real , Y) is the conditional vector extracted from the real text features and the corresponding labels Y.

[0066] Step 2.3: Perform sample augmentation on the sensor signal features. Use the conditional generative adversarial network. Input a part of the real samples F real ∈ F and the random noise z to generate the augmented sensor signal feature F' = G sensor (z, cond(F real, Y)), where cond(F real , Y) is the conditional vector extracted from the real sensor signal features and the corresponding label Y.

[0067] Step 2.4: Concatenate the two augmented samples to form the augmented multimodal feature X′ = [E′, F′].

[0068] Step 2.5: Define the label set Y = {y1, y2,..., y l}, where y i represents the i-th label, and l represents the total number of labels.

[0069] Step 2.6: Construct a label correlation matrix by counting the collinear relationships between labels. The label correlation matrix where count(y i , y j ) represents the number of times the labels y i and y j appear simultaneously, and count(y i ) and count(y j ) represent the number of times y i and y j appear respectively.

[0070] Step 2.7: Optimize the multimodal feature matrix X″ = X′ · R using the label correlation matrix R.

[0071] Step 3: Capture the global feature correlations of multimodal data at time steps and spatial positions through the self-attention mechanism to enhance the comprehensive extraction ability of spatial and temporal features. The specific method is as follows:

[0072] Step 3.1: Model the temporal features in the multimodal feature matrix to capture the dependencies between different time steps.

[0073] Step 3.2: First, perform positional encoding on the optimized multimodal feature matrix to generate the positional encoding matrix pn, and concatenate the two feature matrices to obtain the fused feature matrix X pe .

[0074] Step 3.3: The self-attention mechanism allows the model to consider the information of all time steps at each time step. For the fused feature matrix X pe , perform feature mapping to generate query, key, and value vectors, Q = X pe · W Q , K = X pe · W K , V = X pe · W V ; where W Q , W K , WV They are all weight matrices, which are used to calculate the query (Q), key (K), and value (V) respectively.

[0075] Step 3.4: Calculate self-attention: where Q·K T is the dot product of the query and the key, which is used to calculate the attention weights. The softmax operation ensures that the weights are normalized, and the final output of the attention-weighted features is obtained by taking the dot product with the value V.

[0076] Step 3.5: Perform temporal modeling on the features at each time step through the self-attention mechanism to obtain the spatio-temporal sequence encoding X att = Attention(Q, K, V).

[0077] Step 3.6: Output the temporal encoding X att , the multi-modal features processed by the self-attention mechanism can capture the dependencies of the time series.

[0078] Step 4: Use dilated convolutions to extract local and global features of device faults at multiple scales, and at the same time optimize the dynamic routing algorithm of the capsule network by combining the Bhattacharyya coefficient. The specific method is as follows:

[0079] Step 4.1: Perform dilated convolutions on the temporal encoding X att to extract features at different scales.

[0080] The core idea of dilated convolutions is to change the receptive field size by adjusting the dilation rate r of the convolutional kernel to extract features at different scales. The formula is as follows:

[0081]

[0082] where I represents the elements of the input feature matrix X att for operation with the convolutional kernel, H represents the convolutional kernel, r represents the dilation rate, i and j represent the indices of the input feature I, and m and n represent the indices of the convolutional kernel H.

[0083] Step 4.3: Set different dilation rates r to extract features at different scales.

[0084] Step 4.4: Perform multi-scale feature fusion on the features at different scales obtained after dilated convolutions, X m = Concat(DilatedConv(X att ))), where DilatedConv(X att ) represents the operation of extracting features at different scales under dilated convolutions.

[0085] Step 4.5: Take the fused multi-scale features Xm Convert to the initial capsule vector u h|g = W hg ·X m , where W hg represents the transformation matrix for converting the input feature X m into the capsule prediction vector, u h|g represents the prediction vector from the capsule in the h-th layer to the capsule in the g-th layer.

[0086] Step 4.6: Calculate the Bhattacharyya coefficient between capsules. The Bhattacharyya coefficient measures the similarity between the lower-layer capsules and the higher-layer capsules. The formula is as follows: where v g represents the initial higher-layer capsule vector, u h|g represents the prediction vector of the lower-layer capsule, and BC(v g , u h|g ) represents the similarity between the two capsules.

[0087] Step 4.7: Update the routing weight b between capsules based on the Bhattacharyya coefficient hg , and optimize the connection between capsules. The calculation formula is as follows: b hg (a + 1) = b hg (a) - ln(BC(v g , u h|g ))), where b hg (a) represents the routing weight at the a-th iteration, and ln(BC(v g , u h|g ) represents the weight adjustment value calculated based on the similarity.

[0088] Step 4.8: Convert the routing weight into a routing coefficient through the softmax function, which is used to represent the contribution ratio of each lower-layer capsule to the higher layer. The calculation formula is as follows: Routing coefficient e hg = softmax(b hg ).

[0089] Step 4.9: Weighted sum the lower-layer prediction vectors according to the routing coefficient. The calculation formula is as follows: where s g represents the total input vector of the capsules in the g-th layer. The total input vector s g of the higher-layer capsules is transformed through the non-linear activation function squash to obtain the output of the higher-layer capsules.

[0090] Step 4.10: After multiple rounds of iteration, finally obtain the output of the optimized higher-layer capsules.

[0091] Step 5: Identify the fault type or status of the device by recursively modeling the interaction relationships of multimodal features through a spatio-temporal capsule network. The specific method is as follows:

[0092] Step 5.1: Reconstruct the outputs of all high-level capsules into matrix P.

[0093] Step 5.2: Perform multi-dimensional feature transformation on matrix P, P H = P · T H ,P V = P · T V ,where T H and T V represent the correlation tensors in the horizontal and hammer directions respectively, and P H and P V are the feature representations obtained in the horizontal and vertical directions respectively.

[0094] Step 5.3: Combine the features in the horizontal and vertical directions to generate the final feature representation where δ represents the ReLU activation function.

[0095] Step 5.4: Input the final feature P Z into the classifier network, and obtain the final diagnostic result through the softmax function.

[0096] The above embodiments are only used to illustrate the technical ideas and features of the present invention, aiming to enable professionals familiar with the technology to understand the embodiments of the present invention. The present invention is not limited to the above embodiments, and any equivalent transformation or improvement based on the core idea of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for intelligent device fault diagnosis based on association embedding and multi-scale spatiotemporal capsule network, characterized in that: The steps include: Step 1: Clean and preprocess the signal data and text fault information collected by the equipment sensor; Step 2: Use the conditional generator to generate diversified data based on the existing samples of text data features and sensor signal data features, and form an expanded multimodal feature after splicing. Perform association modeling on the labels, build a label association matrix, and use the label association matrix to optimize the multimodal feature matrix. Step 3: The global feature correlation of the multimodal feature matrix in time steps and spatial positions is captured through the self-attention mechanism to obtain spatiotemporal sequence encoding and enhance the comprehensive extraction capability of spatial and temporal features; Step 4: Use dilated convolution to encode the spatiotemporal sequence to extract local and global features of equipment failures from multiple scales and perform multi-scale fusion. At the same time, combine the Bhattacharyya coefficient to optimize the dynamic routing algorithm of the capsule network. Take the fused multi-scale features as input to obtain the output of the high-level capsule of the dynamic routing algorithm of the optimized capsule network. Step 5: The outputs of all high-level capsules are reorganized into a matrix and multi-dimensional feature transformation is performed to identify the fault type or status of the equipment.

2. The intelligent device fault diagnosis method based on association embedding and multi-scale space-time capsule network according to claim 1 is characterized in that: The specific method of step 1 is: Step 1.1: Define the text data set C = {c1, c2, ..., c N }, where c i represents the i-th device log or alarm record, the sensor signal data set D = {d1, d2, ..., d N }, where d i ∈R T×dd , T is the number of time steps, d d is the sensor feature dimension, and N represents the total number; Step 1.2: For each text c i The cleaned text c i ′, use the embedding model to convert the cleaned text into an embedding vector; Step 1.3: For each sensor signal d i Use short-time Fourier transform (STFT) to extract spectral features f i ; Step 1.4: Use sliding windows to generate feature blocks for text features and sensor signal data features, where E = {E1, E2, ..., E M } represents a text feature block, F = {F1, F2, ..., F M } represents the sensor signal data feature block, and M represents the number of divided time windows; Step 1.5: Align text features E according to time windows i and sensor signal data characteristics F i , for E i and F i Splice to generate multimodal features X i =[E i ,F i ]; Step 1.6: Concatenate the features of each time window and finally output the multimodal feature X = {X1, X2, ...X M }.

3. The intelligent device fault diagnosis method based on association embedding and multi-scale space-time capsule network according to claim 1 is characterized in that: The specific method of step 2 is: Step 2.1: Generate sample expansion for text data features and sensor signal data features; Step 2.2: Expand the sample of text data features, use conditional generative adversarial network, and input some real samples E real ∈E and random noise z, E is the text feature block, generating the extended text feature E′=G text (z,cond(E real ,Y)), where cond(E real ,Y) is the conditional vector extracted from the real text features and the corresponding label Y; Step 2.3: Perform sample expansion on the sensor signal features, use the conditional generative adversarial network, and input some real samples F real ∈F and random noise z, generate the extended sensor signal feature F′=G sensor (z,cond(F real ,Y)), where cond(F real ,Y) is the conditional vector extracted from the real sensor signal features and the corresponding label Y; Step 2.4: Concatenate the two expanded samples to form the expanded multimodal feature X′=[E′,F′]; Step 2.5: Define the label set Y = {y1,y2,...,y l },y i represents the i-th label, and l represents the total number of labels; Step 2.6: Construct a label association matrix by counting the collinear relationships between labels. where count(y i ,y j ) indicates label y i and j The number of times it appears at the same time, count(y i ) and count(y j ) represent y i and j Number of occurrences; Step 2.7: Use the label association matrix R to optimize the multimodal feature matrix X″=X′·R.

4. The intelligent device fault diagnosis method based on association embedding and multi-scale space-time capsule network according to claim 1 is characterized in that: The specific method of step 3 is: Step 3.1: Model the temporal features in the multimodal feature matrix to capture the dependencies between different time steps; Step 3.2: First, position encode the optimized multimodal feature matrix to generate the position encoding matrix pn, and concatenate the two feature matrices to obtain the fused feature matrix X pe ; Step 3.3: The self-attention mechanism allows the model to consider the information of all time steps at each time step. For the fused feature matrix X pe Perform feature mapping to generate query, key, and value vectors, Q = X pe ·W Q , K = X pe ·W K , V = X pe ·W V , where W Q , W K , W V They are all weight matrices, used to calculate query Q, key K and value V respectively; Step 3.4: Calculate self-attention: Among them, Q·K T It is the dot product of the query and the key, which is used to calculate the attention weight. The softmax operation ensures that the weight is normalized. The final output attention weighted feature is obtained by the dot product with the value V; Step 3.5: Use the self-attention mechanism to perform temporal modeling on the features of each time step to obtain the spatiotemporal sequence encoding X att =Attention(Q,K,V), output spatiotemporal sequence code X att .

5. The intelligent device fault diagnosis method based on association embedding and multi-scale space-time capsule network according to claim 1 is characterized in that: The specific method of step 4 is: Step 4.1: Encode the spatiotemporal sequence X att Perform dilated convolution to extract features of different scales; Step 4.2: Change the size of the receptive field by adjusting the dilation rate r of the convolution kernel to extract features of different scales, as shown below: Where I represents the input feature matrix X att The elements of are used to operate with the convolution kernel, H represents the convolution kernel, r represents the dilation rate, i and j represent the index of the input feature I, and m and n represent the index of the convolution kernel H; Step 4.3: Set different expansion rates r to extract features of different scales; Step 4.4: Multi-scale feature fusion is performed on the features of different scales obtained after the dilated convolution, X m =Concat(DilatedConv(X att )), where DilatedConv(X att ) represents the extraction of features of different scales under the dilated convolution operation; Step 4.5: The fused multi-scale features X m Transformed into the initial capsule vector u h|g =W hg ·X m , W hg Represents the transformation matrix, which is used to transform the input feature X m Converted to capsule prediction vector, u h|g Represents the prediction vector from the h-th layer capsule to the g-th layer capsule; Step 4.6: Calculate the Bhattacharyya coefficient between capsules. The Bhattacharyya coefficient measures the similarity between low-level capsules and high-level capsules. The formula is as follows: where v g represents the initial high-level capsule vector, u h|g Represents the prediction vector of the bottom-level capsule, BC(v g ,u h|g ) represents the similarity between two capsules; Step 4.7: Update the routing weight b between capsules based on the Bhattacharyya coefficient hg , optimize the link between capsules, the calculation formula is as follows: b hg (a+1)=b hg (a)-ln(BC(v g ,u h|g )), b hg (a) represents the routing weight at the ath iteration, ln(BC(v g ,u h|g ) represents the weight adjustment value calculated according to the similarity; Step 4.8: The routing weight is converted into a routing coefficient through the softmax function to represent the contribution ratio of each bottom-level capsule to the upper layer. The calculation formula is as follows: Routing coefficient e hg =softmax(b hg ); Step 4.9: Perform weighted summation of the underlying prediction vectors according to the routing coefficients. The calculation formula is as follows: where s g Represents the total input vector of the g-th layer capsule, and the total input vector s of the high-level capsule g The output of the high-level capsule is obtained through the nonlinear activation function squash conversion; Step 4.10: After multiple rounds of iterations, the output of the optimized high-level capsule is finally obtained.

6. The intelligent device fault diagnosis method based on association embedding and multi-scale space-time capsule network according to claim 1 is characterized in that: The specific method of step 5 is: Step 5.1: Reorganize the outputs of all high-level capsules into matrix P; Step 5.2: Matrix P performs multidimensional feature transformation, P H =P·T H , P V =P·T V , where T H and T V Represent the correlation tensor of the horizontal and hammer directions, P H and P V Get feature representations for horizontal and vertical directions respectively; Step 5.3: Combine horizontal and vertical features to generate the final feature representation Where δ represents the ReLU activation function; Step 5.4: The final feature P Z Input into the classifier network and pass through the softmax function to get the final diagnosis result.

7. An intelligent device fault diagnosis device based on associative embedding and multi-scale space-time capsule network, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the intelligent device fault diagnosis method based on associative embedding and multi-scale space-time capsule network is implemented according to any one of claims 1 to 6.