Equipment fault diagnosis method based on deep Brown distance and attention mechanism
Through the fault diagnosis method of depth Brownian distance and attention mechanism, the problem of insufficient diagnostic accuracy of traditional methods in complex fault modes is solved, and adaptive multi-scale feature extraction and high-precision diagnosis in small sample scenarios across working conditions is realized.
Patent Information
- Application Number
- CN202510419914.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-01
AI Technical Summary
Traditional equipment fault diagnosis methods are insufficient in the face of complex, nonlinear and diverse fault modes, and are difficult to effectively handle in small sample scenarios, and traditional convolutional operations are difficult to capture long-distance dependencies and multi-scale fault characteristics.
The fault diagnosis method based on the deep Brownian distance and attention mechanism is adopted, and feature extraction is performed through a weighted multi-scale wide kernel feature dynamic extraction network. Combined with global and local metric modules, the relationship between the support set and the query sample is learned, and the fault diagnosis is performed using the deep Brownian distance and attention mechanism.
It significantly improves the accuracy and reliability of fault diagnosis in small sample scenarios across working conditions, overcomes the limitations of traditional convolution in capturing long-distance dependencies, and optimizes feature extraction efficiency.
Smart Images

Figure CN120408295A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a device fault diagnosis method based on the deep Brown distance and the attention mechanism, belonging to the technical field of fault diagnosis. Background Art
[0002] With the development of intelligent manufacturing technology, higher requirements are put forward for the real-time monitoring and maintenance of mechanical equipment status. During the operation of the device, due to factors such as wear, fatigue, and overload, the performance of the device gradually degrades, and ultimately may lead to failures. Therefore, diagnosing device faults in a timely and effective manner is crucial for ensuring the safety and stability of the device.
[0003] Traditional device fault diagnosis methods mostly rely on empirical models, physics-based models, and signal processing technologies, such as vibration analysis, temperature monitoring, and sound signal analysis. Although these methods have achieved certain success in some simple fault scenarios, when facing complex, non-linear, and diverse fault patterns, the diagnostic accuracy and robustness are often difficult to meet the actual needs. Especially in a complex and changing working environment, device faults may be affected by multiple factors simultaneously, which makes it impossible for traditional methods to handle effectively.
[0004] In recent years, device fault diagnosis methods based on deep learning automatically extract features through models such as convolutional neural networks (CNNs), which have improved the fault diagnosis performance to a certain extent. However, they still face the following technical bottlenecks in actual industrial scenarios:
[0005] 1. Limited feature extraction ability: Traditional convolutional operations are limited by the fixed-scale local receptive field, making it difficult to effectively capture long-distance dependencies and multi-scale fault features in vibration signals, especially insufficient in characterizing early weak faults or complex coupled faults.
[0006] 2. Insufficient generalization in small-sample scenarios: Existing methods mostly rely on a large number of labeled samples to train the model. In actual industrial scenarios, the problem of scarce fault samples across working conditions and devices is prominent. Traditional metric learning methods only focus on the shallow statistical differences between samples, ignoring the high-order correlations of feature distributions, resulting in difficulty in accurately modeling the potential associations between the support set and query samples when samples are scarce, and poor cross-domain diagnosis robustness.
[0007] 3. Single application of the attention mechanism: Although some studies have tried to introduce channel or spatial attention to enhance the weights of key features, there is a lack of collaborative optimization of multi-dimensional attention, failing to fully exploit the global context information of fault features and having insufficient dynamic modeling ability for the dependencies between samples. Summary of the Invention
[0008] The object of the present invention is to overcome the deficiencies in the prior art and provide a device fault diagnosis method based on the deep Brown distance and the attention mechanism. First, it overcomes the limitations of traditional convolution in capturing long-distance dependence relationships and realizes adaptive multi-scale feature extraction. Second, based on the fault diagnosis model, the relationship between the support set samples and the query samples is learned from both local and global dimensions for fault diagnosis, significantly improving the fault diagnosis accuracy and reliability in the small-sample scenario across working conditions.
[0009] To achieve the above object, the present invention is implemented by the following technical solutions:
[0010] The present invention discloses a device fault diagnosis method based on the deep Brown distance and the attention mechanism, including the following steps:
[0011] Obtain the preprocessed fault vibration signal data of the device to be diagnosed;
[0012] According to the fault vibration signal data, generate corresponding time-frequency image data through fast Fourier transform and spectral feature calculation;
[0013] Use the time-frequency image data as query samples, and combine with a preset support set containing multiple types of support samples to construct a small-sample learning task including query samples and the support set;
[0014] According to the small-sample learning task, perform feature extraction based on a preset weighted multi-scale wide kernel feature dynamic extraction network to obtain query sample feature maps and multi-type support sample feature maps;
[0015] According to the query sample feature maps and multi-type support sample feature maps, perform fault diagnosis based on a trained fault diagnosis model to obtain corresponding fault diagnosis results; wherein, the fault diagnosis model includes a global metric module based on the attention mechanism, a local metric module based on the deep Brown distance, and a classifier.
[0016] Further, the obtaining of the fault vibration signal data of the device to be diagnosed includes the following steps:
[0017] Obtain the original fault vibration signal at a preset position of the device to be diagnosed;
[0018] According to the original fault vibration signal, perform cutting based on a preset sliding window to obtain multiple fault vibration signal segments with a fixed length;
[0019] According to the fault vibration signal segments corresponding to all positions, obtain the fault vibration signal data of the device to be diagnosed.
[0020] Further, the generation of the corresponding time-frequency image data includes the following steps:
[0021] For any fault vibration signal segment in the fault vibration signal data, the corresponding complex spectrum is obtained through fast Fourier transform;
[0022] According to the complex spectrum, through power spectrum calculation and logarithmic scale conversion, the corresponding time-frequency feature image is obtained;
[0023] For all time-frequency feature images, normalization and size unification processing are performed to convert and obtain time-frequency image data.
[0024] Furthermore, the preset weighted multi-scale wide kernel feature dynamic extraction network includes a backbone network and a multi-scale feature extraction module,
[0025] The backbone network is used to perform two-dimensional convolution, max pooling, layer normalization, and GELU activation function processing respectively according to any class of support samples in the query sample or support set to obtain a backbone feature matrix;
[0026] The multi-scale feature extraction module includes:
[0027] The receptive field unit is used to perform multiple convolution operations based on different receptive fields according to the backbone feature matrix, and then perform normalization processing to obtain multiple normalized receptive field feature matrices; among them, the multiple convolution operations include large kernel-like convolution, depth convolution, extended depth convolution, and point convolution operations;
[0028] The weighted fusion unit is used to perform multi-scale feature fusion based on the weight sharing mechanism according to the backbone feature matrix and multiple normalized receptive field feature matrices to obtain a fusion feature matrix;
[0029] The feature enhancement unit is used to perform two-dimensional convolution and point convolution processing according to the fusion feature matrix to obtain an enhanced feature matrix;
[0030] The residual connection unit is used to perform residual connection according to the enhanced feature matrix and the fusion feature matrix to obtain a query sample feature map or a support sample feature map corresponding to any class of support samples.
[0031] Furthermore, the global metric module based on the attention mechanism includes:
[0032] The support dual attention unit is used to perform feature extraction based on the channel attention mechanism and the spatial attention mechanism according to any class of support sample feature maps to obtain a support dual attention weighted result;
[0033] The support self-attention unit is used to perform processing based on the self-attention mechanism according to the support dual attention weighted result to obtain a support sample correlation matrix;
[0034] A prototype unit for performing a mean operation based on the support sample correlation matrix to obtain a sample prototype feature matrix;
[0035] A query dual attention unit for performing feature extraction based on the query sample feature map using a channel attention mechanism and a spatial attention mechanism to obtain a query dual attention weighted result;
[0036] A query self-attention unit for processing the query dual attention weighted result based on a self-attention mechanism to obtain a query sample correlation matrix;
[0037] A cross-attention unit for processing the query sample correlation matrix and the sample prototype feature matrix based on a cross-attention mechanism to obtain a global correlation result between the query sample and any class of support samples.
[0038] Further, the support dual attention unit is used for:
[0039] Performing average pooling and max pooling on any class of support sample feature maps in the spatial dimension, and then performing a channel attention weighting operation to obtain a support channel attention weighted result;
[0040] Performing average pooling and max pooling on the support channel attention weighted result in the channel dimension, and then performing a spatial attention weighting operation to obtain a support dual attention weighted result;
[0041] The support self-attention unit is used for:
[0042] Performing global average pooling on the support dual attention weighted result to obtain a support sample feature matrix sequence;
[0043] Performing feature projection and normalization processing on the support sample feature matrix sequence based on a self-attention mechanism to obtain a support sample correlation matrix.
[0044] Further, the query dual attention unit is used for:
[0045] Performing average pooling and max pooling on the query sample feature map in the spatial dimension, and then performing a channel attention weighting operation to obtain a query channel attention weighted result;
[0046] Performing average pooling and max pooling on the query channel attention weighted result in the channel dimension, and then performing a spatial attention weighting operation to obtain a query dual attention weighted result;
[0047] The query self-attention unit is used for:
[0048] Perform global average pooling operation according to the query double-attention weighted result to obtain a query sample feature matrix sequence;
[0049] Based on the query sample feature matrix sequence, perform feature projection and normalization processing based on the self-attention mechanism to obtain a query sample correlation matrix.
[0050] Further, the local metric module based on the deep Brownian distance includes:
[0051] A local support sample unit for calculating the BDC matrix of any class of support sample feature maps and performing a mean operation to obtain a sample prototype BDC matrix corresponding to any class of support samples;
[0052] A local query sample unit for calculating the BDC matrix of the query sample feature map;
[0053] A local correlation unit for performing an inner product operation according to the BDC matrix of the query sample feature map and the sample prototype BDC matrix corresponding to any class of support samples to obtain a local correlation result between the query sample and any class of support samples.
[0054] Further, the calculation steps of the BDC matrix are as follows:
[0055] According to the query sample feature map or any class of support sample feature maps, reshape the feature map tensor shape to obtain a reshaped tensor;
[0056] According to the reshaped tensor, obtain random vectors and calculate the Euclidean distance between the vectors to obtain an Euclidean distance matrix;
[0057] According to the Euclidean distance matrix, perform square root and subtract mean operations to obtain the BDC matrix of the query sample feature map or the BDC matrix of any class of support sample feature maps.
[0058] Further, the classifier is used for:
[0059] Calculate the metric similarity score between the query sample and each class of support samples according to the global correlation result and the local correlation result between the query sample and any class of support samples;
[0060] Take the fault type corresponding to the class of support samples with the highest metric similarity score as the corresponding fault diagnosis result.
[0061] Compared with the prior art, the beneficial effects achieved by the present invention:
[0062] The device fault diagnosis method based on the deep Brown distance and attention mechanism of the present invention, first, feature extraction is performed based on a preset weighted multi-scale wide kernel feature dynamic extraction network, which overcomes the limitations of traditional convolutions in capturing long-distance dependency relationships, realizes adaptive multi-scale feature extraction, and simultaneously optimizes the calculation efficiency of large kernels. Secondly, a fault diagnosis model based on the deep Brown distance and attention mechanism is used for fault diagnosis, learning the relationship between support samples and query samples from both local and global dimensions, and significantly improving the fault diagnosis accuracy and reliability in cross-condition small sample scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flowchart of the device fault diagnosis method based on the deep Brown distance and attention mechanism provided in Embodiment 1 of the present invention;
[0064] Figure 2 is a schematic structural diagram of the weighted multi-scale wide kernel feature dynamic extraction network provided in Embodiment 1 of the present invention;
[0065] Figure 3 is a schematic structural diagram of the global metric module provided in Embodiment 1 of the present invention;
[0066] Figure 4 is a calculation schematic diagram of the support dual attention unit provided in Embodiment 1 of the present invention;
[0067] Figure 5 is a schematic structural diagram of the local metric module provided in Embodiment 1 of the present invention;
[0068] Figure 6 is a calculation schematic diagram of the BDC matrix provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0070] This embodiment provides a device fault diagnosis method based on the deep Brown distance and attention mechanism, including the following steps:
[0071] Obtain the preprocessed fault vibration signal data of the device to be diagnosed;
[0072] According to the fault vibration signal data, corresponding time-frequency image data is generated through fast Fourier transform and spectral feature calculation;
[0073] Use the time-frequency image data as query samples, and combine with a preset support set containing multiple types of support samples to construct a small sample learning task including query samples and the support set;
[0074] Based on the small-sample learning task, feature extraction is performed based on a preset weighted multi-scale wide kernel feature dynamic extraction network to obtain a query sample feature map and multi-class support sample feature maps;
[0075] Based on the query sample feature map and multi-class support sample feature maps, fault diagnosis is performed based on a trained fault diagnosis model to obtain corresponding fault diagnosis results; wherein, the fault diagnosis model includes a global metric module based on an attention mechanism, a local metric module based on a deep Brown distance, and a classifier.
[0076] The technical concept of the present invention is as follows: First, feature extraction is performed based on a preset weighted multi-scale wide kernel feature dynamic extraction network, which overcomes the limitations of traditional convolutions in capturing long-distance dependency relationships, realizes adaptive multi-scale feature extraction, and simultaneously optimizes the computational efficiency of large kernels. Second, fault diagnosis is performed based on a fault diagnosis model based on a deep Brown distance and an attention mechanism, which learns the relationship between the support set and the query sample from both local and global dimensions, and significantly improves the fault diagnosis accuracy and reliability in cross-condition small-sample scenarios.
[0077] As Figure 1 shown, the specific steps are as follows:
[0078] Step 1: Obtain the preprocessed fault vibration signal data of the device to be diagnosed.
[0079] Specifically, it includes the following steps:
[0080] Obtain the original fault vibration signal at a preset position of the device to be diagnosed;
[0081] According to the original fault vibration signal, perform cutting based on a preset sliding window to obtain multiple fault vibration signal segments of a fixed length;
[0082] According to the fault vibration signal segments corresponding to all positions, obtain the fault vibration signal data of the device to be diagnosed.
[0083] In practical applications, acceleration sensors are placed at multiple positions of the device to be diagnosed to collect the original fault vibration signals at the corresponding positions. These signals contain rich device fault information, and the rational use of this information is related to the accuracy of device fault diagnosis.
[0084] The preset sliding window size in this embodiment is 4096.
[0085] Step 2: According to the fault vibration signal data, generate corresponding time-frequency image data through fast Fourier transform and spectrum feature calculation.
[0086] Specifically, it includes the following steps:
[0087] 2.1. For any fault vibration signal segment in the fault vibration signal data, the corresponding complex spectrum is obtained through fast Fourier transform.
[0088] The expression of the transformation function of the fast Fourier transform is as follows:
[0089]
[0090] In the formula, represents the complex spectrum obtained through fast Fourier transform with respect to frequency and time offset ; represents the fault vibration signal segment with respect to time t; represents the window function with respect to time , and the superscript represents the complex conjugate; j represents the imaginary unit; represents the frequency; represents the time offset; represents the differentiation with respect to time t.
[0091] Among them, in the fast Fourier transform, the parameter window size is set to 512 sampling points, and the frame step size is set to 512.
[0092] 2.2. According to the complex spectrum, through power spectrum calculation and logarithmic scale conversion, the corresponding time-frequency feature image is obtained.
[0093] 2.3. For all time-frequency feature images, normalization and size unification processing are performed, adjusted to a unified size of 64×64, and the time-frequency image data is converted. The time-frequency image data is represented by a feature matrix with a shape of (B, 1, 64, 64), where B represents the number of time-frequency feature images, that is, the number of fault vibration signal segments, which is also the number of samples.
[0094] Step 3. Use the time-frequency image data as the query sample, and combine it with the preset support set containing multiple types of support samples to construct a few-shot learning task including the query sample and the support set.
[0095] The few-shot learning task in this embodiment includes a query sample and a support set containing multiple types of support samples. Among them, each type of support sample corresponds to a fault type.
[0096] Step 4. According to the few-shot learning task, based on the preset weighted multi-scale wide kernel feature dynamic extraction network, feature extraction is performed to obtain the query sample feature map and multi-class support sample feature maps.
[0097] As Figure 2 shown, the preset weighted multi-scale wide kernel feature dynamic extraction network includes a backbone network and a multi-scale feature extraction module.
[0098] 4.1, Backbone Network.
[0099] The backbone network is used to perform two-dimensional convolution, max pooling, layer normalization, and GELU activation function processing on either the query sample or the support samples of any fault type in the support set, respectively, to obtain a backbone feature matrix.
[0100] Through this backbone network, a backbone feature matrix with an invariant dimension is extracted.
[0101] 4.2, Multi-scale Feature Extraction Module.
[0102] The multi-scale feature extraction module includes:
[0103] Receptive field units, which are used to perform multiple convolution operations based on different receptive fields on the backbone feature matrix and then perform normalization processing to obtain multiple normalized receptive field feature matrices; among them, the multiple convolution operations include large kernel-like convolution, depth convolution, extended depth convolution, and point convolution operations; these convolution operations can adaptively adjust the receptive field to capture long-range pixel dependence relationships at different scales.
[0104] Weighted fusion units, which are used to perform multi-scale feature fusion on the backbone feature matrix and multiple normalized receptive field feature matrices based on a weight sharing mechanism to obtain a fused feature matrix. Through the weight sharing mechanism, the contribution weights of different receptive field feature matrices in the overall feature expression are dynamically adjusted to achieve adaptive multi-scale feature fusion and generate a comprehensive feature matrix. This step can effectively integrate multi-scale information and improve the robustness of feature expression.
[0105] Feature enhancement units, which are used to perform two-dimensional convolution and point convolution processing on the fused feature matrix to obtain an enhanced feature matrix. Specifically, first use 3×3 two-dimensional convolution to further enhance the local feature expression, and then refine the channel information through point convolution to obtain an enhanced feature matrix.
[0106] Residual connection units, which are used to perform residual connection on the enhanced feature matrix and the fused feature matrix to obtain a query sample feature map or a support sample feature map corresponding to any class of support samples. Specifically, add the enhanced feature matrix and the fused feature matrix to form a residual connection structure to obtain a query sample feature map or a support sample feature map corresponding to any class of support samples.
[0107] Step 5, Based on the query sample feature map and multiple classes of support sample feature maps, perform fault diagnosis using the trained fault diagnosis model to obtain corresponding fault diagnosis results.
[0108] The fault diagnosis model includes a global metric module based on the attention mechanism, a local metric module based on the deep Brown distance, and a classifier.
[0109] 5.1. Global metric module based on attention mechanism.
[0110] As Figure 3 shown, the global metric module based on attention mechanism includes:
[0111] 5.1.1. Support dual attention unit, which is used to perform feature extraction based on the channel attention mechanism and spatial attention mechanism according to the feature map of any class of support samples, and obtain the support dual attention weighted result.
[0112] As Figure 4 shown, according to the feature map of any class of support samples, after performing average pooling and max pooling in the spatial dimension, and then performing channel attention weighting operation, the support channel attention weighted result is obtained; according to the support channel attention weighted result, after performing average pooling operation and max pooling operation in the channel dimension, and then performing spatial attention weighting operation, the support dual attention weighted result is obtained.
[0113] Denote the feature map of any class of support samples as , where represents the number of support samples corresponding to any class of support samples, represents the number of channels, <s represents the height of the feature map, represents the width of the feature map.
[0114] S1. Perform global pooling on the support sample feature map , and perform average pooling in the spatial dimension including the feature map height and the feature map width to obtain the support spatial average pooling result:
[0115]
[0116] In the formula, represents the support spatial average pooling result; represents the spatial average pooling operation; i represents the variable index regarding the feature map height; j represents the variable index regarding the feature map width.
[0117] At the same time, perform max pooling operation in the spatial dimension including the feature map height and the feature map width to obtain the support spatial max pooling result, and the expression is as follows:
[0118]
[0119] In the formula, represents the support spatial max pooling result; represents the spatial max pooling operation.
[0120] S2. Input the results of spatial average pooling and the results of spatial max pooling into the fully connected layer for fully connected processing. The number of neurons in the two fully connected layers is and , where represents the number of channels; represents the spatial hyperparameter with a value of 2. After inputting into the fully connected layer, use the Sigmoid activation function to generate the support channel attention weight , .
[0121] S3. Multiply the support channel attention weight with the support sample feature map to obtain the support channel attention weighted result , . ]
[0122] S4. According to the support channel attention weighted result , perform average pooling operation and max pooling operation respectively on the channel dimension to generate two single-channel features, and then obtain the support channel average pooling result and the support channel max pooling result , , .
[0123] S5. According to the support channel average pooling result and the support channel max pooling result , perform splicing processing on the channel dimension to obtain the support dual-channel result , ; Input the support dual-channel result into the convolutional layer to compress the channel dimension to obtain the compressed support dual-channel result , , and then input it into the Sigmod activation function for normalization to obtain the support spatial attention weight , .
[0124] S6. Multiply the support spatial attention weight with the corresponding support channel attention weight to obtain the support dual-attention weighted result , .
[0125] 5.1.2. Support self-attention unit, which is used to process based on the self-attention mechanism according to the support dual-attention weighted result to obtain the support sample correlation matrix.
[0126] The specific steps are as follows:
[0127] First, based on the result of dual attention weighting , perform global average pooling operation to obtain a sequence of support sample feature matrices , .
[0128] Second, according to the sequence of support sample feature matrices , obtain the attention parameter matrix for calculation , including the support query parameter matrix , the support key parameter matrix and the support value parameter matrix .
[0129] In this embodiment, the attention function adopted is the scaled dot-product attention function, and its mathematical expression is:
[0130]
[0131] In the formula, represents the weighted value obtained after attention calculation; represents the operation of the input with the normalized exponential function; represents the query vector; represents the key vector; represents the value vector; c represents the scaling factor.
[0132] Take the support sample feature in the sequence of support sample feature matrices as the input of the attention function, and project the support sample feature onto the , , matrices to obtain the support sample attention matrix .
[0133] Finally, according to the support sample feature and the support sample attention matrix , perform normalization processing to obtain the support sample correlation matrix , .
[0134] 5.1.3. The prototype unit is used to perform a mean operation according to the support sample correlation matrix to obtain the sample prototype feature matrix.
[0135] The expression of the sample prototype feature matrix corresponding to any class of support samples is as follows:
[0136]
[0137] In the formula, represents the mean operation.
[0138] 5.1.4. The query dual-attention unit is used to perform feature extraction based on the query sample feature map according to the channel attention mechanism and the spatial attention mechanism to obtain the query dual-attention weighted result.
[0139] Specifically, according to the query sample feature map, after performing average pooling and max pooling in the spatial dimension, a channel attention weighting operation is performed to obtain the query channel attention weighted result;
[0140] According to the query channel attention weighted result, after performing average pooling operation and max pooling operation in the channel dimension, a spatial attention weighting operation is performed to obtain the query dual-attention weighted result.
[0141] The principle of this step is the same as that of the support dual-attention unit in step 5.1.1, so it will not be elaborated here.
[0142] 5.1.5. The query self-attention unit is used to process according to the query dual-attention weighted result based on the self-attention mechanism to obtain the query sample correlation matrix.
[0143] First, according to the query dual-attention weighted result, a global average pooling operation is performed to obtain the query sample feature matrix sequence .
[0144] Secondly, the query sample features in the query sample feature matrix sequence are used as the input of the attention function, and the query sample features are projected into the , , matrices to obtain the query sample attention matrix .
[0145] Finally, according to the query sample features and the query sample attention matrix , a normalization process is performed to obtain the query sample correlation matrix , .
[0146] 5.1.6. The cross-attention unit is used to process according to the query sample correlation matrix and the sample prototype feature matrix based on the cross-attention mechanism to obtain the global correlation result between the query sample and any class of support samples.
[0147] Specifically, the projection obtained from the query sample correlation matrix is used as the query component For the sample prototype feature matrix corresponding to the nth type of support sample , the corresponding key component is obtained through projection and the value component .
[0148] According to the query component , the key component and the value component , calculate the correlation matrix between the query sample and the prototype of the corresponding class of support samples , . The correlation matrix is flattened into a vector of length , and input into a fully connected network for a fully connected operation to obtain the global correlation result between the query sample and the nth type of support sample .
[0149] 5.2. Local metric module based on the deep Brownian distance
[0150] As shown in Figure 5 , the local metric module based on the deep Brownian distance includes
[0151] 5.2.1 Local support sample unit, which is used to calculate the BDC matrix of the feature map of any class of support samples and perform a mean operation to obtain the sample prototype BDC matrix corresponding to any class of support samples
[0152] Specifically, as shown in Figure 6 , the calculation steps of the BDC matrix are as follows
[0153] Denote the query sample feature map or the feature map of any class of support samples as the data to be calculated , , where is the number of samples, is the number of channels, is the height of the feature map, is the width of the feature map
[0154] Reshape the feature map tensor in the data to be calculated to obtain the reshaped tensor .
[0155] Then, regard each row and each column in the reshaped tensor as random vectors respectively, calculate the Euclidean distance between the vectors, and obtain the Euclidean distance matrix , where represents the Euclidean distance between the kth row and the lth column
[0156] For this Euclidean distance matrix Take the square root to obtain , and then subtract the row average, column average, and the average of all elements from to obtain the BDC matrix .
[0157] Given the BDC matrix of the feature map of the nth class of support samples, calculate the mean to obtain the sample prototype BDC matrix corresponding to the nth class of support samples , , where represents the number of support samples included in the nth class of support samples represents the BDC matrix of the kth support sample feature map in the nth class of support samples
[0158] 5.2.2. Local query sample unit, used to calculate the BDC matrix of the query sample feature map
[0159] The principle of this step is the same as the calculation principle of the BDC matrix in the local support sample unit in step 5.2.1, so it will not be elaborated here
[0160] 5.2.3. Local correlation unit, used to perform an inner product operation based on the BDC matrix of the query sample feature map and the sample prototype BDC matrix corresponding to any class of support samples to obtain the local correlation result between the query sample and any class of support samples
[0161] 5.3. Classifier
[0162] The classifier is used for
[0163] Calculate the metric similarity score between the query sample and each class of support samples according to the global correlation result and the local correlation result between the query sample and any class of support samples
[0164] Take the fault type corresponding to the class of support samples with the highest metric similarity score as the corresponding fault diagnosis result
[0165] The expression of the classifier is as follows
[0166]
[0167] where represents the fault type corresponding to the nth class of support samples represents the metric similarity score between the query sample and the nth class of support samples represents the local correlation result between the query sample and the nth class of support samples represents the global correlation result between the query sample and the nth class of support samples represents the trade-off parameter, with a range of 0 to 1
[0168] 5.4. Training of the fault diagnosis model
[0169] In addition, this embodiment also provides a method for training the fault diagnosis model, and the specific steps are as follows:
[0170] 5.4.1. Obtain the training set, including the support training set and the query training set.
[0171] First, obtain the data set, which contains multiple fault types, and each fault type corresponds to multiple sample data. Then, form an N-way K-shot few-shot learning task according to the fault type. N represents the fault type, and K represents the high-dimensional feature matrix of K samples extracted from each category to construct the support training set: The high-dimensional feature matrix of the remaining samples in each category constitutes the query training set where represents the high-dimensional feature matrix of the i-th support training set sample, represents the high-dimensional feature matrix of the query training set sample; represents the corresponding fault type label; represents the corresponding fault type label; NK represents the number of samples in the support training set; NQ represents the number of samples in the query training set.
[0172] 5.4.2. According to the support training set and the support training set, train the pre-constructed fault diagnosis model until the preset training termination condition is met, and output the trained fault diagnosis model.
[0173] In this embodiment, different loss functions are adopted for the two metric modules in the fault diagnosis model.
[0174] For the global metric module based on the attention mechanism, the contrastive loss function is adopted, and the expression is as follows:
[0175]
[0176] In the formula, represents the contrastive loss function; represents the fault type corresponding to the n-th class of support samples; represents the global correlation result between the query sample and the n-th class of support samples; represents the true label of the support sample; represents the true label of the query sample.
[0177] For the local metric module based on the deep Brown distance, the cross-entropy loss function is adopted for training, and the expression is as follows:
[0178]
[0179] In the formula, represents the cross-entropy loss function; represents the fault type corresponding to the nth class of support samples; represents the local correlation result between the query sample and the nth class of support samples; represents the true label of the query sample; N represents the number of fault types.
[0180] Therefore, the overall loss function The formula is as follows:
[0181]
[0182] In the formula, represents the balance hyperparameter, .
[0183] First, the present application preprocesses the original fault vibration signal, downsamples and segments it through a sliding window, and converts the time-domain information into a time-frequency image using the fast Fourier transform. On this basis, a weighted multi-scale wide-kernel feature dynamic extraction network is used to obtain a feature map, overcoming the limitations of traditional convolution in capturing long-distance dependence relationships, obtaining a spatially enhanced feature representation, and the extracted features are diverse and rich, realizing adaptive multi-scale feature extraction, while optimizing the calculation efficiency of large kernels.
[0184] The fault diagnosis model in the present application integrates channel attention, spatial attention, and self-attention mechanisms, and combines feature normalization and a multi-layer perceptron to realize deep feature interaction and measurement between the support set and the query sample. A sample prototype is constructed based on the self-attention mechanism, and cross-attention is used to calculate the correlation in the feature space, which can not only automatically focus on the key features with the most diagnostic value, but also better capture the potential associations between samples in the case of scarce samples, and has strong practical value and application prospects.
[0185] In addition, different from the traditional first-order moment Euclidean distance metric method, the present application innovatively introduces a deep Brown distance that combines marginal distribution and joint distribution features in the local metric module, learns the feature representation by measuring the difference between the joint feature function and the product of marginal distributions, naturally quantifies the dependence between two random variables, and thus more accurately reflects the deep association between the support set and the query sample, and more comprehensively measures the sample relationship. This method makes full use of limited samples, learns the relationship between the support set and the query sample from both local and global dimensions, and significantly improves the fault diagnosis accuracy and reliability in the small-sample scenario across working conditions.
[0186] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0187] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.
[0188] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.
[0189] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.
[0190] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A device fault diagnosis method based on the deep Brown distance and the attention mechanism, characterized in that It includes the following steps: Obtain the preprocessed fault vibration signal data of the device to be diagnosed; Based on the fault vibration signal data, generate corresponding time-frequency image data through fast Fourier transform and spectrum feature calculation; Use the time-frequency image data as a query sample, and combine it with a preset support set containing multiple types of support samples to construct a few-shot learning task including the query sample and the support set; According to the few-shot learning task, perform feature extraction based on a preset weighted multi-scale wide kernel feature dynamic extraction network to obtain a query sample feature map and multi-class support sample feature maps; Based on the query sample feature map and multi-class support sample feature maps, perform fault diagnosis based on a trained fault diagnosis model to obtain corresponding fault diagnosis results; wherein, the fault diagnosis model includes a global metric module based on an attention mechanism, a local metric module based on a deep Brownian distance, and a classifier.
2. The device fault diagnosis method based on the deep Brownian distance and the attention mechanism according to claim 1, wherein The obtaining of the fault vibration signal data of the device to be diagnosed includes the following steps: Obtain the original fault vibration signal at a preset position of the device to be diagnosed; Based on the original fault vibration signal, perform cutting based on a preset sliding window to obtain multiple fault vibration signal segments of a fixed length; Based on the fault vibration signal segments corresponding to all positions, obtain the fault vibration signal data of the device to be diagnosed.
3. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 1, characterized in that The generating of the corresponding time-frequency image data includes the following steps: For any one of the fault vibration signal segments in the fault vibration signal data, obtain the corresponding complex spectrum through fast Fourier transform; Based on the complex spectrum, obtain the corresponding time-frequency feature image through power spectrum calculation and logarithmic scale conversion; For all the time-frequency feature images, perform normalization and size unification processing to convert and obtain the time-frequency image data.
4. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 1, characterized in that The preset weighted multi-scale wide kernel feature dynamic extraction network includes a backbone network and a multi-scale feature extraction module, The backbone network is used to perform two-dimensional convolution, max pooling, layer normalization, and GELU activation function processing respectively on any one type of support sample in the query sample or the support set to obtain a backbone feature matrix; The multi-scale feature extraction module includes: A receptive field unit, which is used to perform multiple convolution operations based on different receptive fields on the backbone feature matrix and then perform normalization processing to obtain multiple normalized receptive field feature matrices; wherein, the multiple convolution operations include large kernel-like convolution, depth convolution, extended depth convolution, and point convolution operations; A weighted fusion unit, which is used to perform multi-scale feature fusion on the backbone feature matrix and multiple normalized receptive field feature matrices based on a weight sharing mechanism to obtain a fused feature matrix; A feature enhancement unit, which is used to perform two-dimensional convolution and point convolution processing on the fused feature matrix to obtain an enhanced feature matrix; A residual connection unit, which is used to perform residual connection on the enhanced feature matrix and the fused feature matrix to obtain a query sample feature map or a support sample feature map corresponding to any one type of support sample.
5. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 1, characterized in that The global metric module based on the attention mechanism includes: Support dual attention unit, which is used to perform feature extraction based on any class of support sample feature maps according to the channel attention mechanism and the spatial attention mechanism to obtain the support dual attention weighted result; Support self-attention unit, which is used to process according to the support dual attention weighted result based on the self-attention mechanism to obtain the support sample correlation matrix; Prototype unit, which is used to perform a mean operation according to the support sample correlation matrix to obtain the sample prototype feature matrix; Query dual attention unit, which is used to perform feature extraction based on the query sample feature map according to the channel attention mechanism and the spatial attention mechanism to obtain the query dual attention weighted result; Query self-attention unit, which is used to process according to the query dual attention weighted result based on the self-attention mechanism to obtain the query sample correlation matrix; Cross-attention unit, which is used to process according to the query sample correlation matrix and the sample prototype feature matrix based on the cross-attention mechanism to obtain the global correlation result between the query sample and any class of support samples.
6. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 5, characterized in that The support dual attention unit is used for: According to any class of support sample feature maps, perform average pooling and max pooling in the spatial dimension, and then perform channel attention weighting operation to obtain the support channel attention weighted result; According to the support channel attention weighted result, perform average pooling operation and max pooling operation in the channel dimension, and then perform spatial attention weighting operation to obtain the support dual attention weighted result; The support self-attention unit is used for: According to the support dual attention weighted result, perform global average pooling operation to obtain the support sample feature matrix sequence; According to the support sample feature matrix sequence, perform feature projection and normalization processing based on the self-attention mechanism to obtain the support sample correlation matrix.
7. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 5, characterized in that The query dual attention unit is used for: According to the query sample feature map, perform average pooling and max pooling in the spatial dimension, and then perform channel attention weighting operation to obtain the query channel attention weighted result; According to the query channel attention weighted result, perform average pooling operation and max pooling operation in the channel dimension, and then perform spatial attention weighting operation to obtain the query dual attention weighted result; The query self-attention unit is used for: According to the query dual attention weighted result, perform global average pooling operation to obtain the query sample feature matrix sequence; According to the query sample feature matrix sequence, perform feature projection and normalization processing based on the self-attention mechanism to obtain the query sample correlation matrix.
8. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 1, characterized in that, The local metric module based on the deep Brownian distance includes: Local support sample unit, which is used to calculate the BDC matrix of any class of support sample feature maps and perform a mean operation to obtain the sample prototype BDC matrix corresponding to any class of support samples; Local query sample unit, which is used to calculate the BDC matrix of the query sample feature map; Local correlation unit, which is used to perform an inner product operation according to the BDC matrix of the query sample feature map and the sample prototype BDC matrix corresponding to any class of support samples to obtain the local correlation result between the query sample and any class of support samples.
9. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 8, characterized in that, The calculation steps of the BDC matrix are as follows: According to the query sample feature map or any class of support sample feature maps, perform reshaping on the feature map tensor shape to obtain the reshaped tensor; According to the reshaped tensor, obtain random vectors and calculate the Euclidean distance between the vectors to obtain the Euclidean distance matrix; According to the Euclidean distance matrix, perform square root and mean subtraction operations to obtain the BDC matrix of the query sample feature map or the BDC matrix of any class of support sample feature maps.
10. The device fault diagnosis method based on the deep Brown distance and the attention mechanism according to claim 1, characterized in that, The classifier is used for: According to the global correlation result and local correlation result between the query sample and any class of support samples, calculate the metric similarity score between the query sample and each class of support samples; Take the fault type corresponding to the class of support samples with the highest metric similarity score as the corresponding fault diagnosis result.
Citation Information
Cited By
Green ammonia reactor fault mode identification method and device
CN120687945A
Power fault identification method and system based on small sample target detection
CN121053441A