Vision Transform-based gear sound and vibration signal fusion fault diagnosis method

By adopting the acoustic and vibrating signal fusion method based on Vision Transformer in gear fault diagnosis, the problem of fault diagnosis under speed fluctuation is solved, and the cross-domain fault diagnosis effect with high accuracy and robustness is achieved.

CN120012033AActive Publication Date: 2025-05-16SHANDONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510506363.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively diagnose gear failures under speed fluctuation conditions, especially because Transformer is not comprehensive enough when dealing with multi-dimensional features, resulting in insufficient model accuracy and operating efficiency.

Method used

The gear acoustic and vibration signal fusion fault diagnosis method is adopted based on Vision Transformer. By dividing the acoustic and vibration signals into multiple vector blocks and embedding them, dynamic allocation can learn position encoding, adjust the weight of the position encoding in combination with the position-aware scaling factor, and feature extraction and fusion are performed in an independent ViT module, and finally cross-domain fault diagnosis is used using Mahalanobis distance.

Benefits of technology

By fully utilizing the complementary characteristics of the acoustic and vibrating signals, combined with Vision Transformer and Mahalanobis distance, accurate cross-domain fault diagnosis under speed fluctuation conditions is achieved, which significantly improves diagnostic accuracy and improves the robustness and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012033A_ABST
    Figure CN120012033A_ABST
Patent Text Reader

Abstract

The invention discloses a gear sound and vibration signal fusion fault diagnosis method based on Vision Transform, and relates to the technical field of fault diagnosis of vibration signals and acoustic signals of rotating machinery, and the method comprises the steps: segmenting a sound and vibration signal into a plurality of small blocks, carrying out the embedding based on linear mapping, and generating an input sequence; dynamically distributing a learnable position code for each position of the input sequence, and adjusting the weight of the position code to the input sequence in combination with a position sensing scaling factor; respectively carrying out feature extraction and fusion on the sound signal and the vibration signal in the independent ViT module; fusing the sound signal features and the vibration signal features to obtain splicing features; a Mahalanobis distance is used to measure inter-domain feature distribution difference, and cross-domain fault diagnosis is carried out in combination with a fault classifier and a domain classifier. According to the method, a Vision Transform fusion network is used, and sound and vibration signal features are synchronously extracted and fused; learable position scaling and a double-feature extraction channel self-adaptive technology are combined, and collaborative cross-domain fault diagnosis of the gear is achieved under the rotating speed fluctuation working condition of the rotating machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault diagnosis of rotating machinery vibration signals and acoustic signals, and in particular to a gear acoustic-vibration signal fusion fault diagnosis method based on Vision Transformer. Background Art

[0002] The health status of gears directly affects the stability and reliability of equipment operation. Fault diagnosis methods that use signal processing technology to extract fault features have been widely used in this field. Many researchers use Transformer as a framework for fault research. Transformer has demonstrated powerful long-sequence processing capabilities and parallel computing advantages in the field of natural language processing. The Vision Transformer (ViT) aims to introduce the architectural advantages of Transformer into computer vision, break the long-term dominance of CNN in visual tasks, and explore a new way to process visual data such as images.

[0003] In ViT, each image block is input into the model after being embedded and other operations. By default, all image blocks are treated equally, and the contribution to subsequent calculations is not dynamically adjusted according to the position of the image block in the overall image and the importance of the features it contains. This may result in some image block features that are critical to tasks such as classification and detection not being highlighted, while some relatively unimportant or noisy image block features are equally involved in the calculation, interfering with the model's extraction and utilization of effective features, and reducing the overall accuracy and operating efficiency of the model; secondly, the Transformer's multi-head attention mechanism focuses on feature interactions and relationship capture in the sequence dimension, but does not fully utilize the multi-dimensional features of the data. Summary of the invention

[0004] In order to solve the technical problem in the prior art that it is difficult to effectively diagnose gear faults under speed fluctuations, the present invention discloses a gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer.

[0005] To achieve the above purpose, the present invention adopts the following technical scheme: a gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer, comprising the following steps: S1, dividing the acoustic signal and the vibration signal into multiple vector blocks and embedding them based on linear mapping, generating an acoustic signal input sequence and a vibration signal input sequence respectively; the formula is: ; In the formula, x pis a block in the segmented block sequence, and Z0 is the patch embedding; S2, dynamically assigns a learnable position code to each position of the acoustic signal input sequence and the vibration signal input sequence, and adjusts the weight of the position code in combination with the position-aware scaling factor; S3, extracts and fuses the acoustic signal and vibration signal after adjusting the weights in independent ViT modules respectively; S4, fuses the acoustic signal features and vibration signal features obtained in step S3 to obtain the splicing features; S5, uses the Mahalanobis distance to measure the difference in feature distribution between domains, and combines the fault classifier and the domain classifier for cross-domain fault diagnosis.

[0006] Optionally, in step S2, the generated position index is introduced into the sequence, and the corresponding learnable position encoding vector is obtained by searching the position encoding matrix. During the forward propagation process, the learnable position encoding of the current sequence is first calculated, then added to the input sequence, and finally the adjusted position encoding is multiplied by the position-aware scaling factor to adjust its influence on the input, that is, the weight, to form a new representation. This multiplication operation enables the model to flexibly enhance or weaken the influence of the position encoding on the final output during training, thereby improving the model's ability to adapt to position information in sequence data. Specifically including: S21, generating learnable position encoding, introducing the generated position index into the input sequence, and obtaining the corresponding learnable position encoding vector by searching the position encoding matrix; the formula is: PE(i)=Embedding(i); ; In the formula, is the encoder input sequence, R is a set of real numbers, L is the sequence length, D is the embedding dimension, is the classification head vector, is the linear mapping matrix, P and Respectively represent the number of segments and the feature dimension of each segment, is a segmented block sequence, represents the pth block in the segmented block sequence, PE(i) is a learnable embedding function used to generate the embedding vector of position i, The position encoding matrix is ​​generated by PE(i), Embedding represents an embedding operation or an embedding matrix, which is used to convert discrete index values ​​into continuous vector representations; S22, add the learnable position encoding to the input sequence and multiply it by the position-aware scaling factor, the formula is: ; In the formula, is the input sequence tensor, where B is the batch size, L is the sequence length, is the number of channels, α is the position-aware scaling factor, is the position-aware scaling factor The loss function of the gradient, b is the index variable of the batch size, D is the embedding dimension, d is the index variable of the embedding dimension, and c is the index variable of the number of channels. Represents the output value of the b-th batch, c-th channel, and l-th position.

[0007] Optionally, in step S3, the weighted acoustic signal and vibration signal are first transmitted to the ViT module, i.e., two independent feature extraction channels, each of which extracts the feature representation of the dimension. In this process, each feature extraction channel includes a Z-Pool layer, a convolution layer, a batch normalization layer, and an activation function. Among them, the Z-Pool layer is used to reduce the feature dimension while retaining important information; the convolution layer effectively captures the spatial features in the acoustic vibration signal through local connections and weight sharing; the batch normalization layer helps to accelerate training and improve the stability of the model; and the activation function introduces nonlinearity. The dimension reduction process of the Z-Pool layer is: ; ; ; Where i, j, k are the coordinates of the output features, m, n are the relative coordinates within the pooling window, M and N are the height and width of the pooling window, is the maximum pooling operation, is an average pooling operation, Z-Pool(h) is a feature dimensionality reduction operation, Z represents the operation of concatenating the results of maximum pooling and average pooling, and Pool(h) is a pooling operation on the input sequence vector h. Fusion in this step refers to the feature fusion of acoustic signal and acoustic signal, vibration signal and vibration signal after maximum pooling and average pooling respectively. After the convolutional layer feature extraction, weighting is performed, and the weighted features are combined and fused by element-by-element multiplication. This fusion method can not only effectively integrate information of different dimensions, but also enhance the feature expression ability of the model. The element-by-element multiplication operation can be regarded as a gating mechanism. By controlling the mutual influence of different features, the model can adaptively focus on more important features, thereby improving the overall performance. This design can improve the performance of the model in complex tasks, making it more robust and adaptable.

[0008] Optionally, in step S4, the acoustic signal and vibration signal processed by the convolution layer of step S3 are respectively encoded using M layers of Vision Transformer encoders, each layer of the encoder includes a feedforward network and a multi-head attention mechanism, the feedforward network consists of two fully connected layers, which is used to enable the model to perform nonlinear transformation independently at each position; the multi-head attention is used to enable the model to pay attention to different positions in the sequence, and to capture contextual features and relationships by calculating the similarity between elements in the input sequence. In the output process of the encoder, the output of the multi-head attention is combined with the output of the feedforward network through a residual connection. This residual connection not only helps to alleviate the gradient vanishing problem in the deep network, but also improves the efficiency of information flow; after layer-by-layer processing by the encoder, the network extracts specific feature identifiers and performs layer normalization on these features to improve the convergence speed and stability of the model; finally, the extracted and normalized acoustic signal and vibration signal are horizontally spliced ​​to complete feature fusion; the operating formula of the multi-head attention mechanism and the residual connection is: ; The operation formula of the feedforward network and residual connection is: ; The operation formula of layer normalization is: ; In the formula, is the corresponding sequence, is the internal output sequence of the M-layer encoder, MSA is multi-head attention, MLP is the feedforward network, MN is the layer normalization function, y is the final output sequence after being processed by the layer normalization function MN, is the initial input sequence of the Mth layer encoder, It is The internal output sequence of the layer encoder is taken as In this step, fusion refers to the fusion of features of the acoustic signal and the vibration signal through the multi-head attention mechanism and the feedforward network after the fusion is completed in the ViT module.

[0009] Step S4 also includes dividing each vector of the concatenated features by the Euclidean norm (L2 norm) to make the vector length 1, which is used to eliminate the scale difference between different features and improve the stability and accuracy of the model; the formula is: ; In the formula, U represents a vector and u represents the component of the vector.

[0010] Optionally, in step S5, the fusion feature of the acoustic vibration signal is quantified by Mahalanobis distance, thereby improving the robustness of the domain invariant feature. The calculation formula of Mahalanobis distance is: ; In the formula, is the mean vector, is the covariance matrix, e is the eigenvector of the input, is the difference vector between the vector e and the mean vector μ, is the Mahalanobis distance. The fused extracted features are passed to the domain classifier and the fault classifier. The fault classifier is built by a fully connected layer to establish a mapping relationship between collaborative features and fault categories. Based on the optimization method of Mahalanobis distance, the loss function is defined as: ; In the formula, is category y i The covariance matrix of is a hyperparameter, is the regularization term.

[0011] The beneficial effect of the present invention is that the method of the present invention first divides the acoustic signal and vibration signal input vectors into fixed-size, non-overlapping fragments or patches, and each block is flattened into a vector to constitute the input sequence of the model; before the feature fusion stage, a position-aware scaling factor is assigned to each position of the input sequence, so that the model can dynamically adjust the weight of the feature according to the importance of the position, improve the model's ability to capture multi-dimensional features, and can adaptively assign weights to different elements in the channel and sequence; then, the obtained data is directed to the dual feature extraction channel of the ViT module, each feature extraction channel independently processes the input sequence to extract the feature representation of the dimension, and in the feature fusion stage, the feature levels of the acoustic signal and vibration signal after each fusion are spliced; in order to solve the scale difference between the feature vectors, each element of the spliced ​​feature is divided by its Euclidean norm (L2 norm) so that the length of the vector becomes 1, the scale difference between different features is eliminated, and the stability and accuracy of the model are improved; in order to measure the difference between different data distributions, the Mahalanobis distance of the fused feature is calculated, and in the final output stage of the model, the fused feature is further processed by the domain classifier and the fault classifier to perform accurate classification tasks.

[0012] The present invention makes full use of the complementary characteristics of acoustic and vibration signals and demonstrates excellent domain alignment capabilities. The model uses the Vision Transformer fusion network to simultaneously extract and fuse acoustic and vibration signal features. At the same time, by innovatively combining learnable position scaling with dual feature extraction channel adaptive technology, the interference of speed fluctuations on feature extraction is effectively eliminated, thereby accurately and intelligently realizing coordinated cross-domain fault diagnosis of gears under the condition of rotating machinery speed fluctuations. Experimental results in the field of gear fault diagnosis show that compared with the existing technology, the present invention achieves a significant improvement in diagnostic accuracy and performs better in domain-invariant feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flow chart of a gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer of the present invention. DETAILED DESCRIPTION

[0014] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0015] A gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer, which uses the following components: a learnable position encoding module, which dynamically optimizes the position encoding according to the intrinsic characteristics of the data and improves the feature expression; a position-aware scaling module, which emphasizes key features and suppresses secondary information; a feature extraction module, which extracts the features of the acoustic and vibration signals through dual feature extraction channels, ensuring the comprehensiveness and accuracy of feature extraction; a Mahalanobis distance measurement module, which takes into account the covariance structure between features, so as to more accurately measure the differences between features; a fault classifier, which accurately identifies and classifies various fault types; and a domain classifier, which accurately distinguishes whether the data comes from the source domain or the target domain. The fault diagnosis method process is as follows: Figure 1 As shown, the method includes the following steps: S1, setting different gear health states, manually adjusting the frequency converter to control the motor speed to fluctuate randomly within a certain range, thereby obtaining the speed fluctuation working condition, and collecting the vibration signals and acoustic signals of different faulty gears under the speed fluctuation working condition; dividing the acoustic vibration signal into multiple small fixed-size, non-overlapping fragments or patches, each block is flattened into a vector, and embedded based on linear mapping to generate an input sequence; the formula is: ; In the formula, x p is a block in the segmented block sequence, and Z0 is the patch embedding.

[0016] S2. Dynamically assign a learnable position code to each position of the input sequence, and adjust the strength of the influence of the position code on the input sequence in combination with the position-aware scaling factor; introduce the generated position index into the sequence, and obtain the corresponding learnable position code vector by searching the position code matrix. During the forward propagation process, the learnable position code of the current sequence is first calculated, then added to the input sequence, and finally the adjusted position code is multiplied by the position-aware scaling factor to adjust the strength of its influence on the input, thereby forming a new representation. This multiplication operation enables the model to flexibly enhance or weaken the influence of the position code on the final output during training, and improves the model's ability to adapt to position information in sequence data. Specifically including: S21. Generate a learnable position code, introduce the generated position index into the input sequence, and obtain the corresponding learnable position code vector by searching the position code matrix; the formula is: PE(i)=Embedding(i); ; In the formula, is the encoder input sequence, R is a set of real numbers, L is the sequence length, D is the embedding dimension, is the classification head vector, is the linear mapping matrix, P and Respectively represent the number of segments and the feature dimension of each segment, is a segmented block sequence, represents the pth block in the segmented block sequence, PE(i) is a learnable embedding function used to generate the embedding vector of position i, The position encoding matrix is ​​generated by PE(i), Embedding represents an embedding operation or an embedding matrix, which is used to convert discrete index values ​​into continuous vector representations; S22, add the learnable position encoding to the input sequence and multiply it by the position-aware scaling factor, the formula is: ; In the formula, is the input sequence tensor, where B is the batch size, L is the sequence length, is the number of channels, α is the position-aware scaling factor, is the loss function for the gradient of the position-aware scaling factor α, b is the index variable for the batch size, d is the index variable for the embedding dimension, D is the embedding dimension, c is the index variable for the number of channels, Represents the output value of the b-th batch, c-th channel, and l-th position.

[0017] S3, extract and fuse the acoustic signal and vibration signal after adjusting the weights in independent ViT modules respectively. Each ViT module performs Z-Pool layer dimensionality reduction, convolution layer feature extraction, batch normalization and nonlinear activation operations in turn, and fuses them by element-by-element multiplication. The formula is: ; ; ; Where i, j, k are the coordinates of the output features, m, n are the relative coordinates within the pooling window, M and N are the height and width of the pooling window, is the maximum pooling operation, is an average pooling operation, Z-Pool(h) is a feature dimensionality reduction operation, Z represents the operation of concatenating the results of maximum pooling and average pooling, and Pool(h) is a pooling operation on the input sequence vector h.

[0018] Specifically, the acoustic signal and vibration signal obtained in step S2 are first transmitted to the ViT module, that is, two independent feature extraction channels, each of which extracts the feature representation of the dimension. In this process, each feature extraction channel contains a Z-Pool layer, a convolution layer, a batch normalization layer, and an activation function. Among them, the Z-Pool layer is used to reduce the feature dimension while retaining important information; the convolution layer effectively captures the spatial features in the acoustic vibration signal through local connections and weight sharing; the batch normalization layer helps to accelerate training and improve the stability of the model; the activation function introduces nonlinearity. After the convolution layer feature extraction, the weighted features are combined and fused by element-by-element multiplication. This fusion method can not only effectively integrate information of different dimensions, but also enhance the feature expression ability of the model. The element-by-element multiplication operation can be regarded as a gating mechanism. By controlling the mutual influence of different features, the model can adaptively focus on more important features, thereby improving the overall performance. This design can improve the performance of the model in complex tasks and make it more robust and adaptable.

[0019] S4, the acoustic signal and vibration signal processed by the convolutional layer in step S3 are respectively encoded using M layers of VisionTransformer. Each layer of the encoder contains a feedforward network and a multi-head attention mechanism. The feedforward network consists of two fully connected layers, which is used to enable the model to perform nonlinear transformation independently at each position to enhance the expressiveness of the features; the multi-head attention is used to enable the model to pay attention to different positions in the sequence, and to capture contextual features and relationships by calculating the similarity between the elements in the input sequence. In the output process of the encoder, the output of the multi-head attention is combined with the output of the feedforward network through a residual connection. This residual connection not only helps to alleviate the gradient vanishing problem in the deep network, but also improves the efficiency of information flow; after layer-by-layer processing by the encoder, the network extracts specific feature identifiers and performs layer normalization on these features to improve the convergence speed and stability of the model; finally, the extracted and normalized acoustic signal and vibration signal are horizontally spliced ​​to complete feature fusion; the operating formula of the multi-head attention mechanism and the residual connection is: ; The operation formula of the feedforward network and residual connection is: ; The operation formula of layer normalization is: ; In the formula, is the corresponding sequence, is the internal output sequence of the M-layer encoder, MSA is multi-head attention, MLP is the feedforward network, MN is the layer normalization function, y is the final output sequence after being processed by the layer normalization function MN, is the initial input sequence of the Mth layer encoder, It is The internal output sequence of the layer encoder is taken as In this step, fusion refers to the fusion of features of the acoustic signal and the vibration signal through the multi-head attention mechanism and the feedforward network after the fusion is completed in the ViT module.

[0020] After fusion, each vector of the concatenated features is divided by the Euclidean norm (L2 norm) to make the vector length 1, which is used to eliminate the scale differences between different features and improve the stability and accuracy of the model; the formula is: ; In the formula, U represents a vector and u represents the component of the vector.

[0021] S5. Use Mahalanobis distance to measure the difference in feature distribution between domains, and combine fault classifier and domain classifier for cross-domain fault diagnosis. The Mahalanobis distance is used to quantify the fusion features of acoustic vibration signals, which improves the robustness of domain invariant features. The calculation formula of Mahalanobis distance is: ; Where μ is the mean vector, S is the covariance matrix, is the difference vector between the vector e and the mean vector μ, is the Mahalanobis distance.

[0022] The fused extracted features are passed to the domain classifier and the fault classifier. The fault classifier is built by a fully connected layer to establish a mapping relationship between collaborative features and fault categories. Based on the optimization method of Mahalanobis distance, the loss function is defined as: ; In the formula, is category y i The covariance matrix of is a hyperparameter, is the regularization term.

[0023] The present invention is further described by collecting acoustic and vibration signals and conducting intelligent diagnosis under speed fluctuation on a specially designed gear fault test bench.

[0024] Seven gear health states are considered, namely normal gear, sun gear pitting, sun gear fracture, sun gear wear, planet gear pitting, planet gear fracture and planet gear wear. The sampling frequency is set to 12.8 kHz, and the gear acceleration signals are collected as data sets at four stable working conditions of 1000 r / min, 1500 r / min, 1800 r / min, 2000 r / min, 1000~1500 r / min, 1500~1800 r / min, 1800~2000 r / min and three fluctuating working conditions. Each health state has 100 acoustic samples and 100 vibration signal samples, and each sample contains 2048 sample points. 50% of the data set is used as the training set, and the rest is used as the test set.

[0025] In order to evaluate the effectiveness of the proposed method, three latest cross-domain fault diagnosis methods are used for comparison, including the distance-based deep transfer learning network WD-DTL, the feature-based transfer neural network FTNN, and the dual-channel parallel adversarial network DPAN. Under the six migration tasks of 1000 rpm→1000 rpm~1500 rpm (task 1), 1500 rpm→1000 rpm~1500 rpm (task 2), 1500 rpm→1500 rpm~1800 rpm (task 3), 1800 rpm→1800 rpm~2000 rpm (task 4), 1800 rpm→1800 rpm~2000 rpm (task 5), and 2000 rpm→1800 rpm~2000 rpm (task 6), the diagnostic accuracy of the four methods is compared, where the left side of the arrow is the source domain and the right side of the arrow is the target domain. The comparison results of the diagnostic accuracy of the four methods are shown in Table 1 below:

[0026] Table 1 Comparison of diagnostic accuracy of the four methods

[0027] It can be clearly seen from the experimental results that the diagnostic accuracy of FTNN is significantly lower than that of other comparison methods. In most diagnostic tasks, its accuracy is less than 94%. In sharp contrast, DPAN has a diagnostic accuracy of more than 93% in all tasks, thanks to the time-frequency domain signal fusion technology, which can more accurately identify fault features and thus improve the accuracy of diagnosis. Of particular note is the WD-DTL method, which showed good performance in this comparison, with a maximum accuracy of 94.3%. This outstanding performance may be due to its use of Mahalanobis distance as a metric, which can more efficiently learn the distribution differences between the source domain and the target domain. When WD-DTL processes data under different working conditions, it can better adapt to the feature distribution of the target domain, thereby improving the diagnostic accuracy. However, WD-DTL may incur a higher computational cost when processing complex data. Specifically, DPAN performs very well in task performance, achieving a diagnostic accuracy of more than 95% in Task 1, Task 2, and Task 5. However, the performance of the method of the present invention is better than the other three comparison methods in all tasks, with an average diagnostic accuracy of more than 99%. The method of the present invention fully combines the advantages of multi-channel feature fusion, which not only effectively improves the accuracy of the ViT model, but also pays more attention to the robustness of the model. In addition, the multi-channel feature fusion mechanism has unique advantages. It can better process and utilize information from different signal sources, while providing an intuitive way to understand how the model interacts with features of different dimensions, thereby improving the interpretability of the model. In addition, the present invention also utilizes a position-aware scaling module, which enriches the feature representation by assigning a learnable position code to each position, allowing the model to capture more subtle feature differences. In summary, the method of the present invention has significant and effective performance in dealing with migration fault diagnosis problems under gear speed fluctuation conditions.

[0028] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer, characterized in that: The steps include: S1. Divide the acoustic signal and the vibration signal into multiple vector blocks and embed them based on linear mapping to generate an acoustic signal input sequence and a vibration signal input sequence respectively; the formula is: ; In the formula, x p is a block in the segmented block sequence, and Z0 is the patch embedding; S2, dynamically assigning a learnable position code to each position of the acoustic signal input sequence and the vibration signal input sequence, and adjusting the weight of the position code in combination with the position-aware scaling factor; S3, extracting and fusing the weighted acoustic signal and vibration signal in independent ViT modules respectively; S4, fusing the acoustic signal feature and the vibration signal feature obtained in step S3 to obtain a splicing feature; S5. Use Mahalanobis distance to measure the difference in feature distribution between domains, and combine fault classifier and domain classifier to perform cross-domain fault diagnosis.

2. The gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer according to claim 1 is characterized in that: Step S2 includes: S21. Generate a learnable position code, introduce the generated position index into the input sequence, and obtain the corresponding learnable position code vector by searching the position code matrix; the formula is: PE(i)=Embedding(i); ; In the formula, is the encoder input sequence, R is a set of real numbers, L is the sequence length, D is the embedding dimension, is the classification head vector, is the linear mapping matrix, P and Respectively represent the number of segments and the feature dimension of each segment, is a segmented block sequence, represents the pth block in the segmented block sequence, PE(i) is a learnable embedding function used to generate the embedding vector of position i, The position encoding matrix is ​​generated by PE(i), and Embedding represents an embedding operation or an embedding matrix, which is used to convert discrete index values ​​into continuous vector representations; S22. Add the learnable position encoding to the input sequence and multiply it by the position-aware scaling factor. The formula is: ; In the formula, is the input sequence tensor, where B is the batch size, L is the sequence length, is the number of channels, α is the position-aware scaling factor, is the loss function for the gradient of the position-aware scaling factor α, b is the index variable of the batch size, D is the embedding dimension, d is the index variable of the embedding dimension, c is the index variable of the number of channels, Represents the output value of the b-th batch, c-th channel, and l-th position.

3. The gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer according to claim 1 is characterized in that: In step S3, the acoustic signal and the vibration signal after weight adjustment are sequentially subjected to Z-Pool layer dimensionality reduction, convolution layer feature extraction, batch normalization and nonlinear activation operations, weighted after convolution layer feature extraction, and weighted features are combined and fused by element-by-element multiplication. The Z-Pool layer dimensionality reduction process is as follows: ; ; ; Where i, j, k are the coordinates of the output features, m, n are the relative coordinates within the pooling window, M and N are the height and width of the pooling window, is the maximum pooling operation, is an average pooling operation, Z-Pool(h) is a feature dimensionality reduction operation, Z represents the operation of concatenating the results of maximum pooling and average pooling, and Pool(h) is a pooling operation on the input sequence vector h.

4. The gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer according to claim 1 is characterized in that: In step S4, the acoustic signal and vibration signal processed by the convolution layer in step S3 are respectively encoded using M layers of Vision Transformer. Each layer of the encoder includes a feedforward network and a multi-head attention mechanism. In the output process of the encoder, the output of the multi-head attention is combined with the output of the feedforward network through a residual connection. After layer-by-layer processing by the encoder, the network extracts specific feature identifiers and performs layer normalization on these features. Finally, the extracted and normalized acoustic signal and vibration signal are horizontally spliced ​​to complete feature fusion. process for: ; ; ; In the formula, is the corresponding sequence, is the internal output sequence of the M-layer encoder, MSA is multi-head attention, MLP is the feedforward network, MN is the layer normalization function, y is the final output sequence after being processed by the layer normalization function MN, is the initial input sequence of the Mth layer encoder, It is The internal output sequence of the layer encoder is The input of the layer encoder.

5. The gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer according to claim 1 is characterized in that: Step S4 includes dividing each vector of the concatenated features by the Euclidean norm to make the vector length become 1, which is used to eliminate the scale difference between different features. The formula is: ; Where U represents a vector and u represents the component of the vector.

6. The gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer according to claim 1 is characterized in that: In step S5, the calculation formula of Mahalanobis distance is: ; Where μ is the mean vector, S is the covariance matrix, is the difference vector between the vector e and the mean vector μ, is the Mahalanobis distance.

7. The gear acoustic and vibration signal fusion fault diagnosis method based on Vision Transformer according to claim 6 is characterized in that: The Mahalanobis distance loss function is: ; In the formula, is category y i The covariance matrix of is a hyperparameter, is the regularization term.

Citation Information

Patent Citations

  • Multi-mode Mongolian-Chinese translation method based on cyclic common attention Transform

    CN113657124A

  • Bearing fault diagnosis method based on SDP and visual Transform coding

    CN117009770A

  • Deep learning-based panoramic dental film disease identification and auxiliary diagnosis method and system

    CN117934433A

  • Multi-sensor fusion gear fault diagnosis method

    CN118133150A

  • Multi-class welding spot defect classification method, system and equipment based on Transform and channel interaction and medium

    CN119851043A

Cited By

  • Bearing sound and vibration fusion cross-working-condition fault diagnosis method based on KA-Transform

    CN122132820A

  • A bearing acoustic-vibration fusion cross-condition fault diagnosis method based on KA-Transformer

    CN122132820B