Rotating machine fault diagnosis method based on multi-classifier progressive network
Through the fault diagnosis method of multi-classifier progressive network, GFE, IPEA and PDCE are used to improve feature extraction and expression capabilities, solving the accuracy and stability of rotary mechanical fault diagnosis under limited sample conditions, especially in the case of noise and sample imbalance.
Patent Information
- Application Number
- CN202510276972.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Under limited sample conditions, existing rotary machinery fault diagnosis methods are difficult to effectively extract complex feature information, resulting in insufficient diagnostic accuracy and stability, especially poor performance in the case of noise interference and sample imbalance.
The fault diagnosis method based on multi-classifier progressive network (PNMC) is adopted, including particle feature extraction module (GFE), interactive perception enhancement attention mechanism (IPEA) and progressive dynamic classifier integrated structure (PDCE). Through multi-scale feature extraction, feature interaction perception fusion and dynamic adjustment, feature extraction, expression and learning capabilities are improved.
It significantly improves the accuracy and stability of rotary machinery fault diagnosis, especially in noise and limited sample conditions, and is suitable for fault diagnosis tasks in complex environments.
Smart Images

Figure CN120296589A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for diagnosing faults in rotating machinery, and more particularly to a method for diagnosing faults in rotating machinery based on a Progressive Network with Multiple Classifiers (PNMC). Background Art
[0002] Large rotating machinery, including wind turbines, turbogenerators, and aerospace engines, is the cornerstone of the intelligent industrial era and has promoted the progress of intelligent manufacturing. However, during long-term operation, complex environments and improper operations may cause damage to key components, which not only affects production efficiency but may also lead to serious safety accidents. Therefore, accurately diagnosing faults in key components such as bearings, gears, and rotors is crucial for improving industrial production efficiency and ensuring operational safety.
[0003] Compared with traditional diagnostic methods, deep learning can automatically extract complex features and achieve intelligent fault diagnosis, providing a more efficient diagnostic tool for industrial equipment. Unfortunately, due to high acquisition costs and strict safety requirements, the number of available samples is usually limited, resulting in deep learning having difficulty extracting sufficient feature information, thus affecting the accuracy and stability of diagnosis. Therefore, conducting research on rotating machinery fault diagnosis under limited sample conditions is of great significance.
[0004] Under limited sample conditions, generative models, transfer learning, and few-shot learning have made significant progress in improving fault diagnosis performance. First, generative models can use a limited number of samples to generate highly diverse new samples, providing richer and more diverse feature information for deep learning, thereby significantly enhancing the accuracy and stability of fault diagnosis. Nevertheless, the training process of generative models is relatively complex, and the quality of the generated samples may be uneven, which poses higher requirements for the generalization ability of the diagnostic model. Compared with generative models, transfer learning significantly improves the diagnostic performance of the target domain by using the model or feature representation trained in the source domain. Although transfer learning effectively improves the model performance, its effect depends on the similarity between the target domain and the source domain. When the difference between the two is large, transfer bias may lead to performance degradation. Finally, few-shot learning extracts effective features from limited samples through functional modules and optimization strategies, enhancing the generalization ability of the model and showing significant advantages under limited sample conditions. However, few-shot learning also has certain limitations, such as high design complexity and strict requirements for optimization strategies. Nevertheless, given its accuracy and stability under limited sample conditions, few-shot learning remains the core method of current research, and its potential and challenges are worthy of further exploration.
[0005] Under the condition of limited samples, finite-sample learning enhances the comprehensive utilization ability of fault features from multiple levels by improving feature extraction, strengthening feature expression, and optimizing feature learning, thus significantly improving the accuracy and stability of fault diagnosis. As a key step, multiple methods have been proposed for feature extraction to adapt to complex fault patterns and limited-sample conditions. Specifically, Zhang et al. proposed a compact convolutional neural network that processes vibration signals through a multi-scale feature extraction unit to extract more discriminative features, greatly improving the adaptability to complex fault patterns. Similarly, the MFSFormer proposed by Li et al. extracts multi-scale receptive field features by embedding convolutional layers with different kernel sizes, enhancing the comprehensive expression ability of global and local information and showing excellent stability under noise and limited-sample conditions. In addition, Yao et al. proposed a dual-pooling attention residual network that significantly improves the sensitivity and discriminability of feature extraction by fusing global and local information, thereby improving the accuracy and stability of fault diagnosis. Regarding the long-term dependence relationship of signals, Kumar et al. proposed a multi-size wide-kernel convolutional network that captures long-term dependence relationships through wide-kernel design and combines multi-size convolutional kernels to extract features at different scales, thus showing strong adaptability when processing complex pattern signals. In terms of feature fusion, Zhang et al. proposed a deep learning model based on dual-branch time-frequency fusion that significantly improves the feature extraction ability for complex fault patterns by fusing time-domain and frequency-domain features. Although these methods perform well in feature extraction, there are still problems of insufficient information fusion and interaction between features, which may lead to the loss or redundancy of key information.
[0006] In addition, researchers have proposed various attention mechanisms to improve the model's ability to express key features. Specifically, Hu et al. proposed a multi-scale convolutional neural network that combines multiple attention mechanisms. By integrating channel and spatial attention mechanisms, it improves the ability to identify key features of fault signals. Similarly, Snyder et al. proposed a two-headed transformer that enhances the interaction ability of global and local features by embedding self-attention mechanisms. In addition, the multi-branch multi-scale dynamic convolutional network proposed by Liang et al. combines dynamic attention mechanisms, adaptively adjusts the weights of convolutional kernels to enhance the flexibility and robustness of feature expression, and the multi-branch structure supports the parallel processing of information at different scales, further improving the adaptability to complex signals. In addition, Li et al. developed a novel neural network structure that combines ConvNeXt with a multi-scale dilated attention mechanism to capture multi-scale context information, significantly improving the accuracy and stability of fault diagnosis. However, under noise interference or limited-sample conditions, the effectiveness of attention mechanisms is limited to a certain extent, making it challenging to ensure the stability and generalization ability of the model.
[0007] In terms of feature learning, the research mainly focuses on the design of loss functions and the optimization of regularization mechanisms to improve the learning ability in the scenario of limited samples. In terms of loss function design, Shi et al. proposed an incremental learning loss function adapted to class imbalance, which significantly improved the robustness of industrial flow sample processing. In addition, the meta-learning framework proposed by Li et al. combines intra-class and inter-class optimization, and strengthens the feature learning ability in the few-shot scenario through a new loss function. Wang et al. proposed a new loss function for cross-domain adaptation, which effectively enhanced the feature alignment ability in cross-domain fault diagnosis. Finally, to further improve the generalization performance of the model, Dong et al. optimized the feature space distribution through a dual regularization mechanism, thereby improving the stability of the model in different scenarios. Although the above methods optimize the feature learning process from multiple perspectives, they still face some challenges. Specifically, the generality of the loss function and the balance of the regularization method are the key issues to be considered in model design.
[0008] In summary, in limited-sample learning, the feature extraction module, attention mechanism, and optimization strategy play important roles together. The feature extraction module deeply mines the core information of samples and generates high-quality feature representations. The attention mechanism focuses on key regions and significantly improves the feature expression ability. The optimization strategy further improves the accuracy and stability of the model by strengthening the feature learning ability. Summary of the Invention
[0009] To address the challenges of rotating machinery fault diagnosis under limited samples, the present invention provides a rotating machinery fault diagnosis method based on a progressive network of multi-classifiers, which is used to improve the accuracy and robustness of fault recognition under limited samples. The highlight designs of the proposed PNMC in the present invention include GFE, IPEA, and PDCE, which work together to improve the performance of the model in feature extraction, expression, and learning. Specifically, GFE uses a multi-scale feature extraction (MSFE) module to extract multi-scale features of input samples, combines a feature interaction perception fusion module (FIPF) to capture the interdependencies between features, and then realizes the interaction perception fusion of features, accurately extracts granular feature information, and significantly improves the feature extraction ability of the model. IPEA enhances the information transmission and dependency relationship between features through a masking interaction perception mechanism (MIP), and uses a multi-scale masking multi-head attention mechanism (TMA) to capture local and global information, dynamically adjusts the attention to key information, and improves the feature expression ability of the model. PDCE further improves the feature learning ability of the model by adopting a multi-classifier structure, an integrated loss function, and a dynamic adjustment mechanism, and makes full use of multi-level features.
[0010] The object of the present invention is achieved through the following technical solutions:
[0011] A rotating machinery fault diagnosis method based on a multi-classifier progressive network, comprising the following steps:
[0012] Step 1, build a fault simulation platform and collect vibration signals of rotating machinery through sensors;
[0013] Step 2, use the sliding window technique to divide the collected one-dimensional vibration signals into a training set and a test set;
[0014] Step 3, construct a PNMC, and initialize the weights and biases. Among them, the PNMC includes a granular feature extraction module (Granular feature extraction, GFE), an interaction perception enhanced attention mechanism (Interaction perception enhanced attention mechanism, IPEA), and a progressive dynamic classifier ensemble structure (Progressive dynamic classifier ensemble structure, PDCE). The specific construction steps are as follows:
[0015] Step 31, construct the GFE:
[0016] The GFE consists of a multi-scale feature extraction module (Multi-Scale Feature Extraction, MSFE), a feature interaction perception fusion module (Feature Interaction Perception Fusion, FIPF), an adaptive convolution block (AdaConv), and a residual convolution block (ResConv). Among them: The MSFE consists of a large kernel dilated convolution block LKDConv and a small kernel dilated convolution block SKDConv. LKDConv includes LKDConv a and LKDConv b, and SKDConv includes SKDConv a and SKDConv b; AdaConv includes AdaConv a and AdaConv b; Assuming the input sample is x, after being processed by the MSFE, the output feature is expressed as and are the features output by LKDConv a, LKDConv b, SKDConv a, and SKDConv b respectively. Among them: and include two processing paths, and After feature concatenation in the channel dimension, it is passed to AdaConv a, and After being fused by FIPF, it is transmitted to AdaConv b, and the outputs of the two paths are transmitted to ResConv and added element-wise to generate the fused feature and It includes two processing paths, and After feature concatenation in the channel dimension, it is transmitted to AdaConv a, and After being fused by FIPF, it is transmitted to AdaConv b, and the outputs of the two paths are transmitted to ResConv and added element-wise to generate the fused feature and After feature concatenation in the channel dimension, it is transmitted to IPEA;
[0017] Step 32: Construct IPEA:
[0018] IPEA includes a feature mapping module (FM), a masked interaction perception mechanism (MIP), a multi-scale masked multi-head attention mechanism (TMA), and a skip connection structure, where:
[0019] The FM consists of three groups of 1×1 convolutional layers, which map the input features to multiple spaces while keeping the feature size unchanged. The specific calculation formula is as follows:
[0020] {x α ,x δ ,x ζ}=Conv 1×1 (f LS )
[0021] The MIP consists of a linear transformation, a masking operation, and a superposition fusion, xα and xδ Generate two groups of query, key, and value features through linear transformation and masking operation respectively. The specific calculation formula is as follows:
[0022]
[0023] In the formula, and are respectively xα and xδ The query features generated through linear transformation and masking operation, and are respectively xα and xδ The key features generated through linear transformation and masking operation, V i α and V i δ are respectivelyxα and xδ The value features generated through linear transformation and masking operation, where i represents the index of the attention head, represents the projection matrix, and C represents the number of channels, dhead represents the dimension of each attention head, Mi represents the multi-scale masking matrix; for two sets of query, key, and value features, they are fused through element-wise addition to generate the final result of MIP The specific calculation formula is as follows:
[0024]
[0025] The TMA consists of a multi-scale masking mechanism and a multi-head attention mechanism. The TAM performs an inner product calculation on and and applies masks with different scales to the inner product results of different attention heads, and then obtains the attention distribution through Softmax normalization. The specific calculation formula is as follows:
[0026]
[0027] In the formula, T represents the transpose operation; then, using the attention distribution Ai to perform a weighted sum on V i αδ to obtain the output of a single attention head. The specific calculation formula is as follows:
[0028] head i = A i ·V i αδ
[0029] Finally, the outputs of all attention heads are concatenated and linearly mapped to obtain the final result x αδ of the TAM. The specific calculation formula is as follows:
[0030] x αδ = Cat(head1, head2,..., head h ) · W o
[0031] In the formula, represents the projection matrix;
[0032] Adopting a skip connection structure, add xαδ and x ζ element-wise. The specific calculation formula is as follows:
[0033] x αδζ = x αδ + x ζ
[0034] Step 33: Construct the PDCN:
[0035] The PDCE includes a multi-classifier structure, an integrated loss function, and a dynamic adjustment mechanism, where:
[0036] The multi-classifier structure includes an early classifier, a mid-term classifier, and a late classifier;
[0037] The early classifier is composed of GAP, FC, and Softmax, and learns the basic features of the input samples through the linear combination of shallow features. The output of the early classifier is expressed as:
[0038] y e = Softmax{w e2 ·FC[w e1 ·GAP(f LS ) + b e1 + b e2}
[0039] In the formula, and respectively represent the weights of the first and second layers of the early classifier, and respectively represent the biases of the first and second layers of the early classifier, De represents the output dimension of the first layer of the early classifier;
[0040] The mid-term classifier is composed of GAP, FC, and Softmax, and learns the advanced features of the input samples through the linear combination of middle-level features. The output of the mid-term classifier is expressed as:
[0041] y m = Softmax{w m2 ·FC[w m1 ·GAP(x αδζ ) + b m1 + b m2}
[0042] In the formula, wm1 and wm2 respectively represent the weights of the first and second layers of the mid-term classifier, bm1 and bm2 respectively represent the biases of the first and second layers of the mid-term classifier;
[0043] The late classifier is composed of GAP and Softmax, and can accurately learn the deep features of the input samples. The output of the late classifier is expressed as:
[0044]
[0045] Where exp represents the exponential function, GAP represents the global average pooling operation, K represents the total number of categories, and x R represents the input feature of the late classifier;
[0046] The early classifier adopts the focal loss function, and the calculation formula of the focal loss function is:
[0047] L e = -α t (1 - p t ) γ log(p t )
[0048] Where p t represents the predicted probability of the model for the correct category, α t represents the balance factor, and γ represents the focusing parameter;
[0049] The middle classifier adopts the generalized cross-entropy loss function, and the calculation formula of the generalized cross-entropy loss function is:
[0050]
[0051] Where represents the predicted probability of the model for the true label y, and q is a hyperparameter;
[0052] The late classifier adopts the standard cross-entropy loss function, and the calculation formula of the standard cross-entropy loss is:
[0053]
[0054] Where yk represents the true label of the k-th sample, pk represents the probability that the sample is predicted as the k-th category;
[0055] The calculation formula of the integrated loss function is:
[0056] L = L e + L m + L f
[0057] The calculation formula of the dynamic weight adjustment mechanism is:
[0058]
[0059] Where w e , w m and w f respectively represent the loss weights of the early, middle, and late classifiers, t represents the number of iterations, η represents the learning rate, represents the gradient of the current loss L with respect to the loss weight of the early classifier;
[0060] Step 4: Use the training set to train PNMC to optimize the model parameters, where: During the training process, the calculation formula for the integrated dynamic loss is:
[0061] L = w e ·L e (y e , y true ) + w m ·L m (y m , y true ) + w f ·L f (y f , y true )
[0062] In the formula, ytrue represents the true label;
[0063] Step 5: Conduct a fault diagnosis test through the test set to generate the final diagnosis result.
[0064] Compared with the prior art, the present invention has the following advantages:
[0065] 1. The present invention proposes a method for rotating machinery fault diagnosis based on a multi-classifier progressive network. First, a particle feature extraction module is proposed, which effectively captures particle feature information through multi-scale feature extraction and interactive perception fusion, thereby enhancing the model's ability to perceive tiny fault features. In addition, an interactive perception enhanced attention mechanism is proposed, which dynamically adjusts the model's attention to key information through the interactive perception and adaptive enhancement of multi-space features, thereby significantly improving the expression quality of complex fault features. Finally, a progressive dynamic classifier integration structure is proposed, which effectively integrates multi-level feature information by gradually introducing early, middle, and late classifiers and integrating dynamic loss, thereby enhancing the model's ability to learn fault features.
[0066] 2. The present invention evaluates the performance of the proposed method on the Paderborn dataset, the aero-engine test platform dataset, and the laboratory self-collected dataset respectively. The results show that PNMC performs outstandingly in the fault diagnosis tasks of the three datasets. Especially under the conditions of noise & limited samples, its classification performance and anti-noise performance are better than other methods. Further, PNMC has high practical value in industrial applications and is suitable for fault diagnosis tasks in complex environments.
[0067] 3. The present invention not only has practical application value in the field of rotating machinery fault diagnosis under limited sample conditions, but also provides a new research idea for other classification and diagnosis tasks facing limited sample constraints. Description of the Drawings
[0068] Figure 1 It is a fault diagnosis method based on PNMC.
[0069] Figure 2 It is the structural diagram of GFE.
[0070] Figure 3 It is the structural diagram of MSFE.
[0071] Figure 4 It is the structural diagram of FIPF.
[0072] Figure 5 It is the structural diagram of MIP and TMA in IPEA.
[0073] Figure 6 It is the test platform for the PU dataset.
[0074] Figure 7 It is the test platform for the LabSC dataset.
[0075] Figure 8 It is the test platform for the AETP dataset.
[0076] Figure 9 It is the visualization of the feature classification results of PNMC on TP1, TP2, TP3, and TP4.
[0077] Figure 10 It is the visualization of the feature classification results of PNMC on TL1, TL2, TL3, and TL4.
[0078] Figure 11 It is the visualization of the feature classification results of PNMC on TA1, TA2, TA3, and TA4.
[0079] Figure 12 It is the ablation experiment results on the PU dataset.
[0080] Figure 13 It is the ablation experiment results on the LabSC dataset.
[0081] Figure 14 It is the ablation experiment results on the AETP dataset. Specific implementation manners
[0082] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings, but it is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention shall be covered by the protection scope of the present invention.
[0083] The present invention provides a rotating machinery fault diagnosis method based on a multi-classifier progressive network, as Figure 1As shown, it includes three steps: data collection and preprocessing, model training, and model testing. First, a fault simulation platform is built to collect vibration signals of rotating machinery through sensors, and the one-dimensional signals are segmented into a training set and a test set using the sliding window technique. Second, the training set is used to train the PNMC to optimize the model parameters. Finally, a fault diagnosis test is performed through the test set to generate the final diagnosis result. The specific steps are as follows:
[0084] Step 1: Build a fault simulation platform and collect vibration signals of rotating machinery through sensors.
[0085] Step 2: Use the sliding window technique to divide the collected one-dimensional vibration signals into a training set and a test set.
[0086] Step 3: Construct the PNMC and initialize the weights and biases. Among them, the specific steps to construct the PNMC are as follows:
[0087] Step 31: Construct the GFE:
[0088] As Figure 2 shown, the GFE mainly consists of MSFE, FIPF, AdaConv, and ResConv, which is used to enhance the feature extraction ability of the model. AdaConv includes AdaConv a and AdaConv b, both of which consist of a convolutional layer, batch normalization (BN), and a rectified linear unit (ReLU), and can adaptively change the number of channels of the input features as needed. ResConv consists of a convolutional layer, BN, ReLU, and a skip connection structure, which further realizes the fusion of global and local features and effectively prevents the degradation problem that may occur during the model training process. Assuming the input sample is x, after being processed by MSFE, the output features can be expressed as and are the features output by LKDConv a, LKDConv b, SKDConv a, and SKDConv b respectively. and include two processing paths, and After feature concatenation in the channel dimension, they are passed to AdaConv a, and After being fused by FIPF, they are passed to AdaConv b. The outputs of the two paths are passed to ResConv and added element by element to generate more detailed fused features thus enhancing the model's ability to capture global important information. Similarly, and also include two processing paths, and After feature concatenation in the channel dimension, it is passed to AdaConv a. and After fusion through FIPF, it is passed to AdaConv b. The outputs of the two paths are passed to ResConv and added element-wise to generate more detailed fused features. Thereby improving the model's ability to capture local important information. and After feature concatenation in the channel dimension, it is passed to the feature mapping module (FM).
[0089] As Figure 3 shown, MSFE consists of a large kernel dilated convolution block (LKDConv) and a small kernel dilated convolution block (SKDConv), which can be used to capture multi-scale context information, thereby enhancing the model's ability to understand features at different scales. LKDConv includes LKDConv a and LKDConv b, and SKDConv includes SKDConv a and SKDConv b. Specifically, MSFE effectively captures multi-scale and multi-level feature information by applying dilated convolution blocks with different convolution kernel sizes and dilation rates in parallel. LKDConv is used to capture a larger receptive field and extract features containing more global information; SKDConv focuses on fine-grained feature extraction and enhances the ability to capture local information. The expression of the dilated convolution block is as follows:
[0090]
[0091] In the formula, f(c,j) represents the value of the c-th output channel at position j, L represents the size of the convolution kernel, s represents the stride, d represents the dilation rate, x(j·s + n·d) represents the value of the input sample at position j·s + n·d, and w(c,n) represents the weight of the c-th convolution kernel at position n.
[0092] As Figure 4 shown, FIPF realizes interactive perception fusion by introducing third-domain features and using the interactive perception and adaptive enhancement of multi-spatial features, thereby strengthening the dependence relationship and information interaction between features. Compared with traditional feature fusion methods, FIPF strengthens the complementarity and dependence of different-domain features and significantly improves the model's ability to capture and fuse key information. The process of FIPF is as follows:
[0093] (1) Feature mapping: Assume the input features are fα and fδ , which are respectively mapped through a 1×1 convolutional layer to generate corresponding mapped features. The specific calculation formula is as follows:
[0094]
[0095] Wherein, and are mapping features in different spaces, and Conv 1×1 represents a 1×1 convolution operation. In addition, third-domain feature f αδ = f α + f δ is introduced in the feature domain and mapped through a 1×1 convolution layer to generate third-domain mapping features to further enhance the interaction between features. The specific calculation formula is as follows:
[0096]
[0097] (2) Feature interaction perception: For the mapping features in different domains, perform interaction perception processing and generate corresponding interaction perception features sα , sαδ and sδ through sigmoid. The specific calculation formula is as follows:
[0098]
[0099]
[0100] Wherein, T represents the transpose operation, and d d represents the feature dimension.
[0101] (3) Weight generation: For the interaction perception features in different domains, perform GAP processing respectively to generate global descriptions g α , g αδ and g δ . Subsequently, the global descriptions generate corresponding weights w α , w αδ and w δ through MLP and sigmoid to adjust the importance of features. The specific calculation formula is as follows:
[0102] w α = Sigmoid{MLP[GAP(s α )]} (8)
[0103] w αδ = Sigmoid{MLP[GAP(s αδ )]} (9)
[0104] w δ = Sigmoid{MLP[GAP(s δ )]} (10)
[0105] (4) Weighted and fusion: For the mapping features and in different domains, respectively multiply with the corresponding weights wα , w αδ and w δ are multiplied, and then feature concatenation is performed on the channel dimension to generate the fused feature f αδ . The specific calculation formula is as follows:
[0106]
[0107] In the formula, Cat c represents the feature concatenation operation on the channel dimension.
[0108] and are respectively fused through ResConv to generate the features f L and f S . Finally, the features f L and f S are concatenated on the length dimension to generate the fused particle feature f LS . The specific calculation formula is as follows:
[0109] f LS = Cat l (f L , f S ) (12)
[0110] In the formula, Cat l represents the feature concatenation operation on the length dimension.
[0111] Step 32: Construct the IPEA:
[0112] The IPEA includes FM, MIP, TMA, and skip connection structures, which can enhance the feature expression ability of the model. FM consists of three groups of 1×1 convolutional layers. While keeping the feature size unchanged, it maps the input features to multiple spaces, thereby increasing the diversity and information volume of feature representation. The specific calculation formula is as follows:
[0113] {x α , x δ , x ζ} = Conv 1×1 (f LS ) (13)
[0114] Figure 5 shows the structures of MIP and TMA in the IPEA. MIP consists of linear transformation, masking operation, and stacking fusion, while TMA consists of a multi-scale masking mechanism and a multi-head attention mechanism. MIP and TMA work together to capture multi-level context information, strengthen the interaction between local and global features, enhance the dependence relationship between features, and dynamically adjust the model's attention to key information, thereby improving the accuracy and flexibility of feature representation.xα and xδ generate two sets of query, key, and value features through linear transformation and masking operation respectively. The specific calculation formulas are as follows:
[0115]
[0116] In the formula, and are respectively xα and xδ the query features generated through linear transformation and masking operation, and are respectively xα and xδ the key features generated through linear transformation and masking operation, V i α and V i δ are respectively xα and xδ the value features generated through linear transformation and masking operation. i represents the index of the attention head, represents the projection matrix, C represents the number of channels, d head represents the dimension of each attention head, M i represents the multi-scale masking matrix. The specific calculation formulas are as follows:
[0117] d head = C / h (17)
[0118] M i = M o ·Scale i (18)
[0119] In the formula, h represents the number of attention heads, M o represents the original masking matrix, Scale i represents the scaling factor of the i-th attention head. For the two sets of query, key, and value features, they are fused through element-wise addition to generate the final result of MIP The specific calculation formula is as follows:
[0120]
[0121] TAM calculates the inner product of and and applies masks with different scales to the inner product results of different attention heads, and then obtains the attention distribution through Softmax normalization. The specific calculation formula is as follows:
[0122]
[0123] Wherein, T represents the transpose operation. Then, using the attention distribution Ai to i αδ weighted sum V
[0124] to obtain the output of a single attention head. The specific calculation formula is as follows: i head i = A i αδ ·V
[0125] (21) αδ Finally, concatenate the outputs of all attention heads and obtain the final result x of TAM through linear mapping
[0126] x αδ = Cat(head1, head2, …, head h )·W o (22)
[0127] Wherein, represents the projection matrix.
[0128] Adopt a skip connection structure to xαδ and x ζ add element by element, which not only retains the original features of the input, but also effectively avoids the degradation problem that may occur during the model training process. The specific calculation formula is as follows:
[0129] x αδζ = x αδ + x ζ (23)
[0130] Step 33, construct PDCN:
[0131] The PDCE includes a multi-classifier structure, an integrated loss function, and a dynamic adjustment mechanism. By gradually introducing early, middle, and late classifiers, setting a dedicated loss function for each classifier, and dynamically adjusting the weights, it makes full use of multi-level features and improves the feature learning ability of the model.
[0132] The early, middle, and late classifiers in the multi-classifier structure have different feature perception and processing capabilities, can process and learn the feature information at each level, and provide key decisions, significantly improving the overall classification performance. The early classifier focuses on the processing of shallow features, captures the global trend and basic patterns, and quickly generates a preliminary classification result. The early classifier consists of GAP, FC, and Softmax, and learns the basic features of the input samples through the linear combination of shallow features. The output of the early classifier can be expressed as:
[0133] y e= Softmax{w e2 ·FC[w e1 ·GAP(f LS ) + b e1 + b e2} (24)
[0134] In the formula, and respectively represent the weights of the first and second layers of the early classifier, and respectively represent the biases of the first and second layers of the early classifier, De represents the output dimension of the first layer of the early classifier. The intermediate classifier effectively captures subtle differences and improves the classification accuracy by extracting deeper features and combining more context information. The structure of the intermediate classifier is the same as that of the early classifier, consisting of GAP, FC, and Softmax, and learns the high-level features of the input samples by linearly combining the intermediate features. The output of the intermediate classifier can be expressed as:
[0135] y m = Softmax{w m2 ·FC[w m1 ·GAP(x αδζ ) + b m1 + b m2} (25)
[0136] In the formula, wm1 and wm2 respectively represent the weights of the first and second layers of the intermediate classifier, bm1 and bm2 respectively represent the biases of the first and second layers of the intermediate classifier. The late classifier further improves the accuracy and stability of decision-making by fusing deep features and classification information. The late classifier consists of GAP and Softmax and can accurately learn the deep features of the input samples. The output of the late classifier can be expressed as:
[0137]
[0138] In the formula, exp represents the exponential function, GAP represents the global average pooling operation, K represents the total number of categories, and x R represents the input features of the late classifier.
[0139] The early, intermediate, and late classifiers are respectively set with different loss functions to more efficiently learn their respective specific target features. The early classifier focuses on the learning accuracy of global and shallow features, adopts the focal loss function to reduce the attention to easy-to-classify samples, and strengthens the learning of difficult samples, thereby improving the stability of global classification. The calculation formula of the focal loss function is:
[0140] L e = -α t (1 - p t ) γ log(p t ) (27)
[0141] where p t represents the predicted probability of the model for the correct class, reflecting the confidence of the model in the correct class. α t represents the balance factor, which is used to adjust the weights of positive and negative samples. γ represents the focusing parameter, which is used to reduce the weights of easy-to-classify samples and increase the attention to difficult samples. The intermediate classifier utilizes more context features and adopts the generalized cross-entropy loss function to better focus on key features, improving the global classification accuracy under the conditions of class imbalance and complex samples. The calculation formula of the generalized cross-entropy loss function is:
[0142]
[0143] where represents the predicted probability of the model for the true label y, and q is a hyperparameter. By adjusting the value of q, the loss function of the intermediate classifier can achieve a balance between the sensitivity and stability of the model to samples, thereby effectively reducing the influence of limited samples and noisy samples. The late classifier adopts the standard cross-entropy loss function, focusing on refining deep features, further optimizing the final classification decision, and improving the classification accuracy of the model in complex scenarios. The calculation formula of the standard cross-entropy loss is:
[0144]
[0145] where y k represents the true label of the k-th sample, and p k represents the probability that the sample is predicted as the k-th class. The calculation formula of the integrated loss function is:
[0146] L = L e + L m + L f (30)
[0147] In the early stage of training, the dynamic weight adjustment mechanism gives priority to the loss weights of the early and intermediate classifiers to strengthen the rapid learning of global features. As the training progresses, the mechanism gradually turns to the learning of refined features, gradually increasing the loss weight of the late classifier, thereby optimizing the final classification performance. Specifically, the dynamic weight adjustment mechanism calculates the gradient of the current loss with respect to the weights of each classifier through training feedback, and adjusts the contribution of the losses of each classifier in the gradient update in real time, thereby optimizing the model parameters and improving the classification accuracy and stability. The calculation formula of the dynamic weight adjustment mechanism is:
[0148]
[0149] w f (t) = 1 - w e (t) - w m (t) (33)
[0150] where w e , w m and w f represent the loss weights of the early, middle, and late classifiers respectively, t represents the number of iterations, η represents the learning rate, represents the gradient of the current loss L with respect to the loss weight of the early classifier.
[0151] PDCE gradually introduces the early, middle, and late classifiers, and combines the integrated loss function with the dynamic adjustment mechanism, comprehensively improving the model's learning ability for complex sample features, thus significantly enhancing the accuracy and stability of classification. During the training process, the calculation formula for the integrated dynamic loss is:
[0152] L = w e ·L e (y e , y true ) + w m ·L m (y m , y true ) + w f ·L f (y f , y true ) (34)
[0153] where y true represents the true label.
[0154] Step 4: Use the training set to train PNMC to optimize the model parameters.
[0155] Step 5: Conduct fault diagnosis tests through the test set to generate the final diagnosis results.
[0156] Example:
[0157] To verify the effectiveness and superiority of the present invention, a series of experiments were conducted on the Paderborn University (PU) dataset, the Laboratory Self - Collected (LabSC) dataset, and the Aero - Engine Test Platform (AETP) dataset. The fault diagnosis experiments were carried out on a computing platform equipped with an Intel Xeon Gold 6348 CPU and an A800 GPU. The algorithm development was based on Python 3.8.10, using the PyTorch 2.0.0 and CUDA 11.8 frameworks.
[0158] Dataset description:
[0159] (1) The PU dataset covers various bearing states and is one of the important sources of experimental data in the field of rotating machinery fault diagnosis. As Figure 6 shown, the test bench mainly includes a motor, a torque measurement shaft, a rolling bearing test module, a flywheel, and a load motor. During the data acquisition process, the rotational speed of the drive system was 1500 r / min, the radial force applied to the bearing was 1000 N, the load torque of the transmission system was 0.7 Nm, and the sampling frequency was 64 kHz. According to different fault positions, fault modes, fault levels, and fault combinations, the PU dataset can be divided into ten types, and the specific information is shown in Table 1. The fault positions include the outer ring and the inner ring. The fault modes include pitting (P) and indentation (I). The fault levels are divided into level 1, level 2, and level 3. The fault combinations are divided into single damage (S) and repeated damage (R).
[0160] Table 1 Information of the PU dataset
[0161]
[0162] (2) A laboratory mechanical comprehensive diagnosis simulation platform was built in the present invention, and vibration signals of deep - groove rolling bearings in different states were collected to construct the LabSC dataset to verify the performance of the method proposed in the present invention and support related research. As Figure 7 shown, the test bench mainly includes a drive motor, a motor controller, a tachometer, an accelerometer, a bearing, and a load. During the data acquisition process, the load was 5.25 N, and the sampling frequency was 10 kHz. According to different fault positions and load rotational speeds, the LabSC dataset can be divided into ten types, and the specific information is shown in Table 2. The fault positions include the inner ring, the rolling elements, and the outer ring. The load rotational speeds are 1200 r / min, 1800 r / min, and 2400 r / min respectively.
[0163] Table 2 Information of the LabSC dataset
[0164]
[0165] (3) As Figure 8 shown, the test bench of the AETP dataset mainly includes an improved aero-engine, an electric motor drive system, and a lubricating oil system. During the data acquisition process, the sensor measurement points are arranged at five different positions, where 1 and 2 are displacement sensors, and 3, 4, 5, and 6 are acceleration sensors, and the sampling frequency is 25 kHz for all. It can be seen from the 3D model that the dual-rotor structure consists of a low-pressure (LP) compressor and a high-pressure (HP) compressor. According to different rotational speed levels, fault positions, and fault lengths, the AETP dataset can be divided into ten types, and the specific information is shown in Table 3. The rotational speed levels include low 1 (L1), low 2 (L2), low 3 (L3), high 1 (H1), high 2 (H2), and high 3 (H3). The fault positions include the inner ring and the outer ring, and the fault depth is 0.5 mm for both. The fault length of the outer ring is 0.5 mm. The fault lengths of the inner ring are 0.5 mm and 1.0 mm respectively.
[0166] Table 3 Information of the AETP Dataset
[0167]
[0168]
[0169] Experimental Setup and Evaluation Metrics:
[0170] Table 4 shows the architecture of PNMC. The optimizer used in the experiment is Adam. The selected parameters are based on the results of multiple experiments, aiming to balance the training efficiency and accuracy of the model. To ensure the stability of model training, the dynamic adjustment mechanism is only applied during the first one-third of the training phase, and an adjustment is made every one-tenth of the training phase. In addition, all experiments are carried out under the same hardware configuration and programming environment to exclude performance differences caused by non-model factors, thereby ensuring the credibility and comparability of the experimental results.
[0171] Table 4 Architecture of PNMC
[0172]
[0173] Four experimental scenarios are set up in the present invention, including balanced finite samples, unbalanced finite samples, noise & balanced finite samples, and noise & unbalanced finite samples, aiming to verify the performance of PNMC under finite sample conditions and its ability to cope with noise challenges in practical applications. Table 5 shows the experimental task settings on PU dataset, LabSC dataset, AETP dataset and WTPG dataset. TP1, TL1, TA1 and TW1 are verification experiments under balanced finite sample conditions, TP2, TL2, TA2 and TW2 are verification experiments under unbalanced finite sample conditions, TP3, TL3, TA3 and TW3 are verification experiments under noise & balanced finite sample conditions, and TP4, TL4, TA4 and TW4 are verification experiments under noise & unbalanced finite sample conditions. The samples in PU dataset and WTPG dataset contain 2048 data points, while the samples in LabSC dataset and AETP dataset contain 1024 data points. The number of training samples varies according to different datasets and experimental scenarios, while the number of test samples remains unchanged, fixed at 100 for each class. In the experiment, 0dB additive white Gaussian noise is used as additional noise to simulate the interference in the real scenario, so as to verify the noise suppression ability of PNMC.
[0174] Table 5 Settings of Experimental Tasks
[0175]
[0176]
[0177] Analysis of Experimental Results:
[0178] (1) Figure 9The visualization results of the characteristics of PNMC on TP1, TP2, TP3, and TP4 are shown, presenting the classification decision boundaries and clustering situations of each sample category. It can be seen that despite a small number of misclassifications, PNMC can still effectively map various categories to different spatial regions, demonstrating good diagnostic performance. Specifically, in TP1 and TP2, the diagnostic accuracy of PNMC is close to 100%. However, in TP3 and TP4, affected by the noise & limited sample conditions, there are some misdiagnosis cases in PNMC, but the overall accuracy still remains at a relatively high level. Table 6 summarizes the diagnostic effects of different methods on TP1, TP2, TP3, and TP4. Compared with TP1 and TP2, ResNet18 and ResNet50 perform significantly worse on the more complex TP3 and TP4, showing weaker generalization ability. WDCNN and ESPNet have lower accuracies in all tasks and larger standard deviations, indicating poor diagnostic performance and poor stability. The diagnostic performance of ISCNN and TFIFNet is comparable to that of PNMC in TP1 and TP2, but they are slightly insufficient in the more complex TP3 and TP4. DBTIFF and DHET have higher accuracies, but their performance does not exceed that of PNMC in all four tasks. Overall, the overall performance of PNMC in the four tasks is better than that of other comparison methods, with an average accuracy of 99.57% and a lower standard deviation. The results show that PNMC can effectively cope with the limited sample problem, demonstrate excellent anti-noise ability, and have excellent diagnostic accuracy and stability.
[0179] Table 6 Performance comparison of different methods on TP1, TP2, TP3, and TP4
[0180]
[0181] (2) Figure 10The visualization results of the characteristics of PNMC on TL1, TL2, TL3, and TL4 are shown. The results indicate that PNMC exhibits high performance in all four tasks, and samples of all categories are clearly mapped to different spatial regions. Table 7 summarizes the diagnostic effects of different methods on TL1, TL2, TL3, and TL4. In TL1 and TL2, except for ResNet50 and WDCNN, the accuracy rates of other methods exceed 90%, among which the accuracy rates of ISCNN, DHET, DBTIFF, and PNMC exceed 99%, showing good diagnostic performance. In TL3 and TL4, due to the influence of noise & limited samples, the accuracy rates of all methods decrease. Nevertheless, the diagnostic performance of PNMC in TL3 and TL4 is still the highest, being 97.91% and 99.06% respectively, demonstrating good capabilities in dealing with limited samples and anti-noise. It should be noted that in some tasks, the accuracy rates of some methods are slightly higher than that of PNMC. In TL1, the accuracy rates of DHET and DBTIFF are 99.94% and 99.82% respectively, slightly higher than 99.42% of PNMC. In TL2, the accuracy rates of ISCNN and DHET are 99.83% and 99.80% respectively, slightly higher than 99.63% of PNMC. Nevertheless, PNMC has the best overall diagnostic performance in the four tasks, with an average accuracy rate of 99.00% and a low standard deviation, demonstrating excellent diagnostic performance and stability. The results show that PNMC has significant advantages on the LabSC dataset and is applicable to fault diagnosis under limited samples and noise & limited sample conditions.
[0182] Table 7 Performance comparison of different methods on TL1, TL2, TL3, and TL4
[0183]
[0184] (3) Figure 11The visualization results of the characteristics of PNMC on TA1, TA2, TA3, and TA4 are shown. It can be seen that although there are a small number of misclassifications, PNMC can still effectively map different categories to different spatial regions, demonstrating good diagnostic effects. Specifically, in TA1 and TA2, there is almost no overlap or only a small amount of overlap between different categories, indicating that PNMC can effectively handle the problem of limited samples and has good classification effects. In TA3 and TA4, the noise & limited sample conditions lead to a decrease in the diagnostic accuracy, but PNMC still shows strong diagnostic performance. The classification accuracy of most categories remains high, and only a few categories are misclassified. Table 8 summarizes the performance comparison of different methods on TA1, TA2, TA3, and TA4. It can be seen that PNMC performs particularly outstanding in multiple tasks. Especially in TA3 and TA4, the accuracies reach 99.92% and 98.79% respectively, both exceeding all comparison methods, demonstrating its strong generalization performance under noise & limited sample conditions. It should be noted that in TA1 and TA2, the accuracies of TFIFNet and DBTIFF are slightly higher than that of PNMC. Nevertheless, the average accuracy of PNMC in the four tasks reaches 99.39%, higher than all comparison methods, and the standard deviation is low, demonstrating excellent fault diagnosis accuracy and stability. The results show that PNMC performs outstandingly in the fault diagnosis task of the AETP dataset. Especially under noise & limited sample conditions, its classification performance and anti-noise performance are better than other methods. PNMC has high practical value in industrial applications and is suitable for fault diagnosis tasks in complex environments.
[0185] Table 8 Performance comparison of different methods on TA1, TA2, TA3, and TA4
[0186]
[0187] Ablation analysis:
[0188] To verify the effectiveness of the proposed modules, mechanisms, and structures, ablation experiments were conducted on the PU dataset, LabSC dataset, and AETP dataset respectively. The specific experiments included removing GFE, FIPF, IPEA, and PDCE respectively and observing the changes in the performance of PNMC. The experimental design is as follows. (1) PNMC was used as a comparative experiment, denoted as T1. (2) In the experiment of removing GFE, a dilated convolution block was used to replace GFE, denoted as T2. The dilated convolution block includes a dilated convolution layer, BN, and ReLU. (3) In the experiment of removing FIPF, the feature concatenation method in the channel dimension was used to replace FIPF, denoted as T3. (4) The experiment of removing IPEA was denoted as T4. (5) In the experiment of removing PDCE, only one late classifier was retained, denoted as T5. The ablation experiment results on the PU dataset, LabSC dataset, and AETP dataset are respectively as Figure 12, Figure 13 and Figure 14 as shown
[0189] (1) Ablation of GFE: It can be found that after removing GFE, the accuracy of PNMC significantly decreases in most tasks, especially in the diagnostic tasks under the conditions of noisy & balanced finite samples and noisy & unbalanced finite samples. However, in some tasks, when GFE is removed, PNMC shows a slight improvement instead. Specifically, in TL1 and TL2, the accuracy of PNMC increases by 0.32% and 0.12% respectively after removing GFE. In TA2, the accuracy of PNMC does not change significantly after removing GFE. Overall, after removing GFE, the average accuracy of PNMC in the tasks on the PU dataset, LabSC dataset, and AETP dataset decreases by 7.61%, 2.76%, and 2.15% respectively. The results show that GFE can effectively capture particle feature information through multi-scale feature extraction and interactive perception fusion, which is crucial for most complex tasks and is one of the key modules in PNMC. However, in some relatively simple tasks, removing GFE may help simplify the model and thus slightly improve the performance. Generally speaking, GFE plays a key role in improving the accuracy and stability of PNMC for complex tasks, especially under the conditions of noisy and balanced or noisy and unbalanced finite samples.
[0190] (2) Ablation of FIPF: It can be found that after removing FIPF, the accuracy of PNMC significantly decreases in all tasks. Specifically, after removing FIPF, in the tasks on the PU dataset, LabSC dataset, and AETP dataset, the average accuracy of PNMC decreases by 0.28%, 0.88%, and 0.93% respectively. The results show that FIPF can promote the interactive perception fusion of features, enhance the dependence relationship and information interaction between features, thereby promoting the extraction of particle features and ultimately improving the accuracy and stability of PNMC.
[0191] (3) Ablation of IPEA: It can be found that after removing IPEA, the accuracy of PNMC significantly decreases in most tasks, especially in the diagnostic tasks under the conditions of noise & balanced limited samples and noise & unbalanced limited samples. Specifically, after removing IPEA, in the tasks on the PU dataset, LabSC dataset, and AETP dataset, the average accuracy of PNMC decreases by 0.37%, 1.73%, and 0.40% respectively. The results show that IPEA can dynamically adjust the model's attention to key information through the interactive perception and adaptive enhancement of multi-spatial features, enhance the feature expression ability of PNMC, and thus improve the accuracy and stability of diagnosis. However, in some relatively simple tasks, removing IPEA may help simplify PNMC and thus slightly improve performance. Nevertheless, overall, IPEA still plays an irreplaceable role in improving the accuracy and stability of complex tasks.
[0192] (4) Ablation of PDCE: It can be found that after removing PDCE, the accuracy of PNMC significantly decreases in all tasks, especially in the diagnostic tasks under the conditions of noise & balanced limited samples and noise & unbalanced limited samples. Specifically, after removing PDCE, in the tasks on the PU dataset, LabSC dataset, and AETP dataset, the average accuracy of the model decreases by 0.90%, 0.69%, and 0.24% respectively. The results show that PDCE can fully utilize multi-level features and enhance the feature learning ability of PNMC by gradually introducing early, middle, and late classifiers and integrating dynamic losses, thus significantly improving the accuracy and stability of diagnosis.
[0193] (5) Analysis: The ablation experiment results show that the modules, mechanisms, and structures proposed in the present invention play a key role in improving the diagnostic performance of PNMC, especially under the conditions of noise & limited samples. Although in some tasks, removing a certain innovation may slightly improve or have no significant impact on the accuracy, from the overall performance perspective, the synergistic effect of these innovations greatly improves the diagnostic accuracy and stability of PNMC. In particular, GFE and FIPF play a key role in feature extraction, IPEA plays an important role in feature expression, and PDCE provides key support in feature learning. The synergistic effect of the modules, mechanisms, and structures proposed in the present invention enables PNMC to show high accuracy and stability in the diagnostic tasks under limited samples and noise & limited samples conditions, thus providing an effective solution for complex industrial fault diagnosis tasks.
Claims
1. A method for rotating machinery fault diagnosis based on a multi-classifier progressive network, characterized in that The method includes the following steps: Step 1: Build a fault simulation platform and collect vibration signals of rotating machinery through sensors; Step 2: Use the sliding window technique to divide the collected one-dimensional vibration signals into a training set and a test set; Step 3: Construct a multi-classifier progressive network PNMC, and initialize the weights and biases. The PNMC includes a granular feature extraction module GFE, an interactive perception attention mechanism IPEA, and a progressive dynamic classifier ensemble structure PDCE. The specific construction steps are as follows: Step 31: Construct GFE: The GFE consists of a multi-scale feature extraction module MSFE, a feature interaction perception fusion module FIPF, an adaptive convolution block AdaConv, and a residual convolution block ResConv; Step 32: Construct IPEA: IPEA includes a feature mapping module FM, a masking interaction perception mechanism MIP, a multi-scale masking multi-head attention mechanism TMA, and a skip connection structure; Step 33: Construct PDCN: The PDCE contains a multi-classifier structure, an integrated loss function, and a dynamic adjustment mechanism; Step 4: Use the training set to train the PNMC to optimize the model parameters; Step 5: Conduct a fault diagnosis test through the test set to generate the final diagnosis result.
2. The rotating machinery fault diagnosis method based on the multi-classifier progressive network according to claim 1, characterized in that In step 31, MSFE consists of a large kernel dilated convolution block LKDConv and a small kernel dilated convolution block SKDConv. LKDConv includes LKDConv a and LKDConv b, and SKDConv includes SKDConv a and SKDConv b; AdaConv includes AdaConv a and AdaConv b; assuming the input sample is x, after being processed by MSFE, the output feature is represented as and are the features output by LKDConv a, LKDConv b, SKDConv a, and SKDConv b respectively, where: and include two processing paths, and After feature concatenation in the channel dimension, it is passed to AdaConv a, and After being fused by FIPF, it is passed to AdaConv b. The outputs of the two paths are passed to ResConv and added element-wise to generate the fused feature and include two processing paths, and After feature concatenation in the channel dimension, it is passed to AdaConv a, and After being fused by FIPF, it is passed to AdaConv b. The outputs of the two paths are passed to ResConv and added element-wise to generate the fused feature and After feature concatenation in the channel dimension, it is passed to IPEA.
3. The method for diagnosing faults of rotating machinery based on a multi-classifier progressive network according to claim 1, wherein The structure of the FIPF is as follows: Assume the input feature is f α and f δ , are respectively mapped through a 1×1 convolutional layer to generate corresponding mapped features. The specific calculation formula is as follows: In the formula, and are both mapping features in different spaces. Conv 1×1 represents a 1×1 convolution operation, introducing the third-domain feature f αδ = f α + f δ , and performing mapping through a 1×1 convolution layer to generate the third-domain mapping feature. The specific calculation formula is as follows: Perform interactive perception processing on the mapping features for different domains, and generate corresponding interactive perception features through sigmoid sα , sαδ and sδ , and the specific calculation formula is as follows: In the formula, T represents the transpose operation, dd represents the feature dimension; Perform GAP processing on the interaction perception features for different domains respectively to generate global descriptions gα 、 gαδ and gδ Subsequently, the global descriptions generate corresponding weights through MLP and sigmoid wα 、 wαδ and wδ The specific calculation formula is as follows: w α = Sigmoid{MLP[GAP(s α )]} w αδ = Sigmoid{MLP[GAP(s αδ )]} w δ = Sigmoid{MLP[GAP(s δ )]} Mapping features for different domains and are respectively multiplied by the corresponding weights wα , wαδ and wδ , and then feature concatenation is performed on the channel dimension to generate fused features fαδ . The specific calculation formula is as follows: where, Cat c represents the feature concatenation operation in the channel dimension; and are respectively fused through ResConv to generate features fL and fS . Finally, the features fL and fS are concatenated in the length dimension to generate the fused particle features fLS . The specific calculation formula is as follows: f LS = Cat l (f L , f S ) In the formula, Cat l represents the feature splicing operation in the length dimension.
4. The method for rotating machinery fault diagnosis based on a multi-classifier progressive network according to claim 1, characterized in that In step 32, the FM consists of three groups of 1×1 convolutional layers, which map the input features to multiple spaces while keeping the feature size unchanged. The specific calculation formula is as follows: {x α ,x δ ,x ζ} = Conv 1×1 (f LS ) The MIP consists of a linear transformation, a masking operation, and superposition fusion. xα and xδ Generate two sets of query, key, and value features through linear transformation and masking operation respectively. The specific calculation formulas are as follows: Wherein, and are respectively xα and xδ query features generated through linear transformation and masking operation, and are respectively xα and xδ key features generated through linear transformation and masking operation, V i α and V i δ are respectively xα and xδ value features generated through linear transformation and masking operation, i represents the index of the attention head, represents the projection matrix, C represents the number of channels, dhead represents the dimension of each attention head, Mi represents the multi-scale masking matrix; for two groups of query, key and value features, they are fused through element-wise addition to generate the final result of MIP The specific calculation formula is as follows: The TMA consists of a multi-scale masking mechanism and a multi-head attention mechanism. The TAM performs an inner product calculation on and , applies masks of different scales to the inner product results of different attention heads, and then obtains the attention distribution through Softmax normalization. The specific calculation formula is as follows: where T represents the transpose operation; then, using the attention distribution Ai to weight-sum V i αδ to obtain the output of a single attention head, and the specific calculation formula is as follows: head i = A i ·V i αδ Finally, the outputs of all attention heads are concatenated and the final result \(x\) of TAM is obtained through a linear mapping αδ , and the specific calculation formula is as follows: x αδ = Cat(head1, head2, …, head h )·W o In the formula, represents a projection matrix; Adopt a skip connection structure to add xαδ and x ζ element by element. The specific calculation formula is as follows: x αδζ = x αδ + x ζ 。 5. The method for rotating machinery fault diagnosis based on a multi-classifier progressive network according to claim 1, wherein In step 33, the multi-classifier structure includes an early classifier, a mid-term classifier, and a late classifier, where: The early classifier consists of GAP, FC, and Softmax, and learns the basic features of the input samples through the linear combination of shallow features. The output of the early classifier is expressed as: y e = Softmax{w e2 ·FC[w e1 ·GAP(f LS ) + b e1 + b e2} In the formula, and respectively represent the weights of the first and second layers of the early classifier, and respectively represent the biases of the first and second layers of the early classifier, De represents the output dimension of the first layer of the early classifier; The mid-term classifier consists of GAP, FC, and Softmax, and learns the advanced features of the input samples by linearly combining the mid-level features. The output of the mid-term classifier is expressed as: y m = Softmax{w m2 ·FC[w m1 ·GAP(x αδζ ) + b m1 + b m2} wherein, wm1 and wm2 represent the weights of the first and second layers of the intermediate classifier respectively, bm1 and bm2 represent the biases of the first and second layers of the intermediate classifier respectively; The late classifier consists of GAP and Softmax, and can accurately learn the deep features of the input samples. The output of the late classifier is expressed as: In the formula, exp represents the exponential function, GAP represents the global average pooling operation, K represents the total number of categories, xR represents the input features of the late classifier.
6. The rotating machinery fault diagnosis method based on the multi-classifier progressive network according to claim 5, wherein The early classifier adopts a focal loss function, and the calculation formula of the focal loss function is: L e = -α t (1 - p t ) γ log(p t ) In the formula, pt represents the prediction probability of the model for the correct category, αt represents the balance factor, and γ represents the focusing parameter; The mid-term classifier adopts a generalized cross-entropy loss function, and the calculation formula of the generalized cross-entropy loss function is: Where P y q represents the predicted probability of the model for the true label y, and q is a hyperparameter; The late classifier adopts a standard cross-entropy loss function, and the calculation formula of the standard cross-entropy loss is: wherein, yk represents the true label of the k-th sample, pk represents the probability that the sample is predicted to be the k-th class; The calculation formula of the integrated loss function is: L = L e + L m + L f 。 7. The rotating machinery fault diagnosis method based on the multi-classifier progressive network according to claim 1, characterized in that In step 33, the calculation formula of the dynamic weight adjustment mechanism is: w f (t) = 1 - w e (t) - w m (t) wherein, we , wm and w f respectively represent the loss weights of the early, middle, and late classifiers, t represents the number of iterations, η represents the learning rate, represents the gradient of the current loss L with respect to the loss weight of the early classifier.
8. The rotating machinery fault diagnosis method based on a multi-classifier progressive network according to claim 1, wherein In step 4, during the training process, the calculation formula of the integrated dynamic loss is: L = w e ·L e (y e ,y true ) + w m ·L m (y m ,y true ) + w f ·L f (y f ,y true ) where y true represents the true label.
Citation Information
Patent Citations
Image semantic segmentation method based on multiple classifiers
CN115761229A
Semantic segmentation method based on Transform laser radar point cloud and camera image information fusion
CN117788823A
Gear fault diagnosis method and system based on non-local attention measurement network
CN118171162A
Driving control circuit with fault detection function
CN119148608A
Chinese named entity recognition method based on stacked grid structure information enhancement
CN119227685A