Rotating machinery fault diagnosis method based on multi-classifier progressive network

By employing a multi-classifier progressive network fault diagnosis method, utilizing a particle feature extraction module, an interactive perception-enhanced attention mechanism, and a dynamic classifier ensemble structure, the problem of insufficient feature extraction under limited sample conditions in rotating machinery fault diagnosis is solved, achieving fault diagnosis with high accuracy and stability.

CN120296589BActive Publication Date: 2026-03-31HARBIN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In the diagnosis of rotating machinery faults, existing technologies struggle to effectively extract sufficient feature information under limited sample conditions, resulting in insufficient diagnostic accuracy and stability, especially under complex environments and noise interference.

Method used

A fault diagnosis method based on progressive multi-classifier networks (PNMC) is adopted, including a granular feature extraction module (GFE), an interactive perception enhanced attention mechanism (IPEA), and a progressive dynamic classifier ensemble structure (PDCE). Through multi-scale feature extraction, feature interactive perception fusion, and dynamic adjustment, the feature extraction and learning capabilities are improved.

Benefits of technology

It significantly improves the accuracy and robustness of fault diagnosis for rotating machinery, especially performing well under noisy and limited sample conditions, and is suitable for fault diagnosis tasks in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296589B_ABST
    Figure CN120296589B_ABST
Patent Text Reader

Abstract

The application discloses a rotating machinery fault diagnosis method based on a multi-classifier progressive network. First, a particle feature extraction module is proposed, which effectively captures particle feature information through multi-scale feature extraction and interactive perception fusion, thereby enhancing the model's ability to perceive minor fault features. In addition, an interactive perception enhanced attention mechanism is proposed, which dynamically adjusts the model's attention to key information through interactive perception and adaptive enhancement of multi-space features, thereby significantly improving the expression quality of complex fault features. Finally, a progressive dynamic classifier integration structure is proposed, which effectively integrates multi-level feature information by gradually introducing early, middle and terminal classifiers and integrating dynamic loss, thereby enhancing the model's ability to learn fault features. The application not only has practical application value in the field of rotating machinery fault diagnosis under limited sample conditions, but also provides a new research idea for other classification and diagnosis tasks facing limited sample constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for diagnosing faults in rotating machinery, specifically a method for diagnosing faults in rotating machinery based on a Progressive Network with Multiple Classifiers (PNMC). Background Technology

[0002] Large rotating machinery, including wind turbines, turbofans, and aerospace engines, is the cornerstone of the smart industry era, driving the advancement of intelligent manufacturing. However, during long-term operation, complex environments and improper operation can lead to damage to critical components, affecting not only production efficiency but also potentially causing serious safety accidents. Therefore, accurate diagnosis of faults in critical components such as bearings, gears, and rotors is crucial for improving industrial production efficiency and ensuring operational safety.

[0003] Compared to traditional diagnostic methods, deep learning can automatically extract complex features, enabling intelligent fault diagnosis and providing more efficient diagnostic tools for industrial equipment. Unfortunately, due to high data acquisition costs and stringent safety requirements, the number of available samples is usually limited, making it difficult for deep learning to extract sufficient feature information, thus affecting the accuracy and stability of the diagnosis. Therefore, research on rotating machinery fault diagnosis under limited sample conditions is of great significance.

[0004] Under finite sample conditions, generative models, transfer learning, and finite sample learning have made significant progress in improving fault diagnosis performance. First, generative models can generate highly diverse new samples using a limited number of samples, providing richer and more diverse feature information for deep learning, thus significantly improving the accuracy and stability of fault diagnosis. However, the training process for generative models is complex, and the quality of generated samples may be uneven, placing higher demands on the generalization ability of diagnostic models. Compared to generative models, transfer learning significantly improves the diagnostic performance of the target domain by utilizing models or feature representations trained in the source domain. While transfer learning effectively improves model performance, its effect depends on the similarity between the target and source domains. When the differences are large, transfer bias may lead to performance degradation. Finally, finite sample learning, through functional modules and optimization strategies, extracts effective features from a limited number of samples, enhancing the model's generalization ability and demonstrating significant advantages under finite sample conditions. However, finite sample learning also has certain limitations, such as high design complexity and stringent requirements for optimization strategies. Nevertheless, given its accuracy and stability under finite sample conditions, finite sample learning remains a core method in current research, and its potential and challenges warrant further exploration.

[0005] Under limited sample conditions, limited sample learning enhances the comprehensive utilization of fault features from multiple levels by improving feature extraction, strengthening feature representation, and optimizing feature learning, thereby significantly improving the accuracy and stability of fault diagnosis. Feature extraction, as a key step, has seen the development of various methods to adapt to complex fault modes and limited sample conditions. Specifically, Zhang et al. proposed a compact convolutional neural network that processes vibration signals through multi-scale feature extraction units, thereby extracting more discriminative features and significantly improving adaptability to complex fault modes. Similarly, Li et al. proposed MFSFormer, which extracts multi-scale receptive field features by embedding convolutional layers with convolutional kernels of different sizes, enhancing the comprehensive representation of global and local information and exhibiting excellent stability under noisy and limited sample conditions. Furthermore, Yao et al. proposed a dual-pooling attention residual network that significantly improves the sensitivity and discriminativeness of feature extraction by fusing global and local information, thereby improving the accuracy and stability of fault diagnosis. To address long-term dependencies in signals, Kumar et al. proposed a multi-size wide-kernel convolutional network. This network captures long-term dependencies through a wide-kernel design and extracts features at different scales using multi-size convolutional kernels, thus demonstrating strong adaptability when processing complex pattern signals. Regarding feature fusion, Zhang et al. proposed a deep learning model based on dual-branch time-frequency fusion, which significantly improves feature extraction capabilities for complex fault modes by fusing time-domain and frequency-domain features. Despite these methods' excellent performance in feature extraction, they still suffer from insufficient information fusion and interaction between features, which may lead to the loss or redundancy of crucial information.

[0006] In addition, researchers have proposed various attention mechanisms to enhance the model's ability to express key features. Specifically, Hu et al. proposed a multi-scale convolutional neural network combining multiple attention mechanisms, which improved the ability to identify key features of fault signals by integrating channel and spatial attention mechanisms. Similarly, Snyder et al. proposed a dual-head transformer that enhances the interaction between global and local features by embedding a self-attention mechanism. Furthermore, Liang et al. proposed a multi-branch, multi-scale dynamic convolutional network that combines dynamic attention mechanisms to adaptively adjust the convolutional kernel weights to enhance the flexibility and robustness of feature representation. The multi-branch structure also supports parallel processing of information at different scales, further improving the adaptability to complex signals. Additionally, Li et al. developed a novel neural network structure that combines ConvNeXt with a multi-scale dilated attention mechanism to capture multi-scale contextual information, significantly improving the accuracy and stability of fault diagnosis. However, the effectiveness of attention mechanisms is limited under noisy interference or limited sample conditions, making it challenging to ensure the stability and generalization ability of the model.

[0007] In feature learning, research mainly focuses on loss function design and regularization mechanism optimization to improve learning capabilities in scenarios with limited samples. Regarding loss function design, Shi et al. proposed an incremental learning loss function that adapts to class imbalance, significantly improving the robustness of industrial flow sample processing. Furthermore, Li et al. proposed a meta-learning framework that combines intra-class and inter-class optimization, strengthening feature learning capabilities in scenarios with few samples through a novel loss function. Wang et al. proposed a novel cross-domain adaptive loss function, effectively enhancing feature alignment capabilities in cross-domain fault diagnosis. Finally, to further improve the model's generalization performance, Dong et al. optimized the feature space distribution through a dual regularization mechanism, thereby improving the model's stability in different scenarios. Although the above methods optimize the feature learning process from multiple perspectives, some challenges remain. Specifically, the generalizability of the loss function and the balance of regularization methods are key issues that need to be considered in model design.

[0008] In summary, in finite sample learning, the feature extraction module, attention mechanism, and optimization strategy all play crucial roles. The feature extraction module delves into the core information of the samples, generating high-quality feature representations. The attention mechanism focuses on key regions, significantly improving feature representation capabilities. The optimization strategy further enhances the model's accuracy and stability by strengthening feature learning capabilities. Summary of the Invention

[0009] To address the challenges of rotating machinery fault diagnosis under limited sample conditions, this invention provides a rotating machinery fault diagnosis method based on a multi-classifier progressive network (PNMC), which improves the accuracy and robustness of fault identification under limited sample conditions. The key design features of the proposed PNMC include GFE, IPEA, and PDCE, which work synergistically to enhance the model's performance in feature extraction, representation, and learning. Specifically, GFE utilizes a multi-scale feature extraction (MSFE) module to extract multi-scale features from input samples, and combines this with a feature interaction-aware fusion (FIPF) module to capture the interdependencies between features, thereby achieving interactive feature fusion and accurately extracting granular feature information, significantly improving the model's feature extraction capability. IPEA enhances the information transmission and dependencies between features through a masked interactive awareness (MIP) mechanism, and utilizes a multi-scale masked multi-head attention (TMA) mechanism to capture local and global information, dynamically adjusting the focus on key information to improve the model's feature representation capability. PDCE, by employing a multi-classifier structure, an ensemble loss function, and a dynamic adjustment mechanism, fully utilizes multi-level features to further enhance the model's feature learning capability.

[0010] The objective of this invention is achieved through the following technical solution:

[0011] A method for fault diagnosis of rotating machinery based on a multi-classifier progressive network includes the following steps:

[0012] Step 1: Build a fault simulation platform and collect vibration signals of rotating machinery through sensors;

[0013] Step 2: Use the sliding window technique to divide the acquired one-dimensional vibration signal into a training set and a test set;

[0014] Step 3: Construct PNMC and initialize weights and biases. PNMC includes a Granular feature extraction (GFE) module, an Interaction Perception Enhanced Attention (IPEA) mechanism, and a Progressive Dynamic Classifier Ensemble (PDCE) structure. The specific construction steps are as follows:

[0015] Step 31, Construct GFE:

[0016] The GFE consists of a Multi-Scale Feature Extraction (MSFE) module, a Feature Interaction Perception Fusion (FIPF) module, an adaptive convolutional block (AdaConv), and a residual convolutional block (ResConv). Specifically, the MSFE consists of a large-kernel dilated convolutional block (LKDConv) and a small-kernel dilated convolutional block (SKDConv). LKDConv includes LKDConv a and LKDConv b, SKDConv includes SKDConv a and SKDConv b, and AdaConv includes AdaConv a and AdaConv b. Assuming the input sample is x, after processing by the MSFE, the output features are represented as follows: and The features output by LKDConv a, LKDConv b, SKDConv a, and SKDConv b are respectively: and Includes two processing paths, and After feature concatenation along the channel dimension, the data is passed to AdaConv a. and After fusion via FIPF, the outputs are passed to AdaConv b. The outputs of both paths are then passed to ResConv, where they are added element-wise to generate the fused feature. and Includes two processing paths, and After feature concatenation along the channel dimension, the data is passed to AdaConv a. and After fusion via FIPF, the outputs are passed to AdaConv b. The outputs of both paths are then passed to ResConv, where they are added element-wise to generate the fused feature. and After feature concatenation along the channel dimension, the data is transmitted to IPEA.

[0017] Step 32, Build IPEA:

[0018] IPEA includes a Feature Mapping Module (FM), a Masked Interaction Awareness (MIP) mechanism, a Multi-Scale Masked Multi-Head Attention (TMA) mechanism, and a skip connection structure, wherein:

[0019] The FM consists of three 1×1 convolutional layers, which map the input features to multiple spaces while keeping the feature size constant. The specific calculation formula is as follows:

[0020] {x α ,x δ ,x ζ}=Conv 1×1 (f LS )

[0021] The MIP consists of linear transformation, masking operation, and superposition fusion. xα and xδ Two sets of query, key, and value features are generated after linear transformation and masking operations, respectively. The specific calculation formula is as follows:

[0022]

[0023] In the formula, and They are respectively xα and xδ The query features are generated through linear transformation and masking operations. and They are respectively xα and xδ The key features generated after linear transformation and masking operations, V i α and V i δ They are respectivelyxα and xδ The value feature generated after linear transformation and masking operation, where i represents the index of the attention head. Let C represent the projection matrix, and C represent the number of channels. dhead This represents the dimension of each attention head. Mi This represents a multi-scale masking matrix; for two sets of query, key, and value features, they are fused element-wise by addition to generate the final MIP result. The specific calculation formula is as follows:

[0024]

[0025] The TMA consists of a multi-scale masking mechanism and a multi-head attention mechanism. TMA works by... and Inner product calculation is performed, and masks of different scales are applied to the inner product results of different attention heads. Then, the attention distribution is obtained by Softmax normalization. The specific calculation formula is as follows:

[0026]

[0027] In the formula, T represents the transpose operation; then, attention distribution is used... Ai For V i αδ The output of a single attention head is obtained by performing a weighted summation, and the specific calculation formula is as follows:

[0028] head i =A i ·V i αδ

[0029] Finally, the outputs of all attention heads are concatenated, and the final result x of TAM is obtained through linear mapping. αδ The specific calculation formula is as follows:

[0030] x αδ =Cat(head1,head2,…,head) h )·W o

[0031] In the formula, Represents the projection matrix;

[0032] Using a skip connection structure, xαδ and x ζ The specific formula for adding elements one by one is as follows:

[0033] x αδζ =x αδ +x ζ

[0034] Step 33: Build PDCN:

[0035] The PDCE includes a multi-classifier structure, an ensemble loss function, and a dynamic adjustment mechanism, wherein:

[0036] The multi-classifier structure includes an early classifier, a mid-stage classifier, and a late-stage classifier;

[0037] The early classifier consists of GAP, FC, and Softmax. It learns the basic features of the input sample through a linear combination of shallow features. The output of the early classifier is represented as:

[0038] y e =Softmax{w e2 ·FC[w e1 ·GAP(f LS )+b e1 ]+b e2}

[0039] In the formula, and These represent the weights of the first and second layers of the early classifier, respectively. and These represent the biases of the first and second layers of the early classifier, respectively. De This represents the output dimension of the first layer of the early classifier;

[0040] The intermediate classifier consists of GAP, FC, and Softmax. It learns the high-level features of the input samples by linearly combining the intermediate features. The output of the intermediate classifier is represented as:

[0041] y m =Softmax{w m2 ·FC[w m1 ·GAP(x αδζ )+b m1 ]+b m2}

[0042] In the formula, wm1 and wm2 These represent the weights of the first and second layers of the intermediate classifier, respectively. bm1 and bm2 These represent the biases of the first and second layers of the intermediate classifier, respectively.

[0043] The late classifier consists of GAP and Softmax, and is capable of accurately learning the deep features of the input samples. The output of the late classifier is represented as:

[0044]

[0045] In the formula, exp represents the exponential function, GAP represents the global average pooling operation, K represents the total number of classes, and x R This represents the input features of the late classifier;

[0046] The early classifier uses a focal loss function, which is calculated using the following formula:

[0047] L e =-α t (1-p t ) γ log(p t )

[0048] In the formula, p t α represents the model's predicted probability for the correct class. t γ represents the balance factor, and γ represents the focusing parameter;

[0049] The intermediate classifier uses the generalized cross-entropy loss function, which is calculated using the following formula:

[0050]

[0051] In the formula, q represents the model's predicted probability of the true label y, and q is a hyperparameter;

[0052] The late classifier uses the standard cross-entropy loss function, and the formula for calculating the standard cross-entropy loss is as follows:

[0053]

[0054] In the formula, yk This represents the true label of the k-th sample. pk This represents the probability that a sample is predicted to be in the k-th class;

[0055] The formula for calculating the integration loss function is as follows:

[0056] L = L e +L m +L f

[0057] The calculation formula for the dynamic weight adjustment mechanism is as follows:

[0058]

[0059] In the formula, w e w m and w f These represent the loss weights for the early, mid, and late-stage classifiers, respectively; t represents the number of iterations; and η represents the learning rate. This represents the gradient of the current loss L with respect to the weights of the earlier classifier losses;

[0060] Step 4: Train PNMC using the training set to optimize the modulus parameters, where: During training, the formula for calculating the ensemble dynamic loss is:

[0061] L = w e ·L e (y e ,y true )+w m ·L m (y m ,y true )+w f ·L f (y f ,y true )

[0062] In the formula, ytrue Indicates the true label;

[0063] Step 5: Perform fault diagnosis tests using the test set to generate the final diagnosis results.

[0064] Compared with the prior art, the present invention has the following advantages:

[0065] 1. This invention proposes a rotating machinery fault diagnosis method based on a multi-classifier progressive network. First, a particle feature extraction module is proposed, which effectively captures particle feature information through multi-scale feature extraction and interactive perception fusion, thereby enhancing the model's ability to perceive minute fault features. Furthermore, an interactive perception-enhanced attention mechanism is proposed, which dynamically adjusts the model's focus on key information through interactive perception and adaptive enhancement of multi-spatial features, thus significantly improving the representation quality of complex fault features. Finally, a progressive dynamic classifier ensemble structure is proposed, which effectively integrates multi-level feature information by progressively introducing early, mid-term, and final classifiers and integrating dynamic loss, thereby enhancing the model's ability to learn fault features.

[0066] 2. The performance of the proposed method was evaluated on the Paderborn dataset, the aero-engine test platform dataset, and the laboratory-acquired dataset. The results show that PNMC performs exceptionally well in fault diagnosis tasks on all three datasets, especially under noisy and finite sample conditions, where its classification performance and noise resistance are superior to other methods. Furthermore, PNMC has high practical value in industrial applications and is suitable for fault diagnosis tasks in complex environments.

[0067] 3. This invention not only has practical application value in the field of rotating machinery fault diagnosis under limited sample conditions, but also provides new research ideas for other classification and diagnosis tasks facing limited sample constraints. Attached Figure Description

[0068] Figure 1 This is a fault diagnosis method based on PNMC.

[0069] Figure 2 This is a structural diagram of GFE.

[0070] Figure 3 This is a structural diagram of MSFE.

[0071] Figure 4 This is a structural diagram of FIPF.

[0072] Figure 5 This is a structural diagram of MIP and TMA in IPEA.

[0073] Figure 6 This serves as a testing platform for the PU dataset.

[0074] Figure 7 This serves as a test platform for the LabSC dataset.

[0075] Figure 8 This serves as a test platform for the AETP dataset.

[0076] Figure 9 Visualize the feature classification results of PNMC on TP1, TP2, TP3 and TP4.

[0077] Figure 10 Visualize the feature classification results of PNMC on TL1, TL2, TL3 and TL4.

[0078] Figure 11 Visualize the feature classification results of PNMC on TA1, TA2, TA3 and TA4.

[0079] Figure 12 These are the ablation experiment results from the PU dataset.

[0080] Figure 13 These are the results of ablation experiments on the LabSC dataset.

[0081] Figure 14 These are the ablation experiment results on the AETP dataset. Detailed Implementation

[0082] The technical solution of the present invention will be further described below with reference to the accompanying drawings, but it is not limited thereto. Any modifications or equivalent substitutions to the technical solution of the present invention that do not depart from the spirit and scope of the technical solution of the present invention should be covered within the protection scope of the present invention.

[0083] This invention provides a method for diagnosing rotating machinery faults based on a multi-classifier progressive network, such as... Figure 1As shown, the process includes three steps: data collection and preprocessing, model training, and model testing. First, a fault simulation platform is built, and vibration signals from rotating machinery are collected using sensors. The one-dimensional signal is then divided into training and testing sets using a sliding window technique. Second, the PNMC is trained using the training set to optimize the model parameters. Finally, fault diagnosis tests are performed using the testing set to generate the final diagnostic results. The specific steps are as follows:

[0084] Step 1: Build a fault simulation platform and collect vibration signals of rotating machinery through sensors.

[0085] Step 2: Use the sliding window technique to divide the acquired one-dimensional vibration signal into a training set and a test set.

[0086] Step 3: Construct the PNMC and initialize the weights and biases. The specific steps for constructing the PNMC are as follows:

[0087] Step 31, Construct GFE:

[0088] like Figure 2 As shown, the GFE mainly consists of MSFE, FIPF, AdaConv, and ResConv, used to enhance the model's feature extraction capabilities. AdaConv includes AdaConv a and AdaConv b, both composed of convolutional layers, batch normalization (BN), and rectified linear units (ReLU), capable of adaptively changing the number of channels in the input features as needed. ResConv consists of convolutional layers, BN, ReLU, and skip connection structures, further achieving the fusion of global and local features while effectively preventing degradation problems that may occur during model training. Assuming the input sample is x, the output features after MSFE processing can be represented as... and These are the features output by LKDConv a, LKDConv b, SKDConv a, and SKDConvb, respectively. and Includes two processing paths, and After feature concatenation along the channel dimension, the data is passed to AdaConv a. and After fusion via FIPF, the results are passed to AdaConv b. The outputs of both paths are then passed to ResConv, where they are added element-wise to generate more refined fusion features. This enhances the model's ability to capture globally important information. Similarly, and It also includes two processing paths. and After feature concatenation along the channel dimension, the data is passed to AdaConv a. and After fusion via FIPF, the results are passed to AdaConv b. The outputs of both paths are then passed to ResConv, where they are added element-wise to generate more refined fusion features. This improves the model's ability to capture important local information. and After feature concatenation along the channel dimension, the data is passed to the feature mapping module (FM).

[0089] like Figure 3 As shown, MSFE consists of large-kernel dilated convolutional blocks (LKDConv) and small-kernel dilated convolutional blocks (SKDConv), which can be used to capture multi-scale contextual information, thereby enhancing the model's ability to understand features at different scales. LKDConv includes LKDConv a and LKDConv b, and SKDConv includes SKDConv a and SKDConv b. Specifically, MSFE effectively captures multi-scale and multi-level feature information by applying dilated convolutional blocks with different kernel sizes and dilation rates in parallel. LKDConv is used to capture a larger receptive field and extract features containing more global information; SKDConv focuses on fine-grained feature extraction, enhancing the ability to capture local information. The expression for the dilated convolutional block is as follows:

[0090]

[0091] In the formula, f(c,j) represents the value of the c-th output channel at position j, L represents the size of the convolution kernel, s represents the stride, d represents the dilatation rate, x(j·s+n·d) represents the value of the input sample at position j·s+n·d, and w(c,n) represents the weight of the c-th convolution kernel at position n.

[0092] like Figure 4 As shown, FIPF achieves interactive perceptual fusion by introducing third-domain features and utilizing the interactive perception and adaptive enhancement of multi-space features, thereby strengthening the dependencies and information interaction between features. Compared with traditional feature fusion methods, FIPF enhances the complementarity and dependency of features from different domains, significantly improving the model's ability to capture and fuse key information. The FIPF process is as follows:

[0093] (1) Feature mapping: Assume the input features are fα and fδ Each feature is mapped using a 1×1 convolutional layer to generate corresponding mapped features. The specific calculation formula is as follows:

[0094]

[0095] In the formula, and These are all mapping features in different spaces, Conv 1×1 This represents a 1×1 convolution operation. Furthermore, a third domain feature f is introduced into the feature domain. αδ =f α +f δ The feature is then mapped using a 1×1 convolutional layer to generate a third-domain mapping feature, further enhancing the interaction between features. The specific calculation formula is as follows:

[0096]

[0097] (2) Feature Interaction Perception: For the mapping features of different domains, interaction perception processing is performed, and corresponding interaction perception features are generated through sigmoid. sα , sαδ and sδ The specific calculation formula is as follows:

[0098]

[0099]

[0100] In the formula, T represents the transpose operation, and d d Indicates the feature dimension.

[0101] (3) Weight generation: For the interactive perception features of different domains, GAP processing is performed separately to generate a global description g. α g αδ and g δ Subsequently, the global description generates corresponding weights w using MLP and sigmoid. α w αδ and w δ This is used to adjust the importance of features. The specific calculation formula is as follows:

[0102] w α =Sigmoid{MLP[GAP(s α (8)

[0103] w αδ =Sigmoid{MLP[GAP(s αδ (9)

[0104] w δ =Sigmoid{MLP[GAP(s δ (10)

[0105] (4) Weighting and fusion: mapping features for different domains and , respectively with the corresponding weight wα w αδ and w δ Multiply the features, then concatenate them along the channel dimension to generate a fused feature f. αδ The specific calculation formula is as follows:

[0106]

[0107] In the formula, Cat c This indicates that feature splicing operations are performed along the channel dimension.

[0108] and The features f are generated by fusing them separately using ResConv. L and f S Finally, the feature f is analyzed along the length dimension. L and f S The data is spliced ​​together to generate the fused particle features f. LS The specific calculation formula is as follows:

[0109] f LS =Cat l (f L ,f S (12)

[0110] In the formula, Cat l This indicates that feature concatenation is performed along the length dimension.

[0111] Step 32, Build IPEA:

[0112] The IPEA includes FM, MIP, TMA, and skip connection structures, which enhance the model's feature representation capabilities. FM consists of three 1×1 convolutional layers, mapping input features to multiple spaces while maintaining the feature size, thereby increasing the diversity and information content of feature representations. The specific calculation formula is as follows:

[0113] {x α ,x δ ,x ζ}=Conv 1×1 (f LS (13)

[0114] Figure 5 The structure of MIP and TMA in IPEA is demonstrated. MIP consists of linear transformation, masking operation, and superposition fusion, while TMA consists of multi-scale masking mechanism and multi-head attention mechanism. MIP and TMA work together to capture multi-level contextual information, strengthen the interaction between local and global features, enhance the dependencies between features, and dynamically adjust the model's focus on key information, thereby improving the accuracy and flexibility of feature representation.xα and xδ Two sets of query, key, and value features are generated through linear transformation and masking operations, respectively. The specific calculation formula is as follows:

[0115]

[0116] In the formula, and They are respectively xα and xδ The query features are generated through linear transformation and masking operations. and They are respectively xα and xδ The key features generated after linear transformation and masking operations, V i α and V i δ They are respectively xα and xδ The value feature generated after linear transformation and masking operation, where i represents the index of the attention head. Let C represent the projection matrix, and d represent the number of channels. head M represents the dimension of each attention head. i This represents the multi-scale masking matrix. The specific calculation formula is as follows:

[0117] d head =C / h (17)

[0118] M i =M o ·Scale i (18)

[0119] In the formula, h represents the number of attention heads, and M o Represents the original masking matrix, Scale i This represents the scaling factor for the i-th attention head. For two sets of query, key, and value features, the final MIP result is generated by element-wise addition. The specific calculation formula is as follows:

[0120]

[0121] TAM through the and Inner product calculations are performed, and masks of different scales are applied to the inner product results for different attention heads. Then, Softmax normalization is applied to obtain the attention distribution. The specific calculation formula is as follows:

[0122]

[0123] In the formula, T represents the transpose operation. Next, attention distribution is utilized... Ai For V i αδ The weighted summation yields the output of a single attention head. The specific calculation formula is as follows:

[0124] head i =A i ·V i αδ (twenty one)

[0125] Finally, the outputs of all attention heads are concatenated, and the final result x of TAM is obtained through linear mapping. αδ The specific calculation formula is as follows:

[0126] x αδ =Cat(head1,head2,…,head) h )·W o (twenty two)

[0127] In the formula, This represents the projection matrix.

[0128] Using a skip connection structure, xαδ and x ζ Element-by-element addition not only preserves the original features of the input but also effectively avoids degradation problems that may occur during model training. The specific calculation formula is as follows:

[0129] x αδζ =x αδ +x ζ (twenty three)

[0130] Step 33: Build PDCN:

[0131] The PDCE includes a multi-classifier structure, an ensemble loss function, and a dynamic adjustment mechanism. By gradually introducing early, mid, and late classifiers and setting a dedicated loss function for each classifier, while dynamically adjusting the weights, it fully utilizes multi-level features to enhance the model's feature learning ability.

[0132] In a multi-classifier architecture, early, mid, and late classifiers possess different feature perception and processing capabilities, enabling them to process and learn feature information at various levels and provide key decisions, significantly improving overall classification performance. The early classifier focuses on processing shallow features, capturing global trends and basic patterns to quickly generate preliminary classification results. The early classifier consists of GAP, FC, and Softmax, learning the basic features of the input sample through a linear combination of shallow features. The output of the early classifier can be expressed as:

[0133] y e=Softmax{w e2 ·FC[w e1 ·GAP(f LS )+b e1 ]+b e2} (twenty four)

[0134] In the formula, and These represent the weights of the first and second layers of the early classifier, respectively. and These represent the biases of the first and second layers of the early classifier, respectively. De This represents the output dimension of the first layer of the early classifier. The intermediate classifier extracts deeper features and combines them with more contextual information to effectively capture subtle differences and improve classification accuracy. The intermediate classifier has the same structure as the early classifier, consisting of GAP, FC, and Softmax. It learns high-level features of the input samples by linearly combining intermediate features. The output of the intermediate classifier can be represented as:

[0135] y m =Softmax{w m2 ·FC[w m1 ·GAP(x αδζ )+b m1 ]+b m2} (25)

[0136] In the formula, wm1 and wm2 These represent the weights of the first and second layers of the intermediate classifier, respectively. bm1 and bm2 ...

[0137]

[0138] In the formula, exp represents the exponential function, GAP represents the global average pooling operation, K represents the total number of classes, and x R This represents the input features of the late classifier.

[0139] Early, mid, and late-stage classifiers each employ different loss functions to more efficiently learn their respective specific target features. The early classifier focuses on the accuracy of learning global and shallow features, using a focus loss function to reduce attention to easily classified samples and strengthen the learning of difficult samples, thereby improving the stability of global classification. The formula for calculating the focus loss function is:

[0140] L e =-α t (1-p t ) γ log(p t (27)

[0141] In the formula, p t This represents the model's predicted probability of the correct class, reflecting the model's confidence in the correct class. α t γ represents the balance factor, used to adjust the weights of positive and negative samples. γ represents the focusing parameter, used to reduce the weight of easily classified samples and increase attention to difficult samples. The intermediate classifier utilizes more contextual features and employs the generalized cross-entropy loss function to better focus on key features, improving global classification accuracy under imbalanced and complex sample conditions. The formula for calculating the generalized cross-entropy loss function is:

[0142]

[0143] In the formula, Let q represent the model's predicted probability of the true label y, and q be a hyperparameter. By adjusting the value of q, the loss function of the intermediate classifier can achieve a balance between the model's sensitivity to samples and its stability, thereby effectively reducing the influence of limited and noisy samples. The late classifier uses the standard cross-entropy loss function, focusing on refining deep features to further optimize the final classification decision and improve the model's classification accuracy in complex scenarios. The formula for calculating the standard cross-entropy loss is:

[0144]

[0145] In the formula, y k p represents the true label of the k-th sample. k This represents the probability that a sample is predicted to be in the k-th class. The formula for calculating the ensemble loss function is:

[0146] L = L e +L m +L f (30)

[0147] In the early stages of training, the dynamic weight adjustment mechanism prioritizes the loss weights of early and mid-stage classifiers to enhance the rapid learning of global features. As training progresses, the mechanism gradually shifts towards learning more refined features, progressively increasing the loss weights of late-stage classifiers to optimize the final classification performance. Specifically, the dynamic weight adjustment mechanism calculates the gradient of the current loss relative to the weights of each classifier through training feedback, adjusting the contribution of each classifier's loss to gradient updates in real time, thereby optimizing model parameters and improving classification accuracy and stability. The calculation formula for the dynamic weight adjustment mechanism is:

[0148]

[0149] w f (t)=1-w e (t)-w m (t) (33)

[0150] In the formula, w e w m and w f These represent the loss weights for the early, mid, and late-stage classifiers, respectively; t represents the number of iterations; and η represents the learning rate. This represents the gradient of the current loss L with respect to the weights of the earlier classifier loss.

[0151] PDCE gradually introduces early, mid, and late-stage classifiers, and combines them with an ensemble loss function and dynamic adjustment mechanism to comprehensively improve the model's ability to learn complex sample features, thereby significantly enhancing classification accuracy and stability. During training, the formula for calculating the ensemble dynamic loss is:

[0152] L = w e ·L e (y e ,y true )+w m ·L m (y m ,y true )+w f ·L f (y f ,y true (34)

[0153] In the formula, y true This indicates the actual label.

[0154] Step 4: Train the PNMC using the training set to optimize the model parameters.

[0155] Step 5: Perform fault diagnosis tests using the test set to generate the final diagnosis results.

[0156] Example:

[0157] To verify the effectiveness and superiority of this invention, a series of experiments were conducted on the Paderborn University (PU) dataset, the Laboratory Self-Collected (LabSC) dataset, and the Aero-Engine Test Platform (AETP) dataset. Fault diagnosis experiments were performed on a computing platform equipped with an Intel Xeon Gold 6348 CPU and an A800 GPU. The algorithm was developed using Python 3.8.10, employing PyTorch 2.0.0 and the CUDA 11.8 framework.

[0158] Dataset description:

[0159] (1) The PU dataset covers a variety of bearing conditions and is one of the important experimental data sources in the field of rotating machinery fault diagnosis. For example Figure 6 As shown, the test bench mainly includes a motor, torque measurement shaft, rolling bearing test module, flywheel, and load motor. During data acquisition, the drive system speed was 1500 r / min, the radial force applied to the bearing was 1000 N, the load torque of the transmission system was 0.7 Nm, and the sampling frequency was 64 kHz. Based on the different fault locations, fault modes, fault levels, and fault combinations, the PU dataset can be divided into ten types, as detailed in Table 1. Fault locations include outer and inner rings. Fault modes include pitting (P) and indentation (I). Fault levels are divided into Level 1, Level 2, and Level 3. Fault combinations are divided into single damage (S) and repeated damage (R).

[0160] Table 1 Information about the PU dataset

[0161]

[0162] (2) This invention establishes a comprehensive laboratory mechanical diagnostic simulation platform, collects vibration signals of deep groove rolling bearings under different conditions, and constructs a LabSC dataset to verify the performance of the proposed method and support related research. For example... Figure 7 As shown, the test bench mainly includes a drive motor, motor controller, tachometer, accelerometer, bearings, and load. During data acquisition, the load was 5.25 N, and the sampling frequency was 10 kHz. Based on the fault location and load speed, the LabSC dataset can be divided into ten types, as detailed in Table 2. Fault locations include the inner race, rolling elements, and outer race. The load speeds are 1200 r / min, 1800 r / min, and 2400 r / min, respectively.

[0163] Table 2 Information about the LabSC dataset

[0164]

[0165] (3) Figure 8 As shown, the test bench for the AETP dataset mainly includes an improved aero-engine, an electric motor drive system, and a lubrication system. During data acquisition, sensor measurement points were arranged at five different locations, with displacement sensors at positions 1 and 2, and acceleration sensors at positions 3, 4, 5, and 6. The sampling frequency for all points was 25 kHz. The 3D model shows that the dual-rotor structure consists of a low-pressure (LP) compressor and a high-pressure (HP) compressor. Based on the speed level, fault location, and fault length, the AETP dataset can be divided into ten types, as detailed in Table 3. Speed ​​levels include low 1 (L1), low 2 (L2), low 3 (L3), high 1 (H1), high 2 (H2), and high 3 (H3). Fault locations include both the inner and outer rings, with a fault depth of 0.5 mm. The fault length of the outer ring is 0.5 mm. The fault lengths of the inner rings are 0.5 mm and 1.0 mm, respectively.

[0166] Table 3 Information about the AETP dataset

[0167]

[0168]

[0169] Experimental setup and evaluation metrics:

[0170] Table 4 illustrates the architecture of PNMC. The optimizer used in the experiments was Adam. The selected parameters were based on multiple experimental results, aiming to balance the model's training efficiency and accuracy. To ensure the stability of model training, the dynamic adjustment mechanism was applied only during the first third of the training phase, and adjustments were made every tenth of the training phase. Furthermore, all experiments were conducted under the same hardware configuration and programming environment to eliminate performance differences caused by non-model factors, thereby ensuring the reliability and comparability of the experimental results.

[0171] Table 4 PNMC Architecture

[0172]

[0173] This invention sets up four experimental scenarios: balanced finite sample, unbalanced finite sample, noisy & balanced finite sample, and noisy & unbalanced finite sample, aiming to verify the performance of PNMC under finite sample conditions and its ability to cope with noise challenges in practical applications. Table 5 shows the experimental task settings on the PU dataset, LabSC dataset, AETP dataset, and WTPG dataset. TP1, TL1, TA1, and TW1 are verification experiments under the balanced finite sample condition; TP2, TL2, TA2, and TW2 are verification experiments under the unbalanced finite sample condition; TP3, TL3, TA3, and TW3 are verification experiments under the noisy & balanced finite sample condition; and TP4, TL4, TA4, and TW4 are verification experiments under the noisy & unbalanced finite sample condition. The samples in the PU and WTPG datasets contain 2048 data points, while the samples in the LabSC and AETP datasets contain 1024 data points. The number of training samples varies depending on the dataset and experimental scenario, while the number of test samples remains constant at 100 per class. In the experiment, 0dB additive white Gaussian noise was used as additional noise to simulate interference in real-world scenarios, thereby verifying the noise suppression capability of PNMC.

[0174] Table 5 Experimental Task Settings

[0175]

[0176]

[0177] Analysis of experimental results:

[0178] (1) Figure 9The feature visualization results of PNMC on TP1, TP2, TP3, and TP4 are presented, showing the classification decision boundary and clustering situation for each sample category. It can be seen that despite a small number of misclassifications, PNMC can still effectively map each category to different spatial regions, demonstrating good diagnostic performance. Specifically, in TP1 and TP2, the diagnostic accuracy of PNMC is close to 100%. However, in TP3 and TP4, due to noise and limited sample conditions, PNMC experienced some misdiagnosis, but the overall accuracy remained high. Table 6 summarizes the diagnostic performance of different methods on TP1, TP2, TP3, and TP4. Compared with TP1 and TP2, ResNet18 and ResNet50 performed significantly worse on the more complex TP3 and TP4, showing weaker generalization ability. WDCNN and ESPNet had low accuracy and large standard deviation across all tasks, indicating poor diagnostic performance and stability. ISCNN and TFIFNet had diagnostic performance comparable to PNMC in TP1 and TP2, but were slightly inferior in the more complex TP3 and TP4. DBTIFF and DHET achieved high accuracy, but none surpassed the performance of PNMC across all four tasks. Overall, PNMC outperformed the other comparative methods across all four tasks, achieving an average accuracy of 99.57% with a low standard deviation. These results demonstrate that PNMC effectively addresses the limited sample problem, exhibiting excellent noise resistance and superior diagnostic accuracy and stability.

[0179] Table 6. Performance comparison of different methods on TP1, TP2, TP3, and TP4.

[0180]

[0181] (2) Figure 10The feature visualization results of PNMC on TL1, TL2, TL3, and TL4 are presented. The results show that PNMC exhibits high performance in all four tasks, with samples from all classes clearly mapped to different spatial regions. Table 7 summarizes the diagnostic performance of different methods on TL1, TL2, TL3, and TL4. In TL1 and TL2, except for ResNet50 and WDCNN, the accuracy of other methods exceeds 90%, with ISCNN, DHET, DBTIFF, and PNMC all exceeding 99%, demonstrating good diagnostic performance. In TL3 and TL4, the accuracy of all methods decreases due to noise and limited sample size. Nevertheless, PNMC still achieves the highest diagnostic performance in TL3 and TL4, at 97.91% and 99.06%, respectively, demonstrating good handling of limited samples and robustness to noise. It should be noted that in some tasks, the accuracy of some methods is slightly higher than that of PNMC. In TL1, DHET and DBTIFF achieved accuracies of 99.94% and 99.82%, respectively, slightly higher than PNMC's 99.42%. In TL2, ISCNN and DHET achieved accuracies of 99.83% and 99.80%, respectively, also slightly higher than PNMC's 99.63%. Nevertheless, PNMC demonstrated the best overall diagnostic performance across all four tasks, with an average accuracy of 99.00% and a low standard deviation, exhibiting superior diagnostic performance and stability. The results indicate that PNMC has significant advantages on the LabSC dataset and is suitable for fault diagnosis under finite sample conditions and noisy & finite sample conditions.

[0182] Table 7. Performance comparison of different methods on TL1, TL2, TL3, and TL4.

[0183]

[0184] (3) Figure 11The feature visualization results of PNMC on TA1, TA2, TA3, and TA4 are presented. It can be seen that despite a small number of misclassifications, PNMC effectively maps each category to different spatial regions, demonstrating good diagnostic performance. Specifically, in TA1 and TA2, there is almost no overlap or only a small amount of overlap between different categories, indicating that PNMC can effectively handle the limited sample problem and achieve good classification results. In TA3 and TA4, the noisy and limited sample conditions lead to a decrease in diagnostic accuracy, but PNMC still exhibits strong diagnostic performance, maintaining high classification accuracy for most categories, with only a few categories showing misclassification. Table 8 summarizes the performance comparison of different methods on TA1, TA2, TA3, and TA4. It can be seen that PNMC performs particularly well on multiple tasks, especially in TA3 and TA4, with accuracies reaching 99.92% and 98.79% respectively, both surpassing all compared methods, demonstrating its strong generalization performance under noisy and limited sample conditions. It should be noted that in TA1 and TA2, TFIFNet and DBTIFF have slightly higher accuracies than PNMC. Despite this, PNMC achieved an average accuracy of 99.39% across the four tasks, surpassing all comparable methods, and exhibited a low standard deviation, demonstrating superior fault diagnosis accuracy and stability. The results show that PNMC excels in fault diagnosis tasks on the AETP dataset, particularly under noisy and finite sample conditions, where its classification performance and noise resistance are superior to other methods. PNMC has high practical value in industrial applications and is suitable for fault diagnosis tasks in complex environments.

[0185] Table 8. Performance comparison of different methods on TA1, TA2, TA3, and TA4.

[0186]

[0187] Ablation analysis:

[0188] To verify the effectiveness of the proposed module, mechanism, and structure, ablation experiments were conducted on the PU, LabSC, and AETP datasets. Specific experiments included removing GFE, FIPF, IPEA, and PDCE, and observing the performance changes of PNMC. The experimental design is as follows: (1) PNMC was used as a control experiment, denoted as T1. (2) In the GFE removal experiment, dilated convolutional blocks were used instead of GFE, denoted as T2. Dilated convolutional blocks included dilated convolutional layers, BN, and ReLU. (3) In the FIPF removal experiment, feature concatenation along the channel dimension was used instead of FIPF, denoted as T3. (4) The IPEA removal experiment was denoted as T4. (5) In the PDCE removal experiment, only one late classifier was retained, denoted as T5. The ablation experiment results on the PU, LabSC, and AETP datasets are as follows: Figure 12, Figure 13 and Figure 14 As shown.

[0189] (1) Ablation of GFE: It was found that removing GFE significantly reduced the accuracy of PNMC in most tasks, especially in diagnostic tasks under noisy & balanced finite sample and noisy & unbalanced finite sample conditions. However, in some tasks, removing GFE showed a slight improvement in PNMC. Specifically, in TL1 and TL2, removing GFE improved the accuracy of PNMC by 0.32% and 0.12%, respectively. In TA2, removing GFE did not significantly change the accuracy of PNMC. Overall, removing GFE reduced the average accuracy of PNMC by 7.61%, 2.76%, and 2.15% on tasks on the PU, LabSC, and AETP datasets, respectively. The results indicate that GFE, through multi-scale feature extraction and interactive perception fusion, can effectively capture granular feature information and is crucial for most complex tasks, making it one of the key modules in PNMC. However, in some relatively simple tasks, removing GFE may help simplify the model and thus slightly improve performance. Overall, GFE plays a key role in improving the accuracy and stability of PNMC for complex tasks, especially under finite sample conditions with noise and balance or noise and imbalance.

[0190] (2) FIPF Ablation: It was found that removing FIPF significantly reduced the accuracy of PNMC across all tasks. Specifically, after removing FIPF, the average accuracy of PNMC decreased by 0.28%, 0.88%, and 0.93% on the PU, LabSC, and AETP datasets, respectively. The results indicate that FIPF can promote interactive-aware fusion of features, enhance the dependencies and information interaction between features, thereby promoting the extraction of granular features and ultimately improving the accuracy and stability of PNMC.

[0191] (3) IPEA Ablation: It was found that removing IPEA significantly reduced the accuracy of PNMC in most tasks, especially in diagnostic tasks under noisy & balanced finite sample and noisy & unbalanced finite sample conditions. Specifically, after removing IPEA, the average accuracy of PNMC decreased by 0.37%, 1.73%, and 0.40% on tasks on the PU, LabSC, and AETP datasets, respectively. The results show that IPEA can dynamically adjust the model's focus on key information through interactive perception and adaptive enhancement of multi-spatial features, enhancing the feature representation ability of PNMC and thus improving the accuracy and stability of diagnosis. However, in some relatively simple tasks, removing IPEA may help simplify PNMC and thus slightly improve performance. Nevertheless, overall, IPEA still plays an irreplaceable role in improving the accuracy and stability of complex tasks.

[0192] (4) Ablation of PDCE: It was found that removing PDCE significantly reduced the accuracy of PNMC across all tasks, especially in diagnostic tasks under noisy & balanced finite sample and noisy & unbalanced finite sample conditions. Specifically, after removing PDCE, the average accuracy of the model decreased by 0.90%, 0.69%, and 0.24% on tasks on the PU, LabSC, and AETP datasets, respectively. The results indicate that PDCE, by progressively introducing early, mid, and late classifiers and integrating dynamic loss, can fully utilize multi-level features, enhance the feature learning ability of PNMC, and thus significantly improve the accuracy and stability of diagnosis.

[0193] (5) Analysis: Ablation experiments show that the modules, mechanisms, and structures proposed in this invention play a crucial role in improving the diagnostic performance of PNMC, especially under noisy and finite sample conditions. Although removing a particular innovation may slightly improve or have no significant impact on accuracy in some tasks, the synergistic effect of these innovations greatly enhances the diagnostic accuracy and stability of PNMC in terms of overall performance. In particular, GFE and FIPF play a key role in feature extraction, IPEA plays an important role in feature representation, and PDCE provides key support in feature learning. The synergistic effect of the modules, mechanisms, and structures proposed in this invention enables PNMC to exhibit high accuracy and stability in diagnostic tasks under finite sample and noisy & finite sample conditions, thus providing an effective solution for complex industrial fault diagnosis tasks.

Claims

1. A rotating machinery fault diagnosis method based on a multi-classifier progressive network, characterized by The method comprises the following steps: Step 1, build a fault simulation platform, and collect the vibration signal of the rotating machinery through a sensor; Step 2, divide the collected one-dimensional vibration signal into a training set and a test set by using a sliding window technology; Step 3, construct a progressive network of multiple classifiers (PNMC) and initialize the weight and bias, wherein the PNMC comprises a granular feature extraction module (GFE), an interactive perception attention mechanism (IPEA) and a progressive dynamic classifier integration structure (PDCE), and the specific construction steps are as follows: Step 31, construct the GFE: The GFE is composed of a multi-scale feature extraction module (MSFE), a feature interactive perception fusion module (FIPF), an adaptive convolution block (AdaConv) and a residual convolution block (ResConv), the MSFE is composed of a large-kernel and a small-kernel empty convolution block (LKDConv and SKDConv), the large-kernel and small-kernel empty convolution blocks each comprise two sub-convolution blocks, after the data is input into the MSFE, the two sub-convolution blocks of the large-kernel and small-kernel empty convolution blocks output feature representations respectively, the two feature representations output by the large-kernel and small-kernel empty convolution blocks are spliced in the channel dimension and then transmitted to AdaConv a, fused by the FIPF and then transmitted to AdaConv b, finally, the outputs of AdaConv a and AdaConv b are transmitted to the ResConv to generate fused features, and then the fused features corresponding to the large-kernel and small-kernel empty convolution blocks are spliced and transmitted to the IPEA; Step 32, construct the IPEA: The IPEA comprises a feature mapping module FM, a masking interaction perception mechanism MIP, a multi-scale masking multi-head attention mechanism TMA and a skip connection structure, the FM is composed of three groups of convolutional layers, the MIP is composed of a linear transformation, a masking operation and a superposition fusion, and the TMA is composed of a multi-scale masking mechanism and a multi-head attention mechanism. Step 33, construct the PDCN: The PDCE comprises a multiple classifier structure, an integrated loss function and a dynamic adjustment mechanism, the multiple classifier structure comprises an early classifier, a middle classifier and a late classifier, which learn the shallow features, middle features and deep features respectively, the early classifier, the middle classifier and the late classifier have respective loss functions, and the integrated loss function is the sum of the three loss functions; Step 4, train the PNMC using the training set to optimize the model parameters; Step 5, perform fault diagnosis test through the test set to generate a final diagnosis result.

2. The multi-classifier progressive network-based rotating machinery fault diagnosis method according to claim 1, characterized in that The step 31, MSFE is composed of large kernel hole convolution block LKDConv and small kernel hole convolution block SKDConv, LKDConv includes LKDConv a and LKDConv b, SKDConv includes SKDConv a and SKDConv b; AdaConv includes AdaConv a and AdaConv b; assuming that the input sample is , after MSFE processing, the output feature representation is , , , and are the features output by LKDConv a, LKDConv b, SKDConv a and SKDConv b respectively, wherein: and include two processing paths, and are concatenated in the channel dimension and then passed to AdaConv a, and are fused by FIPF and then passed to AdaConv b, and the outputs of the two paths are passed to ResConv to generate fused features by element-wise addition. and include two processing paths, and are concatenated in the channel dimension and then passed to AdaConv a, and are fused by FIPF and then passed to AdaConv b, and the outputs of the two paths are passed to ResConv to generate fused features by element-wise addition. and are concatenated in the channel dimension and then passed to IPEA.

3. The multi-classifier progressive network-based rotating machinery fault diagnosis method according to claim 1, characterized in that The structure of the FIPF is as follows: Assume the input features are and through respectively The convolutional layer performs mapping to generate corresponding mapped features. The specific calculation formula is as follows: In the formula, and are mapping features of different spaces, represent convolution operation, introduce the third domain feature in the feature domain , and map through the convolution layer to generate the third domain mapping feature, and the specific calculation formula is as follows: For the mapping characteristics of different domains, interactive perception processing is performed, and corresponding interactive perception characteristics are generated through sigmoid , and , and the specific calculation formula is as follows: wherein denotes a transpose operation, denotes a feature dimension; For the interactive perception features of different domains, GAP processing is performed respectively to generate global descriptions, and then the global descriptions generate corresponding weights through MLP and sigmoid , and , the specific calculation formula is as follows: Mapping features for different domains , and are multiplied by corresponding weights , and respectively, and then feature concatenation is performed on the channel dimension to generate fusion features . The specific calculation formula is as follows: In the formula, represents a feature concatenation operation in the channel dimension; and are fused by ResConv respectively to generate features and , finally, the features and are spliced in the length dimension to generate the fused grain features , the specific calculation formula is as follows: In the formula, denotes that the feature concatenation operation is performed in the length dimension.

4. The rotating machinery fault diagnosis method based on multi-classifier progressive network according to claim 3, characterized in that In the step 32, the FM is composed of three groups The convolution layer is composed of three groups, which maps the input features to multiple spaces while keeping the feature size unchanged, and the specific calculation formula is as follows: The MIP is composed of linear transformation, masking operation and superposition fusion, and Two sets of query, key and value features are generated through linear transformation and masking operation respectively, and the specific calculation formula is as follows: wherein and are and query features generated by linear transformation and masking operation, and are and key features generated by linear transformation and masking operation, and are and value features generated by linear transformation and masking operation, denotes the index of attention head, denotes the projection matrix, denotes the number of channels, denotes the dimension of each attention head, denotes the multi-scale masking matrix; for two groups of query, key and value features, the final result of MIP is generated by element-wise addition , and the specific calculation formula is as follows: The TMA is composed of a multi-scale masking mechanism and a multi-head attention mechanism. The TAM performs inner product calculation on and applies different scales of masking to the inner product results of different attention heads, and then obtains an attention distribution through Softmax normalization. The specific calculation formula is as follows: In the formula, This indicates the transpose operation; next, attention distribution is used. right The output of a single attention head is obtained by performing a weighted summation, and the specific calculation formula is as follows: Finally, the outputs of all attention heads are concatenated and passed through a linear mapping to get the final result of TAM The specific calculation formula is as follows: In the formula, denotes a projection matrix; The skip connection structure is adopted, and and are added element by element, and the specific calculation formula is as follows: 。 5. The multi-classifier progressive network-based rotating machinery fault diagnosis method according to claim 4, characterized in that In the step 33, the multiple classifier structure comprises an early classifier, a middle classifier and a late classifier, wherein: The early classifier is composed of a GAP, a FC and a Softmax, learns the basic features of the input sample through linear combination of the shallow features, and the output of the early classifier is expressed as: wherein, and W1and W2denote the weights of the first and second layers of the early classifier, respectively, and B1and B2denote the biases of the first and second layers of the early classifier, respectively, denotes the output dimension of the first layer of the early classifier; The middle classifier is composed of a GAP, a FC and a Softmax, learns the high-level features of the input sample through linear combination of the middle features, and the output of the middle classifier is expressed as: wherein and W1and W2represent the weights of the first and second layers of the mid-level classifier, respectively, and b1and b2represent the biases of the first and second layers of the mid-level classifier, respectively. The late classifier is composed of a GAP and a Softmax, can accurately learn the deep features of the input sample, and the output of the late classifier is expressed as: wherein denotes an exponential function, denotes a global average pooling operation, denotes the total number of classes, denotes the input features of the late classifier.

6. The multi-classifier progressive network-based rotating machinery fault diagnosis method according to claim 5, characterized in that The early classifier adopts a focal loss function, and the calculation formula of the focal loss function is: wherein denotes the predicted probability of the model for the correct class, denotes the balancing factor, denotes the focus parameter; The middle classifier adopts a generalized cross-entropy loss function, and the calculation formula of the generalized cross-entropy loss function is: wherein denotes the predicted probability of the model for the real label , and is a hyperparameter; The late classifier adopts a standard cross-entropy loss function, and the calculation formula of the standard cross-entropy loss is: In the formula, denotes the true label of the th sample, denotes the probability that the sample is predicted to be the th class. The calculation formula of the integrated loss function is: 。 7. The rotating machinery fault diagnosis method based on multi-classifier progressive network according to claim 6, characterized in that In the step 33, the calculation formula of the dynamic weight adjustment mechanism is: wherein, , and denote the loss weights for the early, mid and late classifiers, respectively, denotes the number of iterations, denotes the learning rate, denotes the current loss gradient of the early classifier loss weight.

8. The rotating machinery fault diagnosis method based on multi-classifier progressive network according to claim 7, characterized in that In step 4, in the training process, the calculation formula of the integrated dynamic loss is as follows: In the formula, represents the true label.

Citation Information

Patent Citations

  • Image semantic segmentation method based on multiple classifiers

    CN115761229A

  • Driving control circuit with fault detection function

    CN119148608A