Asynchronous motor infrared image fault diagnosis method based on cross-normal-form feature fusion and small sample learning

By combining ConvNeXt and SwinTransformer networks and using self-space adaptive fusion module (SSAFM) for feature fusion, the problem of data scarcity and insufficient feature extraction capabilities in asynchronous motor infrared image fault diagnosis is solved, and efficient and robust fault diagnosis effect is achieved.

CN120219760APending Publication Date: 2025-06-27NORTH CHINA ELECTRIC POWER UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510294571.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art faces the problems of data scarcity and insufficient feature extraction capabilities in the infrared image fault diagnosis of asynchronous motors, resulting in insufficient accuracy and robustness of diagnostic results.

Method used

A method of fault diagnosis of infrared image of asynchronous motors based on cross-paradigm feature fusion and small sample learning is proposed. By combining ConvNeXt and SwinTransformer networks, feature fusion is used to use self-space adaptive fusion module (SSAFM) to optimize model performance using transfer learning and data augmentation technology.

Benefits of technology

It effectively reduces the demand for computing resources, enhances noise immunity, simplifies data synchronization problems, and retains more key information, thereby improving the interpretability and robustness of diagnostic results, and significantly improving the accuracy of fault diagnosis under small sample conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219760A_ABST
    Figure CN120219760A_ABST
Patent Text Reader

Abstract

The invention relates to an asynchronous motor infrared image fault diagnosis method based on cross-normal form feature fusion and small sample learning, is suitable for asynchronous motor fault diagnosis under a small sample condition, and belongs to the technical field of electrical equipment fault diagnosis. According to the method, two heterogeneous networks of ConvNeXt and Swin Transform are combined, local features and global features of an infrared image are extracted respectively, and efficient feature fusion is realized through a self-space adaptive fusion module (SSAFM). The SSAFM module further enhances the feature expression ability by using self-attention and space attention mechanisms, and significantly improves the accuracy and robustness of fault diagnosis. According to the method, under the condition that only one real image is used for training in each class, 95.14% classification precision can be obtained on a real test set, and the method is obviously superior to a single model and other advanced classification models in the prior art. An innovative solution is provided for asynchronous motor infrared image fault diagnosis, and the method has wide engineering application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of asynchronous motor fault detection, and in particular to an asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and few-shot learning. Background Art

[0002] As the core power source of rotating machinery, asynchronous motors play a crucial role in industrial fields such as mechanical manufacturing, power generation, and transportation. However, long-term continuous operation is likely to cause various faults, such as stator short-circuit faults, rotor locked-rotor faults, and cooling fan faults. If these faults are not detected and processed in a timely manner, they not only pose a threat to production safety but may also seriously affect equipment efficiency and personnel safety. Therefore, studying efficient asynchronous motor fault diagnosis technologies is of great significance for ensuring the stable operation of mechanical equipment.

[0003] Traditional asynchronous motor fault diagnosis methods mainly include electrical signal analysis, vibration signal analysis, and infrared thermography analysis, etc. Among them, infrared thermography analysis has become a research hotspot due to its non-contact measurement, high spatial resolution, and wide applicability. By capturing the surface temperature distribution of the motor, this technology can monitor the operating state without internal sensors and accurately capture changes in thermal patterns with high sensitivity, providing important support for early warning and accurate diagnosis of faults. With the rapid development of deep learning technologies, especially the excellent performance of Convolutional Neural Network (CNN) and Transformer networks in the fields of image classification and pattern recognition, the application of infrared thermography analysis in the fault diagnosis of electrical equipment has been significantly expanded.

[0004] However, the performance of deep learning models highly depends on large-scale labeled datasets. In the actual industrial environment, due to equipment limitations and environmental complexity, obtaining high-quality infrared thermal image data is both time-consuming and expensive. Therefore, few-shot learning methods are particularly important in the field of asynchronous motor infrared image fault diagnosis. In recent years, few-shot learning methods such as Transfer Learning, Generative Adversarial Network (GAN), and Meta-Learning have received extensive attention due to their efficient learning ability in the case of scarce data.

[0005] In the field of rotating machinery health monitoring, feature fusion technology helps improve the accuracy of fault diagnosis by integrating different sensor data or feature extraction methods. Feature fusion can be divided into three levels: data-level fusion preserves the most complete information, but has high requirements for computing resources and is sensitive to noise; feature-level fusion improves efficiency and interpretability, but may lose some information during the feature extraction process; decision-level fusion enhances robustness, but is prone to amplifying the deviation of a single model, and has limited optimization space and weak interpretability. Summary of the Invention

[0006] To overcome the deficiencies of traditional feature fusion methods, the present invention proposes an asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and few-shot learning to solve the problems of data scarcity and insufficient feature extraction ability in the prior art. In the fields of deep learning and computer vision, "cross-paradigm" usually refers to the synergistic effect and integration of different computing frameworks, modeling ideas, or representation mechanisms. Each paradigm has unique feature extraction strategies and information processing methods, which can capture different levels, scopes, or types of patterns and semantic information. This method effectively reduces the computing resource requirements, enhances the anti-noise ability, simplifies the data synchronization problem, and retains more key information, thus improving the interpretability and robustness of the diagnostic results.

[0007] The technical solution adopted in this application is as follows:

[0008] The present invention proposes an asynchronous motor infrared image fault diagnosis model based on cross-paradigm feature fusion and few-shot learning, which specifically includes the following steps:

[0009] (1) Feature extraction: Use the ConvNeXt network to extract the local features of the asynchronous motor infrared image, and use the SwinTransformer network to extract the global features. The ConvNeXt network can capture the detailed information such as the texture and edges of the image, while the Swin Transformer models the long-range dependencies of the image through a hierarchical attention mechanism to capture the global context information.

[0010] (2) Feature fusion: Fuse the features extracted by ConvNeXt and SwinTransformer through the Self-Spatial Adaptive Fusion Module (SSAFM). The SSAFM module combines the self-attention and spatial attention mechanisms to dynamically adjust the weights of the local and global features and generate more discriminative fusion features.

[0011] (3) Classification decision: Input the fused features into the classifier, and realize the classification of fault categories through adaptive average pooling and fully connected layers. The classifier adopts a weighted cross-entropy loss function and label smoothing regularization to improve the generalization ability of the model.

[0012] (4) Small-sample training strategy: During the training process, only 1 real image of each type of fault is taken for training, and data augmentation technology is used to generate a pseudo-validation set to optimize the model hyperparameters. Through the transfer learning strategy, pre-trained ConvNeXt and Swin Transformer models are utilized to reduce the problem of insufficient training under small-sample conditions.

[0013] The most significant advantage of this application is:

[0014] (1) Cross-paradigm feature fusion: By combining ConvNeXt and Swin Transformer, the efficient collaborative utilization of local and global features is achieved, significantly enhancing the feature expression ability.

[0015] (2) Adaptation to small-sample scenarios: Through data augmentation and transfer learning techniques, the problem of data scarcity is overcome, enabling efficient learning under small-sample conditions;

[0016] (3) Efficient module design: The self-spatial adaptive fusion module is adopted to reduce the computational complexity while ensuring the model performance;

[0017] (4) Wide application prospects: It is applicable to the asynchronous motor fault diagnosis task in industrial environments and can be extended to the infrared image fault diagnosis of other electrical equipment. Brief Description of the Drawings

[0018] Figure 1 is the structure diagram of the proposed model;

[0019] Figure 2 is the structure diagram of the backbone network;

[0020] Figure 3 is the structure diagram of SSAFM;

[0021] Figure 4 is the real sample diagram of the training set;

[0022] Figure 5 is the accuracy graph of repeated experiments for 1 training image;

[0023] Figure 6 is the confusion matrix graph with an accuracy of 95.14%;

[0024] Figure 7 is the receiver operating characteristic curve graph with an accuracy of 95.14%;

[0025] Figure 8 is the confusion matrix graph of minor faults and no-load conditions. Detailed Implementation Manner

[0026] The following further describes this application with reference to the drawings:

[0027] 1. Fault Diagnosis Model Based on Cross-Paradigm Feature Fusion and Few-Shot Learning

[0028] (1) Model Architecture

[0029] The present invention proposes a model architecture that fuses two heterogeneous feature extraction networks, namely a cross-paradigm feature fusion method based on a convolutional neural network (CNN) and a Transformer. By combining the ConvNeXt and Swin Transformer networks, the model fully exploits the advantages of local and global image features to more efficiently complete the classification task. The present invention overcomes the limitations of traditional fixed-weight fusion methods (such as simple summation or concatenation) by fusing the final features of the two networks.

[0030] To achieve flexible cross-paradigm feature fusion, the present invention designs the SSAFM module, which can dynamically adjust the weights of the input features during the training process, thereby enhancing the ability to capture diverse features. This module further combines the self-attention and spatial attention mechanisms to optimize the information interaction between channels and improve the effect of the cooperation between local and global features.

[0031] The overall architecture is as Figure 1 shown, including a backbone network (ConvNeXt and Swin Transformer) and a fusion module (SSAFM). The backbone network is pre-trained based on a large-scale dataset, and the SSAFM module focuses on efficiently fusing the extracted features to complete the final classification task.

[0032] The image classification process is divided into the following four steps:

[0033] a. Data augmentation: Perform various data augmentations (such as translation, affine transformation, color jitter, etc.) on the training images to increase sample diversity and optimize the model generalization performance, and at the same time generate a pseudo-validation set;

[0034] b. Feature extraction: Input the augmented images into the Swin Transformer Tiny and ConvNeXt Tiny networks respectively to extract global features and local features;

[0035] c. Feature fusion: Combine the features of the two networks through the SSAFM module, and use the self-attention and spatial attention mechanisms to capture cooperative information to generate a feature representation with stronger discrimination ability;

[0036] d. Classification decision: Input the fused features into the classification head, and achieve efficient discrimination of categories through pooling operations.

[0037] The above design enables the model to exhibit significant advantages in feature expression and classification performance, providing a reliable solution for the fault diagnosis of asynchronous motor infrared images in the case of few samples.

[0038] (2) Backbone Network

[0039] The proposed model integrates ConvNeXt and Swin Transformer in the feature extraction module, constructing an efficient architecture that takes into account both local and global feature extraction. The ConvNeXt network focuses on capturing local features in images and can extract detailed information such as texture and edges, as shown in Figure 2 (a). Compared with traditional convolutional neural networks, ConvNeXt enhances the ability to capture local information while maintaining high computational efficiency and scalability.

[0040] On the other hand, as the core module for global feature extraction, Swin Transformer effectively models the long-range dependencies of images through a hierarchical attention mechanism, making it suitable for capturing global context information and semantic features. Its architecture is shown in Figure 2 (b). This clearly defined division of labor enables ConvNeXt and Swin Transformer to complement each other in feature extraction, organically combining local and global features.

[0041] Through this fusion method, the model can more comprehensively understand the complex features of images, thereby providing more discriminative feature representations for downstream tasks.

[0042] (3) Cross-Paradigm Feature Fusion

[0043] SSAFM is an efficient fusion module designed specifically for integrating the local features extracted by ConvNeXt and the global features extracted by Swin Transformer. Combining self-attention and spatial attention mechanisms, SSAFM not only realizes the interactive fusion of local and global features but also overcomes the limitations of traditional feature fusion methods (such as simple concatenation or fixed weighting methods) in dynamic relationship modeling, thus generating more discriminative cross-paradigm fusion features. Its structure is shown in Figure 3 as follows.

[0044] The input of SSAFM includes the local features extracted by the ConvNeXt network and the global features extracted by the Swin Transformer network

[0045] F L = M ConvNeXt (X) (1)

[0046] F G = M Swin (X) (2)

[0047] Where: B is the batch size; C is the number of channels; H and W are the height and width of the feature map; N = H × W is the length of the flattened sequence; X is the model input, usually the original image.

[0048] To operate in the self-attention module, the input local features are flattened into a three-dimensional sequence form:

[0049] F′ L = reshape(F L , B × N × C) (3)

[0050] Where: reshape(·) is the flattening operation.

[0051] The two types of features are added pointwise to achieve preliminary fusion:

[0052] F = F′ L + F G (4)

[0053] Where:

[0054] The preliminarily fused feature F is fed into the self-attention module for dynamic interaction between local and global features. First, query (Q), key (K), and value (V) matrices are generated using linear transformations:

[0055] Q = FW Q , K = FW K , V = FW V (5)

[0056] Where: is the linear weight matrix for projection; d k = C / h is the dimension of each attention head; h is the number of attention heads.

[0057] The similarity between the query and the key is calculated through the dot product, and scaling and normalization operations are performed to generate the attention weights:

[0058]

[0059] Where: is the attention weight matrix; is the scaling factor used to prevent the dot product value from being too large.

[0060] The attention weights A are used to perform weighted summation on the value matrix to generate the updated feature representation:

[0061] F attn = AV (7)

[0062] Where:

[0063] Subsequently, the feature dimension is restored through the fully connected layer, and residual connection and normalization operations are applied:

[0064] F proj = F attn W proj (8)

[0065] F res = F + F proj (9)

[0066] F norm = LayerNorm(F res ) (10)

[0067] Where:

[0068] This mechanism can dynamically adjust the importance of local and global features, effectively capturing the long-term and short-term dependencies between them.

[0069] The features after self-attention optimization are reshaped into a four-dimensional tensor:

[0070] F 4D = reshape(F norm , B×C×H×W) (11)

[0071] Next, the spatial attention mechanism extracts the global feature intensity and significant region features from the channel dimension through average pooling and max pooling operations:

[0072]

[0073] Where: b is the batch index; c is the channel index; i, j are the spatial positions on the feature map.

[0074] After concatenating the results of the two poolings in the channel dimension, a spatial attention map is generated using lightweight convolution:

[0075] A spatial = σ(Conv(concat(F avg , F max ))) (14)

[0076] Where: is the spatial attention map; σ is the Sigmoid activation function.

[0077] The spatial attention map is used to weight the input features point by point, thus completing the enhancement of key regions and the suppression of background noise:

[0078] F fused (b, c, i, j) = A spatial (b, 1, i, j) ⊙ F4D (b, c, i, j) (15)

[0079] In the formula: ⊙ represents point-by-point multiplication.

[0080] Overall, SSAFM realizes the in-depth interaction between local and global features through the self-attention mechanism and optimizes the spatial distribution of feature representations through the spatial attention mechanism. The self-attention mechanism can dynamically model the long-term and short-term dependence relationships between features and adjust the importance of features; the spatial attention mechanism further enhances the identification ability of features by enhancing the significant regions and suppressing background noise. Compared with traditional feature fusion methods, SSAFM achieves a good balance between computational efficiency and feature expression ability, providing an efficient and robust solution for cross-paradigm feature fusion.

[0081] (4) Pooling and Classification

[0082] Pooling and classification are the key links for a deep learning model to go from feature extraction to final prediction. In this model, the fused features are first compressed into global statistical features by adaptive average pooling. Adaptive pooling can dynamically adjust the output size according to the size of the input features, thereby enhancing the adaptability of the model to different inputs. This process extracts the global information of each channel, providing a concise and efficient representation for the classification task.

[0083] Subsequently, the pooled features are flattened into a one-dimensional vector and passed to the classifier module. The classifier consists of two fully connected layers and a non-linear activation function, and incorporates the Dropout mechanism to improve the generalization ability. The first fully connected layer is used for dimensionality reduction and extraction of high-level semantic features, and at the same time introduces non-linearity through the activation function, enabling the model to learn complex decision boundaries. The second fully connected layer maps the features to the classification space to generate class prediction scores. Dropout reduces the model's dependence on specific features by randomly discarding some features during training, effectively alleviating the overfitting problem.

[0084] Overall, the pooling and classification module completes the transformation from visual features to class prediction through global feature extraction and efficient feature mapping, significantly improving the generalization ability and classification performance of the model.

[0085] (5) Small Sample Training Strategy

[0086] The present invention proposes a cross-paradigm feature fusion model that combines local features and global features, which is specifically used for the asynchronous motor infrared image classification task. The model consists of two core parts: First, local fine-grained features are extracted through a pre-trained ConvNext Tiny network; Second, a pre-trained Swin Transformer Tiny network is used to obtain global context information. When the model loads the publicly available pre-trained weights, the relevant parameters of the classification head are removed to reduce the bias in the transfer task. At the same time, the parameters of the feature extraction part are not frozen, allowing for adaptive adjustment during training to optimize the feature representation.

[0087] To improve the model performance, a weighted cross-entropy loss function (Weighted Cross-Entropy Loss) is adopted, which has good performance in the case of unbalanced class distributions. At the same time, label smoothing regularization (Label Smoothing Regularization) is introduced, and a smoothing coefficient of 0.1 is used to replace some hard label values, thereby reducing the sensitivity to noisy data. The AdamW optimizer is selected, with an initial learning rate set to 1e-4 and a weight decay coefficient of 1e-4. To improve the robustness of the optimization process, a piecewise learning rate scheduling strategy is designed: in the first 10 epochs, linear growth is used for learning rate warm-up to help the model converge quickly; subsequently, the cosine annealing strategy is adopted to gradually reduce the learning rate, which finally drops to 10% of the initial value.

[0088] To address the problem of insufficient data, a variety of data augmentation methods are designed to alleviate overfitting, including center cropping, random horizontal flipping, random vertical flipping, and random rotation. The augmented images are uniformly adjusted to 224×224 pixels and normalized according to the requirements of the pre-trained model. Only center cropping and normalization are applied to the pseudo-validation set and the test set to ensure the objectivity of the evaluation results.

[0089] During the training process, after each epoch, the performance change of the model is evaluated through the accuracy of the pseudo-validation set, and an early stopping mechanism is introduced to prevent overfitting. When the validation accuracy has not improved for 15 consecutive epochs, the training automatically terminates.

[0090] 2. Experimental Verification

[0091] (1) Dataset

[0092] This dataset was provided by Najafi et al. from Babol Noshirvani University of Technology in Iran. It is an infrared thermal image dataset dedicated to the condition monitoring of induction motors. The dataset covers various operating states of induction motors, including 8 different degrees of stator winding short-circuit faults, rotor locked-rotor faults, cooling fan faults, and no-load states. The thermal images were collected in a laboratory with an ambient temperature of 23 °C using a Dali-tech T4 / T8 infrared thermal imager on a workbench. The structure of the dataset is shown in Table 1.

[0093] Table 1 Dataset Structure

[0094]

[0095] In the experimental design, one real image was randomly selected from each category as the basic data for the training set (as Figure 4 shown), and slight data augmentation was performed on it. The augmentation operations include translation, affine transformation, elastic transformation, superpixel perturbation, color jitter, adding Gaussian noise, and blurring, etc., to improve the generalization ability of the model and alleviate the overfitting problem of small samples. Each basic image generated 20 extended images through augmentation, 10 of which were added to the training set, and 10 were used as a pseudo-validation set to optimize the model. All real images not used for training in the dataset were used as the test set to evaluate the model performance.

[0096] (2) Experimental Results

[0097] a. Training Results of 1 Real Image

[0098] To comprehensively evaluate the model performance, the present invention uses multiple metrics for analysis, including Accuracy, Recall (Macro-Average), Precision (Macro-Average), F1 Score (Macro-Average), and Area Under the Receiver Operating Characteristic Curve (ROC AUC, OVO). Accuracy is used to measure the overall correctness of the model's predictions; Recall reflects the model's ability to identify positive examples, and macro-average is used to calculate to comprehensively evaluate the performance of various categories; Precision represents the proportion of samples predicted as positive examples that are actually positive examples, and is also measured by macro-average; the F1 Score evaluates the comprehensive performance of the model through the harmonic mean of Precision and Recall; while the ROC AUC metric describes the model's ability to distinguish positive and negative samples through the "one-versus-one" method. The combination of these metrics provides a reliable basis for the comprehensive evaluation of the model performance. In addition, to further analyze the classification performance, a Confusion Matrix is plotted to visually display the distribution of classification errors, which helps to deeply understand the performance differences of the model in various categories.

[0099] Under the condition of training with 1 real image per classification, to ensure the stability and reliability of the evaluation results, the fault diagnosis experiment of asynchronous motor infrared images was repeated 5 times, and the change of Precision was statistically analyzed. As Figure 5 shown, the maximum Precision of the model reached 95.14%, and the Precision generally remained above 90%, indicating that the model has high consistency and robustness in small-sample classification tasks.

[0100] The evaluation metrics of the 5 repeated experiments are summarized in Table 2. The Confusion Matrix corresponding to the experiment with the highest Precision (95.14%) is as Figure 6 shown, and the Receiver Operating Characteristic Curve is as Figure 7 shown.

[0101] Table 2 Evaluation Metrics of Repeated Experiments

[0102]

[0103] As can be seen from Table 2, the model shows high prediction stability in terms of Precision, ranging from 0.9261 to 0.9514. Among them, the precision of the fifth experiment reaches the highest value of 0.9514, indicating its optimal performance in the reliability of positive example sample prediction. The ROC AUC indicators are all maintained at a high level of 0.9956 to 0.9973, close to 1, indicating that the model performs extremely well in the ability to distinguish between positive and negative samples. In particular, the ROC AUC value of Experiment 4 reaches the highest point of 0.9973.

[0104] Comprehensively analyzed, the experimental results verify the excellent performance and robustness of the model in small-sample classification tasks. In particular, Experiment 5 shows the most prominent performance in various indicators, providing important theoretical support and practical reference for the research and application of small-sample classification tasks.

[0105] Judging from the results of the confusion matrix, the classification accuracy of the model for the vast majority of fault categories exceeds 90%. Among them, the classification accuracies of A&B50 class, A&C30 class, A50 class, and Fan class reach 100%, fully demonstrating the excellent performance of the model in specific categories. However, for A10 class, Noload class, and categories with highly overlapping features (such as A&C&B30 class and A&C&B10 class), the classification accuracy decreases slightly, which is mainly due to misclassification caused by the similarity between different category features.

[0106] Overall, the model can efficiently identify different faults and no-load states of asynchronous motors and shows high classification accuracy in most scenarios, making it suitable for the diagnostic tasks in complex fault scenarios and providing strong support for the practical application of asynchronous motor fault diagnosis technology.

[0107] b. Training results with 3 and 5 real images

[0108] When training with 3 and 5 real images, 5 experiments were repeated respectively, and the summary of evaluation indicators is shown in Table 3.

[0109] Table 3 Evaluation indicators for 3 and 5 repeated experiments

[0110]

[0111] As can be seen from Table 3, when the number of training samples increases from 1 to 3 and then expands to 5, the classification performance of the model is significantly improved. This improvement is reflected in multiple metrics such as accuracy, recall, F1-score, and precision. Especially when the training samples increase to 5, the performance of Experiment 5 reaches the best level, with an accuracy of up to 0.9618, an F1-score of 0.9661, and the ROC AUC value approaching perfection at 0.9996. These results indicate that appropriately increasing the number of training samples can effectively enhance the robustness and comprehensive performance of the model in small-sample classification tasks.

[0112] The experimental results further verify the applicability and expansion potential of the model in small-sample classification tasks, providing a reliable solution for the fault diagnosis of asynchronous motor infrared images and small-sample scenarios in other similar fields. At the same time, these results also provide an important reference direction for model optimization and performance improvement in related fields.

[0113] c. Experimental results of slight faults and no-load state

[0114] To evaluate the diagnostic ability of the model in the state of slight faults and no-load, the present invention conducted multiple repeated experiments on 10 types of faults of A&C&B, 10 types of faults of A&C, 10 types of faults of A, and the Noload state (i.e., 10% of the short-circuited turns per phase of the stator and no-load state). The evaluation metrics of the experimental results are listed in Table 4, and the confusion matrix is as Figure 8 shown.

[0115] Table 4 shows that although there are certain fluctuations in the classification performance during the experiment, the average precision remains stable at about 90%, indicating that this method can be used as an effective auxiliary means for diagnosing slight faults of asynchronous motors. The experimental results show a significant performance improvement: the accuracy increases from 0.8291 in Experiment 1 to 0.9231 in Experiment 5, the F1-score increases from 0.8206 to 0.9283, the precision increases from 0.8684 to 0.9288, and the ROC AUC always remains at a high level (>0.95), reaching the highest value of 0.9939 in Experiment 5. This indicates that the model has strong robustness and discrimination ability in dealing with the classification tasks of slight faults and no-load state. Especially in Experiment 5, all metrics show the best performance.

[0116] Table 4 Evaluation metrics of slight faults and no-load state

[0117]

[0118] From Figure 7From the confusion matrix analysis, the classification accuracy of the A&C&B10 class and the No-load class reached 100%, indicating that the model can accurately extract features and achieve reliable classification when processing images of 10% short-circuit turns in the stator windings of phases A, C, and B and no-load conditions. However, for the A&C10 class and the A10 class, the classification performance is relatively low, with 5 and 4 images misclassified into each other's classes respectively, and the accuracy rates are 83.3% and 87.9% respectively. This misclassification phenomenon is mainly due to the high similarity of the fault features of the two classes, which increases the difficulty of distinguishing by the model.

[0119] Generally speaking, the experimental results verify the applicability and stability of the model in the diagnosis tasks of mild faults and no-load conditions. The model can accurately distinguish most fault classes and achieve high classification accuracy on the main classes. This will provide more reliable support for the practical application of asynchronous motor fault diagnosis technology and lay a solid foundation for industrial equipment maintenance.

[0120] (3) Comparative verification

[0121] To verify the advantages of the model of the present invention in small-sample classification tasks, ConvNeXt, SwinTransformer, DenseNet, EfficientNet, EfficientNetV2, MobileViT, RegNet, ShuffleNet, and Vision Transformer were selected as comparative models, and under the same training strategy, only 1 real image per class was used for training. The experimental results are shown in Table 5, comprehensively demonstrating the significant advantages of the model of the present invention in multiple evaluation metrics.

[0122] Table 5 Evaluation metrics of the comparative experiment

[0123]

[0124] As can be seen from Table 5, the model of the present invention achieved the best performance in all metrics: the accuracy rate was 0.9269, the F1 score was 0.9452, the precision and recall rates were 0.9514 and 0.9441 respectively, and the ROC AUC was as high as 0.9969. This indicates that the model of the present invention can effectively overcome the problem of feature extraction under small-sample conditions and demonstrate excellent classification ability and robustness. Compared with other models, the model of the present invention not only has obvious advantages in accuracy rate and comprehensive performance, but also is close to optimal in the ability to distinguish positive and negative samples (ROC AUC is close to 1), which is particularly prominent.

[0125] In summary, the efficiency and robustness of the model of the present invention in the small-sample classification task have been fully verified. It can extract more accurate features from a very small number of training samples and achieve excellent classification results. This provides important technical support and theoretical reference for the asynchronous motor fault diagnosis under small-sample conditions.

Claims

1. An asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning, characterized in that: The following steps are involved: (1) Data augmentation: Perform various data augmentation operations on training images, including translation, affine transformation, color jitter, etc., to increase sample diversity and optimize model generalization performance, while generating a pseudo validation set; (2) Feature extraction: The enhanced image is input into the Swin Transformer Tiny and ConvNeXt Tiny networks respectively to extract global features and local features; (3) Feature fusion: The features of the two networks are combined through the self-spatial adaptive fusion module (SSAFM), and the self-attention and spatial attention mechanisms are used to capture collaborative information to generate feature representations with stronger discrimination capabilities; (4) Classification decision: The fused features are input into the classification head, and the pooling operation is used to achieve efficient classification of categories.

2. According to claim 1, an asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning is characterized in that: The self-space adaptive fusion module (SSAFM) comprises the following steps: (1) Input local features and global features, which are extracted by ConvNeXt and Swin Transformer networks respectively; (2) Dynamically interact local and global features through the self-attention mechanism to generate updated feature representations; (3) Extract global feature strength and salient regional features from the channel dimension through the spatial attention mechanism to generate a spatial attention map; (4) The spatial attention map is used to weight the input features point by point to enhance the key areas and suppress the background noise.

3. According to claim 2, an asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning is characterized in that: The specific steps of the self-attention mechanism include: (1) Generate query (Q), key (K), and value (V) matrices through linear transformation; (2) Calculate the similarity between the query and the key through the dot product, and perform scaling and normalization operations to generate attention weights; (3) Use the attention weights to perform weighted summation on the value matrix to generate an updated feature representation; (4) Restore the feature dimension through the fully connected layer and apply residual connection and normalization operations.

4. According to claim 2, an asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning is characterized in that: The specific steps of the spatial attention mechanism include: (1) Through average pooling and maximum pooling operations, the global feature intensity and salient area features are extracted from the channel dimension; (2) After concatenating the two pooling results in the channel dimension, a spatial attention map is generated using lightweight convolution; (3) The spatial attention map is used to weight the input features point by point to enhance the key areas and suppress the background noise.

5. The asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning according to claim 1 is characterized in that: The classification decision step includes: (1) Compress spatial information into global statistical features through adaptive average pooling; (2) Flatten the pooled features into a one-dimensional vector and input it into the classifier module; (3) The classifier module consists of two fully connected layers and a nonlinear activation function, combined with the Dropout mechanism to improve generalization ability.

6. According to claim 1, the asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning is characterized in that: The small sample training strategy includes: (1) Use weighted cross entropy loss function for model training; (2) Label smoothing regularization is used to replace some hard label values ​​with a smoothing coefficient of 0.1 to reduce sensitivity to noisy data; (3) The optimizer uses AdamW, the initial learning rate is 1e-4, and the weight decay coefficient is 1e-4; (4) Design a segmented learning rate scheduling strategy. The learning rate is warmed up by linear growth in the first 10 epochs, and then the cosine annealing strategy is used to gradually reduce the learning rate.

7. The asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning according to claim 1 is characterized in that: The model is trained and tested on an infrared image dataset of asynchronous motors containing 10 fault categories and no-load states. Only one real image is used for training for each category, and a pseudo validation set is generated through data augmentation to optimize hyperparameters.

8. The asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning according to claim 1 is characterized in that: The classification accuracy of the model on the real test set reached 95.14%, significantly better than ConvNeXt, Swin Transformer and other advanced classification models.

9. The asynchronous motor infrared image fault diagnosis method based on cross-paradigm feature fusion and small sample learning according to claim 1 is characterized in that: The method is applicable to various fault diagnosis of asynchronous motors, including stator faults, rotor faults and cooling fan faults, and has good engineering applicability and robustness.