Small sample surface defect detection method based on multistage feature enhancement

Through contrast feature extraction, multi-level feature interaction and sliding multi-level feature fusion modules, the feature extraction difficulties and loss problems in surface defect detection under small sample conditions are solved, and more efficient feature fusion and detection accuracy are achieved.

CN120495823APending Publication Date: 2025-08-15HEBEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510606925.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In industrial production, surface defect detection faces the problems of data scarcity and high labeling costs, which leads to a significant decline in detection performance of traditional deep learning methods under small sample conditions, difficulty in feature extraction, and a semantic gap in feature loss and multi-level feature fusion in deep networks.

Method used

The contrast feature extraction module, multi-level feature interaction module and sliding multi-level feature fusion module are adopted to extract and optimize multi-scale local features through technologies such as multi-branch parallel architecture, heterogeneous convolution, multi-head self-attention and channel shuffling, and capture long-distance dependencies and enhance feature diversity.

Benefits of technology

It significantly improves the surface defect detection accuracy under small sample conditions, can effectively extract and fuse multi-layer features, alleviate feature loss problem, and improves the model's detection performance under small samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495823A_ABST
    Figure CN120495823A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample surface defect detection method based on multi-level feature enhancement. The method comprises the following steps: S1, constructing a method framework of a contrast feature extraction module, a multi-level feature interaction module and a sliding multi-level feature fusion module; s2, extracting and integrating context features and Gaussian saliency features of the defect target; s3, the extracted multi-scale local features are optimized, and the long-distance dependency relationship of the target in the complex background is captured from the high-level features; and S4, combining the features with different resolutions to serve as a supplement of single-layer feature information to ensure the diversity of fusion features. Feature extraction is guaranteed, and the problem of feature loss in a deep network is relieved; and the diversity of fusion features is ensured by fusing more feature layers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of industrial surface defect detection and deep learning technology, and in particular to a small sample surface defect detection method based on multi-level feature enhancement. Background Art

[0002] Surface defects caused by uncertainties in the industrial production process not only directly affect the appearance quality and mechanical properties of the product, but may also cause functional failure or even the scrapping of the entire product. This quality risk makes surface defect detection technology a key link in ensuring product yield and optimizing production costs. However, in actual industrial scenarios, defect samples are inherently scarce and have high labeling costs, which poses a fundamental challenge to traditional deep learning methods: supervised learning-based detection models are difficult to achieve sufficient feature learning under limited labeled data conditions, resulting in a significant decrease in detection performance. Small sample learning, as a new learning paradigm, can make full use of a small amount of sample data for learning and training, solving the problem of poor defect detection results of traditional deep learning methods caused by insufficient and low-quality labeled data.

[0003] Algorithms that learn from a small number of samples are called few-shot learning. Their goal is to use a small amount of sample data to learn new tasks. Through techniques such as metric learning, transfer learning, and meta-learning, efficient defect detection can be achieved under limited data conditions. Metric learning projects samples into a feature space by designing a mapping function, narrowing the distance between similar samples. Transfer learning uses pre-trained models with abundant labeled data from the source domain to transfer the learned knowledge to target domain tasks. By fine-tuning the model parameters, rapid learning is achieved under limited target domain samples. Meta-learning methods train on different tasks to enable the model to learn parameters or optimization strategies suitable for new tasks, thereby improving the model's generalization ability under small sample conditions.

[0004] In the advancement of intelligent industrial quality inspection, small-sample learning technology offers a new research path for addressing the scarcity of defect data. For example, the deep metric learning framework constructed in the article [Kim MS, Park T, Park PG. Classification of steel surface defects using convolutional neural network with few images [C] / / 2019 12th Asian control conference (ASCC). IEEE, 2019: 1398-1401.] employs a parameter-sharing convolutional feature encoder and effectively addresses the feature separability issue for steel surface defects through a joint optimization mechanism using the cosine similarity metric and a triplet loss function. Experiments on the NEU steel surface defect benchmark database demonstrate that this architecture achieves multi-class classification accuracy of 85.1%-86.5% with only 5-10 annotated samples, validating the generalization capability of metric learning in small-sample scenarios. In terms of optimizing the transfer learning framework, the article [Yang Zhiming, Huang Tianlun, Xue Yan, et al. High-precision detection method for injection molded product defects under unbalanced sample conditions [J]. Automation and Information Engineering, 2024, 45(06): 53-58+72.] proposed a dynamic sample distribution reconstruction strategy, which enhances the model's sensitivity to defect representation through a feature space recalibration mechanism. The two-stage training paradigm they designed establishes a universal feature base in the pre-training phase and implements adaptive feature distillation in the fine-tuning phase, improving model recognition performance by 19.7% under limited sample conditions. It is worth noting that the article [Shen Zhenyu, Zhu Changming, Wang Zhe. YOLOv3 target detection model based on MAML algorithm [J]. Journal of East China University of Science and Technology (Natural Science Edition), 2022, 48(01): 112-119.] deeply integrates the meta-learning mechanism with YOLOv3 (You Only Look Once, YOLO) at the target detection network architecture level, and constructs a two-stage parameter optimization architecture: the outer loop establishes meta-knowledge priors, and the inner loop implements a meta-task-driven parameter update strategy. Experimental data show that this improved scheme increases the average detection accuracy to 87.6% while maintaining a 2.3-fold increase in model convergence speed, providing a new technical paradigm for small-sample industrial quality inspection. In summary, although the small-sample method provides ideas for defect detection with a small amount of labeled data information, the existing methods still have three key limitations at the feature expression level: First, traditional convolution is difficult to capture the local contrast features of defects; second, the long-range dependency modeling of deep features is insufficient; and third, there is a semantic gap in multi-level feature fusion. Therefore, how to extract enough discernible features from a limited small sample is still a key issue that needs to be solved in current research.

[0005] To address these issues, this paper proposes a small-sample surface defect detection method based on multi-level feature enhancement. This method first proposes a contrast feature extraction module to fully extract local information features, contextual information, and contrast features. Secondly, a multi-level feature extraction module is further proposed to capture long-range dependencies. Finally, a sliding multi-level feature interaction module is proposed to fully interact features between adjacent levels, making the semantics more consistent and significantly improving the model's detection accuracy under small sample conditions. Summary of the Invention

[0006] The purpose of the present invention is to provide a small sample surface defect detection method based on multi-level feature enhancement, which includes three parts: a contrast feature extraction module, a multi-level feature interaction module and a sliding multi-level feature interaction module. Among them, the contrast feature extraction module extracts and integrates the contextual features and Gaussian saliency features of the defect target to solve the problem of difficult feature extraction. The multi-level feature interaction module optimizes the extracted multi-scale local features, captures the long-distance dependencies of the target in the complex background from the high-level features, and alleviates the problem of feature loss in the deep network. The sliding multi-level feature interaction module combines features of different resolutions as a supplement to the single-layer feature information, and ensures the diversity of the fused features by fusing more feature layers.

[0007] The technical solution adopted by the present invention is:

[0008] A small sample surface defect detection method based on multi-level feature enhancement includes the following steps:

[0009] S1: Constructing the method framework of contrast feature extraction module, multi-level feature interaction module and sliding multi-level feature fusion module;

[0010] S2: A multi-branch parallel architecture is used in the contrast feature extraction module. Three heterogeneous convolution operations are used to achieve multi-level feature extraction of the target. The module focuses on integrating the contextual features of the target object and the saliency features based on the Gaussian distribution to address the difficulty of feature extraction.

[0011] S3: Optimize the extracted multi-scale local features in the multi-level feature interaction module. This allows high-level features to capture long-range dependencies of objects in complex backgrounds, alleviating the problem of feature loss in deep networks.

[0012] S4: In the sliding multi-level feature fusion module, features of different resolutions are combined to supplement the single-layer feature information. The diversity of the fused features is ensured by fusing more feature layers. While retaining the original layer features, multi-scale features are added, and information exchange is enhanced through channel shuffling. The supplementary features are composed of features from other layers in a nonlinear manner. This not only adds multi-scale features while retaining the original features, but also enhances the information exchange between the supplementary features and the original features.

[0013] Furthermore, in step S2, the contrast feature extraction module adopts a multi-branch parallel architecture to achieve multi-level feature extraction of the target through three heterogeneous convolution operations, and focuses on integrating the context-related features of the target object and the saliency features based on Gaussian distribution to solve the problem of difficult feature extraction. The process includes the following steps:

[0014] 1-1) In the feature extraction process, a standard 3×3 convolution kernel is first used to capture the fine-grained features of the target, including visual attributes such as shape, structure, and texture. This process is similar to the feature extraction mechanism of traditional detection networks.

[0015] 1-2) Secondly, by introducing a dilated convolution operation with a dilation rate of 1, the module can effectively extract local contextual feature information, expanding the receptive field while accurately capturing artifact features around the defect. The introduction of this local contextual information significantly enhances the target's feature representation and plays a significant role in improving detection performance.

[0016] 1-3) Finally, the module uses Gaussian Saliency Convolution (G) to specifically extract salient features of small-scale targets. This design effectively solves the problem of feature loss when detecting defects in small and medium-sized targets.

[0017] Furthermore, in step S3, the extracted multi-scale local features are optimized in the multi-level feature interaction module. In this way, the long-range dependencies of the target in the complex background are captured from the high-level features to alleviate the problem of feature loss in the deep network. The process includes the following steps:

[0018] 2-1) Specifically, MFIM contains (k-1) ViT branches, all of which have the same structure. CFEM is used as the feature embedding module to extract multi-scale local features. Feature F i Flattened into a 2D block in is the number of blocks, (P i , P i) is the resolution of each block. After position embedding, we get Obtained Where n is the number of tokens, and

[0019] Embedded token E (i) Divided into m heads in Then enter the multi-head self-attention module (MSA) to obtain the interaction This article defines these processes as:

[0020]

[0021] Among them, LN represents layer normalization. In each head, the multi-head self-attention module MSA defines three trainable weight matrices to transform the query Q (i) , key K (i) Sum V (i) .Then, is input into the MLP module to obtain the final The result of MLP can be expressed as:

[0022]

[0023] in

[0024] 2-2) Next, Reintegrated into Feature V i , where P i =H k =W K The result of the ViT branch can be expressed as:

[0025]

[0026] Furthermore, in step S4, in the sliding multi-level feature fusion module, features of different resolutions are combined as a supplement to the single-layer feature information, and the diversity of the fused features is ensured by fusing more feature layers. SMIM supplements multi-scale features while retaining the original layer features, and enhances information exchange through channel shuffle. The supplementary features are composed of features from other layers in a nonlinear manner. It not only adds multi-scale features while retaining the original features, but also enhances the information exchange between the supplementary features and the original features. The process includes the following steps:

[0027] 3-1) First, generate supplementary features. Use the features X generated by the backbone network in four different stages ori ={Xori |1≤i≤4} as the original features. An adaptive weighting function is then used to generate supplementary features for the other three layers, excluding layer i. Secondly, the original features and the supplementary features of layer i are shuffled from a channel perspective. Finally, a convolutional layer is used to adjust the number of channels in the shuffled features to match that of the original features of layer i.

[0028] The calculation formula of the SMIM module is as follows:

[0029]

[0030] in, is the convolutional feature extracted by the backbone network at layer i; represents the complementary features of the i-th layer; CAT represents the concatenation layer; CH_Shuffle represents channel shuffle; W1 represents a 1×1 convolutional layer; represents the i-th layer features enhanced by the SMIM module.

[0031] 3-2) The specific generation process of the supplementary features is as follows:

[0032] First, the feature maps of different layers are adjusted to the same size and number of channels as the i-th layer features using up-sampling, down-sampling, and 1×1 convolution operations; secondly, the adjusted three-layer features are spliced along the channel to form multi-scale original supplementary features for the i-th layer supplementary information; finally, the channel attention mechanism is used to capture the spatial correlation between features and enhance the feature representation ability generated by the network. Specifically, global average pooling (GAP) is performed on the original supplementary features of the i-th layer to obtain one-dimensional context features, and then two fully connected layers and sigmoid activation functions are used to obtain the adaptive fusion weights of the original supplementary features of the i-th layer. Finally, the weights are scaled through the normalization layer, and the features are reweighted and fused to obtain the final supplementary features of the i-th layer.

[0033] 3-3) Supplementary features The calculation formula is as follows:

[0034]

[0035] in: represents the supplementary features of the i-th layer; Represents the i-th layer convolutional feature map extracted by the backbone network; F is a nonlinear function used to perform channel-level scaling on multi-level features; λ is a hyperparameter used to adjust the balance between original features and supplementary features; represents multiplication, ∑ represents addition; T is a function that adjusts the feature maps of different stages to the same size and number of channels as the feature maps of the i-th layer; is the original supplementary feature of the i-th layer generated by the original feature of the n-th layer; GAP represents the global average pooling operation; fc represents the fully connected layer; δ represents the ReLU activation function; σ represents the Sigmoid activation function; Softmax represents the normalized exponential function.

[0036] The channel shuffle operation concatenates the original features and supplementary features along the channel and then shuffles them along the channel to enhance the information flow between them. Specifically, the original features and supplementary features are staggered along the channel to achieve better feature fusion.

[0037] The beneficial effects of adopting the above technical solution are:

[0038] The purpose of the present invention is to provide a small sample surface defect detection method based on multi-level feature enhancement, which includes three parts: a contrast feature extraction module, a multi-level feature interaction module, and a sliding multi-level feature fusion module. The contrast feature extraction module extracts and integrates the contextual features and Gaussian saliency features of the defect target, solving the problem of difficult feature extraction; the multi-level feature interaction module optimizes the extracted multi-scale local features, captures the long-range dependencies of the target in a complex background from the high-level features, and alleviates the problem of feature loss in deep networks; the sliding multi-level feature fusion module combines features of different resolutions as a supplement to the single-layer feature information, and ensures the diversity of the fused features by fusing more feature layers. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A framework diagram of a small sample surface defect detection method based on multi-level feature enhancement;

[0040] Figure 2 Diagram of contrast feature extraction module;

[0041] Figure 3 Sliding multi-level feature fusion module diagram;

[0042] Figure 4 NEU-DET data graph;

[0043] Figure 5 PVEL-AD dataset;

[0044] Figure 6 Small sample task division;

[0045] Figure 7 Comparison of mAP results of different methods (NEU-DET);

[0046] Figure 8 Comparison of Precision results of different methods (NEU-DET);

[0047] Figure 9 Comparison of Recall results of different methods (NEU-DET);

[0048] Figure 10 Comparison of F1 results of different methods (NEU-DET);

[0049] Figure 11 Comparison of mAP results of different methods (PVEL-AD);

[0050] Figure 12 Comparison of Precision results of different methods (PVEL-AD);

[0051] Figure 13 Comparison of Recall results of different methods (PVEL-AD);

[0052] Figure 14 Comparison of F1 results of different methods (PVEL-AD);

[0053] Figure 15 Visualization of detection results of NEU-DET dataset;

[0054] Figure 16 Visual detection results of the PVEL-AD dataset;

[0055] Figure 17 Line graph of mAP results of ablation experiment (NEU-DET);

[0056] Figure 18 Line chart of ablation experiment Precision results (NEU-DET);

[0057] Figure 19 Line graph of Recall results of ablation experiment (NEU-DET);

[0058] Figure 20 Line graph of F1 results of ablation experiment (NEU-DET);

[0059] Figure 21 Ablation experiment mAP result line chart (EVPL-AD);

[0060] Figure 22 Line chart of ablation experiment Precision results (EVPL-AD);

[0061] Figure 23 Line chart of recall results of ablation experiment (EVPL-AD);

[0062] Figure 24 Line graph of F1 results of ablation experiment (EVPL-AD);

[0063] Figure 25 grad_cam visualization of the SMFE-Net feature layer on the NEU-DET dataset;

[0064] Figure 26 grad_cam visualization of the SMFE-Net feature layer on the EVPL-AD dataset; DETAILED DESCRIPTION

[0065] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0066] The present invention takes surface defect detection under small sample conditions as the background, uses deep learning technology as the carrier, and takes contrast feature extraction module, multi-level feature interaction module and sliding multi-level feature fusion module as the main method framework. It proposes a small sample surface defect detection method based on multi-level feature enhancement. The framework is as follows Figure 1 As shown, the following steps are included:

[0067] S1: Constructing the method framework of contrast feature extraction module, multi-level feature interaction module and sliding multi-level feature fusion module;

[0068] S2: In the contrast feature extraction module, a multi-branch parallel architecture is adopted to achieve multi-level feature extraction of the target through three heterogeneous convolution operations, and focuses on integrating the contextual features of the target object and the saliency features based on Gaussian distribution to solve the problem of difficult feature extraction. The method is as follows Figure 2 As shown;

[0069] In the contrast feature extraction module, a multi-branch parallel architecture is used to achieve multi-level feature extraction of the target through three heterogeneous convolution operations. It focuses on integrating the contextual features of the target object and the saliency features based on Gaussian distribution to solve the problem of difficult feature extraction. The process includes the following steps:

[0070] 1-1) In the feature extraction process, a standard 3×3 convolution kernel is first used to capture the fine-grained features of the target, including visual attributes such as shape, structure, and texture. This process is similar to the feature extraction mechanism of traditional detection networks.

[0071] 1-2) Secondly, by introducing a dilated convolution operation with a dilation rate of 1, the module can effectively extract local contextual feature information, expanding the receptive field while accurately capturing artifact features around the defect. The introduction of this local contextual information significantly enhances the target's feature representation and plays a significant role in improving detection performance.

[0072] 1-3) Finally, the module uses Gaussian Saliency Convolution (G) to specifically extract salient features of small-scale targets. This design effectively solves the problem of feature loss when detecting defects in small and medium-sized targets.

[0073] S3: Optimize the extracted multi-scale local features in the multi-level feature interaction module. This allows high-level features to capture long-range dependencies of objects in complex backgrounds, alleviating the problem of feature loss in deep networks.

[0074] 2-1) Specifically, MFIM contains (k-1) ViT branches, all of which have the same structure. CFEM is used as the feature embedding module to extract multi-scale local features. Feature F i Flattened into a 2D block in is the number of blocks, (P i , P i ) is the resolution of each block. After position embedding, we get Obtained Where n is the number of tokens, and

[0075] Embedded token E (i) Divided into m heads in Then enter the multi-head self-attention module (MSA) to obtain the interaction This article defines these processes as:

[0076]

[0077] Among them, LN represents layer normalization. In each head, the multi-head self-attention module MSA defines three trainable weight matrices to transform the query Q (i) , key K (i) Sum V (i) .Then, is input into the MLP module to obtain the final The result of MLP can be expressed as:

[0078]

[0079] in

[0080] 2-2) Next, Reintegrated into Feature V i , where P i =Hk =W K The result of the ViT branch can be expressed as:

[0081]

[0082] S4: In the sliding multi-level feature fusion module, features of different resolutions are combined to supplement the single-layer feature information. The diversity of the fused features is ensured by fusing more feature layers. While retaining the original layer features, multi-scale features are added, and information exchange is enhanced through channel shuffling. The supplementary features are composed of features from other layers in a nonlinear manner. This not only adds multi-scale features while retaining the original features, but also enhances the information exchange between the supplementary features and the original features.

[0083] SMIM combines features of different resolutions to supplement single-layer feature information and ensures the diversity of fused features by fusing more feature layers. While preserving the original layer features, it supplements multi-scale features and enhances information exchange through channel shuffle. Supplementary features are composed of features from other layers in a nonlinear manner. This adds multi-scale features while preserving the original features and enhances information exchange between the supplementary features and the original features. The process includes the following steps:

[0084] 3-1) First, generate supplementary features. Use the features X generated by the backbone network in four different stages ori ={X ori |1≤i≤4} as the original features. An adaptive weighting function is then used to generate supplementary features for the other three layers, excluding layer i. Secondly, the original features and the supplementary features of layer i are shuffled from a channel perspective. Finally, a convolutional layer is used to adjust the number of channels in the shuffled features to match that of the original features of layer i.

[0085] The calculation formula of the SMIM module is as follows:

[0086]

[0087] in, is the convolutional feature extracted by the backbone network at layer i; represents the complementary features of the i-th layer; CAT represents the concatenation layer; CH_Shuffle represents channel shuffle; W1 represents a 1×1 convolutional layer; represents the i-th layer features enhanced by the SMIM module.

[0088] 3-2) The specific generation process of the supplementary features is as follows:

[0089] First, the feature maps of different layers are adjusted to the same size and number of channels as the i-th layer features using up-sampling, down-sampling, and 1×1 convolution operations; secondly, the adjusted three-layer features are spliced along the channel to form multi-scale original supplementary features for the i-th layer supplementary information; finally, the channel attention mechanism is used to capture the spatial correlation between features and enhance the feature representation ability generated by the network. Specifically, global average pooling (GAP) is performed on the original supplementary features of the i-th layer to obtain one-dimensional context features, and then two fully connected layers and sigmoid activation functions are used to obtain the adaptive fusion weights of the original supplementary features of the i-th layer. Finally, the weights are scaled through the normalization layer, and the features are reweighted and fused to obtain the final supplementary features of the i-th layer.

[0090] 3-3) Supplementary features The calculation formula is as follows:

[0091]

[0092] in: represents the supplementary features of the i-th layer; Represents the i-th layer convolutional feature map extracted by the backbone network; F is a nonlinear function used to perform channel-level scaling on multi-level features; λ is a hyperparameter used to adjust the balance between original features and supplementary features; represents multiplication, ∑ represents addition; T is a function that adjusts the feature maps of different stages to the same size and number of channels as the feature maps of the i-th layer; is the original supplementary feature of the i-th layer generated by the original feature of the n-th layer; GAP represents the global average pooling operation; fc represents the fully connected layer; δ represents the ReLU activation function; σ represents the Sigmoid activation function; Softmax represents the normalized exponential function.

[0093] The channel shuffle operation concatenates the original features and supplementary features along the channel and then shuffles them along the channel to enhance the information flow between them. Specifically, the original features and supplementary features are staggered along the channel to achieve better feature fusion.

[0094] Based on the above steps, the present invention effectively solves the problem that industrial surface defect detection under small sample conditions is difficult to effectively detect unknown classes, and proposes a small sample surface defect detection method based on multi-level feature enhancement. The method includes three parts. The first part extracts and integrates the contextual features and Gaussian saliency features of the defect target through a contrast feature extraction module to solve the problem of difficult feature extraction; the second part optimizes the extracted multi-scale local features through a multi-level feature interaction module, captures the long-distance dependency of the target in a complex background from high-level features, and alleviates the problem of feature loss in deep networks; the third part combines features of different resolutions through a sliding multi-level feature fusion module as a supplement to the single-layer feature information, and ensures the diversity of fused features by fusing more feature layers.

[0095] In order to verify the surface defect detection performance of the method in this chapter under small sample conditions, the mAP value, PR value and F1 score were tested under four different sample sizes using the small sample evaluation method on the surface defect dataset of Northeastern University and the photovoltaic cell defect anomaly detection dataset jointly released by Hebei University of Technology and Beihang University, demonstrating the effectiveness of the SMFE-Net model. It was also compared with the latest small sample target detection methods IFDD, FSDD, MLDA-YOLOv7, Faster R-CNN-H-BFC, FSSDD and Meta-FSDet in recent years to verify the superiority of SMFE-Net.

[0096] Experimental verification of the surface defect open set detection method under small sample conditions of the present invention:

[0097] 1. Test environment

[0098] The system was Windows 11, with a Core(TM) i5-13400 CPU and a GeForce RGTX 3060 graphics card with 12GB of video memory and 32GB of RAM. The experimental software was developed using Python 3.8, the PyTorch 1.12.1 deep learning framework, Torch 1.12.1, and TorchVision 0.13.1.

[0099] 2. Test verification

[0100] Experimental results and analysis on the Northeastern University Surface Defect Detection Dataset (NEU-DET) and the photovoltaic cell defect anomaly detection dataset jointly released by Hebei University of Technology and Beihang University.

[0101] (1) Dataset description

[0102] Dataset 1: The NEU-DET steel strip surface defect dataset collects six typical surface defects of hot-rolled steel strip, including rolling scale (RS), spots (Pa), cracks (Cr), pits (PS), inclusions (In), and scratches (Sc). The database contains 1,800 grayscale images, with 300 samples for each defect. The original resolution of each image is 200 × 200 pixels.

[0103] Dataset 2: Photovoltaic Cell Defect Anomaly Detection Dataset (PVEL-AD) jointly released by Hebei University of Technology and Beihang University: This dataset consists of 36,543 high-resolution images (1024x1024), collecting one type of defect-free image and 12 different categories of abnormal defect images. These include: linear crack, finger interruption, black core, thick line, star crack, corner, fragment, scratch, vertical dislocation, horizontal dislocation, and short circuit. It includes 11,351 non-abnormal samples, 21,044 defect samples with box annotations, and 4,148 near-infrared image defect samples with category annotations. Since some types of defects rarely occur in the manufacturing process of photovoltaic modules, the small number of samples is not conducive to model training. Therefore, this paper uses six defect categories: linear cracks, broken grids, black cores, thick lines, star-shaped cracks, and vertical dislocations for experiments.

[0104] Existing research on small-sample steel surface defect detection divides the NEU-DET dataset into a base class Cbase (RS, In, Sc) and a novel class Cnovel (Cr, Pa, Ps). This section uses the same task division format to divide the PVEL-AD dataset, selecting cracks, black cores, and broken grids as base classes, and star-shaped cracks, thick lines, and vertical dislocations as novel classes. 1,200 images are randomly selected as the training set and 300 images as the validation set. Figure 6 shown.

[0105] (2) Implementation details and evaluation indicators

[0106] The model's initial learning rate was 0.01, and a cosine learning rate decay strategy was used for updates with a decay exponent of 0.8. Stochastic gradient descent was used to update network parameters, with a weight decay of 0.0005 and a momentum of 0.937. The input image size was 640x640, the batch size was set to 16, and the number of training iterations was 300. An early stopping mechanism was introduced to prevent overfitting. Mosaic enhancement was used for image preprocessing.

[0107] To demonstrate the effectiveness of this method, we used four metrics: precision (P), recall (R), mean average precision (mAP), and F1 score to compare and verify the effectiveness of the model. The calculation formula is shown below.

[0108]

[0109] In the performance evaluation of classification models, the four basic indicators of the confusion matrix provide an important basis for model evaluation. Specifically, the True Positive (TP) indicator reflects the number of correctly identified positive samples, reflecting the model's ability to identify the target category. False Positive (FP) characterizes the number of negative samples that are mistakenly judged as positive samples. This indicator directly reflects the model's misjudgment rate. True Negative (TN) indicates the number of correctly classified negative samples, reflecting the model's recognition accuracy for non-target categories. False Negative (FN) records the number of positive samples that are mistakenly classified as negative samples. This indicator is closely related to the model's missed detection rate. Recall is the recall rate, which indicates the proportion of correctly identified samples among all positive samples. mAP refers to the average AP value of all categories of defects. AP refers to the area of the curve under the precision-recall curve. The formulas for AP and mAP are as follows. The higher the mAP value, the better the comprehensive detection performance of the model in all categories. The specific formulas are as follows:

[0110]

[0111] Where N is the number of categories.

[0112] (3) Comparative experiment and result analysis

[0113] In order to verify the surface defect detection performance of the method in this chapter under small sample conditions, the mAP value, PR value and F1 score were tested on two datasets under four different sample sizes using the small sample evaluation method to demonstrate the effectiveness of the SMFE-Net model. It was also compared with the latest small sample target detection methods IFDD, FSDD, MLDA-YOLOv7, Faster R-CNN-H-BFC, FSSDD and Meta-FSDet in recent years to verify the superiority of SMFE-Net.

[0114] Comparison method introduction:

[0115] IFDD: By pre-training the target task data, the model's generalization to new types of data is improved, and domain generalization and noise regularization strategies are proposed to improve detection accuracy.

[0116] FSDD: A small sample steel defect detection network based on the improved YOLOv3, replacing Darknet53 with Resnet34, and introducing same-scale fusion to enhance its feature extraction capability for deep categories.

[0117] MLDA-YOLOv7: This algorithm combines the meta-learning algorithm MAML with the YOLOv7 network, designs the SPP module to enhance the recognition of small target defects, introduces a regional adaptive optimizer to improve model generalization, and uses label smoothing technology to prevent model overfitting, thereby improving detection capabilities.

[0118] Faster R-CNN-H-BFC: It proposes a hallucination network based on multi-layer perceptron to learn shared features, replaces the VGG16 backbone network with ResNet50, introduces a feature pyramid network to achieve multi-scale feature fusion, and finally adds an attention mechanism to enhance the model's feature extraction capabilities.

[0119] FSSDD: Through multi-scale semantic enhancement representation and mask category information mapping, it effectively utilizes information to guide the multi-head detector and mask category representation module, achieving accurate detection of steel surface defects under conditions of very few labeled samples.

[0120] Meta-FSDet: By designing a representative prototype vector extractor, a cosine similarity-based saliency highlighting network, and a reconstruction candidate region network, it extracts and reconstructs effective defect features under limited labeled samples, achieving accurate detection and positioning of photovoltaic module defects.

[0121] Table 1 and Figures 7-10The comparison results of this model and six comparison models on the NEU-DET dataset are shown. The bolded parts indicate the best performance under the corresponding indicators. It can be seen that IFDD has the worst performance. The reason is that the pre-training mechanism it uses still requires a certain amount of data to support model training, which makes it difficult to adapt to defect detection tasks in scenarios with fewer samples. FSDD enhances feature extraction for deep categories by fusing same-scale features, thereby improving detection performance. MLDA-YOLOv7 enhances the feature extraction capability for small target defects, helping the model extract fine-grained features, but the overall improvement is not significant in small sample sizes because not all defects are small target defects. Faster R-CNN-H-BFC performs multi-scale feature fusion and uses attention to further enhance the model's feature extraction capabilities, significantly improving detection accuracy under different sample sizes compared to previous models. FSSDD utilizes multi-scale semantic enhancement and mask category information mapping to effectively guide the multi-head detector and mask category representation modules, achieving accurate detection of steel surface defects with minimal labeled samples. For example, the mAP values for k = 2 and k = 3 are 48.8% and 51.4%, respectively. This demonstrates that FSSDD has significant advantages in extracting deep categorical features and is better suited for defect detection tasks in very few-sample scenarios. Meta-FSDet achieves a mAP of 58.3% for k = 10, surpassing IFDD's 50.3% and FSDD's 52.0%. This indicates that through prototype vector extraction and a saliency highlighting network, Meta-FSDet effectively extracts and reconstructs defect features, thereby improving detection accuracy. Compared to the FSSDD model, which shows significant performance improvements, our SMFE-Net model is essentially on par with it in very few-sample conditions with k = 2 and k = 3. However, with further increases in the sample size, SMFE-Net shows a sustained improvement, reaching its optimal overall performance at k = 5. This shows that SMFE-Net has a stronger ability in extracting and utilizing key features from limited samples.

[0122] Table 1 Comparison of Model Performance on the NEU-DET Dataset

[0123]

[0124]

[0125] Table 2 and Figures 11-14Comparison results of our model with four comparison models on the PVEL-AD dataset are shown. The bolded results in the table indicate the best performance under the corresponding metric. Due to the high complexity of photovoltaic cell defect detection, the diverse defect types, and their uneven distribution, the overall detection performance of our model is slightly lower than that of the NEU-DET dataset. Our SMFE-Net model shows performance decreases of 1.9%, 1.7%, 1.6%, and 1.7% for sample sizes of 2, 3, 5, and 10, respectively. The smallest decrease is achieved with a sample size of 5, demonstrating that SMFE-Net is more suitable for detection tasks in small sample sizes. Furthermore, the minimal decrease in overall accuracy for sample sizes of 2, 3, and 10 demonstrates the excellent generalization performance of our model across diverse data distributions.

[0126] Table 2 Comparative experiments of models on PVEL-AD dataset

[0127]

[0128]

[0129] To further verify the six-category defect detection performance of SMFE-Net on the NEU-DET and EVPL-AD datasets, Figures 15-16 The detection effect of SMFE-Net and the comparison method under the condition of k=5 is demonstrated. The experimental results show that SMFE-Net has a significant advantage in multi-type defect detection. The experimental results on the NEU-DET dataset show that Figure 3 As shown in Figure 15, SMFE-Net demonstrates excellent performance in multiple detection scenarios. Specifically, in column (b), SMFE-Net successfully achieves complete detection of small inclusions with low contrast against the background, while IFDD and FSDD can only detect local areas due to the lack of a multi-scale feature fusion mechanism; MLDA-YOLOv7 misses detections due to insufficient prototype vector extraction. In column (c), SMFE-Net accurately fits the boundaries of irregular patches, while IFDD and FSDD, due to the limited receptive field of the backbone network, result in adjacent patches being mistakenly merged and detected. In addition, in column (e), SMFE-Net also demonstrates good detection results for densely distributed small target defects.

[0130] Experimental results on the EVPL-AD dataset. Figure 16The results further demonstrate the superiority of SMFE-Net. In column (a), SMFE-Net performs best in the linear crack detection task, accurately detecting thin, discontinuous cracks and maintaining high detection accuracy even when the crack-background contrast is low. Because broken grid and thick line detection rely on high-resolution features and background suppression capabilities, in columns (c) and (e), SMFE-Net's cross-layer feature enhancement strategy significantly outperforms the traditional attention mechanism, demonstrating superior detection results.

[0131] In summary, SMFE-Net performs well across all six defect detection tasks on the NEU-DET and EVPL-AD datasets. In particular, its detection accuracy significantly outperforms other comparison methods for slender defects, small target defects, and defects with low background contrast. This result demonstrates that SMFE-Net has a stronger ability to extract and utilize key features from limited samples, making it better suited for defect detection tasks in small sample sizes.

[0132] (4) Ablation experiment and result analysis

[0133] In order to evaluate the effectiveness of each module in the SMFE-Net model, this section conducts ablation experiments on the NEU-DET dataset and the EVPL-AD dataset, and uses the mAP value for comprehensive evaluation. The experimental parameter settings are the same as described in 3.3.4. A total of three benchmark models are designed, SMFE-Net-C, SMFE-Net-M, and SMFE-Net-S. The SMFE-Net-C model replaces CFEM with an ordinary convolutional layer while keeping other experimental indicators unchanged. The SMFE-Net-M model represents the SMFE-Net network without MFIM. The SMFE-Net-S model means that the SMIM module is removed, and the model only uses the feature information of this layer. The experimental results are shown in Tables 3 to 4 and Figures 17-24 shown.

[0134] Results from the NEU-DET dataset show that, for k = 2, 3, 5, and 10, removing the modules in this chapter leads to a decrease in the overall model's Precision, Recall, and F1 scores. SMFE-Net-C's performance significantly deteriorates, with mAP values dropping by 12.2%, 14.4%, 12%, and 12.7%, respectively. This demonstrates that CFEM plays a key role in extracting local, contextual, and contrast features, effectively constraining the network's feature learning and enhancing its ability to fully learn the characteristic representation of defect targets. SMFE-Net-M's mAP values drop by 8.9%, 10.5%, 8.3%, and 9.1%, respectively. This is due to the reduced use of long-range information, which results in the loss of some deep network features, leading to performance degradation. SMFE-Net-S's mAP values drop by 7.1%, 7.8%, 5.6%, and 6.5%, respectively. This is because the SMIF module contains multi-level features, which can simultaneously extract fine-grained defect features from shallow features and semantic information contained in deep features. This information helps the network directly identify clearer boundaries between background and target areas and provide better localization. SMFE-Net outperforms the ablation model in all k-shot settings, demonstrating the effectiveness of the CFEM, MFIM, and SMIM modules. Notably, when the number of training samples k>5, the performance curve of SMFE-Net shows a superlinear growth trend, indicating that this method has better feature expansion capabilities as the amount of data increases.

[0135] Table 3 Ablation study results on NEU-DET dataset

[0136]

[0137] Results from the EVPL-AD dataset show that the overall performance degradation of the ablation model is slightly greater than that of NEU-DET. This is due to the greater diversity and uneven distribution of defect types in the EVPL-AD dataset, but the overall trends are consistent. When the sample size is small (k = 2 and 3), the mAP of SMFE-Net-C decreases by 12.2% and 14.4%. This is because photovoltaic defects have high background similarity, and the model relies solely on single-scale features, resulting in a higher false detection rate for defects with similar backgrounds. When k = 3, the mAP of SMFE-Net-M decreases by 10.5%. This is because the removal of MFIM prevents the model from modeling long-range dependencies, resulting in a decrease in detection of continuous defects and overall performance. When k = 5 and 10, the mAP of SMFE-Net-S decreases by 5.6% and 6.5%, respectively. This indicates that a larger sample size can partially compensate for the model's feature representation ability, but the decrease is greater when the sample size is smaller (k = 2 and 3), demonstrating the more effective feature representation of the SMIM module under small sample sizes.

[0138] Table 4 Ablation study results on EVPL-AD dataset

[0139]

[0140] In summary, ablation experiments show that the CFEM, MFIM, and SMIM modules in the SMFE-Net model all improve overall model performance. The CFEM module enhances feature extraction capabilities, the MFIM module models long-range dependencies, and the SMIM module integrates multi-level feature fusion to improve model performance across a wide range of sample sizes. The SMIM module's effectiveness is particularly evident under small sample sizes.

[0141] (5) Feature layer visualization analysis

[0142] In order to better demonstrate the learning ability of SMFE-Net for defect features, this section visualizes the grad_cam feature maps of the first four layers of the six types of defect features of the two data sets, as shown in the following example: Figures 25-26 shown.

[0143] Figures 25-26 The visualization results of the SMFE-Net feature layer on the NEU-DET dataset and EVPL-AD dataset. Figure 24 From left to right, they are cracks, inclusions, spots, pits, rolling scales, and scratches. Figure 25From left to right, they are linear cracks, star-shaped cracks, broken grids, black cores, thick lines, and vertical error defects. From top to bottom, they are the original image and the visualization results of the feature maps of the first four layers of our SMFE-Net network. In the visualization analysis of model performance, by comparing the visual representation of the feature maps, the following conclusions can be drawn: First, the feature maps generated by our model show higher brightness values, larger coverage areas, and more uniform distribution characteristics in the defect area. This phenomenon fully verifies the model's superiority in defect feature extraction. Second, the noise in the background area of the feature map is significantly reduced. This result shows that the model has stronger background suppression capabilities. More importantly, our model can focus on specific areas where defects may occur, such as edges prone to scratches, while maintaining global uniformity. This feature significantly improves the accuracy and reliability of defect detection. The above shows that our SMFE-Net model can better aggregate defect features and achieve more accurate detection.

[0144] To address the challenges of insufficient sample size and difficult feature extraction in small-sample surface defect detection tasks, a small-sample surface defect detection method based on multi-level feature enhancement, SMFE-Net, was proposed. This method significantly improves the model's detection performance under limited sample conditions by designing three components: a contrast feature extraction module (CFEM), a multi-level feature interaction module (MFIM), and a sliding multi-level feature fusion module (SMIM). The first component, the contrast feature extraction module, combines fine-grained features, local context features, and Gaussian saliency features to address the difficulty in feature extraction caused by the irregularity and diversity of industrial surface defects, thereby enhancing the model's ability to express defect targets. The second component, the multi-level feature interaction module, captures the long-range dependencies of targets in complex backgrounds, alleviating the problem of feature loss in deep networks and further improving the model's detection accuracy. The third component, the sliding multi-level feature fusion module, fuses multi-scale features and enhances information interaction between features, ensuring the diversity of fused features and enabling the model to more accurately locate and identify defect targets. Experimental results demonstrate that SMFE-Net achieves optimal detection performance across multiple k-shot settings on both the NEU-DET and PVEL-AD datasets, significantly outperforming existing small-sample target detection methods. Furthermore, ablation experiments validate the effectiveness of the CFEM, MFIM, and SMIM modules; removing any of these modules significantly degrades model performance. Visualization analysis further demonstrates that SMFE-Net can more accurately capture defect features while also exhibiting stronger background suppression capabilities. In summary, the SMFE-Net approach proposed in this chapter demonstrates outstanding performance in small-sample surface defect detection tasks, providing an effective solution to the challenges of insufficient sample size and difficult feature extraction in industrial scenarios.

[0145] The above examples of the present invention are described in detail, but the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A small sample surface defect detection method based on multi-level feature enhancement, characterized by: Using deep learning technology to detect surface defects in small samples includes the following steps: S1: Constructing the method framework of contrast feature extraction module, multi-level feature interaction module and sliding multi-level feature fusion module; S2: A multi-branch parallel architecture is used in the contrast feature extraction module to achieve multi-level feature extraction of the target through three heterogeneous convolution operations, focusing on integrating the contextual features of the target object and the saliency features based on the Gaussian distribution; S3: Optimizes the extracted multi-scale local features in the multi-level feature interaction module, captures the long-range dependencies of targets in complex backgrounds from high-level features, and alleviates feature loss in deep networks; S4: In the sliding multi-level feature fusion module, features of different resolutions are combined as a supplement to the single-layer feature information, and the diversity of the fused features is ensured by fusing more feature layers. While retaining the original layer features, multi-scale features are supplemented, and information exchange is enhanced through channel shuffling. The supplementary features are composed of features from other layers in a nonlinear way. While retaining the original features, multi-scale features are added to enhance the information exchange between the supplementary features and the original features.

2. The small sample surface defect detection method based on multi-level feature enhancement according to claim 1 is characterized in that: In step S2, the contrast feature extraction module adopts a multi-branch parallel architecture and realizes multi-level feature extraction of the target through three heterogeneous convolution operations, focusing on integrating the contextual features of the target object and the saliency features based on Gaussian distribution to solve the feature extraction problem. The process includes the following steps: 1-1) In the feature extraction process, a standard 3×3 convolution kernel is first used to capture the fine-grained features of the target, including visual attributes such as shape, structure, and texture; 1-2) By introducing a dilated convolution operation with a dilation rate of 1, local context feature information is extracted, which not only expands the receptive field but also captures artifact features around the defect; 1-3) Finally, the module extracts the salient features of small-scale targets through Gaussian saliency convolution.

3. The small sample surface defect detection method based on multi-level feature enhancement according to claim 1 is characterized in that: Step S3: In the multi-level feature interaction module, the extracted multi-scale local features are optimized. This process includes the following steps: 2-1) Specifically, MFIM contains (k-1) ViT branches, all of which have the same structure, and uses CFEM as the feature embedding module to extract multi-scale local features. (i∈{1, 2, ..., k}), feature F i Flattened into a 2D block where N i =(H i W i / P i 2 ) is the number of blocks, (P i , P i ) is the resolution of each block, and after position embedding, we get Obtained Where n is the number of tokens, and Embedded token E (i) Divided into m heads in Then enter the multi-head self-attention module to obtain the interaction This article defines these processes as: Among them, LN represents layer normalization. In each head, the multi-head self-attention module MSA defines three trainable weight matrices to transform the query Q (i) , key K (i) Sum V (i) ,Then, is input into the MLP module to obtain the final The result of MLP can be expressed as: in 2-2) Next, Reintegrated into Feature V i , where P i =H k =W K , the result of the ViT branch can be expressed as:

4. The method for detecting small sample surface defects based on multi-level feature enhancement according to claim 1, characterized in that: The step S4 includes the following steps: 3-1) First, generate supplementary features using the features X generated by the backbone network in four different stages ori ={X ori |1≤i≤4} as the original feature input, and generate the supplementary features of the other three layers except the i-th layer through the adaptive weighting function. Secondly, the original features and the supplementary features of the i-th layer are shuffled from the channel perspective. Finally, the convolution layer is used to adjust the number of channels of the shuffled features to make them consistent with the number of channels of the original features of the i-th layer. The calculation formula of the SMIM module is as follows: in, is the convolutional feature extracted by the backbone network at layer i; represents the complementary features of the i-th layer; CAT represents the concatenation layer; CH_Shuffle represents channel shuffle; W1 represents a 1×1 convolutional layer; represents the i-th layer features enhanced by the SMIM module; 3-2) The specific generation process of the supplementary features is as follows: First, the up-sampling, down-sampling and 1×1 convolution operations are used to adjust the feature maps of different layers to the same size and number of channels as the i-th layer features; secondly, the adjusted three-layer features are spliced along the channel to form multi-scale original supplementary features for the i-th layer supplementary information; finally, the spatial correlation between features is captured through the channel attention mechanism to enhance the feature representation ability generated by the network. Specifically, the original supplementary features of the i-th layer are globally averaged pooled to obtain one-dimensional context features, and then two fully connected layers and sigmoid activation functions are used to obtain the adaptive fusion weights of the i-th layer original supplementary features. Finally, the weights are scaled through the normalization layer, and the features are reweighted and fused to obtain the final supplementary features of the i-th layer. 3-3) Supplementary features The calculation formula is as follows: in: represents the supplementary features of the i-th layer; Represents the i-th layer convolutional feature map extracted by the backbone network; F is a nonlinear function used to perform channel-level scaling on multi-level features; λ is a hyperparameter used to adjust the balance between original features and supplementary features; represents multiplication, ∑ represents addition; T is a function that adjusts the feature maps of different stages to the same size and number of channels as the feature maps of the i-th layer; is the original supplementary feature of the i-th layer generated by the original feature of the n-th layer; GAP represents the global average pooling operation; fc represents the fully connected layer; δ represents the ReLU activation function; σ represents the Sigmoid activation function; Softmax represents the normalized exponential function; The channel shuffling operation splices the original features and the supplementary features along the channel, and then shuffles the original features and the supplementary features on the channel to enhance the information flow between the two; that is, the original features and the supplementary features are staggered along the channel to achieve feature fusion.