Small sample lightweight flaw detection method based on improved YOLOv11n
By improving the YOLOv11n network, combining CycleGAN, MDSA, HS-FPNs, and using DWConv and HetConv for lightweight design, the problem of unstable detection and high computational volume in complex environments is solved, and efficient and real-time defect detection is achieved.
Patent Information
- Application Number
- CN202510420839.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing defect detection algorithms are unstable in complex and changeable industrial environments, which are difficult to meet the needs of high precision and high reliability, and the calculation volume is large, making it difficult to achieve efficient real-time detection on the production line.
The improved YOLOv11n network is adopted, combined with CycleGAN and traditional data augmentation technology to expand samples, introduce multi-scale dynamic collaborative attention (MDSA) and high-level screening feature pyramid networks (HS-FPNs), and use deep convolution (DWConv) and heterogeneous kernel convolution (HetConv) for lightweight design to optimize network structure.
On the premise of ensuring detection accuracy, the calculation amount and hardware dependence are significantly reduced, and efficient defect detection is achieved, with a reduced parameter amount of 49.5%, a reduced calculation amount of 96.0%, and a 3-fold increase in inference speed, which is suitable for real-time deployment of edge devices.
Smart Images

Figure CN120339227A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of clothing detection and deep learning networks, and specifically relates to a small-sample lightweight defect detection method based on an improved YOLOv11n network. Background Art
[0002] In the highly competitive manufacturing industry, defect detection, as a core link in quality control, directly affects the market competitiveness of products. Once defective products enter the market, it not only increases after-sales costs but also may damage the brand reputation. Therefore, improving the efficiency and accuracy of defect detection is crucial for the manufacturing industry.
[0003] Traditional manual detection methods rely on workers' experience, are inefficient, and are easily affected by subjective factors, making it difficult to meet the requirements of large-scale production. Detection methods based on template matching have high precision in specific scenarios, but their generalization ability is limited, making it difficult to adapt to newly emerging defect types, and the maintenance cost is relatively high. With the development of deep learning, convolutional neural networks have gradually been applied to defect detection, which can automatically extract image features and improve detection accuracy. However, classic CNN architectures still have problems with unstable detection in complex scenarios and are difficult to meet the high real-time requirements of industrial production.
[0004] In recent years, the YOLO series of algorithms have received wide attention in the field of object detection due to their fast detection ability and high accuracy. YOLOv11 further optimizes the backbone network and feature fusion method, improving the multi-scale feature extraction ability and detection robustness, and performing excellently in general object detection tasks. However, there is little research on applying existing YOLO to industrial defect detection, and this task faces many challenges, such as diverse defects and unbalanced sample types, variable clothing materials and colors, and variable lighting conditions, resulting in the detection effect of existing methods still being less than satisfactory.
[0005] At the same time, there are also problems in obtaining defect samples. On the one hand, it is difficult to meet the demand for the sample size of defect detection, and the number of samples is relatively scarce; on the other hand, the types of defects are rich and diverse, including various categories such as breakage, deformation, and color abnormality, and the sample distribution of various defects shows a serious imbalance. This further has a negative impact on the accuracy and generalization ability of detection. The limitations of samples also further exacerbate the difficulties faced by the algorithm in practical applications.
[0006] From the algorithm level, the current challenges are even more severe. The target is placed at different angles in the image, and the complex texture on the object surface will also interfere with the detection. Coupled with the different sizes and complex types of defects, this poses higher requirements for the recognition accuracy, maintainability, and reliability of the algorithm. Existing algorithms are difficult to stably extract accurate defect features under complex and changeable conditions, resulting in unstable detection results and unable to meet the requirements of high precision and high reliability in actual production. These algorithmic problems not only affect the accuracy and stability of detection but also pose a huge challenge to the detection efficiency in actual application scenarios. And in actual production lines, detection efficiency is a key factor determining production scale and cost. Due to the higher requirements for the lightweight and real-time nature of the algorithm due to commercialization needs, the production line needs to run the detection algorithm under online and high-speed conditions to meet the rhythm of large-scale production. However, the existing deep learning Yolo algorithm often has a large computational amount and a relatively complex model, making it difficult to achieve fast inference while ensuring detection accuracy and unable to meet the strict requirements of the production line for detection speed. This makes it difficult to achieve efficient real-time detection in actual applications and limits the large-scale application of deep learning algorithms on the production line. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a small-sample lightweight defect detection method based on an improved YOLOv11n network, for fast and effective detection of defects based on the training and lightweight deployment of small-sample clothing images.
[0008] To solve the above technical problems, the present invention provides a small-sample lightweight defect detection method based on an improved YOLOv11n network, including the process of collecting clothing images with trademark parts, preprocessing the collected images and then inputting them into the offline-trained CZ-YOLO network to generate trademark defect prediction candidate boxes and confidence levels.
[0009] The CZ-YOLO network is improved based on the Yolov11n network. The backbone of the CZ-YOLO network includes a series of Conv modules, a structure of two-layer DWConv modules + Het-C3k2 module, an MDSA module, a structure of two-layer DWConv modules + Het-C3k2 module, an SPPF module, and a C2PSA module, for outputting three different-scale features S2, feature S3, and feature S4; the neck of the CZ-YOLO network includes an HS-FPNs network and a Het-C3k2 module. After the features S2, S3, and S4 pass through the HS-FPNs network, the output features N2, N3, and N4 are each processed by a Het-C3k2 module and then used as the inputs of the three detection heads respectively; the detection head of the CZ-YOLO network uses the detection head of the Yolov11n network.
[0010] Improvement of the small-sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0011] The Het-C3k2 module is specifically:
[0012] The Het-C3k2 module is an improvement of the C3k2 module. First, replace the ordinary convolution in the Bottleneck of the C3k2 module in the original Yolov11n network with heterogeneous kernel convolution, and remove the original skip connection, thereby constructing HetBottleneck. Then, replace the Bottleneck in C3k of the original Yolov11n network with HetBottleneck to obtain HetC3k;
[0013] Then, keep the structure of the C3k2 module in the original Yolov11n network unchanged when C3k = False as the structure of the Het-C3k2 module when C3k = False; when C3k = True, replace C3k in the C3k2 module of the original Yolov11n network with HetC3k as the structure of the Het-Cek2 module when C3k = True.
[0014] Further improvement of the small-sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0015] The MDSA module includes multi-scale channel attention, multi-scale spatial attention, and a dynamic collaboration layer. The output feature F generates the feature CFF through multi-scale channel attention. The feature CFF obtains the spatial attention weight W through multi-scale spatial attention s , and the feature CFF obtains the channel attention weight W through the dynamic collaboration layer c , then use two learnable dynamic weights for the spatial attention weight W s and the channel attention weight W c to perform multiplication and then weighted fusion, and then calculate the overall attention weight W through Sigmoid activation and act on the input feature CFF as the output of the MDSA module.
[0016] Further improvement of the small-sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0017] The multi-scale channel attention includes two stages: feature extraction and channel feature fusion. In the feature extraction stage, the input feature F obtains the global feature F GAP and the global feature F GMP; The input feature F is respectively passed through local average pooling + global average pooling and local max pooling + global max pooling to obtain local feature F LAP+GAP and local feature F LMP+GMP ;
[0018] The channel feature fusion stage is as follows:
[0019] F1 = F GAP + F LAP+GAP (1)
[0020] F2 = F GNP + F LMP+GMP (2)
[0021] Then, feature F1 and feature F2 are concatenated along the channel dimension and passed through a channel weight calculation gate to generate channel attention weights, and the generated channel attention weights are applied to the input feature F to obtain feature CFF;
[0022] The channel weight calculation gate is composed of a fully connected layer, a ReLU activation function, a fully connected layer, a convolutional layer, and a sigmoid activation function connected in series.
[0023] As a further improvement of the small sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0024] The multi-scale spatial attention is as follows: After the feature CFF is respectively subjected to global average pooling and global max pooling along the channel, they are concatenated in the channel dimension, and then a convolutional layer is used to obtain the spatial attention weight W s .
[0025] As a further improvement of the small sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0026] The dynamic collaboration layer is as follows: The feature CFF is subjected to feature extraction through depthwise separable convolution, and global average pooling is performed along the channel to obtain the channel attention weight W c .
[0027] As a further improvement of the small sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0028] The HS-FPNs are as follows: The features S2, S3, and S4 output by the backbone respectively pass through the CA module and the convolutional module to obtain features P2, P3, and P4. While feature P4 is directly output as feature N4, feature N4 and feature P3 pass through the improved SFF module to obtain feature N3, and feature N4 is deconvolved and then passes through the improved SFF module with feature P2 to obtain feature N2.
[0029] As a further improvement of the small-sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0030] The improved SFF module is as follows: For the input high-level feature f high Perform a transposed convolution operation to obtain the upsampled high-level feature f high_up , and for the high-level feature f high_up Obtain the channel weight ω through the CA module, and then use the channel weight ω to process the input low-level feature f low Complete channel-by-channel weighted screening to obtain the attentively screened low-level feature Then, fuse the upsampled high-level feature f high_up with the attentively screened low-level feature as the output feature of the improved SFF module.
[0031] As a further improvement of the small-sample lightweight defect detection method based on the improved YOLOv11n network of the present invention:
[0032] The offline training of the CZ-YOLO network is small-sample offline training:
[0033] (1) Dataset collection: Collect clothing pictures with trademarks, where the number is no more than 80; in the computer, select trademark defect pictures from the collected pictures and annotate them with LabelImg. Pictures with trademark defects are labeled as 'NG', and pictures of the L trademark are labeled as 'L' as the original dataset;
[0034] (2) Dataset augmentation: Use the CycleGAN network and traditional data augmentation methods to augment the number of pictures in the original dataset respectively, and then randomly divide the augmented images and the original images into a training set, a validation set, and a test set according to the ratio of 8:1:1;
[0035] (3) Input the labeled pictures in the training set into the CZ-YOLO network for training, calculate the loss function value and perform backpropagation to iteratively optimize the model parameters. End the training after reaching the preset number of epochs; put the labeled test set pictures into the trained model, generate candidate prediction boxes, and calculate the intersection over union and performance metrics of the candidate prediction boxes and the labeled boxes to evaluate the model.
[0036] The beneficial effects of the present invention are mainly reflected in:
[0037] 1. During the training stage of the network of the present invention, by introducing the CycleGAN generative adversarial network, the sample diversity is augmented through unsupervised learning. At the same time, by rotating, introducing noise, changing brightness and contrast to simulate trademarks under different angles and lighting conditions, the data complexity is enhanced, and the model's ability to adapt to different image qualities is improved;
[0038] 2. The MDSA attention mechanism proposed in the present invention can adaptively focus on the key areas of trademark images and enhance the ability to extract defect features. In the face of interference such as the reflection of clothing materials and changes in illumination, this mechanism uses dual attention in space and channels for dynamic calculation, accurately captures defect features, optimizes the model's detection ability and the enhancement effect of robustness to complex backgrounds. It demonstrates the balance characteristics in the three dimensions of accuracy-speed-complexity in industrial inspection scenarios. While maintaining the detection accuracy comparable to mainstream algorithms, it achieves a significant performance improvement through the trinity optimization strategy of "lightweight design-feature enhancement-efficient inference", which is superior to YOLOv11 and YOLOv8.
[0039] 3. The present invention optimizes the network structure to reduce the model complexity and optimizes the Neck structure with the HS-FPNs module. This module can fuse different scale information of the image, ensuring the model accuracy while significantly reducing the computational amount. At the same time, DWconv is used to replace some traditional convolutions, maintaining the feature extraction ability while reducing the computational amount. In addition, HetConv is introduced to improve the C3k2 module, enhancing the model's expressive ability while reducing the computational amount and parameters.
[0040] The verification results of the lightweight design show that the CZ-YOLO network has achieved a key performance breakthrough on the premise of ensuring the detection accuracy: the number of parameters is only 1.302×106, a 49.5% reduction compared to YOLOv11 and a 96.0% reduction compared to RTDETR; the computational amount is reduced from 6.3 GFLOPs of YOLOv11 to 4.5 GFLOPs. At the same time, the inference speed is 3 times that of YOLOv11 and 2.58 times that of YOLOv8 respectively. This lightweight design lays a solid foundation for the real-time deployment of the model on edge devices, significantly reducing the hardware dependence while maintaining the detection ability, and providing an innovative solution for the high-precision, low-power, and fast-response requirements in the industrial inspection field. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The following further details the specific embodiments of the present invention with reference to the accompanying drawings.
[0042] Figure 1 It is a schematic structural diagram of the CZ-YOLO network of the present invention;
[0043] Figure 2 It is a schematic structural diagram of the CycleGAN network;
[0044] Figure 3 It is a schematic structural diagram of the multi-scale channel attention of the present invention;
[0045] Figure 4 It is a schematic structural diagram of the multi-scale spatial attention and dynamic collaboration layer of the present invention;
[0046] Figure 5 Schematic diagram of the improved high-level screening feature pyramid network of the present invention;
[0047] Figure 6 Schematic diagram of the improved selective feature fusion of the present invention;
[0048] Figure 7 Network structure diagram of depth convolution;
[0049] Figure 8 Schematic diagram of the structural comparison between heterogeneous kernel convolution and standard convolution;
[0050] Figure 9 Delay comparison chart of three filters: depth convolution + pointwise convolution, grouped convolution + pointwise convolution, and heterogeneous kernel convolution;
[0051] Figure 10 Example of defective pictures in the dataset of the present invention;
[0052] Figure 11 Diagram showing the detection effect of the present invention. Detailed implementation manners
[0053] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto:
[0054] Embodiment 1, A small-sample lightweight defect detection method based on the improved YOLOv11n network.
[0055] In order to solve problems such as insufficient defective samples, unstable performance in complex industrial environments, and low detection efficiency, the present invention constructs a defect detection network based on YOLOv11n: the CZ-YOLO network. Then, a small number of defective images at the production site are collected. While introducing the CycleGAN network to expand the dataset of defective images, the defective images are processed and expanded in combination with traditional image enhancement techniques. Then, an online usable CZ-YOLO network is obtained through offline training for online detection of defects at the production site. The specific process is as follows:
[0056] 1. Construct the CZ-YOLO network
[0057] 1.1. Multi-scale dynamic collaborative attention (MDSA)
[0058] In the defect detection task, efficient feature extraction and enhancement are crucial for improving detection performance. To this end, the present invention proposes a new attention mechanism based on the design idea of CBAM: multi-scale dynamic collaborative attention (MDSA). MDSA mainly consists of multi-scale channel attention (MSCA), multi-scale spatial attention (MSSA), and dynamic collaborative layer (DSL). It inherits the spatial attention module of CBAM and the idea of cascaded calculation of channel attention combined with spatial attention. Its biggest feature is that it can capture key information at different scales and adaptively adjust the weight ratio of channel and spatial attention to achieve more accurate object detection.
[0059] In the channel attention mechanism, due to the differences in the importance of feature channels at different scales for the detection task, MDSA adopts a combination of global pooling and local pooling to fully extract multi-scale feature information. The MDSA module can capture local defect details while retaining sufficient context information to distinguish complex background interference, and will not significantly increase the computational amount. MDSA enhances defect feature extraction through three stages: First, the multi-scale channel attention (MSCA) captures global semantics and local detail features. Then, the feature CFF (Channel-refined feature) generated by MSCA will generate a spatial weight map through pooling and convolution in the multi-scale spatial attention (MSSA), focusing on the shape and position changes of the defect area. Finally, the spatial weight map is adaptively adjusted by the dynamic collaborative layer (DSL) through two dynamic weights (Dynamic weight) to adjust the feature CFF (Channel-refined feature) generated by MSCA after passing through channel attention and the spatial attention weight ratio generated by MSSA after extracting features through depthwise separable convolution (DSConv), improving the model's ability to extract trademark features.
[0060] The multi-scale channel attention (MSCA) includes two stages: feature extraction and channel feature fusion. In the feature extraction stage, the input feature F obtains channel statistical information through four different scale pooling methods: global average pooling (GAP), local average pooling (LAP), global max pooling (GMP), and local max pooling (LMP): using global average pooling (GAP) to extract global channel information to obtain the global feature F GAP , using global max pooling (GMP) to extract global channel information to obtain the global feature F GMP ; using local average pooling (LAP) + global average pooling (GAP) to obtain the local feature F LAP+GAP , using local max pooling (LMP) + global max pooling (GMP) to obtain the local feature F LMP+GMP。The combination of global pooling and local pooling can focus on the feature distribution of local regions while obtaining the overall target features, thereby enhancing the perception ability of detailed information.
[0061] In the channel feature fusion stage, addition operations are respectively performed on the two global features and two local features obtained in the feature extraction stage.
[0062] F1 = F GAP + F LAP+GAP (1)
[0063] F2 = F GMP + F LMP+GMP (2)
[0064] Then, the features F1 and F2 after the addition operation are concatenated (CAT) along the channel dimension and then passed through a Channel Weight Gate (CWG) for information interaction between channels to generate channel attention weights. The generated channel attention weights are applied to the input feature F to obtain the feature CFF (Channel-refined feature) after the action of multi-scale channel attention (MSCA). The Channel Weight Gate (CWG) is composed of a fully connected layer, a ReLU activation function, a fully connected layer, a convolutional layer, and a sigmoid activation function in series. The MSCA network structure is as Figure 3 shown. This mechanism can effectively improve the model's attention to the defective area and strengthen the representation ability of the defective area at different scales.
[0065] Multi-scale spatial attention (MSSA) is used to enhance the model's focusing ability on the target area. Especially when dealing with the shape and position changes of the defective area, it can effectively enhance the extraction of key features. In this part, the spatial attention module in CBAM is adopted. First, after performing global average pooling (GAP) and global max pooling (GMP) on the feature CFF output by MSCA along the channels respectively, they are concatenated (CAT) in the channel dimension to construct the multi-scale spatial feature F s denoted. Then, convolution (Conv) is used to extract information from the multi-scale spatial feature F s to obtain the spatial attention weight W s .
[0066] To further enhance the representation ability of the attention mechanism, MDSA designed a Dynamic Synergy Layer (DSL) to adaptively adjust the weight ratio of channel attention and spatial attention. Specifically, first, the depthwise separable convolution (DSConv) is used to perform re-feature extraction on the Channel-refined feature (CFF), and global average pooling (GAP) is performed along the channels to obtain the channel attention weight W of the re-featured extraction. c . Subsequently, combined with the spatial attention weight W s , two learnable dynamic weights are used to perform multiplicative weighted fusion on the channels to obtain the overall attention weight W, and then through the Sigmoid activation calculation and acting on the input feature CFF to achieve precise attention to the target area. The MSSA+DSL network structure is as Figure 4 shown.
[0067] The depthwise separable convolution is used because it can effectively separate spatial information and channel information. Its independent channel convolution process helps to enhance the spatial feature extraction ability, and at the same time has the characteristics of light weight, reducing the computational complexity, so as to achieve efficient attention modeling.
[0068] 1.2. Improved High-Level Screening Feature Pyramid Network (HS-FPNs)
[0069] To solve the problem of multi-scale feature processing in defect detection, based on the High-Level Screening Feature Pyramid Network (HS-FPN), this invention modifies the network structure and the SFF module structure to propose an improved High-Level Screening Feature Pyramid Network (HS-FPNs). By using the feature screening and cross-scale fusion mechanism, it significantly enhances the multi-scale feature representation ability for defects of different sizes. Through the dual-channel attention mechanism, it dynamically adjusts the weight distribution of features at different scales, reducing FLOPs while ensuring the detection stability in scenarios with large industrial size differences.
[0070] The network structure of the improved High-Level Screening Feature Pyramid Network (HS-FPNs) is as Figure 5 shown. Specifically:
[0071] The features S2 output by the second layer Het-C3k2 of the backbone, the features S3 output by the third layer Het-C3k2, and the features S4 output by the C2PSA each pass through the CA module and the 1×1 convolution module to obtain the features P2, P3, and P4. While the feature P4 is directly output as the feature N4 in the feature fusion stage, the feature N4 and the feature P3 pass through the improved SFF module to obtain the output of the feature N3 in the feature fusion stage. The feature N4 is deconvolved and then passes through the improved SFF module with the feature P2 to obtain the output of the feature N2 in the feature fusion stage.
[0072] Subsequently, features N2, N3, and N4 are respectively processed by the Het-C3k2 module and used as the input of the detection head Head.
[0073] The HS-FPNs network introduces a channel attention mechanism (CA module) in the feature selection module. Through the combined operation of global average pooling and global max pooling, a feature mask with channel weights is generated. This mask can adaptively filter the effective information in low-level features and suppress redundant noise. In the dimension matching module, 1×1 convolution is used to unify features of different scales to 256 channels. In this module, the original HS-FPN used the method of adding the data after being processed by the CA module to the data before being processed by the CA module. The present invention removes the addition step after passing through the CA module in the original network structure, simplifies this module, and is used to reduce the computational amount and meet the requirements of the Yolov11n network.
[0074] In the feature fusion stage, the HS-FPNs improves the SFF module in the original HS-FPN. The improved SFF module network structure is as Figure 6 shown. Specifically, first, the high-level feature f high is upsampled to the same spatial dimension as the low-level feature f low through transposed convolution. Subsequently, the high-level feature f high_up is converted into attention weights through the CA module, and the low-level feature f low is weighted and screened channel by channel. In the original HS-FPN, the SFF module adjusts the scale through transposed convolution and bilinear interpolation operations to match the size of the low-level feature. The present invention removes its bilinear interpolation operation and adds a deconvolution (ConvTranspose) module outside the improved SFF module of the present invention to perform size matching of the semantic information of the high-level feature and the detailed information of the low-scale feature. This operation is shown in Figure 5 . This operation optimizes the SFF module structure and reduces the model's computational amount while ensuring the model's accuracy. Finally, the screened low-level feature and the high-level feature are added element by element to form a fused feature (f out ). This fusion strategy effectively retains the detailed information of the low-level feature and injects the semantic information of the high-level feature, enabling the model to capture the structural features of small-sized defects and the context information of large-sized defects. The specific principles of each module will be introduced below:
[0075] The feature selection module filters the multi-scale features (S2, S3, S4) output by the backbone network based on the channel attention (CA) mechanism, as shown in the feature selection module in Figure 5 :
[0076] The multi-scale features (S2, S3, S4) output by the backbone network are respectively input into the CA module. The CA module extracts the statistical information in the channel dimension through global average pooling (Avg Pool) and global max pooling (Max Pool), and the formulas are shown as (3) and (4):
[0077]
[0078] f max = max i,j S i (i,j,:) (4)
[0079] where S i are the multi-scale features (S2, S3, S4).
[0080] After concatenating the results of the two poolings, a channel attention weight ω is generated through a fully connected layer and a Sigmoid activation function, and the formula is as follows:
[0081] ω = σ(MLP([f avg ; f max )) (5)
[0082] The input features are weighted in the channel using the weight to obtain the filtered features:
[0083] f ca = Si ⊙ ω (6)
[0084] Finally, the number of channels is adjusted through a 1*1 convolution, and the output feature selection results (P2, P3, P4) are obtained, that is, equation (7):
[0085] P i = Conv 1×1 (f ca ) (7)
[0086] The improved selective feature fusion (SFF) module focuses on the efficient fusion of high-level and low-level features to enhance the expression ability of defect features, as Figure 6 shown:
[0087] Upsampling of high-level features: First, a transposed convolution (ConvTranspose) operation is performed on the high-level feature f high to adjust its spatial size to match the low-level feature f low , as shown in (8):
[0088] f high_up = ConvTranspose(f high ) (8)
[0089] Attention screening of low-level features: The high-level feature f high_upInput Channel Attention (CA) module. The CA module extracts channel statistics through global average pooling and global maximum pooling, generates channel weights ω through a fully connected layer and Sigmoid activation, and then uses the channel weights ω to perform channel-weighted screening on the input low-level feature f low Complete the channel-weighted screening of the low-level feature:
[0090] f low_ca = f low ⊙ ω (9)
[0091] Feature fusion output: Fuse the upsampled high-level feature f high_up with the low-level feature after attention screening to finally output the fused feature f out :
[0092] f out = f low_ca + f high_up (10)
[0093] HS-FPNs demonstrate multiple technical advantages in defect detection: In terms of attention, through the dual-channel attention mechanism and selective feature fusion strategy, it dynamically adjusts the weight distribution of features at different scales, effectively handles size differences, and enhances the response of tiny defect features. Moreover, its channel attention module can suppress background interference and improve the feature expression ability in complex scenarios; Not only that, its lightweight design, through channel dimension matching and dynamic feature routing, significantly reduces computational redundancy while maintaining detection accuracy, achieving a balance between high-precision and high-efficiency defect detection.
[0094] 1.3, Lightweight Network
[0095] In an actual production line scenario, real-time performance is a crucial metric. To further improve the practicality and detection efficiency of the model, the present invention uses two lightweight convolutional modules (Depth-Wise Convolution (DWConv) and Heterogeneous Kernel Convolution (HetConv)) to reduce the model parameters, and adopts the DWConv module to replace the convolutional (Conv) modules except the first one in the backbone. The Het-C3k2 module of the present invention is constructed by Heterogeneous Kernel Convolution (HetConv) to replace the original C3k2 module. Specifically, the Het-C3k2 module first replaces the ordinary convolution (Conv) in the Bottleneck of the original C3k2 module with Heterogeneous Kernel Convolution (HetConv) and removes the original skip connection, thereby constructing the HetBottleneck of the present invention. The Bottleneck in the original C3k is replaced by HetBottleneck to obtain HetC3k. Then, when C3k = False, the structure of the original C3k2 module remains unchanged and serves as the structure of the Het-C3k2 module of the present invention when C3k = False. When C3k = True, the C3k in the original C3k2 module is replaced by HetC3k of the present invention, which serves as the structure of the Het-Cek2 module of the present invention when C3k = True.
[0096] In the field of defect detection, the advantages of DWConv and HetConv complement each other. By retaining the independence between channels, DWConv focuses computing resources on the spatial dimension and can deeply explore local spatial patterns in the input feature map. HetConv excels in multi-scale feature fusion. Its 3×3 kernel has strong local perception ability and can accurately capture local texture details at the edges. The 1×1 kernel is good at integrating global semantic information such as overall offset and splicing misalignment of trademarks. This multi-scale feature fusion mechanism enables HetConv to effectively identify composite defects and further enhances the ability to recognize complex defects. The comparison between HetConv and ordinary convolution is as Figure 8 shown.
[0097] Depth-Wise Convolution (DWConv) achieves efficient feature extraction through a clever channel separation strategy. Its specific structure is as Figure 7 shown. It applies a 3x3 convolutional kernel to each input channel independently for convolution operation, and the number of output channels is the same as the number of input channels. This design decomposes the originally complex cross-channel convolution into multiple independent single-channel convolutions, greatly reducing unnecessary computations.
[0098] From the perspective of computational complexity, the computational complexity of DWConv can be expressed by formula (11):
[0099] FLOPsDW =D o ×D o ×M×K×K (11)
[0100] Where Do represents the size of the output feature map, M is the number of input channels, and K is the size of the convolution kernel (here K = 3). The calculation formula of the traditional convolution is shown in (12):
[0101] FLOPs Conv =D o ×D o ×M×N×K 2 (12)
[0102] By comparison, it can be clearly found that the amount of calculation of DWConv is only 1 / N of that of traditional convolution (N is the number of output channels). For example, when the number of input channels M = 64, the number of output channels N = 128, the output feature map size Do = 32, and the convolution kernel size K = 3, the amount of calculation of traditional convolution is 32x32x64x128x3x3 floating-point operations, while the amount of calculation of DWConv is only 32x32x64x128x3 floating-point operations, which is a huge difference.
[0103] DWConv's design of preserving the independence between channels has many significant advantages. It focuses computing resources on the spatial dimension, allowing the convolution kernel to more deeply explore the local spatial patterns in the input feature map. In the field of image processing, these local spatial patterns cover a wealth of information, such as texture and edges. Texture information can reflect the microstructural characteristics of the image surface; while edge information can outline the contours of the object, helping the model to accurately locate and distinguish different objects. Traditional convolution operations may not be able to accurately capture these subtle defect information due to the limitations of their calculation methods. DWConv, with its efficient ability to capture local spatial patterns, can keenly identify these subtle defects, thereby improving the accuracy and reliability of defect detection.
[0104] HetConv adopts a unique dynamic kernel combination strategy. Each convolution kernel combines sub-kernels of different sizes, and the most common one is a combination of 3×3 and 1×1 kernels. The parameter P is used to flexibly control the proportion of kernels of different sizes. For example, when P=4, it means that 25% of the kernels are 3×3 and 75% of the kernels are 1×1. This dynamic kernel combination breaks the limitation of the singleness of the traditional convolution kernel, allowing the convolution operation to adaptively adjust the kernel usage ratio according to different task requirements and data characteristics, thereby achieving a better balance between computational efficiency and feature extraction capabilities.
[0105] Looking at the computational complexity, due to its fixed computational pattern, the computational complexity of traditional convolution is relatively high. The computational complexity formula of HetConv is (13):
[0106]
[0107] Among them, Do represents the size of the output feature map, M is the number of input channels, N is the number of output channels, and K is the size of the convolution kernel. When K = 3 and P = 4, the computational complexity of HetConv is only 31% of that of traditional convolution, which means it can achieve a 3X to 8X reduction in FLOPs. This significant reduction in computational complexity is of great significance in the actual model training and inference processes, which can significantly reduce the consumption of computational resources and improve the running speed of the model.
[0108] Figure 9 clearly shows the latency comparison between the HetConv (heterogeneous kernel convolution) filter and other efficient convolution filters. DWC+PWC (Depthwise Convolution+Pointwise Convolution) and GWC+PWC (Groupwise Convolution+Pointwise Convolution) adopt a two-level cascaded structure, and this serial processing method will result in a latency of 1. HetConv cleverly avoids the latency caused by serial processing by integrating heterogeneous kernels into a single layer. This zero-latency feature is very crucial in practical applications, especially in scenarios with high real-time requirements, which can ensure that the model quickly gives detection results.
[0109] 1.4, CZ-YOLO Network
[0110] Combining Steps 1.1 - 1.3, the CZ-YOLO network of the present invention is constructed based on the Yolov11n network, and the structure is as Figure 1 shown, including a backbone, a neck, and a detection head.
[0111] The backbone of the Yolov11n network includes a Conv module, a structure of four consecutive Conv modules + C3k2 module, an SPPF module, and a C2PSA module connected in sequence. First, the present invention lightweightens the backbone of the Yolov11n network. The Conv modules in the structure of four Conv modules + C3k2 module are all replaced with DWConv modules, and the C3k2 module is replaced with the Het-C3k2 module of the present invention (i.e., the structure of four Conv modules + C3k2 module is replaced with the structure of DWConv module + Het-C3k2 module). Then, the multi-scale dynamic collaborative attention (MDSA) module of the present invention is added between the second and third layer structures. The backbone of the CZ-YOLO network outputs three feature maps of different scales, namely feature S2, feature S2, and feature S4, as the input to the neck. The neck of the CZ-YOLO network uses an improved high-level screening feature pyramid network (HS-FPNs) to replace the traditional neck data processing part in the Yolov11n network except for the C3k2 module. After passing through the improved high-level screening feature pyramid network (HS-FPNs), feature S2, feature S2, and feature S4 output feature N2, feature N3, and feature N4. Then, feature N2, feature N3, and feature N4 are each processed by a Het-C3k2 module, thus forming a complete neck structure. The detection head of the CZ-YOLO network uses the original detection head of the Yolov11n network.
[0112] 2. Model Training
[0113] 2.1 Data Acquisition
[0114] The dataset is collected from a fabric production factory in Zhejiang Province. The dataset images are all taken by an industrial CCD camera. To achieve real-time online detection, the picture acquisition pixel is set to 1440×1080, and the collected clothing pictures all contain trademarks. In the computer, from the collected clothing pictures, pictures with trademark defects are selected and labeled by LabelImg to make the dataset. Among them, pictures with trademark defects are labeled as 'NG', and pictures of L trademarks are labeled as 'L'. The dataset pictures are as Figure 10 shown.
[0115] 2.2 Dataset Expansion
[0116] In view of the relatively scarce number of defect detection samples and the distribution of the number of extreme defect types, the present invention combines CycleGAN and traditional data augmentation techniques to expand the dataset. First, CycleGAN is used to expand the original data collected in step 2.1 by 3 times, and then traditional data augmentation is used to make the original data collected in step 2.1 reach 15 times the initial number in the case of a small sample dataset with only 20 clothing trademark defect pictures in the original data.
[0117] (1) CycleGAN Network
[0118] As a breakthrough generative adversarial network in the field of image style transfer, the core innovation of CycleGAN lies in achieving cross-domain style conversion through training with unpaired data. Different from traditional GAN models (such as pix2pix) that rely on strictly paired training samples, CycleGAN constructs a dual generator-discriminator architecture by introducing cycle-consistency loss, realizing bidirectional domain mapping under unsupervised conditions. Its overall structure is as shown in Figure 2 . The CycleGAN model uses the first generator G(x→y) to convert the input X-domain image into a Y-domain image; then, through the second generator F(y→x), the generated Y-domain image is converted back into an X-domain image. For Y-domain images, the model performs a similar process. First, the input Y-domain image is converted into an X-domain image through the generator F(y→x), and then the generated X-domain image is converted into a Y-domain image through the generator G(x→y).
[0119] The input image undergoes two conversions and is then compared with the original input image to obtain the cycle-consistency loss. The calculation method of the CycleGAN cycle-consistency loss function is shown in Equation (14):
[0120]
[0121] The adversarial loss consists of two parts corresponding to two independent GAN models, and the formulas are shown in (15) and (16):
[0122]
[0123] The total loss calculation formula for the entire cyclic GAN network is:[[]]
[0124] L(G,F,D x ,D y )=L GAN (G,D y ,X,Y)+L GAN (F,D x ,Y,X+λL c (G,F) (17)
[0125] In the formula: X and Y represent data domains X and Y respectively; x and y represent sample data in the two data domains; G represents the mapping function from the X domain to the Y domain; F represents the mapping function from the Y domain to the X domain; Dx and Dy are discriminant networks; λ is the weight value controlling the cycle-consistency network.
[0126] The goal of the cycle consistency loss function is to make the input image as similar as possible to the images generated twice, so as to ensure a certain consistency between the original image and the transformed image, making the algorithm not rely on a dataset with strict correspondence relationships.
[0127] (2) Traditional image enhancement techniques
[0128] Traditional data augmentation techniques effectively improve data diversity through multi-dimensional image transformation strategies. Specifically, random rotation within a certain range is used to simulate the multi-angle placement of defects, Gaussian noise is superimposed to simulate acquisition interference, and the brightness coefficient and contrast coefficient are dynamically adjusted to cover complex lighting conditions. These transformations are combined with a certain probability to generate enhanced samples with composite features such as wrinkles, occlusions, and slight deformations. These methods construct a virtual sample library containing industrial scene features such as angular deviation, lighting changes, and noise interference, making the training data distribution closer to the actual production environment, significantly reducing the overfitting risk of the model to specific shooting conditions, providing more challenging feature learning materials for deep models, and effectively improving the robustness and generalization ability of detection algorithms in complex industrial scenarios.
[0129] (3) Construction of the training set and test set
[0130] All the images in the dataset obtained by expanding CycleGAN, the dataset obtained by traditional data augmentation, and the original dataset are randomly divided into a training set, a validation set, and a test set in the ratio of 8:1:1.
[0131] 2.3 Offline training and testing process
[0132] The preset learning rate for offline training is 0.01, the number of workers is set to 4, the batch size is set to 16, and the number of epochs is set to 150. The labeled images in the training set are input into the CZ-YOLO network for training. During the training process, the original loss function of the yolov11n network is used to calculate the loss function value and backpropagate to iteratively optimize the model parameters. Training ends after reaching the preset number of epochs;
[0133] In the testing process, the labeled test set images are put into the trained model to generate candidate prediction boxes. The intersection over union (IoU) and performance metrics between the candidate prediction boxes and the labeled boxes are calculated to evaluate the model. When the IoU is greater than 0.5, it is a positive sample; when it is less than 0.5, it is a negative sample. The mean Average Precision (mAP) is used as the evaluation metric, and mAP0.5 = 98.0 is obtained, thus obtaining a CZ-YOLO network that can be used online.
[0134] 3. Online use
[0135] Use an industrial CCD camera to collect clothing images with trademark parts. Preprocess the collected images, and then input the preprocessed images into the online usable CZ-YOLO network obtained in step 2.3. Perform inference by loading the model weight file to generate trademark defect prediction candidate boxes. Remove low-confidence candidate boxes by setting a confidence threshold, and then use non-maximum suppression (NMS) for screening to output clothing anomaly detection images with trademark defects.
[0136] Experiment
[0137] 1. The dataset of Example 1 is used in the experiment.
[0138] 2. Evaluation Metrics
[0139] The experimental evaluation metrics include the number of parameters, floating-point operations per second (FLOPs), mean average precision (mAP), and frames per second (FPS).
[0140] 3. Data Augmentation Verification Experiment
[0141] To verify the effectiveness of the data augmentation strategy of combining CycleGAN with traditional data augmentation techniques to expand the dataset in the present invention, a control experiment is carried out on the CZ-YOLO network of the present invention. The results show that this strategy can significantly improve the detection performance under different original datasets. The dataset is based on the size of the original dataset (20 / 50 / 80 samples), and three groups of control experiments are carried out respectively:
[0142] ① represents the original dataset;
[0143] ② represents generating twice the number of data of the original dataset using CycleGAN + the original dataset;
[0144] ③ On the basis of ②, traditional data augmentation techniques are added, including rotation, Gaussian noise, changing brightness and contrast.
[0145] Table 1: Performance of Data Augmentation in the CZ-YOLO Network
[0146]
[0147]
[0148] The experimental results show that the strategy of CycleGAN + traditional data augmentation shows a significant improvement in universality in the CZ-YOLO network:
[0149] When there are 20 small samples, the CZ-YOLO network is improved from 61.4% to 82.8% (+21.4%), verifying that using CycleGAN can better solve the problem of insufficient basic data for small samples. In the cases of 50 samples and 80 samples, there is still a large improvement, but the improvement amplitude gradually decreases as the basic dataset increases.
[0150] After adding traditional data augmentation, the comparison with only using CycleGAN also shows improvements, and the improvement amplitude also decreases as the basic dataset increases. This verifies that using traditional data augmentation can also play a role in expanding the data. However, it should be noted that the improvement amplitude of mAP@.5:.9 decreases when the number of samples increases from 50 to 80. We speculate that overfitting occurs after the dataset is expanded.
[0151] From the comparison in the table, we can conclude that in the small-sample defect detection task, the hybrid augmentation strategy has significant advantages. CycleGAN is responsible for expanding the coverage of defect categories, while traditional augmentation enhances the model's sensitivity to subtle defect features, thus providing a reliable solution for industrial-level small-sample defect detection.
[0152] 4. Ablation Experiment
[0153] Verify the effectiveness of each improved module through ablation experiments, and the results are shown in Table 2:
[0154] Table 2 Results of CZ-YOLO Ablation Experiment
[0155]
[0156] Among them, yolov11n is the basic network; yolov11n+HS-FPNs means using yolov11n as the basic network and the neck of the CZ-YOLO network for the neck; yolov11n+HS-FPNs+MDSA means adding a multi-scale dynamic collaborative attention (MDSA) module between the Conv module + C3k2 module structures of the second and third layers of the backbone on the basis of yolov11n+HS-FPNs; yolov11n+HS-FPNs+DWconv+Hetconv means replacing the structure of the four-layer Conv module + C3k2 module of the backbone with a DWConv module + Het-C3k2 module structure on the basis of yolov11n+HS-FPNs; ours is the complete version of the CZ-YOLO network.
[0157] Experiments show that the synergistic effect of HS-FPNs, MDSA and hybrid convolution is the key to improving the model performance and lightweight, and the specific analysis is as follows:
[0158] The joint optimization of the HS-FPNs module and the hybrid convolutional architecture achieves a significant lightweight effect while maintaining the model detection performance. The HS-FPNs module reduces the model's mAP@.5:.95 from 80.3% to 79.2%, but reduces the model's parameters by 28.5% (1.846×106), reduces Flops by 11.1% (5.6GFLOPs), and increases the inference speed by 13.3% (87.7FPS). This shows that HS-FPNs reduces redundant calculations through selective feature fusion, and the loss of some detail features caused by the lightweight design is not obvious. After replacing the traditional convolution with the "DWConv+HetConv" hybrid convolutional architecture, the model parameters dropped by 51.7% (1.247×106), Flops dropped by 33.3% (4.2GFLOPs), and the inference speed soared to 302.6FPS. However, mAP@0.5 dropped by 0.2% and mAP@0.5:0.95 dropped by 1%, indicating that pure convolution optimization would sacrifice some semantic features. Nevertheless, the combined effect of hybrid convolution and HS-FPNs enabled the model to maintain 97.4% mAP@0.5, indicating that the lightweight design did not undermine the effectiveness of feature fusion.
[0159] The MDSA attention mechanism plays a key compensatory role in the improvement of this model. It shows significant effectiveness in both the feature loss caused by HS-FPNs lightweighting and the feature weakening caused by mixed convolution. Adding the HS-FPNs module alone will cause some detail features to be lost. After the MDSA attention mechanism is introduced, mAP@.5:.95 rises to 79.8%, which fully verifies that MDSA can effectively compensate for the feature loss caused by HS-FPNs lightweighting through dynamic feature focusing, reflecting the synergistic effect of "feature screening-attention enhancement". Moreover, the final model (Ours) integrates the MDSA attention mechanism, which enables the model to achieve a 0.6% increase in mAP@0.5 on the basis of HS-FPNs+mixed convolution, reaching 98.0%, which is better than all ablation variants. This shows that MDSA effectively compensates for the feature weakening problem caused by lightweight convolution through a multi-scale collaborative attention mechanism.
[0160] 4. Comparative experiment
[0161] The lightweight improved model proposed in this invention achieves significant improvement in computing efficiency through innovative network structure optimization while maintaining detection accuracy. The performance comparison data with the existing network is shown in Table 3:
[0162] Table 3 Performance comparison of mainstream target detection algorithms
[0163]
[0164]
[0165] The CZ-YOLO network exhibits a balance of accuracy, speed, and complexity in industrial detection scenarios. Experimental data show that while maintaining detection accuracy comparable to mainstream algorithms, the model achieves significant performance improvements through the "lightweight design-feature enhancement-efficient reasoning" trinity optimization strategy. In terms of key indicators of target detection, the CZ-YOLO network achieved an mAP@0.5 index of 98.0%, only 0.4 percentage points behind the current optimal RTDETR algorithm, and significantly better than YOLOv11 (97.6%) and YOLOv8 (98.0%). On the stricter mAP@0.5:0.95 comprehensive index, the CZ-YOLO network maintained a stable performance of 80.0%, only 0.1% lower than RTDETR and YOLOv11, verifying the robustness enhancement effect of its MDSA attention mechanism on complex backgrounds.
[0166] The lightweight design verification results show that the CZ-YOLO network has achieved key performance breakthroughs while ensuring detection accuracy through the coordinated optimization of the DWConv+HetConv hybrid convolution architecture and the HS-FPNs lightweight feature fusion module: the number of parameters is only 1.302×106, which is 49.5% less than YOLOv11 (2.583×106) and 96.0% less than RTDETR (32.810×106); the amount of calculation is reduced from 6.3GFLOPs of YOLOv11 to 4.5GFLOPs, which is only 16.6% of EfficientDET. At the same time, the inference speed reaches 303.4FPS, which is 3 times that of YOLOv11 (100.0FPS) and 2.58 times that of YOLOv8 (117.6FPS). This lightweight design lays a solid foundation for the real-time deployment of the model on edge devices, significantly reduces hardware dependence while maintaining detection capabilities, and provides an innovative solution for the high-precision, low-power, and fast-response requirements in the field of industrial detection.
[0167] Figure 11 This is an example diagram of the detection effect of using the CZ-YOLO network of the present invention on trademark defects in the production process. It can be seen that the CZ-YOLO network proposed in the present invention has a very high confidence score in the identification of trademark anomalies and can complete the task of real-time detection.
[0168] Finally, it should be noted that the above examples are only some specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A small-sample lightweight defect detection method based on improved YOLOv11n, characterized in that: The specific process is as follows: collect the clothing image with the trademark part, and after preprocessing the collected image, input it into the offline trained CZ-YOLO network to generate the trademark defect prediction candidate box and confidence level. The CZ-YOLO network is improved based on the Yolov11n network. The backbone of the CZ-YOLO network includes a series of Conv modules, a structure of two-layer DWConv modules + Het-C3k2 module, an MDSA module, a structure of two-layer DWConv modules + Het-C3k2 module, an SPPF module, and a C2PSA module, which are used to output three different scale features S2, feature S3, and feature S4. The neck of the CZ-YOLO network includes an HS-FPNs network and a Het-C3k2 module. After the features S2, feature S3, and feature S4 pass through the HS-FPNs network, the output features N2, feature N3, and feature N4 are each processed by a Het-C3k2 module and then used as the inputs of the three detection heads respectively. The detection head of the CZ-YOLO network uses the detection head of the Yolov11n network.
2. The small sample lightweight defect detection method based on the improved YOLOv11n according to claim 1, characterized in that: The Het-C3k2 module is specifically: The Het-C3k2 module is an improvement of the C3k2 module. First, replace the ordinary convolution in the Bottleneck of the C3k2 module in the original Yolov11n network with a heterogeneous kernel convolution, and remove the original skip connection to construct a HetBottleneck. Then, replace the Bottleneck in the C3k of the original Yolov11n network with HetBottleneck to obtain HetC3k. Then, keep the structure of the C3k2 module in the original Yolov11n network unchanged when C3k = False as the structure of the Het-C3k2 module when C3k = False. When C3k = True, replace the C3k in the C3k2 module of the original Yolov11n network with HetC3k as the structure of the Het-Cek2 module when C3k = True.
3. The small sample lightweight defect detection method based on the improved YOLOv11n according to claim 2, characterized in that: The MDSA module includes multi-scale channel attention, multi-scale spatial attention, and a dynamic collaboration layer. The output feature F generates the feature CFF through multi-scale channel attention. The feature CFF obtains the spatial attention weight W through multi-scale spatial attention s , and the feature CFF obtains the channel attention weight W through the dynamic collaboration layer c . Then, two learnable dynamic weights are used to multiply the spatial attention weight W s and the channel attention weight W c , and then weighted fusion is performed by addition. Then, the overall attention weight W is calculated through Sigmoid activation and acts on the input feature CFF as the output of the MDSA module 4. The small sample lightweight defect detection method based on the improved YOLOv11n according to claim 3, characterized in that: The multi-scale channel attention includes two stages: feature extraction and channel feature fusion. In the feature extraction stage, the input feature F is respectively subjected to global average pooling and global max pooling to obtain the global feature F GAP and the global feature F GMP ; the input feature F is respectively subjected to local average pooling + global average pooling and local max pooling + global max pooling to obtain the local feature F LAP+GAP and the local feature F LMP+GMP ; The channel feature fusion stage is as follows: F1 = F GAP + F LAP+GAP (1) F2 = F GMP +F LMP+GMP (2) Then, splice the feature F1 and feature F2 along the channel dimension and generate the channel attention weight through the channel weight calculation gate. Apply the generated channel attention weight to the input feature F to obtain the feature CFF. The channel weight calculation gate is composed of a fully connected layer, a ReLU activation function, a fully connected layer, a convolutional layer, and a sigmoid activation function in series.
5. The small sample lightweight defect detection method based on the improved YOLOv11n according to claim 4, characterized in that: The multi-scale spatial attention is as follows: after the feature CFF is subjected to global average pooling and global maximum pooling along channels respectively, they are concatenated in the channel dimension, and then a convolution is used to obtain the spatial attention weight W s .
6. The small-sample lightweight defect detection method based on the improved YOLOv11n according to claim 5, wherein: The dynamic collaboration layer is as follows: the feature CFF is subjected to feature extraction through depthwise separable convolution, and global average pooling is performed along the channels to obtain the channel attention weight W c .
7. The small-sample lightweight defect detection method based on the improved YOLOv11n according to claim 6, wherein: The HS-FPNs are as follows: the feature S2, feature S3, and feature S4 output by the backbone each pass through a CA module and a convolutional module to obtain the feature P2, feature P3, and feature P4. While the feature P4 is directly output as the feature N4, the feature N4 and the feature P3 pass through an improved SFF module to obtain the feature N3, and the feature N4 passes through deconvolution and then passes through an improved SFF module with the feature P2 to obtain the feature N2.
8. The small-sample lightweight defect detection method based on the improved YOLOv11n according to claim 7, wherein: The improved SFF module is as follows: for the input high-level feature f high perform a transposed convolution operation to obtain the upsampled high-level feature f high_up , and for the high-level feature f high_up obtain the channel weight ω through the CA module, and then use the channel weight ω to perform per-channel weighted screening on the input low-level feature f low to complete per-channel weighted screening to obtain the attention-screened low-level feature Then, fuse the upsampled high-level feature f high_up with the attention-screened low-level feature as the output feature of the improved SFF module.
9. The small-sample lightweight defect detection method based on the improved YOLOv11n according to claim 8, wherein: The offline training of the CZ-YOLO network is offline training with a small number of samples: (1) Dataset collection: Collect clothing pictures with trademarks, the number of which is no more than 80; in the computer, select trademark defect pictures from the collected pictures and annotate them with LabelImg. The pictures with trademark defects are labeled as 'NG', and the pictures of the L trademark are labeled as 'L' as the original dataset. (2) Dataset augmentation: Use the CycleGAN network and traditional data augmentation methods to augment the number of pictures in the original dataset respectively, and then randomly divide the images obtained by dataset augmentation and the original images into a training set, a validation set, and a test set according to the ratio of 8:1:
1. (3) Input the labeled pictures in the training set into the CZ-YOLO network for training, calculate the loss function value and backpropagate to iteratively optimize the model parameters, and end the training after reaching the preset number of epochs; put the labeled test set pictures into the trained model, generate candidate prediction boxes, and calculate the intersection over union and performance indicators of the candidate prediction boxes and the labeled boxes to evaluate the model.
Citation Information
Cited By
Multi-core feature representation learning system and method for enhancing remote sensing image based on content retrieval
CN120853029A
Textile product defect identification method based on improved YOLOv11
CN121353859A
Serum sample identification and detection system and method based on improved YOLO11n network
CN121366341A