Small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n

By introducing a method of channel cascade feature fusion and enhancement in the YOLOv8n detector, the problems of difficulty in detecting small target defects and large background interference in industrial defect detection are solved, and higher detection accuracy and efficiency are achieved.

CN120071085APending Publication Date: 2025-05-30TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146356.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In industrial defect detection, there are problems such as difficulty in detecting small target defects and large background interference, resulting in low detection accuracy.

Method used

Using the detection method of channel cascade feature fusion and enhancement based on YOLOv8n, the feature fusion and extraction capabilities are enhanced and background interference is reduced by introducing the spatial channel deep fusion attention module and the channel cascade fusion attention upsampling module.

Benefits of technology

It effectively improves the accuracy and efficiency of small-target defect detection, reduces background interference, and enhances the network's ability to capture small-target defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071085A_ABST
    Figure CN120071085A_ABST
Patent Text Reader

Abstract

The invention discloses a small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n, and relates to the technical field of artificial intelligence, and the method mainly comprises the following steps: in a backbone part, carrying out the feature representation of an image through a space channel depth fusion attention module, and extracting feature maps of different scales; in the check part, a channel cascade fusion attention up-sampling module is utilized to perform multi-scale feature fusion and enhancement so as to compensate for feature differences among feature maps with different scales; and in the head part, target detection is performed on the obtained fusion features by using the check part, and the target category, position and confidence are obtained. Therefore, by adopting the small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n, the semantic information of the feature map can be efficiently extracted, meanwhile, the original image information is better reserved, and the influence of a complex background on small target defect detection is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n. Background Art

[0002] In modern industrial production, electronic product manufacturing, aerospace and other fields, product quality control is of crucial importance. As a key link in quality control, defect detection can timely detect defects on the surface or inside of products, such as cracks, holes, scratches, impurities, etc., avoid defective products from entering the market, and reduce potential safety hazards and economic losses.

[0003] Traditional defect detection methods mainly rely on manual visual inspection and rule-based machine vision methods. Manual visual inspection is inefficient and its accuracy is easily affected by subjective factors, while rule-based machine vision methods have poor adaptability to complex defect patterns and changing environments.

[0004] With the development of deep learning, it has demonstrated excellent performance in the field of defect detection. For example, two-stage detectors based on R-CNN, Fast R-CNN, and Faster R-CNN have shown high accuracy on the COCO dataset. Subsequently, single-stage detectors such as YOLOX, YOLOv7, and YOLOv8 emerged, providing higher detection accuracy while maintaining a satisfactory detection speed, gradually replacing the two-stage detectors led by R-CNN. Currently, transformers are very popular, and DETR, as a pioneer of transformers in the field of object detection, has become the first end-to-end detector. However, problems such as its slow training speed and low small target detection accuracy have become the main reasons why DETR-like models cannot become the mainstream detectors in the industrial field.

[0005] In addition, for the industrial defect detection field, due to the particularity of its data, there are still certain limitations in using existing basic detectors to detect industrial defects. For example, problems such as similar defect styles, small defect ratios, and high similarity between defects and the background will all lead to poor performance of classical object detectors in industrial defect detection. To improve the accuracy of defect detectors, researchers have proposed adding attention mechanisms, such as convolutional block attention module (CBAM) and coordinate attention module (CA), etc., to enable the model to focus on more relevant features; and have also introduced improved feature pyramid networks and upsampling and other technologies for enhancing feature fusion and feature representation. Moreover, deep learning networks improved based on the YOLO series have been successfully applied to detect defects in steel, concrete, and other materials.

[0006] Although deep learning-based defect detection methods have been applied in actual industries, industrial defect detection still faces many challenges: (1) Difficulty in detecting small target defects: Due to problems such as small defect size and low resolution, how to accurately and efficiently detect small target defects has always been a difficult problem being tackled in the field of industrial defect detection. (2) Large background interference: Since the shapes and colors of various defects are extremely similar to the background, it increases the interference of the background on the detection of small target defects. Therefore, it is necessary to provide an industrial defect detection method to further solve the problems existing in the process of extracting small target defects. Summary of the Invention

[0007] The purpose of the present invention is to provide a small target defect detection method based on channel cascaded feature fusion and enhancement of YOLOv8n, which can reduce the interference caused by the background and achieve the detection of small target defects, so as to solve the problem of low detection accuracy due to the very small size of metal surface defects and the high similarity between defect patterns and the background.

[0008] To achieve the above purpose, the present invention provides a small target defect detection method based on channel cascaded feature fusion and enhancement of YOLOv8n, including:

[0009] Based on the YOLO v8n network framework, a spatial channel depth fusion attention module and a channel cascaded fusion attention upsampling module are introduced to construct a network based on channel cascaded feature fusion and enhancement, including three parts: backbone, neck, and head;

[0010] In the backbone part, a spatial channel depth fusion attention module is designed. By introducing the CA attention mechanism, long-range dependencies are enhanced, and spatial and channel attention are fused to obtain an enhanced feature map. Feature maps of different scales are extracted to obtain the enhanced feature map and extract feature maps of different scales, enabling the network to better capture small target defects;

[0011] In the neck part, a channel cascaded fusion attention upsampling module is designed. Through cascaded fusion convolution, feature semantic information is extracted fully and efficiently, while the original image features are retained. The augmented channels are used to supplement the spatial information to make up for the missing spatial information, providing high-quality images for the fusion of feature maps of different scales. Then, the feature maps of different scales are upsampled and channel-connected with the corresponding-sized feature maps in the backbone to obtain the fused feature map, thereby making up for the feature difference between feature maps of different scales;

[0012] In the head part, target detection is performed using the fused feature map obtained from the neck part.

[0013] Preferably, the spatial channel depth fusion attention module includes a parallel channel attention mechanism and a spatial attention mechanism, which respectively obtain a channel attention feature map and a spatial attention feature map, and then connect them in channels to obtain an enhanced feature map.

[0014] Preferably, the spatial attention mechanism includes:

[0015] First, split the input feature map along the spatial dimension into two tensors, and perform convolution and non-linear activation operations on them respectively to obtain corresponding attention weights;

[0016] Then, multiply the attention weights by the input feature map to obtain a spatial attention feature map.

[0017] Preferably, the channel attention mechanism includes:

[0018] First, process the input feature map using global max pooling and global average pooling respectively, and obtain corresponding attention weights after passing through the MLP layer;

[0019] Next, add the obtained attention weights and then pass through an activation function to obtain channel attention weights;

[0020] Then, multiply the channel attention weights by the input feature map to obtain a channel attention feature map.

[0021] Preferably, the channel concatenation fusion attention upsampling module includes performing feature extraction on the input feature map of the module to different degrees, and connecting the obtained feature maps of different degrees with the input feature map in channels to obtain a first feature map; at the same time, performing channel amplification on the input feature map of the module to obtain a second feature map; then, adding the first feature map and the second feature map so that the fused feature map is fused with the input feature map of the module again, and then passing through a channel attention mechanism, a Pixel Shuffle operation, and a CA attention mechanism in sequence to obtain a channel attention feature map.

[0022] Preferably, the channel attention mechanism is used to obtain the importance of each channel; the Pixel Shuffle operation is used to rearrange each pixel on the feature map; the CA attention mechanism is used to obtain the relationship between each pixel.

[0023] Therefore, the present invention adopts the above small target defect detection method based on channel concatenation feature fusion and enhancement of YOLOv8n, and has the following technical effects:

[0024] (1) In the spatial-channel depth fusion attention module, while retaining the channel attention in CBAM, the CA attention is introduced to form parallel channels. The depth fusion of the two attentions complements each other, enabling the network to better capture small target defects and reduce the interference caused by the background.

[0025] (2) In the channel cascaded fusion attention upsampling module, with sufficient cascaded fusion convolutions, the network can efficiently extract the semantic information of the feature map while better retaining the features of the original image. At the same time, the increased number of channels enables the network to make full use of the channel information, thus compensating for the missing spatial information and providing high-quality images for the subsequent fusion of feature maps at different scales.

[0026] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0027] Figure 1 is a schematic diagram of a channel cascaded feature fusion and enhancement network based on YOLOv8n;

[0028] Figure 2 is a schematic diagram of the spatial-channel depth fusion attention module;

[0029] Figure 3 is a schematic diagram of the channel cascaded fusion attention upsampling module. Detailed Embodiment

[0030] The present invention can be more specifically explained through the following embodiments. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following embodiments.

[0031] In this embodiment, based on the YOLO v8n network, an improved YOLO-CCFE model is proposed, that is, a channel cascaded feature fusion and enhancement network, to improve the accuracy and detection speed of industrial surface defect detection and facilitate the deployment of the model.

[0032] As Figure 1 shown, the YOLO-CCFE model includes three parts: backbone, neck, and head. Among them, the backbone part is used for feature extraction, the neck part is used for extracting features at different scales and feature fusion, and the head part is used for target detection.

[0033] In the backbone part, a spatial-channel depth fusion attention (SCDF) module is introduced, which is placed at the beginning of the intermediate fusion part and the end of the backbone respectively. It can enhance the feature extraction ability of the backbone part and enable the subsequent feature fusion process to adaptively capture the local context in the multi-scale receptive field.

[0034] As shown Figure 2 in the figure, the SCDF module adopts a hybrid attention similar to CBAM, but different from the cascaded fusion of CBAM, SCDF uses a parallel fusion method, reducing the mutual influence between channel and spatial attention.

[0035] In terms of channels, the same channel attention as in CBAM is adopted. Through two different pooling methods, key information of different scales is obtained. At the same time, CA attention is introduced in terms of space. By pooling and splicing in two different directions of X and Y, the weight values of different pixel points in different channels are calculated, thus solving the problem of being unable to capture long-range dependencies in CBAM. Then, the two attention results are weighted and summed. After convolution, normalization, and activation functions, the residual idea is introduced and fused with the original feature map, which can extract key features while supplementing the original features.

[0036] Specifically, given the input feature map f in , through two parallel spatial attention and channel attention, as follows:

[0037] In spatial attention, first, two spatial ranges (H, 1) and (1, W) of the pooling kernel are used to encode each channel along the horizontal and vertical coordinates respectively. The two feature maps are concatenated along the channels (Concat), and then convolution (Conv) is used to reduce the number of channels to achieve the purpose of reducing the number of parameters. After normalization (BN) and non-linear activation function (NL), the feature map f is obtained. The specific expression is as follows:

[0039] In the formula, [·, ·] represents the concatenation operation along the spatial dimension, G 1 is the convolution function, δ is the non-linear activation function, c represents the corresponding number of channels, and r is the reduction ratio controlling the block size;

[0040] After that, the input feature map is split along the spatial dimension to obtain two separate tensors f h ∈R C / r×H and f w ∈R C / r×W , with lengths H and W respectively; then two convolution transforms G h , G w are used to transform the two split tensors into tensors with the same channel number as the input feature map respectively. The specific expression is as follows:

[0041] g h =σ(G h (f h ));

[0042] g w =σ(Gh (f w ));

[0043] Wherein, σ is the sigmoid activation function, and g h and g w are the corresponding attention weights.

[0044] Multiply the obtained attention weights by the input feature f in to obtain the spatial attention feature map f a , and the specific expression is as follows:

[0046] Wherein,

[0047] In the channel attention, for the input feature map f in , use global max pooling (GMP) and global average pooling (GAP) respectively, and obtain two attention weights through the MLP layer respectively. Then add them up and pass through the activation function to obtain the final channel attention weight g, and the specific expression is as follows:

[0048] g = σ(MLP(AvgPool(f in )) + MLP(MaxPool(f in )));

[0049] Wherein, AvgPool and MaxPool are global average pooling and global max pooling respectively.

[0050] Multiply the channel attention weight by the input feature map f in to obtain the channel attention feature map f b .

[0051] Subsequently, add the two attention-processed feature maps f a and f b , and after convolution, normalization and activation function processing, add it to the input feature f in to supplement the original image features, and finally obtain the enhanced feature map.

[0052] In the neck part, a brand-new channel concatenation fusion attention upsampling module (CCFA) is proposed to replace the original nearest neighbor interpolation upsampling method, and splice the feature maps after cascade fusion convolution to increase the number of channels, so as to make full use of channel information to make up for the missing spatial information and provide high-quality images for the subsequent fusion of feature maps at different scales.

[0053] Among them, upsampling refers to the process of increasing the spatial resolution of an image to increase the size of the image while trying to maintain the image quality. This process is crucial for subsequent tasks such as classification, detection, and segmentation.

[0054] Currently, the most widely used method in existing methods is the mathematical linear interpolation method. This method interpolates new pixel points using existing pixel points and obtains a relatively stable image without adding any parameters. However, for small target defects, simply relying on individual pixels to obtain a new image through mathematical calculations cannot well support subsequent detection tasks.

[0055] Therefore, for deep learning, adopting a suitable upsampling method is an important way to improve small target detection. Among existing methods, CARAFE uses sub-pixel convolution to perform pixel rearrangement on the feature map after using convolution to expand the number of channels, obtains an upsampling kernel, and then reassembles features in a predefined area. However, CARAFE mainly emphasizes how to perform upsampling using the upsampling kernel after channel expansion, but only uses convolution to directly expand channels during channel expansion and does not fully utilize the feature map information.

[0056] For this reason, this embodiment improves on the basis of the CARAFE module and proposes a channel cascaded fusion attention upsampling module, aiming to study how to use the feature map for channel expansion, as Figure 3 shown. On the one hand, convolution is used to extract features from feature maps at different levels and intersperse and fuse the original feature maps to obtain key information at different levels. At the same time, the feature maps obtained through multiple feature extractions are used as the expanded channels, solving the problems of insufficient information utilization and information redundancy in directly expanding channels using convolution. On the other hand, the expanded number of channels can fully utilize channel information, thereby compensating for the missing spatial information, and introducing channel attention and spatial attention after sub-pixel convolution to better extract features, providing high-quality images for the fusion of feature maps at different scales in the future.

[0057] Specifically, given the input feature map f in , after feature extraction through convolution, it is added to the original feature map for full fusion, and this operation is repeated 3 times to obtain feature maps g 1 , g 2 and g 3 with different degrees of feature extraction. The specific expressions are as follows:

[0058] g 1 = G 1 (f in ) + f in ;

[0059] g 2 = G 2 (g 1 ) + f in ;

[0060] g 3 = G3 (g 2 ) + f in ;

[0061] In the formula, G 1 , G 2 , G 3 are all convolution functions that contain only one convolution kernel. Since it is added to the original image after each convolution, rich original image information is included in each feature map, ensuring that the feature map does not deviate from the original image during the continuous feature extraction process.

[0062] Connect the three convolved feature maps g 1 , g 2 , g 3 and the input feature map f in through channel connection (Concat) to form a feature map f a of 4C×H×W, as follows:

[0063] f a = [f in ; g 1 ; g 2 ; g 3 ;

[0064] In the formula, [·;·] represents the splicing operation along the channel dimension.

[0065] Meanwhile, directly amplify the channels of the input feature map f in through convolution to become a feature map of 4C×H×W. Add the two feature maps to fuse the original image information again.

[0066] Since the order of the feature maps of each channel is fixed and the importance of the 4C channels is unknown, it is necessary to introduce channel attention first, so that the network can obtain the importance of each channel in the feature map, enhance the importance of the feature channels useful for the current task, and suppress the feature channels that are not very useful for the current task, so that the network can focus on the feature channels with high weights. Then, perform the Pixel Shuffle operation on the feature map after channel attention to rearrange each pixel on the feature map. Similarly, since the arrangement method is fixed, it is necessary to introduce the CA attention mechanism (CoordinateAttention) so that the network can obtain the importance of each pixel in the feature map and can capture the long-range dependencies across channels to obtain the relationship between each pixel, so that the network can focus on the regional features with high weights. The specific expression is as follows:

[0067] f out = CA(PS(CHA(f a + G(f in )))));

[0068] where f out is the output feature map of C×2H×2W, CHA is the channel attention, and PS is the Pixel Shuffle operation.

[0069] Generally speaking, the channel cascade fusion attention upsampling network utilizes the idea of sub-pixel convolution. If the feature map of C×H×W is upsampled to the feature map of C×2H×2W, first, the number of channels of the feature map of C×H×W is expanded to 4 times the original, that is, from C×H×W to 4C×H×W, and then the 4 channels of each pixel on the feature map are rearranged into a 2×2 area. So far, the upsampling operation is completed, making the original feature map of C×H×W become the feature map of C×2H×2W.

[0070] Therefore, the small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n adopted by the present invention can efficiently extract the semantic information of the feature map, while better retaining the original image information, effectively solving the problems of small defect size and low resolution, and reducing the influence of complex backgrounds on small target defect detection.

[0071] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n, characterized in that: include: Based on the YOLO v8n network framework, the spatial channel deep fusion attention module and the channel cascade fusion attention upsampling module are introduced to build a channel cascade feature fusion and enhancement network, including backbone, neck and head parts; In the backbone part, a spatial channel deep fusion attention module is designed. By introducing the CA attention mechanism, long-distance dependencies are enhanced, and spatial and channel attention are fused to obtain enhanced feature maps and extract feature maps of different scales. In the neck part, a channel cascade fusion attention upsampling module is designed to extract feature semantic information through cascade fusion convolution while retaining the original image features. The amplification channel is used to supplement the spatial information, and then the feature maps of different scales are upsampled and connected with the feature maps of the corresponding size in the backbone through channels to obtain the fused feature maps. In the head part, the fused feature map obtained in the neck part is used for target detection.

2. The small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n according to claim 1 is characterized in that: The spatial channel deep fusion attention module includes the use of parallel channel attention mechanism and spatial attention mechanism to obtain channel attention feature map and spatial attention feature map respectively, and then connect them through channels to obtain enhanced feature map.

3. The small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n according to claim 2 is characterized in that: Spatial attention mechanism, including: First, each channel is encoded along the horizontal and vertical coordinates respectively, the channels are concatenated, and the channels are reduced by convolution, and then normalized and nonlinearly activated. Next, it is split into two tensors along the spatial dimension, and convolution and nonlinear activation operations are performed on them respectively to obtain the corresponding attention weights; Then, the obtained attention weights are multiplied with the input feature map to obtain the spatial attention feature map.

4. The small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n according to claim 2 is characterized in that: Channel attention mechanism, including: First, the input feature map is processed using global maximum pooling and global average pooling respectively, and the corresponding attention weights are obtained after each passing through the MLP layer; Next, the obtained attention weights are added and passed through the activation function to obtain the channel attention weight; Then, the channel attention weight is multiplied by the input feature map to obtain the channel attention feature map.

5. The small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n according to claim 1 is characterized in that: The channel cascade fusion attention upsampling module includes performing different degrees of feature extraction on the input feature map of the module, and channel-connecting the obtained feature maps of different degrees with the input feature map to obtain a first feature map; at the same time, channel amplifying the input feature map of the module to obtain a second feature map; then, adding the first feature map and the second feature map, so that the fused feature map is fused again with the input feature map of the module, and then successively undergoing a channel attention mechanism, a Pixel Shuffle operation and a CA attention mechanism to obtain a channel attention feature map.

6. The small target defect detection method based on channel cascade feature fusion and enhancement of YOLOv8n according to claim 5 is characterized in that: The channel attention mechanism is used to obtain the importance of each channel; the pixel shuffle operation is used to rearrange each pixel on the feature map; the CA attention mechanism is used to obtain the relationship between each pixel.