Construction method and application of small sample chip surface defect detection model based on improved YOLOv8

By improving the YOLOv8 model, using GAN to generate amplified images, and introducing the MSCAM and Head sections into the Neck layer to use the focus loss function, the small sample and category imbalance in chip surface defect detection is solved, and the generalization ability and detection accuracy of the model are improved.

CN120070343APending Publication Date: 2025-05-30JIANGSU UNIV
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510101449.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing deep learning models face small sample problems and category imbalance in chip surface defect detection, resulting in limited model generalization ability and detection performance.

Method used

Using the improved YOLOv8 model, amplified images are generated by generating adversarial networks (GANs), increasing the richness of training data, and introducing a multi-scale convolutional attention module (MSCAM) in the Neck layer, using the focus loss function in the head part.

Benefits of technology

It effectively improves the generalization ability of the model in small sample learning, enhances the detection accuracy of defects in complex backgrounds, and optimizes the detection performance of a few defect categories, solving the problem of category imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure BDA0005254299480000031
    Figure BDA0005254299480000031
  • Figure BDA0005254299480000072
    Figure BDA0005254299480000072
Patent Text Reader

Abstract

The invention discloses an improved YOLOv8-based small sample chip surface defect detection model construction method and application. The method comprises the following steps: acquiring an original image of a chip surface defect; inputting the preprocessed original image and the random noise into a GAN (Generative Adversarial Network), and obtaining an amplified image through the GAN; respectively marking chip surface defects in the original image and the amplified image to obtain a chip surface defect data set; a multi-scale convolution attention module MSCAM is introduced into the Neck part of a YOLOv8 network model, the perception ability of the model for different scale features can be enhanced through introduction of the MSCAM, the defect detection precision under the complex background is improved, meanwhile, a classification loss function of the Head part is modified into a focus loss function, and then an improved YOLOv8 target detection model is constructed. The improved YOLOv8 target detection model can be used for detecting the surface defects of the chip, and the problems of small samples and unbalanced categories can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and particularly relates to a method for constructing a small-sample chip surface defect detection model based on improved YOLOv8 and its application. Background Art

[0002] As the core component of modern electronic devices, chips play a crucial role in many fields. With the continuous expansion and deepening of chip applications, the requirements for chip quality and reliability are also increasing day by day. Therefore, chip quality inspection and defect classification have become key research directions.

[0003] However, due to the complex chip structure and diverse defect types, traditional detection methods face many challenges. Although there are methods in the prior art that use deep learning for chip defect detection, deep learning models require a sufficient number of data samples during training. The cost of obtaining labeled data is high, and it is often difficult to collect enough defect samples. This leads to the small-sample problem when training deep learning models, limiting the generalization ability and detection performance of the models. Due to manufacturing process or manufacturing environment problems, there is also a class imbalance in chip surface defects, resulting in existing deep learning models being biased towards the majority class during training and prediction and ignoring the minority class. This greatly limits the development of chip surface defect detection technology. Summary of the Invention

[0004] In order to solve the deficiencies in the prior art, the present application proposes a method for constructing a small-sample chip surface defect detection model based on improved YOLOv8 and its application, aiming to effectively solve the small-sample problem and the class imbalance problem.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for constructing a small-sample chip surface defect detection model based on improved YOLOv8, comprising the following steps:

[0007] Step 1, collect the original images of chip surface defects;

[0008] Step 2, input the preprocessed original images and random noise into the generative adversarial network GAN to obtain amplified images by the generative adversarial network;

[0009] Step 3, label the chip surface defects in the original images and the amplified images respectively to obtain a chip surface defect dataset;

[0010] Step 4, construct an improved YOLOv8 object detection model. The improved YOLOv8 object detection model includes a Backbone layer, a Neck layer, and a Head layer. Among them, the Backbone layer is sequentially connected by the first CBL network structure, the second CBL network structure, the first C2f network structure, the third CBL network structure, the second C2f network structure, the fourth CBL network structure, the third C2f network structure, the fifth CBL network structure, the fourth C2f network structure, and the SPPF network structure. The second C2f network structure, the third C2f network structure, and the SPPF network structure respectively output the first feature map P1, the second feature map P2, and the third feature map P. The Neck layer is sequentially connected by the first upsampling, the first Concat network structure, the fifth C2f network structure, the second upsampling, the second Concat network structure, the sixth C2f network structure, the first MSCAM multi-scale convolutional attention module, the sixth CBL network structure, the third Concat network structure, the seventh C2f network structure, the second MSCAM multi-scale convolutional attention module, the seventh CBL network structure, the fourth Concat network structure, the eighth C2f network structure, and the third MSCAM multi-scale convolutional attention module. Among them, the first upsampling and the fourth Concat network structure are both connected to the SPPF network structure and receive the first feature map P1. The first Concat network structure is connected to the fifth C2f network structure and receives the output second feature map P2. The second Concat network structure is connected to the second C2f network structure and receives the output third feature map P3. The output features of the Neck layer are respectively output by the first MSCAM multi-scale convolutional attention module, the second MSCAM multi-scale convolutional attention module, and the third MSCAM multi-scale convolutional attention module.

[0011] The Head layer receives the input of the Neck layer and outputs the detection result. The Head layer uses the focal loss function.

[0012] Step 5, after completing the construction of the improved YOLOv8 object detection model, it is necessary to train, validate, and test the built improved YOLOv8 object detection model.

[0013] Furthermore, the MSCAM multi-scale convolutional attention module includes a channel attention block CAB, a spatial attention block SAB, and a multi-scale convolutional block MSCB connected in sequence.

[0014] Furthermore, in the channel attention block CAB, global max pooling and global average pooling are respectively used to process the feature maps to extract the most important features of each feature map. Subsequently, convolutions are respectively used to reduce the number of channels to 1 / 16 of the original, and then ReLU is used as the activation function to improve the speed and performance of the neural network. Then, convolutions are used to restore the number of its channels. Finally, after concatenation, the sigmoid activation function is used to apply weights to the original feature map to achieve channel weighting.

[0015] Furthermore, in the spatial attention block SAB, max pooling and average pooling are respectively performed on the input feature map to focus on local features. After concatenation of the max pooling and average pooling processing structures, a large kernel convolution of 7×7 is used to emphasize these local features. Finally, the sigmoid function is used to calculate the weights, and the spatial attention map is applied to the original feature map.

[0016] Furthermore, in the multi-scale convolution block MSCB, first, a 1×1 convolution is used to expand the number of channels to twice the original. Then, the BN and ReLU6 activation functions are used to improve the performance of the network. Then, MSDC multi-scale depth convolution is used to capture multi-scale and multi-resolution context information. Finally, a 1×1 convolution and a BN layer are used to restore the number of channels.

[0017] Furthermore, the MSDC multi-scale depth convolution is used to capture multi-scale context information. Different-scale features are captured through three parallel branches. Each branch expands the receptive field through depth convolution, BN, and ReLU6 operations to enhance the ability to sense diverse targets. Finally, the three parallel branches are subjected to channel shuffle to rearrange the channels of the feature map.

[0018] Furthermore, the focal loss function in the Head part is expressed as:

[0019] L FL (y, p) = -α t (1 - p t ) γ log(p t )

[0020]

[0021] where y is the true value, p is the predicted value, p t is the predicted probability of the model for the correct class, and α t is a balance factor used to adjust the influence of different classes, and γ is the focusing factor.

[0022] Furthermore, preprocessing is performed on the original images collected in step 1. The preprocessing includes letterBox and normalization.

[0023] Further, the Labelimg tool is used to label the defects on the chip surface.

[0024] A method for detecting defects on the chip surface inputs the image of the chip surface to be detected into an improved YOLOv8 object detection model to identify the defects on the chip surface in the image. The beneficial effects of the present invention are as follows:

[0025] (1) The present invention uses a generative adversarial network to obtain amplified images, which not only increases the richness of training data but also improves the generalization ability of the model in few-shot learning, thus effectively alleviating the challenges brought by data scarcity; the original images and amplified images are labeled to obtain a dataset of chip surface defects, providing high-quality supervision signals for the training of the model.

[0026] (2) The present invention introduces a multi-scale convolutional attention module (MSCAM) in the Neck part of the YOLOv8 network model. The introduction of MSCAM can enhance the model's perception ability of features at different scales and improve the detection accuracy of defects under complex backgrounds.

[0027] (3) The present invention modifies the classification loss function to a focal loss function in the Head part of the YOLOv8 network model. This loss function optimizes the learning process of the model by dynamically adjusting the attention degree to easy and difficult samples, thereby improving the detection performance for a small number of defect categories and further effectively solving the problem of class imbalance. Description of the Drawings

[0028] Figure 1 is the workflow of the present invention

[0029] Figure 2 is the flowchart of the generative adversarial network amplified image of the present invention

[0030] Figure 3 is the structural diagram of the improved YOLOv8 model constructed by the present invention

[0031] Figure 4 is the schematic diagram of the MSCAM multi-scale convolutional attention module

[0032] Figure 5 is the schematic diagram of the CAB channel attention block

[0033] Figure 6 is the schematic diagram of the SAB spatial attention block

[0034] Figure 7 is the schematic diagram of the MSCB multi-scale convolutional block

[0035] Figure 8 is the schematic diagram of the MSDC multi-scale depth convolutional block Detailed Embodiments

[0036] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the present invention and are not intended to limit the present invention.

[0037] Combined with the attached Figure 1-8 , a method for constructing a small-sample chip surface defect detection model based on improved YOLOv8 in this application includes the following steps:

[0038] Step 1: Collect the original images of chip surface defects and perform preprocessing; specifically as follows:

[0039] Step 1.1: Use an industrial camera to obtain images with chip surface defects.

[0040] Step 1.2: Perform preprocessing on the images in the dataset, specifically including operations such as letterBox and normalization. The purpose of letterBox is to convert the original image size to the network input size, that is, 640*640, for subsequent batch processing and calculation. Normalization is to scale the pixel values from the range [0, 255] to the range [0, 1] or [-1, 1]. This helps to accelerate the convergence process, reduce the differences between different features, improve the stability and efficiency of training, and at the same time reduce the sensitivity of the model to the input data distribution.

[0041] Step 2: Input the preprocessed original images and random noise into a generative adversarial network (GAN) to obtain amplified images by the generative adversarial network;

[0042] More specifically, the structure of the generative adversarial network can refer to Figure 2 , including a discriminator D and a generator G. Among them, the generator G can generate corresponding images according to the input random noise and original images. The set of generated images z is denoted as G(z); the discriminator D judges whether the input image is an original image or a generated image. The set of real images x is denoted as D(x);

[0043] More specifically, in the generative adversarial network, the loss function is an important indicator for evaluating the discriminator D and the generator G. The generator G aims to generate images that are indistinguishable from real ones to deceive the discriminator D, and the discriminator D aims to accurately identify whether the input image is a real image or a generated image. The loss function of the generative adversarial network describes the game relationship between the generator G and the discriminator D. The loss function is usually expressed as:

[0044] m G inm D axV(D,G)=E x~P(x) [logD(x)]+E z~P(z) [log(1-D(G(z)))]

[0045] Among them, E x~P(x) represents the expectation of the real image x and its corresponding distribution, and E z~P(z) represents the expectation of the generated image z and its corresponding distribution. P(x) is the distribution of the real image x, and P(z) is the distribution of the generated image z. The loss function aims to make D as large as possible and G as small as possible through the game between the generator and the discriminator.

[0046] More specifically, as Figure 1 shown, the process of the generative adversarial network for amplifying image samples includes: introducing a real label while inputting noise, generating samples by the generator G. After the generated samples are recognized by the discriminator D, an initial threshold β of confidence is artificially set. Only when the confidence of the generated samples exceeds this threshold will they be expanded into the data set. Moreover, a weight coefficient is also introduced for the expanded part of the data. When the samples of this category are scarce, the weight coefficient is increased to improve the detection effect of small-sample categories. When the samples of this category are not scarce, the weight coefficient is reduced. Here, the weight coefficient is a hyperparameter used to control the relative contribution of the generated data to training.

[0047] Step 3: Label the chip surface defects in the original image and the amplified image respectively to obtain a chip surface defect data set;

[0048] More specifically, use the Labelimg tool to label the chip surface defects, that is, use a horizontal box to mark the chip surface defects and save them in the YOLO format.

[0049] Step 4: Construct an improved YOLOv8 object detection model: The improved YOLOv8 object detection model constructed in the present invention replaces all CBS modules in the Backbone and Neck parts of YOLOv8 with CBL modules. And introduce the MSCAM attention mechanism in the Neck part of the YOLOv8 neural network model to enhance the model's perception ability of features at different scales. And modify the classification loss to the focal loss function to further effectively solve the problem of class imbalance. The network structure of the improved YOLOv8 object detection model of the present invention is as Figure 3 shown, including a Backbone layer, a Neck layer, and a Head layer connected in sequence.

[0050] More specifically, the Backbone layer consists of 5 CBL network structures, 4 C2f network structures and one SPPF network structure; the specific connection relationship is as follows: the first CBL network structure, the second CBL network structure, the first C2f network structure, the third CBL network structure, the second C2f network structure, the fourth CBL network structure, the third C2f network structure, the fifth CBL network structure, the fourth C2f network structure and the SPPF network structure are connected in sequence. Among them, the first CBL network structure is connected to the Input layer, and the image is input into the object detection model by the Input layer. The second C2f network structure, the third C2f network structure and the SPPF network structure output the first feature map P1, the second feature map P2 and the third feature map P3 respectively.

[0051] More specifically, the three feature maps output by the Backbone part are: P1 = (80, 80, 256), P2 = (40, 40, 512), P3 = (20, 20, 512).

[0052] More specifically, the Neck layer consists of 2 CBL network structures, 4 C2f network structures, 2 upsamplings, 4 Concat network structures, and 3 MSCAM multi-scale convolutional attention modules; the specific connection relationship is as follows: the first upsampling, the first Concat network structure, the fifth C2f network structure, the second upsampling, the second Concat network structure, the sixth C2f network structure, the first MSCAM multi-scale convolutional attention module, the sixth CBL network structure, the third Concat network structure, the seventh C2f network structure, the second MSCAM multi-scale convolutional attention module, the seventh CBL network structure, the fourth Concat network structure, the eighth C2f network structure, and the third MSCAM multi-scale convolutional attention module are connected in sequence; among them, the first upsampling and the fourth Concat network structure are both connected to the SPPF network structure to receive the first feature map P1; the first Concat network structure is connected to the fifth C2f network structure to receive the output second feature map P2; the second Concat network structure is connected to the second C2f network structure to receive the output third feature map P3; the output features of the Neck part are respectively output by the first MSCAM multi-scale convolutional attention module, the second MSCAM multi-scale convolutional attention module, and the third MSCAM multi-scale convolutional attention module. More specifically, refer to Figure 4 The MSCAM multi-scale convolutional attention module includes a channel attention block CAB, a spatial attention block SAB and a multi-scale convolutional block MSCB connected in sequence.

[0053] More specifically, the channel attention block CAB can enhance the attention to key channels, thereby emphasizing more relevant features while suppressing less useful features. As Figure 5 shown, in the channel attention block CAB, first, global max pooling and global average pooling are respectively used to process the feature map to extract the most important features of each feature map. Then, convolution is respectively used to reduce the number of channels to 1 / 16 of the original, and then ReLU is used as the activation function to improve the speed and performance of the neural network. Next, convolution is used to restore the number of channels. Finally, after concatenation, the sigmoid activation function is used to apply weights to the original feature map, thereby achieving channel weighting.

[0054] More specifically, the spatial attention block SAB is mainly used for the regional evaluation of the feature map to help the model focus on key regions. As Figure 6 shown, in the spatial attention block SAB, max pooling and average pooling are respectively performed on the input feature map to focus on local features. Then, after concatenating the max pooling and average pooling processing structures, a large kernel convolution of 7*7 is used to emphasize these local features. Finally, the sigmoid function is used to calculate the weights, and the spatial attention map is applied to the original feature map.

[0055] More specifically, the multi-scale convolution block MSCB enables the network to effectively capture diverse features of objects through parallel convolution and feature fusion, thereby improving the overall performance of the model. As Figure 7 shown, in the MSCB, first, a 1×1 convolution is used to expand the number of channels to twice the original, then BN and ReLU6 activation functions are used to improve the performance of the network. Next, MSDC multi-scale depth convolution is used to capture multi-scale and multi-resolution context information. Finally, a 1×1 convolution and a BN layer are used to restore the number of channels.

[0056] More specifically, the MSDC multi-scale depth convolution is used to capture multi-scale context information. It captures features of different scales through three parallel branches. Each branch expands the receptive field through depth convolution, BN, and ReLU6 operations, enhancing the ability to sense diverse targets. Finally, the three parallel branches are subjected to channel shuffle to rearrange the channels of the feature map, achieving the effect of promoting information interaction between different channels.

[0057] More specifically, the working steps of the Neck part are as follows: The feature maps P1, P2, and P3 output by the feature extraction network Backbone are respectively input into the Neck part, so that feature maps of multiple scales are fused. The feature map P3 is upsampled and fused with P2 to obtain the feature map Q1. Then, Q1 is fused with P1 after passing through a C2f and an upsampling to obtain Q2. Q2 passes through the MSCAM multi-scale convolutional attention module to obtain T1. Then, T1 is fused with Q1 after passing through a CBL module to obtain Q3. Q3 passes through a C2f and the MSCAM multi-scale convolutional attention module to obtain T2. Then, T2 is fused with P3 after passing through a CBL module to obtain Q4. Finally, Q4 passes through a C2f and the MSCAM multi-scale convolutional attention module to obtain T3. At this time, the three feature maps at different levels of T1, T2, and T3 are the outputs of the entire Neck.

[0058] More specifically, the feature maps T1, T2, and T3 are respectively input into the detection head in the Head part for prediction. The classification branch in the detection head is used to detect the chip defect categories, while the regression branch is used to detect the positions of the chip defects.

[0059] More specifically, the classification loss of the detect detection head in the Head part is modified to the focal loss function FocalLoss, and its formula is as follows:

[0060] L FL (y,p)=-α t (1-p t ) γ log(p t )

[0061] where,

[0062] where, y is the true value, p is the predicted value, and p t is the predicted probability of the model for the correct class. α t is a balancing factor used to adjust the influence of different classes. γ is a modulating factor used to adjust the attention of the loss function. It can improve the model performance when dealing with the problem of class imbalance. By enhancing the weights of difficult-to-classify samples, it helps the model better learn the minority classes, thereby improving the overall detection effect.

[0063] More specifically, the CBL modules in the Backbone part and the Neck part are composed of Conv2d, BN layers, and activation layers. Among them, the activation function in the activation layer is the Leaky ReLU activation function. Leaky ReLU outputs a small linear value for negative inputs instead of completely outputting zero, which can avoid the occurrence of the "dead neuron" phenomenon. By maintaining a small output for negative inputs, Leaky ReLU can enhance the ability of feature learning, especially in deep neural networks. In addition, the sparse activation characteristic of Leaky ReLU helps reduce the risk of overfitting, thereby improving the generalization ability of the model. The expression of the Leaky ReLU function is as follows:

[0064]

[0065] Step 5. After the construction of the improved YOLOv8 object detection model is completed, it is necessary to train, validate, and test the built improved YOLOv8 object detection model.

[0066] Step 5.1. Divide the chip surface defect dataset obtained in Step 3 according to the ratio of (training set + validation set): test set = 9:1, and training set: validation set = 9:1.

[0067] Step 5.2. Use the divided dataset to train the improved YOLOv8 object detection model constructed in Step 4.

[0068] Before training, the epoch is set to 300. The backbone feature extraction network is frozen in the first 50 epochs and the initial learning rate is set to 0.001. When training the model, the preprocessed training data is input into the YOLOv8 model. During the training process, the optimizer updates the weights of the model through the backpropagation algorithm to minimize the loss function.

[0069] Step 5.3. Use the divided dataset to validate the improved YOLOv8 object detection model constructed in Step 4.

[0070] Set the model weights to best.pt after the training in Step 5.2 is completed, and use the test set divided in Step 5.1 to validate the object detection model.

[0071] Based on the improved YOLOv8 object detection model constructed by the above method, the present invention also proposes a method for detecting chip surface defects. The surface image of the chip to be detected is input into the improved YOLOv8 object detection model to realize the recognition of chip surface defects in the image. The above embodiments are only used to illustrate the design concept and characteristics of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A method for constructing a small sample chip surface defect detection model based on improved YOLOv8, characterized in that: The following steps are involved: Step 1, collecting original images of chip surface defects; Step 2, inputting the preprocessed original image and random noise into the Generative Adversarial Network (GAN), and obtaining the augmented image by the Generative Adversarial Network; Step 3, annotating the chip surface defects in the original image and the amplified image respectively to obtain a chip surface defect dataset; Step 4, construct an improved YOLOv8 target detection model, the improved YOLOv8 target detection model includes a Backbone layer, a Neck layer, and a Head layer; wherein the Backbone layer is connected in sequence by a first CBL network structure, a second CBL network structure, a first C2f network structure, a third CBL network structure, a second C2f network structure, a fourth CBL network structure, a third C2f network structure, a fifth CBL network structure, a fourth C2f network structure and an SPPF network structure, and the second C2f network structure, the third C2f network structure and the SPPF network structure output a first feature map P1, a second feature map P2 and a third feature map P respectively; the Neck layer is composed of a first upsampling, a first Concat network structure, a fifth C2f network structure, a second upsampling, a second Concat network structure, a sixth C2f network structure, The first MSCAM multi-scale convolutional attention module, the sixth CBL network structure, the third Concat network structure, the seventh C2f network structure, the second MSCAM multi-scale convolutional attention module, the seventh CBL network structure, the fourth Concat network structure, the eighth C2f network structure, and the third MSCAM multi-scale convolutional attention module are connected in sequence; wherein, the first upsampling and the fourth Concat network structure are connected with the SPPF network structure to receive the first feature map P1; the first Concat network structure is connected with the fifth C2f network structure to receive the output second feature map P2; the second Concat network structure is connected with the second C2f network structure to receive the output third feature map P3; the output features of the Neck layer are output by the first MSCAM multi-scale convolutional attention module, the second MSCAM multi-scale convolutional attention module, and the third MSCAM multi-scale convolutional attention module respectively; The Head layer receives the input of the Neck layer and outputs the detection result. The Head layer adopts the focal loss function; Step 5: After completing the construction of the improved YOLOv8 target detection model, it is necessary to train, verify and test the built improved YOLOv8 target detection model.

2. According to claim 1, a method for constructing a small sample chip surface defect detection model based on improved YOLOv8 is characterized in that: The MSCAM multi-scale convolutional attention module includes a channel attention block CAB, a spatial attention block SAB and a multi-scale convolution block MSCB which are connected in sequence.

3. According to claim 2, a method for constructing a small sample chip surface defect detection model based on improved YOLOv8 is characterized in that: In the channel attention block CAB, the feature map is processed using global maximum pooling and global average pooling respectively to extract the most important features of each feature map; Then, convolution is used to reduce the number of channels to 1 / 16 of the original number, and then ReLU is used as the activation function to improve the speed and performance of the neural network. Convolution is then used to restore the number of channels. Finally, after splicing, the sigmoid activation function is used to apply the weights to the original feature map to achieve channel weighting.

4. The method for constructing a small sample chip surface defect detection model based on improved YOLOv8 according to claim 2, characterized in that: In the spatial attention block SAB, the input feature map is subjected to maximum pooling and average pooling respectively to focus on local features. After the maximum pooling and average pooling processing structures are concatenated, a 7*7 large kernel convolution is used to emphasize these local features. Finally, the sigmoid function is used to calculate the weights and the spatial attention map is applied to the original feature map.

5. The method for constructing a small sample chip surface defect detection model based on improved YOLOv8 according to claim 2, characterized in that: In the multi-scale convolution block MSCB, a 1×1 convolution is first used to expand the number of channels to twice the original number, then BN and ReLU6 activation functions are used to improve the performance of the network, and then MSDC multi-scale deep convolution is used to capture multi-scale and multi-resolution contextual information; finally, a 1×1 convolution and BN layer are used to restore the number of channels.

6. The method for constructing a small sample chip surface defect detection model based on improved YOLOv8 according to claim 5, characterized in that: The MSDC multi-scale deep convolution is used to capture multi-scale context information. Features of different scales are captured through three parallel branches. Each branch expands the receptive field through deep convolution, BN and ReLU6 operations to enhance the perception of diverse targets. Finally, the three parallel branches are shuffled to rearrange the channels of the feature map.

7. The method for constructing a small sample chip surface defect detection model based on improved YOLOv8 according to claim 1, characterized in that: The focal loss function of the Head part is expressed as: L FL (y,p)=-α t (1-p t ) γ log(p t ) in, Among them, y is the true value, p is the predicted value, and p t is the model’s predicted probability for the correct category, α t is a balancing factor used to adjust the influence of different categories, and γ is a focusing factor.

8. The method for constructing a small sample chip surface defect detection model based on improved YOLOv8 according to claim 1, characterized in that: The original image collected in step 1 is preprocessed, and the preprocessing includes letterBox and normalization processing.

9. The method for constructing a small sample chip surface defect detection model based on improved YOLOv8 according to claim 1, characterized in that: Use Labelimg tool to mark chip surface defects.

10. A chip surface defect detection method, characterized in that: The chip surface image to be detected is input into the improved YOLOv8 target detection model constructed by the small sample chip surface defect detection model construction method based on improved YOLOv8 as described in claim 1 to identify the chip surface defects in the image.

Citation Information

Patent Citations

  • Small sample rolling bearing fault diagnosis method based on convolutional transformer generative adversarial network

    CN115859142A

  • Small sample chip appearance defect detection method and detection system based on improved YOLOv5

    CN116228740A

  • Strip steel surface small target defect detection method and system based on super-resolution and YOLOv8

    CN116630301A

  • PCB surface defect detection method based on improved YOLOv8 algorithm

    CN117422681A

  • PCB image defect detection method based on improved YOLOv8 network model

    CN117455889A