Lightweight model-based chip appearance defect detection method, system and terminal

By adopting the improved YOLOv8 network model in chip appearance defect detection, combining the convolution and attention fusion module, efficient fusion attention C2f module and CA attention mechanism module, the problems of low manual detection efficiency and computer vision technology affected by motion blur are solved, and efficient and accurate chip appearance defect detection is achieved.

CN120182207APending Publication Date: 2025-06-20SIEN (QINGDAO) INTEGRATED CIRCUITS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510252211.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, manual detection chip appearance defects are inefficient and costly, and computer vision technology is affected by motion blur during the detection process, resulting in a decrease in detection accuracy.

Method used

Using a chip appearance defect detection method based on lightweight model, images are preprocessed through motion fuzzy functions, data sets marked with defect type tags are constructed, and improved YOLOv8 network model is trained to build a lightweight detection model. The improved model introduces convolution and attention fusion modules in the backbone layer feature extraction network, designs efficient attention C2f modules in the neck feature fusion network, and embeds the CA attention mechanism module, and finally designs a lightweight detection head in the detection network.

Benefits of technology

The model parameter quantity and inference floating point number are reduced, and both accuracy and lightweight are taken into account, which solves the problem of blurred and loss of chip appearance defects caused by motion blur, and improves the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182207A_ABST
    Figure CN120182207A_ABST
Patent Text Reader

Abstract

The invention provides a chip appearance defect detection method, system and terminal based on a lightweight model, and the method comprises the steps: firstly carrying out the preprocessing of a chip appearance defect image through a motion blur function, constructing a chip appearance defect image data set, training an improved YOLOv8 network model through the chip appearance defect image data set, and constructing a detection model, and performing chip appearance defect type detection. According to the invention, a YOLOv8 network model is improved, a convolution and attention fusion module is introduced into a trunk layer feature extraction part, a lightweight module is designed, and an efficient fusion attention C2f module is designed in a neck feature fusion network. Meanwhile, a CA attention mechanism module and a trunk layer convolution and attention fusion module are embedded to carry out multi-attention mechanism fusion, and a lightweight detection head is self-designed in the detection network. According to the method, the model parameter quantity and the reasoning floating-point number are reduced, and both precision and light weight are considered. And meanwhile, the problems of fuzzy and lost appearance defect characteristics of the chip caused by motion blur are solved, and the accuracy and reliability of detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semiconductor integrated circuits, and particularly to a chip appearance defect detection method, system and terminal based on a lightweight model. Background Art

[0002] With the acceleration of the global semiconductor integrated circuit process, the production volume of chips has increased sharply. However, there will be contact between products during the chip production process and the transportation process of finished chips, resulting in problems of appearance defects in the chips. Appearance defects include stains, scratches, pin defects, etc. Traditional chip appearance defect detection mainly relies on manual inspection. However, manual inspection has problems of low efficiency and high detection cost.

[0003] Computer vision technology is increasingly used in the intelligent detection and recognition process of chip appearance defects. However, there are many factors affecting the accuracy and efficiency of intelligent detection of chip appearance defects. Existing chip appearance defect detection equipment has strict requirements for the lightweight of the detection model. In the actual working environment, the detection process will be affected by motion blur, resulting in the blurring and loss of surface features of chip appearance defects. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a chip appearance defect detection method, system and terminal based on a lightweight model, which is used to solve the technical problems of low detection efficiency and high detection cost of manual detection of chip appearance defects in the prior art, and the reduction of detection accuracy due to the influence of motion blur in the detection process of using computer vision technology to detect chip appearance defects.

[0005] To achieve the above purpose and other related purposes, the present invention provides a chip appearance defect detection method based on a lightweight model. The method includes: preprocessing each collected chip appearance defect image using a motion blur function, and constructing a chip appearance defect image data set labeled with defect type labels; training an improved YOLOv8 network model using the chip appearance defect image data set to construct a lightweight chip appearance defect detection model. Among them, the improved YOLOv8 network model includes: a backbone feature extraction network, a neck feature fusion network, and a detection network; a convolution and attention fusion module is introduced in the backbone feature extraction network and a lightweight C2f module is designed; an efficient fusion attention C2f module is designed in the neck feature fusion network, and a CA attention mechanism module is embedded for multi-attention mechanism fusion with the convolution and attention fusion module; a lightweight detection head is designed in the detection network; based on the lightweight chip appearance defect detection model, obtaining the corresponding chip appearance defect type result according to the chip appearance defect image to be detected.

[0006] In an embodiment of the present invention, the backbone layer feature extraction network is constructed as follows: all C2f modules in the backbone layer feature extraction network of the original YOLOv8 network model are replaced with the designed lightweight C2f module, and the convolution and attention fusion module is connected after the pyramid pooling layer.

[0007] In an embodiment of the present invention, a Ghost bottleneck layer module is introduced into the designed lightweight C2f module; wherein, the Ghost bottleneck layer module includes: a Ghost convolutional layer, a batch normalization layer, and a SiLU activation function.

[0008] In an embodiment of the present invention, the convolution and attention fusion module includes: a global branch and a local branch; wherein, the global branch contains four branches, and the input chip appearance defect image enters the second, third, and fourth branches respectively after passing through a 1×1 convolution. The second and third branches are multiplied after passing through a 3×3 convolution, and then processed by the Softmax function to form the first chip appearance defect attention feature map; the fourth branch is multiplied by the first chip appearance defect attention feature map after passing through a 3×3 convolution to form the second chip appearance defect attention feature map; the second chip appearance defect attention feature map is added to the input chip appearance defect image of the first branch after passing through a 1×1 convolution to output the third chip appearance defect attention feature map; the local branch passes the input image through a 1×1 convolution and a channel transformation, and then passes through a 3×3 convolution to form the first appearance defect feature map; the convolution and attention fusion module adds the third chip appearance defect attention feature map output by the global branch and the first appearance defect feature map output by the local branch to output the chip appearance defect image feature.

[0009] In an embodiment of the present invention, the construction method of the neck feature fusion network includes: replacing all C2f modules in the neck feature fusion network of the original YOLOv8 network model with the designed efficient fusion attention C2f module, and embedding a CA attention mechanism module between the efficient fusion attention C2f module and the lightweight detection head.

[0010] In an embodiment of the present invention, an efficient fusion attention module is introduced into the efficient fusion attention C2f module; wherein, the efficient fusion attention module is used to process the input chip appearance defect image features through a 1*1 convolution and a Sigmoid activation function to form first chip appearance defect features, and perform a dot product operation on the first chip appearance defect features and the input chip appearance defect image features to form second chip appearance defect features; the input chip appearance defect image features are processed through global average pooling, a 1*1 convolution and a Sigmoid activation function to form third chip appearance defect features, and perform a dot product operation on the third chip appearance defect features and the input chip appearance defect image features to form fourth chip appearance defect features; the second chip appearance defect features and the fourth chip appearance defect features are added to output the final chip appearance defect features.

[0011] In an embodiment of the present invention, the lightweight detection head consists of two branches; wherein, the first branch consists of a convolutional layer and a regression loss layer; the second branch consists of a convolutional layer and a class loss layer.

[0012] In an embodiment of the present invention, the Powerful-IoU loss function is used as the loss function when training the improved YOLOv8 network model.

[0013] To achieve the above object and other related objects, the present invention provides a chip appearance defect detection system based on a lightweight model, the system includes: a preprocessing module, configured to preprocess each collected chip appearance defect image by using a motion blur function and construct a chip appearance defect image dataset labeled with defect type labels; a model construction module, connected to the preprocessing module, configured to train an improved YOLOv8 network model by using the chip appearance defect image dataset to construct a lightweight chip appearance defect detection model; wherein, the improved YOLOv8 network model includes: a backbone feature extraction network, a neck feature fusion network, and a detection network; a convolution and attention fusion module is introduced in the backbone feature extraction network and a lightweight C2f module is designed; an efficient fusion attention C2f module is designed in the neck feature fusion network, and a CA attention mechanism module is embedded to perform multi-attention mechanism fusion with the convolution and attention fusion module; a lightweight detection head is designed in the detection network; a defect detection module, connected to the model construction module, based on the lightweight chip appearance defect detection model, obtains the corresponding chip appearance defect type result according to the chip appearance defect image to be detected.

[0014] To achieve the above and other related objectives, the present invention provides an electronic terminal, including: one or more memories and one or more processors; the one or more memories are used for storing computer programs; the one or more processors are connected to the memories and are used for running the computer programs to execute the chip appearance defect detection method based on the lightweight model.

[0015] As described above, the present invention is a chip appearance defect detection method, system and terminal based on a lightweight model, having the following beneficial effects: The present invention first preprocesses the chip appearance defect image by applying a motion blur function to construct a chip appearance defect image dataset, and then uses the chip appearance defect image dataset to train an improved YOLOv8 network model to construct a detection model for detecting the types of chip appearance defects. The present invention improves the YOLOv8 network model, introduces a convolution and attention fusion module in the backbone layer feature extraction part and designs a lightweight module, designs an efficient fusion attention C2f module in the neck feature fusion network, and at the same time embeds a CA attention mechanism module to perform multi-attention mechanism fusion with the backbone layer convolution and attention fusion module, and designs a lightweight detection head in the detection network by itself. The present invention not only reduces the model parameter quantity and inference floating point number, taking into account both accuracy and lightweight. At the same time, it solves the problems of fuzzy and lost chip appearance defect features caused by motion blur, and improves the accuracy and reliability of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It shows a schematic flow chart of the chip appearance defect detection method based on the lightweight model in an embodiment of the present invention.

[0017] Figure 2 It shows a schematic structural diagram of the improved YOLOv8 network model in an embodiment of the present invention.

[0018] Figure 3 It shows a schematic structural diagram of the lightweight C2f module in an embodiment of the present invention.

[0019] Figure 4 It shows a schematic structural diagram of the convolution and attention fusion module in an embodiment of the present invention.

[0020] Figure 5 It shows a schematic structural diagram of the efficient fusion attention module in an embodiment of the present invention.

[0021] Figure 6 It shows a schematic structural diagram of the efficient fusion attention C2f module in an embodiment of the present invention.

[0022] Figure 7 It shows a schematic structural diagram of the CA attention mechanism module in an embodiment of the present invention.

[0023] Figure 8 It shows a schematic structural diagram of the lightweight detection head in an embodiment of the present invention.

[0024] Figure 9 It shows a schematic flow diagram of the chip appearance defect detection method based on the lightweight model in an embodiment of the present invention.

[0025] Figure 10 It shows a schematic structural diagram of the chip appearance defect detection system based on the lightweight model in an embodiment of the present invention.

[0026] Figure 11 It shows a schematic structural diagram of the electronic terminal in an embodiment of the present invention. Detailed implementation manners

[0027] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0028] It should be noted that in the following description, reference is made to the accompanying drawings, which describe several embodiments of the present invention. It should be understood that other embodiments can also be used, and mechanical composition, structure, electrical, and operational changes can be made without departing from the spirit and scope of the present invention. The following detailed description should not be considered restrictive, and the scope of the embodiments of the present invention is only defined by the claims of the published patent. The terms used here are only for describing specific embodiments and are not intended to limit the present invention. Spatially related terms, such as "upper", "lower", "left", "right", "below", "beneath", "lower part", "above", "upper part", etc., can be used in the text to facilitate the description of the relationship between one element or feature shown in the figure and another element or feature.

[0029] Throughout the specification, when it is said that a part is "connected" to another part, this includes not only the case of "direct connection", but also the case of "indirect connection" with other elements placed in between. In addition, when it is said that a certain part "includes" a certain constituent element, unless there is a particularly contrary record, it does not exclude other constituent elements, but means that other constituent elements can also be included.

[0030] The first, second, third, etc. terms mentioned herein are used to illustrate various parts, components, regions, layers, and / or segments, but are not limited thereto. These terms are only used to distinguish a part, component, region, layer, or segment from other parts, components, regions, layers, or segments. Therefore, the first part, component, region, layer, or segment described below may refer to the second part, component, region, layer, or segment without departing from the scope of the present invention.

[0031] Furthermore, as used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should be further understood that the terms "comprising", "including" indicate the presence of the stated features, operations, elements, components, items, kinds, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms "or" and "and / or" used herein are interpreted as inclusive, meaning any one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B, and C". An exception to this definition only occurs when the combination of elements, functions, or operations is inherently mutually exclusive in some way.

[0032] In view of the advantages of the YOLOv8 network model, such as high detection accuracy, fast efficiency, and easy deployment, the present invention provides a method for detecting chip appearance defects based on a lightweight model. First, a motion blur function is applied to preprocess the chip appearance defect image to construct a chip appearance defect image dataset. Then, the improved YOLOv8 network model is trained using the chip appearance defect image dataset to construct a detection model for detecting the types of chip appearance defects. The present invention improves the YOLOv8 network model by introducing a convolution and attention fusion module in the backbone layer feature extraction part and designing a lightweight module, designing an efficient fusion attention C2f module in the neck feature fusion network, and simultaneously embedding a CA attention mechanism module for multi-attention mechanism fusion with the backbone layer convolution and attention fusion module. A lightweight detection head is designed in the detection network. The present invention not only reduces the model parameter quantity and inference floating-point number, taking into account both accuracy and lightweight. At the same time, it solves the problem of fuzzy and lost features of chip appearance defects caused by motion blur, improving the accuracy and reliability of detection.

[0033] The following will be described in detail with reference to the accompanying drawings for the embodiments of the present invention, so that those skilled in the technical field of the present invention can easily implement it. The present invention can be embodied in many different forms and is not limited to the embodiments described herein.

[0034] As Figure 1Schematic flowchart showing a method for detecting chip appearance defects based on a lightweight model in an embodiment of the present invention.

[0035] The method includes:

[0036] Step S1: Preprocess each collected chip appearance defect image using a motion blur function, and construct a chip appearance defect image dataset labeled with defect type tags.

[0037] In one embodiment, first, collect multiple chip appearance defect images involving various defect categories, including but not limited to smudges, scratches, pin defects, etc. Then, in order to better reflect the actual chip appearance defect detection scenario and improve the robustness and generalization ability of the model in actual applications, preprocess each collected chip appearance defect image using a motion blur function;

[0038] The applied motion blur function is:

[0039]

[0040] In the formula: L is the blur length, is the blur angle or direction, and x, y are the pixel positions of the chip appearance defect dataset image.

[0041] Through this function, we can effectively apply different degrees of blur effects to the image, thus simulating the detection challenges in real scenarios and further improving the performance of the model in actual applications.

[0042] After the preprocessing is completed, label the preprocessed chip appearance defect images with defect type tags according to their corresponding defect types, and divide the labeled chip appearance defect images into a training set, a validation set, and a test set to construct a chip appearance defect image dataset. Preferably, use LabelImg to label the defect type tags of the preprocessed chip appearance defect images and make them into a VOC format dataset. And divide them into a training set, a validation set, and a test set according to the set data volume ratio.

[0043] Step S2: Use the chip appearance defect image dataset to train an improved YOLOv8 network model to construct a lightweight chip appearance defect detection model.

[0044] Specifically, use the training set, validation set, and test set in the chip appearance defect image dataset to train, validate, and test the improved YOLOv8 network model. After the model is trained and validated iteratively, an optimal weight is obtained. Load the obtained optimal weight into the improved YOLOv8 network model to obtain a lightweight chip appearance defect detection model.

[0045] To better describe the improvement of the YOLOv8 network model, the specific structure of the improved YOLOv8 network model will be described in combination with the following specific embodiments.

[0046] As Figure 2 , the improved YOLOv8 network model includes: a backbone feature extraction network, a neck feature fusion network, and a detection network; a convolution and attention fusion module is introduced in the backbone feature extraction network and a lightweight C2f module is designed; an efficient fusion attention C2f module is designed in the neck feature fusion network, and a CA attention mechanism module is embedded to perform multi-attention mechanism fusion with the convolution and attention fusion module; a lightweight detection head is designed in the detection network.

[0047] In one embodiment, in the original YOLOv8 network model, the backbone feature extraction network is usually composed of a series of convolutional layers, C2f modules, and pooling layers. In the improved YOLOv8 network model, the backbone feature extraction network is composed of a series of convolutional layers, a newly designed lightweight C2f module, pooling layers, and an additionally introduced convolution and attention fusion module.

[0048] Based on the backbone feature extraction network of the original YOLOv8 network model, all C2f modules in the original backbone feature extraction network are replaced with the designed lightweight C2f module, and the convolution and attention fusion module is connected after the pyramid pooling layer of the original backbone feature extraction network, so as to construct the backbone feature extraction network in the improved YOLOv8 network model.

[0049] Specifically, the backbone feature extraction network of the improved YOLOv8 network model has 12 layers. The first layer is the input image layer for inputting the chip appearance defect image, the second layer is convolutional layer 1, the third layer is convolutional layer 2, the fourth layer is lightweight C2f module 1, the fifth layer is convolutional layer 3, the sixth layer is lightweight C2f module 2, and the chip appearance defect feature Feature1 is output. The seventh layer is convolutional layer 4, the eighth layer is lightweight C2f module 3, and the chip appearance defect feature Feature2 is output. The ninth layer is convolutional layer 5, the tenth layer is lightweight C2f module 4, the eleventh spatial pyramid pooling layer, and the twelfth layer is the convolution and attention fusion module, and the chip appearance defect feature Feature3 is output.

[0050] In one embodiment, a Ghost bottleneck layer module is introduced in the designed lightweight C2f module; wherein, the Ghost bottleneck layer module includes: a Ghost convolutional layer, a batch normalization layer, and a SiLU activation function.

[0051] Specifically, the difference between the lightweight C2f module in the backbone feature extraction network of the improved YOLOv8 network model and the C2f module in the backbone feature extraction network of the original YOLOv8 network is that the C2f module in the backbone feature extraction network of the original YOLOv8 network uses a traditional bottleneck module, which is generally composed of multiple conventional convolutional layers combined in a specific structure. The improved lightweight C2f module adopts a Ghost bottleneck layer module. The Ghost bottleneck layer module includes: a Ghost convolutional layer, a batch normalization layer, and a SiLU activation function. Through a unique structural combination, the Ghost bottleneck layer module significantly reduces the computational amount and memory occupancy while ensuring the feature extraction ability. This design enables the model to process image data more efficiently while maintaining a high detection performance.

[0052] It should be noted that according to the requirements of the actual task, multiple Ghost bottleneck layer modules can be introduced into the lightweight C2f module. Taking Figure 3 the introduction of two Ghost bottleneck layer modules as an example, the stacking of multiple modules can further enhance the feature extraction ability.

[0053] In one embodiment, the convolution and attention fusion module introduced in the backbone feature extraction network includes: a global branch and a local branch; it aims to process the input image through different branches, extract features at different levels, and finally fuse these features to output the chip appearance defect image features, helping the model to more accurately detect chip appearance defects.

[0054] Among them, as Figure 4, the global branch contains four branches, and the input chip appearance defect images are respectively subjected to feature extraction and transformation through the four branches; the input chip appearance defect images first pass through a 1*1 convolution, and the function of this step is to adjust the number of channels of the input chip appearance defect images, reduce the subsequent calculation amount, and at the same time perform a preliminary linear combination of the features. Then, the convolved images enter the second branch, the third branch, and the fourth branch respectively. Both the second branch and the third branch perform 3*3 convolution operations, which can extract the local feature information of the images. The results of the convolution of these two branches are subjected to a dot product operation, and the dot product can highlight the parts where the feature values are both large in the two branches and suppress the parts with small feature values. After that, it is processed by the Softmax function, and the Softmax function converts the dot product result into a probability distribution, so that the value of each element is between 0 and 1, and the sum of all elements is 1, thus forming the first chip appearance defect attention feature map. The fourth branch also performs 3*3 convolution, and after extracting the features, it performs a dot product with the first chip appearance defect attention feature map. This step combines the features extracted by the fourth branch and the attention information, highlights the features at important positions, and forms the second chip appearance defect attention feature map. The second chip appearance defect attention feature map passes through a 1*1 convolution to further adjust the number of channels and perform a linear combination of the features. Then it is added to the input chip appearance defect image of the first branch, integrating the attention information into the original image features, and outputting the third chip appearance defect attention feature map.

[0055] As Figure 4 , the local branch mainly focuses on the local features of the image. The input chip appearance defect image first passes through a 1*1 convolution to adjust the number of channels, then undergoes a channel transformation, and finally passes through a 3*3*3 convolution to extract local features, forming the first appearance defect feature map.

[0056] The convolution and attention fusion module adds the third chip appearance defect attention feature map output by the global branch and the first appearance defect feature map output by the local branch. This fusion method combines the global attention information and the local feature information, so that the output chip appearance defect image features contain both the important information of the whole image and the local detailed features, which helps to improve the detection accuracy of the model for chip appearance defects.

[0057] In one embodiment, in the original YOLOv8 network model, the neck feature fusion network is usually composed of a series of convolutional layers, C2f modules, sampling layers, and connection layers. In the improved YOLOv8 network model, the neck feature fusion network is composed of a series of convolutional layers, a newly designed efficient fusion attention C2f module, sampling layers, connection layers, and a newly introduced CA attention mechanism module.

[0058] Based on the neck feature fusion network of the original YOLOv8 network model, replace all C2f modules in the neck feature fusion network of the original YOLOv8 network model with the designed efficient fusion attention C2f module, and embed a CA attention mechanism module between the efficient fusion attention C2f module and the detection head, then the neck feature fusion network of the improved YOLOv8 network model can be constructed.

[0059] Specifically, the chip appearance defect feature Feature3 output by the convolution and attention fusion module is upsampled and then concatenated with the output feature Feature2 to form the output chip appearance defect feature Feature4. The output chip appearance defect feature Feature4 is output as the chip appearance defect feature Feature5 after passing through the efficient fusion attention C2f module 1, then is upsampled, and then concatenated with the output chip appearance defect feature Feature1 to form the output chip appearance defect feature Feature6. The output chip appearance defect feature Feature6 enters the lightweight detection head 1 (small target detection layer) of the detection network after passing through the efficient fusion attention C2f module 2 and the CA attention mechanism module 1. At the same time, the output chip appearance defect feature Feature6 is concatenated with the output chip appearance defect feature Feature5 after passing through the efficient fusion attention C2f module 2 and convolution 6 to form the output chip appearance defect feature Feature7. The output chip appearance defect feature Feature7 enters the lightweight detection head 2 (medium target detection layer) of the detection network after passing through the efficient fusion attention C2f module 3 and the CA attention mechanism block 2. At the same time, the output chip appearance defect feature Feature7 is concatenated with the output chip appearance defect feature Feature3 after passing through the efficient fusion attention C2f module 3 and convolution 7 to form the output chip appearance defect feature Feature8. The output chip appearance defect feature Feature8 enters the lightweight detection head (large target detection layer) of the detection network after passing through the efficient fusion attention C2f module 4 and the CA attention mechanism block 3. Among them, Convolution 1 to Convolution 7 are all composed of ordinary convolution (Conv2d), BN batch normalization layer, and SiLU activation function.

[0060] In one embodiment, an efficient fusion attention module is introduced into the efficient fusion attention C2f module. Specifically, the difference between the efficient fusion attention C2f module of the neck feature fusion network of the improved YOLOv8 network model and the C2f module of the YOLOv8 network model of the original YOLOv8 network is that the C2f module in the backbone feature extraction network of the original YOLOv8 network uses a traditional bottleneck module, which is generally composed of multiple convolutional layers combined in a specific structure. The improved lightweight C2f module adopts a newly designed efficient fusion attention module. This module can use the attention mechanism to perform more detailed weighting and fusion of features during the feature fusion process of the traditional C2f module, thereby enhancing the network's ability to capture chip appearance defect features. Compared with the original C2f module, it can process and fuse features more effectively and improve the performance of the model in the chip appearance defect detection task.

[0061] Among them, as Figure 5 , the efficient fusion attention module is used to form the first chip appearance defect feature after processing the input chip appearance defect image feature through a 1*1 convolution and a Sigmoid activation function, and perform a dot product operation between the first chip appearance defect feature and the input chip appearance defect image feature to form the second chip appearance defect feature; the input chip appearance defect image feature is processed through global average pooling, a 1*1 convolution and a Sigmoid activation function to form the third chip appearance defect feature, and a dot product operation is performed between the third chip appearance defect feature and the input chip appearance defect image feature to form the fourth chip appearance defect feature; the second chip appearance defect feature and the fourth chip appearance defect feature are added to output the chip appearance defect feature.

[0062] It should be noted that according to the requirements of the actual task, multiple efficient fusion attention modules can be introduced into the efficient fusion attention C2f module. Taking Figure 6 the introduction of two efficient fusion attention modules as an example, it can further enhance the model's ability to capture image features, especially when dealing with complex or blurred chip appearance defect images, it can significantly improve the detection accuracy and robustness.

[0063] In one embodiment, a CA attention mechanism module is embedded between the efficient fusion attention C2f module and the detection head. By embedding the CA attention mechanism module between the efficient fusion attention C2f module and the detection head, and performing multi-attention mechanism fusion with the convolution and attention fusion module in the improved YOLOv8 network model.

[0064] As Figure 7, the CA attention mechanism module first takes into account the feature information of the chip appearance defect image features in the width and height directions, then divides them in the height and width directions respectively, then performs Concat connection and convolution in the spatial dimension, then performs batch normalization processing and non-linear functions. The image feature information after batch normalization and non-linear functions is then processed through convolution transformation and Sigmoid activation function to obtain the attention weights in the height and width directions. Finally, through weighted multiplication calculation on the input chip appearance defect feature map, the attention weights in the height and width directions are combined into a weight matrix.

[0065] Since the chip appearance defects have a specific spatial distribution pattern, the CA attention mechanism module can focus on the feature information of the feature map in the width and height directions. By generating spatial attention weights, the model can more accurately capture the specific position of the defects in the image, thereby improving the ability to capture spatial features. The CA attention mechanism module is fused with the convolution and attention fusion module in the improved YOLOv8 network model. The convolution and attention fusion module focuses on feature enhancement and information fusion in the channel dimension, while the CA attention mechanism module focuses on feature capture in the spatial dimension. The two are combined to process and enhance the chip appearance defect image features from different angles.

[0066] In one embodiment, as Figure 2 , 3 lightweight detection heads are designed in the detection network;

[0067] As Figure 8 , the lightweight detection head consists of two branches;

[0068] The first branch consists of a convolution layer and a regression loss layer; the convolution layer is a commonly used feature extraction layer in neural networks. In the first branch, the convolution layer further extracts and transforms the input feature map. It performs convolution operations by sliding the convolution kernel on the feature map, and can capture the local feature information in the feature map. By adjusting parameters such as the size, number, and stride of the convolution kernel, features of different scales and levels can be extracted, providing a more representative feature representation for subsequent regression tasks. The regression loss layer is mainly used to process the position information of the target. In chip appearance defect detection, it is necessary to accurately locate the position of the defect. The regression loss layer will predict the bounding box information of the defect, such as the center coordinates, width, and height of the bounding box, based on the features output by the convolution layer. Then, the loss between the predicted value and the true value is calculated.

[0069] The second branch consists of a convolutional layer and a class loss layer; similar to the convolutional layer of the first branch, the convolutional layer of the second branch also performs feature extraction and transformation on the input feature map. It extracts feature information related to the defect type, and through different combinations of convolutional kernels, it mines the features that can distinguish different types of chip appearance defects. The class loss layer is used to process the classification information of the target. In chip appearance defect detection, it is necessary to determine which type the defect belongs to, such as scratches, cracks, holes, etc. The class loss layer will classify and predict the defects based on the features output by the convolutional layer. Commonly used classification loss functions include cross-entropy loss, etc. By calculating the loss between the predicted class probability distribution and the true class label, the accuracy of classification is measured, and the model is guided to learn more effective classification features to improve the accuracy of defect classification.

[0070] This lightweight detection head focuses on the localization regression of the target through the first branch, while the second branch focuses on the classification of the target. This clear division of labor enables the model to process different types of information more targeted, avoiding the interference that may occur when simultaneously processing localization and classification tasks in the same module, and enabling each task to be processed more precisely. The convolutional layers in the two branches can extract features according to their respective task requirements. Since the two branches process different tasks respectively, the structure of each branch can be relatively simplified, avoiding the design of an overly complex integrated module. This lightweight design can reduce the number of model parameters and computational volume, reduce the demand for computing resources, improve the running efficiency of the model, and make it more suitable for deployment on resource-constrained devices.

[0071] In one embodiment, when training the improved YOLOv8 network model, the loss function uses the Powerful-IoU loss function. Powerful-IoU is an improved version of the IoU (Intersection over Union) loss function, aiming to better handle the prediction problem of bounding boxes, especially for those defect detection tasks with complex shapes and boundaries. Powerful-IoU usually combines multiple loss components, such as CIoU (Complete IoU), DIoU (Distance IoU), and GIoU (Generalized IoU), to provide a more comprehensive loss metric. Using the Powerful-IoU loss function when training the improved YOLOv8 network model in the present invention can improve the detection accuracy and convergence speed of chip appearance defects.

[0072] Step S3: Based on the lightweight chip appearance defect detection model, obtain the corresponding chip appearance defect type result according to the chip appearance defect image to be detected.

[0073] In one embodiment, inputting the chip appearance defect image to be detected into the lightweight chip appearance defect detection model can obtain the corresponding chip appearance defect type result.

[0074] To better describe the above chip appearance defect detection method based on the lightweight model, the following specific embodiments are now combined for illustration.

[0075] Embodiment 1: A chip appearance defect detection method based on a lightweight model. Figure 9 This is a schematic flowchart of the chip appearance defect detection method based on the lightweight model in this embodiment.

[0076] The method includes:

[0077] Step 1: Collect images of the chip appearance defect dataset. The defects include three types: stains, scratches, and pin defects. There are a total of 4000 image samples in the chip appearance defect dataset. Apply the motion blur function to preprocess the images of the chip appearance defect dataset. The motion blur function:

[0078]

[0079] In the formula: L is the blur length, is the blur angle or direction, and x, y are the pixel positions of the images in the chip appearance defect dataset.

[0080] Step 2: Use LabelImg to label the preprocessed chip appearance defect dataset, make it into a VOC format dataset, and divide it into a training set, a validation set, and a test set according to the ratio of 8:1:1. After division, there are 3200 chip appearance defect images in the training set, 400 chip appearance defect images in the validation set, and 400 chip appearance defect images in the test set.

[0081] Step 3: Design a lightweight chip appearance defect detection model. Introduce a convolution and attention fusion module (CAFM) into the backbone layer feature extraction network of the model and design a lightweight module; design an efficient fusion attention C2f module in the neck feature fusion network, and at the same time embed the CA attention mechanism module to perform multi-attention mechanism fusion with the convolution and attention fusion module in the backbone layer; in the detection head network, design a lightweight detection head by itself.

[0082] Step 4: Apply the designed lightweight chip appearance defect detection model for training and validation to obtain the optimal weights. Meanwhile, during the training process, the Powerful-IoU loss function is adopted to improve the detection accuracy and convergence speed of chip appearance defects. This is completed in an experimental environment with the Windows 10 operating system, NVIDIA GeForce RTX 3060Ti GPU, Intel(R) Core(TM) i5-10400F CPU@2.90GHz CPU, CUDA 11.6, cuDNN 8.4.1.50, and Python 3.9. During the training process, the number of training epochs is set to 300, and the batch size is 16.

[0083] Step 5: Load the obtained optimal weights into the improved lightweight chip appearance defect detection model, place the product chips for detection, and the detection interface outputs the results of the chip appearance defect types.

[0084] Similar to the principle of the above embodiment, the present invention provides a chip appearance defect detection system based on a lightweight model.

[0085] The following provides specific embodiments in conjunction with the accompanying drawings:

[0086] As Figure 10 shows a schematic structural diagram of a chip appearance defect detection system based on a lightweight model in an embodiment of the present invention.

[0087] The system includes:

[0088] A preprocessing module 1, which is used to preprocess each collected chip appearance defect image by using a motion blur function and construct a chip appearance defect image dataset labeled with defect type tags;

[0089] A model construction module 2, connected to the preprocessing module, which is used to train an improved YOLOv8 network model by using the chip appearance defect image dataset to construct a lightweight chip appearance defect detection model; among them, the improved YOLOv8 network model includes: a backbone layer feature extraction network, a neck feature fusion network, and a detection network; a convolution and attention fusion module is introduced in the backbone layer feature extraction network, and a lightweight C2f module is designed; an efficient fusion attention C2f module is designed in the neck feature fusion network, and a CA attention mechanism module is embedded for multi-attention mechanism fusion with the convolution and attention fusion module; a lightweight detection head is designed in the detection network.

[0090] A defect detection module 3, connected to the model construction module 2, based on the lightweight chip appearance defect detection model, obtains the corresponding chip appearance defect type results according to the chip appearance defect image to be detected.

[0091] Since the implementation principle of the chip appearance defect detection system based on the lightweight model has been described in the foregoing embodiments, it will not be repeated here.

[0092] The chip appearance defect detection method based on the lightweight model provided by the embodiments of the present invention can be implemented on the terminal side or the server side. In terms of the hardware structure of the electronic terminal, please refer to Figure 11 , which is an optional hardware structure diagram of the electronic terminal 1000 provided by the embodiments of the present invention. The terminal 1000 can be a mobile phone, a computer device, a tablet device, a personal digital processing device, a factory background processing device, etc. The terminal 1000 includes: at least one processor 1001, a memory 1002, at least one network interface 10010, and a user interface 1009. Each component in the device is coupled together through a bus system 1005. It can be understood that the bus system 1005 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1005 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 11 all kinds of buses are labeled as the bus system.

[0093] Among them, the user interface 1009 can include a display, a keyboard, a mouse, a trackball, a click gun, a button, a button, a touchpad, or a touch screen, etc.

[0094] It can be understood that the memory 1002 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory). The memory described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable categories of memory.

[0095] The memory 1002 in the embodiments of the present invention is used to store various types of data to support the operation of the terminal 1000. Examples of such data include: any executable programs for operating on the terminal 1000, such as the operating system 10021 and application programs 10022; the operating system 10021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs 10022 can include various application programs, such as a MediaPlayer, a Browser, etc., for implementing various application services. The method for detecting chip appearance defects based on a lightweight model provided by the embodiments of the present invention can be included in the application programs 10022.

[0096] The method disclosed in the above embodiments of the present invention can be applied to the processor 1001 or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1001 or instructions in software form. The above-mentioned processor 1001 can be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1001 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 1001 can be a microprocessor or any conventional processor, etc. Combining the steps of the accessory optimization method provided by the embodiments of the present invention can be directly embodied as being completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.

[0097] In an exemplary embodiment, the terminal 1000 can be one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuit), DSPs, programmable logic devices (PLDs, ProgrammableLogic Device), complex programmable logic devices (CPLDs, Complex Programmable LogicDevice) for executing the foregoing method.

[0098] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0099] In the embodiments provided in the present application, the computer-readable and writable storage medium may include read-only memory, random access memory, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, USB flash drives, portable hard drives, or any other medium that can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. Additionally, any connection can be appropriately referred to as a computer-readable medium. For example, if instructions are sent from a website, server, or other remote source using coaxial cables, fiber optic cables, twisted pairs, digital subscriber lines (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cables, fiber optic cables, twisted pairs, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable and writable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are intended to refer to non-transient, tangible storage media. As used in the application, magnetic disks and optical discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks typically replicate data magnetically, while optical discs use lasers to optically replicate data.

[0100] In summary, for the chip appearance defect detection method, system, and terminal based on a lightweight model of the present invention, first, a motion blur function is applied to preprocess the chip appearance defect image to construct a chip appearance defect image dataset. Then, the improved YOLOv8 network model is trained using the chip appearance defect image dataset to construct a detection model for detecting the types of chip appearance defects. The present invention improves the YOLOv8 network model by introducing a convolution and attention fusion module in the backbone layer feature extraction part and designing a lightweight module, designing an efficient fusion attention C2f module in the neck feature fusion network, and at the same time embedding a CA attention mechanism module to perform multi-attention mechanism fusion with the backbone layer convolution and attention fusion module, and self-designing a lightweight detection head in the detection network. The present invention not only reduces the model parameter quantity and inference floating-point numbers, taking into account both accuracy and lightweight. At the same time, it solves the problem of fuzzy and lost chip appearance defect features caused by motion blur, improving the accuracy and reliability of detection. Therefore, the present invention effectively overcomes various drawbacks in the prior art and has high industrial utilization value.

[0101] The above embodiments are only used to exemplarily illustrate the principles and effects of the present invention, rather than to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A chip appearance defect detection method based on a lightweight model, characterized in that: The method comprises: Use motion blur function to preprocess each chip appearance defect image collected, and build a chip appearance defect image dataset annotated with defect type labels; The chip appearance defect image dataset is used to train an improved YOLOv8 network model to construct a lightweight chip appearance defect detection model; wherein the improved YOLOv8 network model includes: a backbone layer feature extraction network, a neck feature fusion network and a detection network; a convolution and attention fusion module is introduced into the backbone layer feature extraction network and a lightweight C2f module is designed; an efficient fusion attention C2f module is designed in the neck feature fusion network, and a CA attention mechanism module is embedded to perform multi-attention mechanism fusion with the convolution and attention fusion module; a lightweight detection head is designed in the detection network; Based on the lightweight chip appearance defect detection model, a corresponding chip appearance defect type result is obtained according to the chip appearance defect image to be detected.

2. The chip appearance defect detection method based on a lightweight model according to claim 1, characterized in that: The backbone layer feature extraction network is constructed by replacing all C2f modules in the backbone layer feature extraction network of the original YOLOv8 network model with the designed lightweight C2f module, and connecting the convolution and attention fusion module after the pyramid pooling layer.

3. The chip appearance defect detection method based on a lightweight model according to claim 2, characterized in that: The designed lightweight C2f module introduces a Ghost bottleneck layer module; wherein the Ghost bottleneck layer module includes: Ghost convolution layer, batch normalization layer, and SiLU activation function.

4. The chip appearance defect detection method based on a lightweight model according to claim 2, characterized in that: The convolution and attention fusion module includes: a global branch and a local branch; wherein the global branch contains four branches, the input chip appearance defect image enters the second, third and fourth branches respectively after 1*1 convolution, the second and third branches are dot-multiplied after 3*3 convolution, and then processed by the Softmax function to form a first chip appearance defect attention feature map; the fourth branch is dot-multiplied with the first chip appearance defect attention feature map after 3*3 convolution to form a second chip appearance defect attention feature map; the second chip appearance defect attention feature map is added with the input chip appearance defect image of the first branch after 1*1 convolution to output a third chip appearance defect attention feature map; the local branch performs 1*1 convolution and channel transformation on the input image, and then performs 3*3 convolution to form a first appearance defect feature map; the convolution and attention fusion module adds the third chip appearance defect attention feature map output by the global branch and the first appearance defect feature map output by the local branch, and outputs the chip appearance defect image feature.

5. The chip appearance defect detection method based on a lightweight model according to claim 1, characterized in that: The neck feature fusion network is constructed by replacing all C2f modules in the neck feature fusion network of the original YOLOv8 network model with the designed efficient fusion attention C2f module, and embedding a CA attention mechanism module between the efficient fusion attention C2f module and the lightweight detection head.

6. The chip appearance defect detection method based on a lightweight model according to claim 5, characterized in that: An efficient fusion attention module is introduced into the efficient fusion attention C2f module; wherein the efficient fusion attention module is used to process the input chip appearance defect image feature through 1*1 convolution and Sigmoid activation function to form a first chip appearance defect feature, perform a dot product operation on the first chip appearance defect feature and the input chip appearance defect image feature to form a second chip appearance defect feature; process the input chip appearance defect image feature through global average pooling, 1*1 convolution and Sigmoid activation function to form a third chip appearance defect feature, perform a dot product operation on the third chip appearance defect feature and the input chip appearance defect image feature to form a fourth chip appearance defect feature; add the second chip appearance defect feature and the fourth chip appearance defect feature, and output the final chip appearance defect feature.

7. The chip appearance defect detection method based on a lightweight model according to claim 1, characterized in that: The lightweight detection head consists of two branches; the first branch consists of a convolution layer and a regression loss layer; the second branch consists of a convolution layer and a category loss layer.

8. The chip appearance defect detection method based on a lightweight model according to claim 1, characterized in that: The loss function uses the Powerful-IoU loss function when training the improved YOLOv8 network model.

9. A chip appearance defect detection system based on a lightweight model, characterized in that: The system comprises: A preprocessing module is used to preprocess each acquired chip appearance defect image using a motion blur function, and to construct a chip appearance defect image dataset annotated with defect type labels; A model building module, connected to the preprocessing module, is used to train an improved YOLOv8 network model using a chip appearance defect image dataset to build a lightweight chip appearance defect detection model; wherein the improved YOLOv8 network model includes: a backbone layer feature extraction network, a neck feature fusion network, and a detection network; a convolution and attention fusion module is introduced into the backbone layer feature extraction network and a lightweight C2f module is designed; an efficient fusion attention C2f module is designed in the neck feature fusion network, and a CA attention mechanism module is embedded to perform multi-attention mechanism fusion with the convolution and attention fusion module; a lightweight detection head is designed in the detection network; The defect detection module is connected to the model building module, and based on the lightweight chip appearance defect detection model, obtains the corresponding chip appearance defect type result according to the chip appearance defect image to be detected.

10. An electronic terminal, characterized in that: include: one or more memories and one or more processors; The one or more memories are used to store computer programs; The one or more processors, connected to the memory, are configured to run the computer program to perform the method as claimed in any one of claims 1 to 8.