Chip surface defect detection method based on improved YOLOv8

By improving the YOLOv8 model, combining data preprocessing, feature focus diffusion pyramid network, Ghost-HGNetV2 backbone network, Inner-WIoU loss function and SE attention mechanism, the existing chip defect detection technology is solved, and efficient and accurate chip defect detection is achieved.

CN120198388APending Publication Date: 2025-06-24TIANJIN UNIV OF TECH & EDUCATION (TEACHER DEV CENT OF CHINA VOCATIONAL TRAINING & GUIDANCE)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510270496.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing chip defect detection technology has problems such as low efficiency, low accuracy, high cost, high labor intensity and inconsistent standards, which is difficult to meet the market's demand for high-quality and reliable chips.

Method used

A chip surface defect detection method based on improved YOLOv8 is proposed, and the detection accuracy and efficiency are improved through data preprocessing, the introduction of feature focus diffusion pyramid network (FDPN) module, the use of Ghost-HGNetV2 as the backbone network, the replacement of the border regression loss function as the Inner-WIoU loss function, and the addition of the SE attention mechanism.

Benefits of technology

It significantly improves the detection accuracy and efficiency of tiny defects on the chip surface, reduces the false detection rate and missed detection rate, improves the robustness and generalization ability of detection, and is suitable for automated detection in the semiconductor manufacturing industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198388A_ABST
    Figure CN120198388A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chip surface defect detection, in particular to a chip surface defect detection method based on improved YOLOv8. The detection method mainly comprises the following steps: step 1, collecting data; 2, data preprocessing; 3, introducing a feature focusing diffusion pyramid network module; step 4, the Ghost-HGNetV2 is used as a backbone network of the YOLOv8; step 5, replacing a frame regression loss CIoU loss function with an Inner-WIoU loss function; step 6, adding an SE attention mechanism; and step 7, outputting a result. And 8, optimizing the model. According to the method, a new training strategy and an optimization algorithm are introduced into the YOLOv8, the training effect and generalization ability of the model are improved, the model is more robust when processing small objects, dense scenes and complex backgrounds, and powerful technical support is provided for quality control and production efficiency improvement of the chip manufacturing industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip surface defect detection, and specifically relates to a chip surface defect detection method based on improved YOLOv8. Background Technique

[0002] Chip defect detection is a key link in the semiconductor manufacturing process, which is directly related to the quality and reliability of products. During the chip production process, due to small changes in various factors such as technology, materials, environment, and process parameters, various defects may occur on the chip, such as particles, scratches, cracks, protrusions, misalignments, missing parts, etched rust spots, excessive electroplating, different colors, and damaged metal wires. These defects not only affect the yield of products, but also may cause unstable chip performance, shortened lifespan, and even safety problems.

[0003] Traditional chip defect detection methods are mainly manual visual inspection, but this method has disadvantages such as low efficiency, low accuracy, high cost, high labor intensity, and inconsistent standards, and has gradually been replaced by automatic detection technologies. In automatic detection technologies, machine vision technology has been widely studied and applied due to its advantages such as high efficiency, high accuracy, high reliability, non-contact, and strong objectivity. Machine vision technology can measure, track, and identify chips through a camera instead of human eyes, and can achieve automatic detection of chip surface defects. In recent years, with the development of deep learning technology, especially the successful application of deep learning models represented by convolutional neural networks in the field of computer vision, it has provided a new development direction for chip defect detection. Deep learning models can be trained on a large number of defect images marked with labels to learn the characteristics of defects, so as to achieve automatic recognition and classification of chip surface defects. This method can not only improve the accuracy and efficiency of detection, but also adapt to different types and forms of defects, and has stronger generalization ability.

[0004] The main problems considered in the present invention:

[0005] (1) Improve product quality and reliability: Through chip defect detection, defects generated during the chip production process can be discovered and repaired in a timely manner, thereby improving the quality and reliability of products and extending the service life of products.

[0006] (2) Reduce production costs: Effective defect detection can reduce the scrap rate caused by defects and lower production costs. At the same time, through early defect detection, the process can be adjusted in a timely manner during the production process to avoid unnecessary rework and remaking, and improve production efficiency.

[0007] (3) Meet market demands: As consumers' requirements for the performance and reliability of electronic products are increasing, chip defect detection has become a necessary means to meet market demands. Only through strict defect detection can it be ensured that products can meet market demands, thereby maintaining the competitiveness of enterprises. Summary of the Invention

[0008] Currently, the development and application of existing chip defect detection technologies are of great significance for improving product quality, reducing production costs, and meeting market demands. The present invention proposes a chip surface defect detection algorithm based on improved YOLOv8.

[0009] A chip surface defect detection method based on improved YOLOv8 mainly includes the following steps:

[0010] Step 1: Collect data. First, collect a large amount of chip surface image data, which should contain various possible defect types, such as scratches, stains, cracks, etc., for subsequent processing.

[0011] Step 2: Data preprocessing. Perform preprocessing on the collected image data, including operations such as rotation, flipping, adding noise, and changing brightness, to achieve the effect of data augmentation and improve the generalization ability of the model.

[0012] Step 3: Introduce the Feature Focus Diffusion Pyramid Network (FDPN) module. This module can not only enhance the context information of each scale feature but also optimize the multi-scale fusion of feature maps.

[0013] Step 4: Use Ghost-HGNetV2 as the backbone network of YOLOv8. The Ghost-HGNetV2 module enables the model to better capture the details of small targets through multi-scale feature fusion technology and detail information retention mechanism. At the same time, the model can maintain high efficiency.

[0014] Step 5: Replace the bounding box regression loss CIoU loss function with the Inner-WIoU loss function. Inner-WIoU can not only improve the localization accuracy but also enhance the robustness of the model, providing a more accurate and reliable solution for object detection tasks.

[0015] Step 6: Add the SE attention mechanism. The SE attention mechanism enables the model to better capture the details of small targets and improve the effect of multi-scale feature fusion.

[0016] Step 7: Result output. Finally, output the position, category, and confidence of chip surface defects for subsequent quality control and defect analysis.

[0017] Step 8: Model optimization. According to the detection results, the model is iteratively optimized to improve the generalization ability and robustness of the detection algorithm.

[0018] Through the above steps, this method can effectively improve the detection accuracy and efficiency of micro defects on the chip surface, providing strong technical support for the semiconductor manufacturing industry.

[0019] Furthermore, step 3 includes the following content. The FDPN module ensures effective object detection on feature maps of different scales, enabling the transfer of rich context information between detection scales, thereby improving detection accuracy. By methods such as feature focusing and feature diffusion, context information is increased to achieve feature extraction and fusion at multiple scales with the goal of enhancing the accuracy and robustness of small object detection. The FDPN module is the Feature Focusing and Diffusing Pyramid Network module, and its specific steps are as follows:

[0020] (1) Obtain the P3, P4, and P5 feature maps of different scales in the module.

[0021] (2) The module generates an attention feature map through a single feature fusion operation on the P3 layer. This feature map is processed through a convolutional layer and downsampled to generate the feature map of the P4 layer. The newly generated P4 layer feature map is concatenated with the previous P4 layer feature map.

[0022] (3) The concatenated P4 layer feature map is further processed through a composite feature processing layer to enhance feature representation.

[0023] The module restores the P4 layer feature map to the P3 layer resolution through an upsampling operation and concatenates it with another feature map of the P3 layer.

[0024] (4) In the P3 layer, the concatenated feature map is processed again through the composite feature processing layer. At this time, the module generates the second attention feature map and fuses the feature maps of the P3 and P4 layers.

[0025] (5) The fused feature map is processed through a convolutional layer and downsampled in the P4 layer to generate the feature map of the P5 layer, which is concatenated with the previously processed P5 layer feature map.

[0026] (6) The concatenated P5 layer feature map is optimized through the composite feature processing layer.

[0027] (7) The module upsamples the attention feature map of the P4 layer to restore it to the P3 layer resolution and concatenates it with the feature map of the P3 layer.

[0028] (8) The concatenated P3 layer feature map is optimized through the composite feature processing layer.

[0029] (9) The module fuses the optimized feature maps of the P3, P4, and P5 layers and outputs the final detection results on each resolution layer (P3, P4, P5) through the Detect layer.

[0030] Furthermore, step 4 includes the following content: using the Ghost-HGNetV2 module as the backbone network of YOLOv8. The Ghost-HGNetV2 module is an improvement of the HGNetV2 module, combining the lightweight characteristics of GhostNet and the advanced feature extraction ability of HGNetV2, thus providing stronger feature extraction ability while maintaining a low computational complexity. The Ghost-HGNetV2 module reduces redundant calculations and the number of parameters in the network by introducing the Ghost module, enabling the model to retain rich feature information while ensuring efficiency. Through this improvement, small object detection can capture sufficient detailed information in a smaller feature map area. The flowchart of the Ghost-HGNetV2 module is as Figure 2 shown.

[0031] The step process of the Ghost-HGNetV2 module is as follows:

[0032] (1) Initialize the HGStem module, which is responsible for processing the input image and generating the initial feature map. At this time, the resolution of the feature map is 1 / 4 of the original image;

[0033] (2) Enter the first stage and reuse the Ghost-HGBlock module six times. Each Ghost-HGBlock module contains 3 HG units to enhance the feature extraction ability. This process is denoted as the Ghost-HGBlock-6 process; the input channels of the first stage are 48, and the output channels are 128. The module obtained in the first stage is denoted as the Ghost-HGBlock (1) module;

[0034] (3) Perform downsampling through the depthwise separable convolution module to reduce the resolution of the feature map to 1 / 8 of the original image and adjust the number of channels to 128;

[0035] (4) Enter the second stage, and the network reuses the Ghost-HGBlock (1) module six times. Each module contains 3 HG units, denoted as the Ghost-HGBlock (1) -6 process; the input channels of the second stage are 96, and the output channels are 512. The module obtained in the second stage is denoted as the Ghost-HGBlock (2) module;

[0036] (5) Downsample through the depthwise separable convolution module, reducing the resolution of the feature map to 1 / 16 of the original image and adjusting the number of channels to 512;

[0037] (6) Enter the third stage, and the network repeatedly uses the Ghost-HGBlock (2) module six times, where each module contains 1 HG unit, denoted as Ghost-HGBlock (2) - 6 process, repeatedly use the Ghost-HGBlock (2) - 6 process three times continuously; the input channels of the third stage are 192, and the output channels are 1024. The module obtained in the third stage is denoted as Ghost-HGBlock (3) module;

[0038] (7) Downsample through the depthwise separable convolution module, reducing the resolution of the feature map to 1 / 32 of the original image and adjusting the number of channels to 1024;

[0039] (8) Enter the fourth stage, and the network repeatedly uses the Ghost-HGBlock (3) module six times, where each module contains 1 HG unit, denoted as Ghost-HGBlock (3) - 6 process; the input channels of the fourth stage are 384, and the output channels are 2048. The module obtained in the fourth stage is denoted as Ghost-HGBlock (4) module;

[0040] (9) Finally, add the SPPF module. The SPPF module can enhance the multi-scale representation of the feature map, capture context information at different scales, further enhance the representation ability of the feature map, and enable the model to improve the detection accuracy and adaptability to different scenarios.

[0041] Further, step 5 includes the following content, the improvement of the bounding box regression loss CIoU.

[0042] In the YOLOv8 bounding box regression loss function, the CIoU loss function is replaced by the Inner-WIoU loss function. Inner-WIoU is a weighted IoU loss function, and its core idea is to improve the accuracy of small object detection by adjusting the weights in the IoU calculation process. Specifically, the Inner-WIoU loss function introduces a weight coefficient to adjust the importance of each part in the IoU calculation. In this way, during the training process, the model not only focuses on the overall overlapping area but also pays more attention to the internal region of the target box, thereby improving the localization accuracy. The Inner-WIoU loss function has obvious advantages compared with traditional WIoU loss function, CIoU loss function and other loss functions: (1) The Inner-WIoU loss function provides a more accurate overlapping metric. It focuses on the overlapping inside the target region, which can better handle the details of the target region in the bounding box. When the detection scenario is small object detection, the Inner-WIoU loss function can capture more subtle localization errors, thereby providing more accurate regression feedback and improving the detection accuracy and recall rate. In contrast, the traditional WIoU loss function may not fully reflect the internal overlapping of the target region, thus affecting the localization accuracy. (2) The Inner-WIoU loss function can better handle the problems of target position and scale changes. It helps to improve the detection performance of small objects. Through weighted processing, the Inner-WIoU loss function can reduce the negative impact of position and scale changes on the detection results. It is more sensitive to the overlapping of small target regions and can better adapt to targets of different sizes and scales, thereby improving the localization accuracy of the model in applications and significantly improving the detection performance of small objects.

[0043] The calculation formula of the Inner-WIoU loss function is as follows:

[0044]

[0045] In the formula, Area-I is the area inside the intersection region of the predicted box and the ground truth box; Area-U is the area of the union region of the predicted box and the ground truth box.

[0046] The calculation formula of the Inner-WIoU loss function is as follows:

[0047] L Inner-WIoU = 1 - Inner-WIoU(2)

[0048] After replacing the CIoU function in the bounding box regression loss of YOLOv8 with the Inner-WIoU function, the calculation formula of the total bounding box regression loss of YOLOv8 is as follows:

[0049] L bbox = λ Inner-WIoU ·LInner-WIoU +λ DFL ·L DFL (3)

[0050] The total loss of YOLOv8 includes bounding box regression loss, classification loss, and confidence loss. The calculation formulas for classification loss and confidence loss are as follows:

[0051] L cls =-α(1 - p t ) γ log(p t )(4)

[0052] L conf =-(ylog(p)+(1 - y)log(1 - p))(5)

[0053] Based on formulas (3), (4), and (5), the calculation formula for the total loss is as follows:

[0054] L total =λ bbox ·L bbox +λ cls ·L cls +λ conf ·L conf (6)

[0055] Overall, the introduction of the Inner - WIoU function is an effective supplement to the traditional CIoU loss, especially in small object detection tasks. Through the weighting mechanism, Inner - WIoU not only improves the localization accuracy but also enhances the robustness of the model, providing a more accurate and reliable solution for object detection tasks.

[0056] Furthermore, step 6 includes the following content: the introduction of the SE attention mechanism.

[0057] YOLOv8 introduces the SE attention mechanism, which is a technique to enhance the model performance by adaptively adjusting the importance of feature maps. The network diagram of the SE attention mechanism is as Figure 3 shown. The SE attention mechanism enhances the feature expression ability of the model through three main steps:

[0058] (1) The SE attention mechanism performs global information aggregation on the feature maps of each channel through the "squeeze" operation. The feature maps are compressed into a single - channel descriptor through global average pooling, capturing the global context information of each channel.

[0059] (2) Through the "excitation" operation, the SE attention mechanism weights these channel descriptors to readjust the weight of each channel, thereby highlighting important features and suppressing unimportant features, effectively enhancing the expression of useful information in the feature map while suppressing the interference of irrelevant information, thus improving the feature learning ability of the model.

[0060] (3) Weight the normalized weights obtained above to the features of each channel.

[0061] After introducing the SE attention mechanism, the deficiencies of traditional YOLOv8 in small target detection are solved, and the feature learning ability and detection accuracy of the model are significantly improved. (1) By adaptively adjusting the weights of feature channels through the SE attention mechanism, the ability to capture small target features is enhanced. Since the feature information of small targets in images is often sparse, traditional networks may have difficulty fully extracting this information. The SE attention mechanism, through weighted adjustment, enables the features of small targets to be better expressed in the network, thus improving the detection accuracy. (2) The SE attention mechanism improves the effect of feature fusion. Traditional YOLOv8 may have problems of information loss during multi-scale feature fusion, while the SE attention mechanism can ensure the effective fusion and utilization of features at different scales by dynamically adjusting the importance of feature maps. (3) The SE mechanism can also improve the robustness of YOLOv8 in complex scenarios. By weighted adjustment of the feature map, the SE mechanism can effectively reduce the influence of background noise and interference and enhance the detection ability of targets.

[0062] Beneficial effects:

[0063] (1) Improve detection accuracy: By introducing and improving various modules (such as Ghost-HGNetV2 module, FDPN module, etc.) and optimizing algorithms, the improved YOLOv8 model can more accurately capture the defect features on the chip surface and improve the detection accuracy.

[0064] (2) Reduce the false detection rate and missed detection rate: While maintaining high detection accuracy, the improved model reduces the false detection rate and missed detection rate. This benefits from the model's better understanding of context information and the ability to extract multi-scale features. Through strategies such as optimizing the loss function and adding attention mechanisms, the model can more accurately locate the defect positions and reduce false detections and missed detections.

[0065] (3) Improve detection efficiency: Although the model has been improved and optimized in structure, the computational cost is not significantly increased. Instead, by introducing improved modules and reducing unnecessary computational operations, the improved model improves the detection efficiency while maintaining high detection accuracy.

[0066] (4) High degree of automation: The improved YOLOv8 model realizes a highly automated detection of chip surface defects. Simply input the chip image to be detected into the model, and the detection results can be quickly obtained.

[0067] (5) Strong robustness and generalization ability: The improved model shows good detection performance on different types of chips and defects. This benefits from the model's ability to understand context information and extract multi-scale features. At the same time, through strategies such as data augmentation and model tuning, the robustness and generalization ability of the model have been further improved.

[0068] (6) Easy to deploy and integrate: The improved YOLOv8 model has a small model size and low computational complexity, which makes it easier to deploy and integrate into existing production lines and detection systems. Brief Description of the Drawings

[0069] Figure 1 It is the overall flowchart of the present invention.

[0070] Figure 2 It is the flowchart of the Ghost-HGNetV2 module.

[0071] Figure 3 It is the network diagram of the SE attention mechanism.

[0072] Figure 4 It is the sample diagram of the defect types of the chip surface defect detection data of the present invention.

[0073] Figure 5 It is the performance comparison diagram between the present invention and the original YOLOv8 model. Detailed Embodiments

[0074] Example 1:

[0075] The experiment of the present invention is based on the Ubuntu 20.04.6 LTS operating system, with an Intel Core i9-12900K processor; the GPU is NVIDIA GeForce RTX3090*2, the memory is 128G, and the hard disk is 6TB. The software environment is: the deep learning framework of Python3.8, CUDA 11.8, and PyTorch 2.0.0. During the experiment, the training batch size is 16, the initial learning rate is 0.01, the number of iterations is 150, the input image size is 640×640, the optimizer is SGD, and the momentum is 0.9.

[0076] A method for detecting chip surface defects based on improved YOLOv8 mainly includes the following steps:

[0077] Step 1: Collect data. First, collect a large amount of image data of the chip surface. This data should contain various possible defect types, such as scratches, stains, cracks, etc., for subsequent processing.

[0078] Step 2: Data preprocessing. Perform preprocessing on the collected image data, including operations such as rotation, flipping, adding noise, and changing brightness, to achieve the effect of data augmentation and improve the generalization ability of the model.

[0079] Step 3: Introduce the Feature Focused Diffusion Pyramid Network (FDPN) module. This module can not only enhance the context information of each scale feature but also optimize the multi-scale fusion of feature maps.

[0080] Step 4: Use Ghost-HGNetV2 as the backbone network of YOLOv8. The Ghost-HGNetV2 module enables the model to better capture the details of small targets through multi-scale feature fusion technology and detail information retention mechanism. At the same time, the model can maintain high efficiency.

[0081] Step 5: Replace the bounding box regression loss CIoU loss function with the Inner-WIoU loss function. Inner-WIoU can not only improve the localization accuracy but also enhance the robustness of the model, providing a more accurate and reliable solution for the object detection task.

[0082] Step 6: Add the SE attention mechanism. The SE attention mechanism enables the model to better capture the details of small targets and improve the effect of multi-scale feature fusion.

[0083] Step 7: Result output. Finally, output the location, category, and confidence of the chip surface defects for subsequent quality control and defect analysis.

[0084] Step 8: Model optimization. According to the detection results, iteratively optimize the model to improve the generalization ability and robustness of the detection algorithm.

[0085] Through the above steps, this method can effectively improve the detection accuracy and efficiency of micro-defects on the chip surface, providing strong technical support for the semiconductor manufacturing industry.

[0086] Furthermore, Step 3 includes the following content. The FDPN module ensures effective object detection on feature maps of different scales, enables the transfer of rich context information between different detection scales, thereby improving the detection accuracy. By methods such as feature focusing and feature diffusion, the context information is increased to achieve feature extraction and fusion at multiple scales with the goal of improving the accuracy and robustness of small target detection. The FDPN module is the Feature Focused Diffusion Pyramid Network module, and its specific steps are as follows:

[0087] (1) Obtain P3, P4, and P5 feature maps of different scales in the acquisition module.

[0088] (2) The module generates an attention feature map through a feature fusion operation on the P3 layer. This feature map is processed by a convolutional layer and downsampled to generate the feature map of the P4 layer. The newly generated P4 layer feature map is concatenated with the previous P4 layer feature map.

[0089] (3) The concatenated P4 layer feature map is further processed by a composite feature processing layer to enhance the feature representation.

[0090] The module restores the P4 layer feature map to the P3 layer resolution through an upsampling operation and concatenates it with another feature map of the P3 layer.

[0091] (4) In the P3 layer, the concatenated feature map is processed again by the composite feature processing layer. At this time, the module generates the second attention feature map and fuses the feature maps of the P3 and P4 layers.

[0092] (5) The fused feature map is processed by a convolutional layer and downsampled in the P4 layer to generate the feature map of the P5 layer, which is concatenated with the previously processed P5 layer feature map.

[0093] (6) The concatenated P5 layer feature map is optimized by the composite feature processing layer.

[0094] (7) The module upsamples the attention feature map of the P4 layer to restore it to the P3 layer resolution and concatenates it with the feature map of the P3 layer.

[0095] (8) The concatenated P3 layer feature map is optimized by the composite feature processing layer.

[0096] (9) The module fuses the optimized P3, P4, and P5 layer feature maps and outputs the final detection results at each resolution layer (P3, P4, P5) through the detection layer (Detect).

[0097] Furthermore, step 4 includes the following content: using the Ghost-HGNetV2 module as the backbone network of YOLOv8. The Ghost-HGNetV2 module is an improvement of the HGNetV2 module, combining the lightweight characteristics of GhostNet and the advanced feature extraction ability of HGNetV2, thus providing stronger feature extraction ability while maintaining a relatively low computational complexity. By introducing the Ghost module, the Ghost-HGNetV2 module reduces the redundant calculations and the number of parameters in the network, enabling the model to retain rich feature information while ensuring high efficiency. Through this improvement, small object detection can capture sufficient detailed information in a smaller feature map area. The flowchart of the Ghost-HGNetV2 module is as Figure 2 shown.

[0098] The step process of the Ghost-HGNetV2 module is as follows:

[0099] (1) Initialize the HGStem module, which is responsible for processing the input image and generating the initial feature map. At this time, the resolution of the feature map is 1 / 4 of the original image;

[0100] (2) Enter the first stage and reuse the Ghost-HGBlock module six times. Each Ghost-HGBlock module contains 3 HG units to enhance the feature extraction ability. This process is denoted as the Ghost-HGBlock-6 process; the input channels of the first stage are 48, and the output channels are 128. The module obtained in the first stage is denoted as the Ghost-HGBlock (1) module;

[0101] (3) Perform downsampling through the depthwise separable convolution module to reduce the resolution of the feature map to 1 / 8 of the original image and adjust the number of channels to 128;

[0102] (4) Enter the second stage, and the network reuses the Ghost-HGBlock (1) module six times. Each module contains 3 HG units, denoted as the Ghost-HGBlock (1) -6 process; the input channels of the second stage are 96, and the output channels are 512. The module obtained in the second stage is denoted as the Ghost-HGBlock (2) module;

[0103] (5) Perform downsampling through the depthwise separable convolution module to reduce the resolution of the feature map to 1 / 16 of the original image and adjust the number of channels to 512;

[0104] (6) Enter the third stage, and the network reuses the Ghost-HGBlock (2)The module is repeated six times, and each module contains 1 HG unit, denoted as Ghost-HGBlock (2) -6 process, repeating Ghost-HGBlock three times continuously (2) -6 process; the input channels of the third stage are 192, and the output channels are 1024. The module obtained in the third stage is denoted as Ghost-HGBlock (3) module;

[0105] (7) Downsample through the depthwise separable convolution module, reducing the resolution of the feature map to 1 / 32 of the original image and adjusting the number of channels to 1024;

[0106] (8) Enter the fourth stage, and the network repeatedly uses Ghost-HGBlock (3) The module is repeated six times, and each module contains 1 HG unit, denoted as Ghost-HGBlock (3) -6 process; the input channels of the fourth stage are 384, and the output channels are 2048. The module obtained in the fourth stage is denoted as Ghost-HGBlock (4) module;

[0107] (9) Finally, add the SPPF module. The SPPF module can enhance the multi-scale representation of the feature map, capture context information at different scales, further enhance the representation ability of the feature map, and enable the model to improve the detection accuracy and adaptability to different scenarios.

[0108] In the YOLOv8 bounding box regression loss function, the CIoU loss function is replaced by the Inner-WIoU loss function. Inner-WIoU is a weighted IoU loss function. Its core idea is to adjust the weights in the IoU calculation process to improve the accuracy of small object detection. Specifically, the Inner-WIoU loss function introduces a weight coefficient to adjust the importance of each part in the IoU calculation. In this way, during the training process, the model not only focuses on the overall overlapping area but also pays more attention to the internal area of the target box, thereby improving the localization accuracy. The Inner-WIoU loss function has obvious advantages compared with traditional WIoU loss function, CIoU loss function, and other loss functions: (1) The Inner-WIoU loss function provides a more accurate overlapping metric. It focuses on the internal overlap of the target area, which can better handle the details of the target area in the bounding box. When the detection scenario is small object detection, the Inner-WIoU loss function can capture more subtle localization errors, thereby providing more accurate regression feedback and improving the detection accuracy and recall rate. In contrast, the traditional WIoU loss function may not fully reflect the internal overlap of the target area, thus affecting the localization accuracy. (2) The Inner-WIoU loss function can better handle the problems of target position and scale changes. It helps to improve the detection performance of small objects. Through weighted processing, the Inner-WIoU loss function can reduce the negative impact of position and scale changes on the detection results. It is more sensitive to the overlap of small target areas and can better adapt to targets of different sizes and scales, thereby improving the localization accuracy of the model in applications and significantly improving the detection performance of small objects.

[0109] The calculation formula of the Inner-WIoU loss function is as follows:

[0110]

[0111] In the formula, Area-I is the area inside the intersection region of the predicted box and the ground truth box; Area-U is the area of the union region of the predicted box and the ground truth box.

[0112] The calculation formula of the Inner-WIoU loss function is as follows:

[0113] L Inner-WIoU = 1 - Inner-WIoU(2)

[0114] After replacing the CIoU function in the bounding box regression loss of YOLOv8 with the Inner-WIoU function, the calculation formula of the total bounding box regression loss of YOLOv8 is as follows:

[0115] L bbox = λ Inner-WIoU·L Inner-WIoU + λ DFL ·L DFL (3)

[0116] The total loss of YOLOv8 includes bounding box regression loss, classification loss, and confidence loss. The calculation formulas for classification loss and confidence loss are as follows:

[0117] L cls = -α(1 - p t ) γ log(p t ) (4)

[0118] L conf = -(y log(p) + (1 - y) log(1 - p)) (5)

[0119] Based on formulas (3), (4), and (5), the calculation formula for the total loss is as follows:

[0120] L total = λ bbox ·L bbox + λ cls ·L cls + λ conf ·L conf (6)

[0121] Overall, the introduction of the Inner-WIoU function is an effective complement to the traditional CIoU loss, especially in small object detection tasks. Through the weighting mechanism, Inner-WIoU not only improves the localization accuracy but also enhances the robustness of the model, providing a more accurate and reliable solution for object detection tasks.

[0122] Furthermore, step 6 includes the following content: the introduction of the SE attention mechanism.

[0123] YOLOv8 introduces the SE attention mechanism, which is a technique that enhances the model's performance by adaptively adjusting the importance of feature maps. The network diagram of the SE attention mechanism is as Figure 3 shown. The SE attention mechanism enhances the model's feature expression ability through three main steps:

[0124] (1) The SE attention mechanism performs global information aggregation on the feature maps of each channel through a "squeeze" operation. By using global average pooling, the feature maps are compressed into a single-channel descriptor, capturing the global context information of each channel.

[0125] (2) Through the "incentive" operation, the SE attention mechanism weights these channel descriptors to readjust the weight of each channel, thereby highlighting important features and suppressing unimportant features, effectively enhancing the expression of useful information in the feature map while suppressing the interference of irrelevant information, thus improving the feature learning ability of the model.

[0126] (3) Weight the normalized weights obtained above to the features of each channel.

[0127] Comparative Example 1:

[0128] The original YOLOv8 model, different from Example 1 in that steps (3), (4), (5), and (6) are omitted, specifically refer to No. 1 in Table 1.

[0129] Comparative Example 2:

[0130] The technical solution is similar to that of Example 1, the difference is only that steps (4), (5), and (6) are omitted, specifically refer to No. 2 in Table 1.

[0131] Comparative Example 3:

[0132] The technical solution is similar to that of Example 1, the difference is only that steps (3), (5), and (6) are omitted, specifically refer to No. 3 in Table 1.

[0133] Comparative Example 4:

[0134] The technical solution is similar to that of Example 1, the difference is only that steps (3), (4), and (6) are omitted, specifically refer to No. 4 in Table 1.

[0135] Comparative Example 5:

[0136] The technical solution is similar to that of Example 1, the difference is only that steps (3), (4), and (5) are omitted, specifically refer to No. 5 in Table 1.

[0137] Comparative Examples 6 - 11:

[0138] The technical solution is similar to that of Example 1, the difference is only that two of the steps (3), (4), (5), and (6) are omitted, specifically refer to Nos. 6 - 11 in Table 1.

[0139] Comparative Examples 12 - 15:

[0140] The technical solution is similar to that of Example 1, the difference is only that one of the steps (3), (4), (5), and (6) is omitted, specifically refer to Nos. 12 - 15 in Table 1.

[0141] Table 1 is a comparison chart of ablation experiments for the improved YOLOv8 algorithm. Among them, "A" represents using Ghost-HGNetV2 as the backbone network of YOLOv8, "B" represents introducing the FDPN module into the YOLOv8 model, "C" represents replacing the bounding box regression loss CIoU loss function in YOLOv8 with the Inner-WIoU loss function, and "D" represents introducing the SE attention mechanism into the YOLOv8 model. "√" indicates that the module is introduced into the YOLOv8 model, and no annotation means it is not introduced. P (Precision): Among all detected targets, the ratio of correctly detected targets. It measures the accuracy of the model's detection. R (Recall): Among all real targets, the ratio of correctly detected targets. It measures the comprehensiveness of the model's detection. mAP@50 (mean Average Precision at 50% IoU, mean average precision at 50% intersection over union): The mean average precision when the IoU threshold is set to 50%. This metric is a commonly used performance evaluation metric in object detection and is used to measure the average detection performance of the model across multiple classes. mAP@50:95 (mean Average Precision at 50% to 95% IoU, mean average precision at 50% to 95% intersection over union): Calculate the average precision for all thresholds in the range of IoU thresholds from 50% to 95%, and then calculate their average. This metric can more comprehensively evaluate the performance of the model at different IoU thresholds. Params (Number of parameters): The total number of all learnable parameters in the model. This metric can be used to measure the complexity and size of the model. GFLOPs (Giga Floating-point Operations): A metric used to evaluate the computational complexity of the model, representing the number of floating-point operations required for the model to perform one forward pass, with the unit of billions. This metric helps us understand the running efficiency of the model on hardware.

[0142] Table 2 is a comparison chart of the detection accuracy values of various types of defects before and after the improvement of the improved YOLOv8 algorithm. Among them, "wr" represents contamination defects, "qx" represents wire missing defects, "lg" represents missing fixation defects, "jlqx" represents crystal grain tilt defects, and "wx" represents wire skew defects.

[0143] Table 1

[0144]

[0145] In Table 1, serial number 1 is the traditional YOLOv8 algorithm with poor comprehensive performance, and none of the indicators reached the expected effect. Serial number 2 introduced the FDPN module into the traditional YOLOv8, aiming to improve the model's ability to process multi-scale features. The accuracy only increased by 2.8%, but Params and GFLOPs increased. Serial number 3 used Ghost-HGNetV2 as the backbone network of YOLOv8, making the model lightweight. The accuracy increased by 3.2%, the recall rate increased by 0.3%, and the mAP@0.5 and mAP@0.5:0.95 values increased by 0.1% and 0.3%, respectively. However, Params and GFLOPs effectively decreased by 0.7M and 0.3G, achieving both model lightweight and performance improvement. Serial number 4 replaced the bounding box regression loss CIoU loss function of the YOLOv8 model with the InnerWIoU loss function to improve the localization accuracy, and the accuracy effectively increased by 1.6%. Serial number 5 added the SE attention mechanism to the YOLOv8 model. While capturing the details of small targets, it improved the effect of multi-scale feature fusion. Without changing Params and GFLOPs, the accuracy value increased by 4.2% significantly, the recall rate increased by 1.7%, and the mAP@0.5 and mAP@0.5:0.95 values increased by 1.1% and 0.4%, respectively. Serial numbers 6 - 15 conducted experiments by combining the four improvement methods in permutations and combinations. The experimental results showed that some combinations of improvement methods would further enhance the model's performance. For example, in the technical solution of serial number 9, the FDPN module was introduced into the YOLOv8 model, and at the same time, the bounding box regression loss CIoU loss function in YOLOv8 was replaced with the Inner-WIoU loss function. The comprehensive performance of the obtained model was better than the technical solutions of serial numbers 3 and 4. At the same time, there were also combinations of improvement methods that did not bring beneficial effects and even produced negative effects. For example, in the technical solution of serial number 11, the bounding box regression loss CIoU loss function in YOLOv8 was replaced with the Inner-WIoU loss function, and at the same time, the SE attention mechanism was introduced into the YOLOv8 model. The comprehensive performance of the obtained model was significantly worse than the technical solutions of serial numbers 4 and 5. It can be seen that not all combinations of improvement methods can bring optimization effects compared to the previous group. The interaction and combination of improvement methods are crucial, and the experimental results cannot be simply improved by combination. Considering the four improvement schemes, although Params increased by 0.9M and GFLOPs increased by 4.1G compared with the traditional YOLOv8 model, the accuracy increased by 9.1% significantly, the recall rate increased by 4.4%, and the mAP@0.5 and mAP@0.5:0.95 values increased by 6.3% and 2.3%, respectively. Through the above analysis, it can be verified that the improvement methods are feasible and effective in improving the accuracy of model defect detection.

[0146] Table 2

[0147]

[0148]

[0149] Through the analysis of Table 2, the detection accuracy of various types of defects on the improved model has been significantly improved. Specifically, the accuracies of general defects (missing lines, missing fixation, grain tilt, and crooked lines) have increased by 3.5%, 11.1%, 7.3%, and 4.3% respectively. The accuracy of small target contamination defects has increased significantly by 19.5%. This significant progress fully verifies that the improved model has achieved remarkable results in capturing the defect characteristics on the chip surface, especially for small target defects.

[0150] Figure 1 Figure 11 is the overall flowchart of YOLOv8. The model first goes through the data collection stage to ensure the diversity and quality of the training data. Subsequently, the data undergoes data preprocessing steps, including cleaning, annotation, and augmentation, to improve the generalization ability of the model. In the model optimization and improvement section, YOLOv8 integrates four improvement methods to enhance its detection accuracy and efficiency. After a series of training iterations, the model outputs the training results, demonstrating the performance of the model at each stage and providing a basis for further optimization. Finally, in the model optimization stage, through careful hyperparameter tuning and structural fine-tuning, the stability and accuracy of the model are further improved.

[0151] Figure 2 Figure 12 is the flowchart of the Ghost-HGNetV2 module. The advantages of the Ghost-HGNetV2 module lie in its improved feature extraction and fusion capabilities, which help to make up for the deficiencies of YOLOv8 in small target detection. (1) Lightweight: The GhostConv module reduces the amount of computation and the number of parameters, making the entire network lighter, with lower inference time and memory occupancy. This helps to reduce the computational burden of the model and improve the detection speed. (2) Adaptability: The Ghost-HGNetV2 module combines convolutional operations and special structural designs to extract richer feature information and process input images of different sizes and shapes. It can better adapt to defect data of different sizes in the chip surface defect detection task, improving the generalization ability and detection accuracy of the model.

[0152] Figure 3It is the network diagram of the SE attention mechanism. The SE attention mechanism first performs a "Squeeze" operation, which globally averages the pooling of the spatial dimension of the feature map, compressing the two-dimensional feature map of each channel into a real number, which represents the response intensity of the channel in the global range; secondly, through the "Excitation" operation, two fully connected layers are used to learn the dependencies between channels and generate a weight coefficient for each channel. These weight coefficients are then rescaled back to the original channel dimension and multiplied by the input feature map through a multiplication operation, thereby enhancing the important feature channels and suppressing the unimportant feature channels while maintaining spatial information. The SE attention mechanism can adaptively adjust the response of the feature channels, improve the representation ability of the network, and without increasing too much computational complexity.

[0153] Figure 4 This is the sample diagram of the data defect types for chip surface defect detection by improving YOLOv8 of the present invention. It includes defects such as contamination, grain tilt, missing solidification, missing wire, and crooked wire.

[0154] From Figure 5 It can be seen that the model (Ours) of the present invention shows a more stable and optimized performance on the recall rate, precision, mean average precision, and loss value curves. This significant improvement not only reflects the robustness of the model training but also confirms its superiority over the original YOLOv8 in multiple key analysis dimensions:

[0155] (1) In terms of the smoothness of the curve, the improved model significantly reduces fluctuations and oscillations, indicating that it is more stable during the training process, can learn data features more effectively, and avoids unnecessary performance fluctuations.

[0156] (2) In terms of the performance of peaks and valleys, the improved model reaches higher peaks on the recall rate and precision curves, and at the same time reaches deeper valleys on the loss value curve, which reflects the significant progress of the model in achieving optimal performance and minimizing errors.

[0157] (3) Observing the position of the convergence point, the improved model shows a faster convergence speed, which means that it can reach a higher performance level after fewer training iterations, thereby improving the training efficiency and time cost effectiveness.

[0158] (4) In terms of the final values of the performance metrics, the improved model has achieved a significant improvement in the mean average precision, comprehensively reflecting the detection performance of the model in various categories.

[0159] Generally speaking, the stable performance of the improved model on the recall rate, precision, and mean average precision curves effectively proves its superiority over the traditional YOLOv8 model in terms of performance, stability, and generalization ability through multi-dimensional analysis.

Claims

1. A chip surface defect detection method based on improved YOLOv8, characterized in that: The following steps are involved: Step 1: Collect data: Collect a large amount of chip surface image data; Step 2: Data preprocessing: preprocess the collected image data; Step 3: Introduce the feature-focused diffusion pyramid network FDPN module; Step 4: Use Ghost-HGNetV2 as the backbone network of YOLOV8; Step 5: Replace the bounding box regression loss CIoU loss function with the Inner-WIoU loss function; Step 6: Add SE attention mechanism; Step 7: Result output: The location, category and confidence level of the chip surface defects are finally output for subsequent quality control and defect analysis; Step 8: Model optimization: Based on the detection results, the model is iteratively optimized to improve the generalization ability and robustness of the detection algorithm.

2. The chip surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The data collected in step 1 include various defect types, and the defect type includes at least one of scratches, stains, and cracks.

3. The chip surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The preprocessing process of step 2 includes rotating, flipping, adding noise, and / or changing the brightness of the collected image.

4. The chip surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The step 3 is as follows: The FDPN module is a feature-focused diffusion pyramid network module, and its specific steps are as follows: (1) Obtain P3, P4, and P5 feature maps of different scales in the module; (2) The module generates a focused feature map through a feature fusion operation on the P3 layer, which is processed by a convolution layer and downsampled to generate a feature map of the P4 layer. The newly generated P4 layer feature map is concatenated with the previous P4 layer feature map. (3) The spliced ​​P4 layer feature map is further processed through a composite feature processing layer to enhance the feature representation; the module restores the P4 layer feature map to the P3 layer resolution through an upsampling operation and splices it with another feature map of the P3 layer; (4) In the P3 layer, the spliced ​​feature map is processed again by the composite feature processing layer. The module will generate the feature map for the second time and fuse the feature maps of the P3 and P4 layers. (5) The fused feature map is processed by convolution layer and downsampling in P4 layer to generate the feature map of P5 layer, which is then concatenated with the previously processed feature map of P5 layer; (6) The spliced ​​P5 layer feature map is optimized through the composite feature processing layer; (7) The module upsamples the attention feature map of the P4 layer to restore it to the resolution of the P3 layer and concatenates it with the feature map of the P3 layer; (8) The spliced ​​P3 layer feature map is optimized through the composite feature processing layer; (9) The module fuses the optimized feature maps of the P3, P4, and P5 layers, and outputs the final detection results at each resolution layer through the detection layer.

5. The chip surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The step 4 is specifically as follows: The steps of the Ghost-HGNetV2 module are as follows: (1) Initialize the HGStem module, which is responsible for processing the input image and generating the initial feature map. The resolution of the feature map is 1 / 4 of the original image. (2) Entering the first stage, the Ghost-HGBlock module is reused six times, where each Ghost-HGBlock module contains 3 HG units to enhance the feature extraction capability. This process is recorded as the Ghost-HGBlock-6 process; the input channel of the first stage is 48, the output channel is 128, and the module obtained in the first stage is recorded as Ghost-HGBlock (1) Modules; (3) Downsampling is performed through a depth-wise separable convolution module to reduce the resolution of the feature map to 1 / 8 of the original image and adjust the number of channels to 128; (4) Entering the second stage, the network reuses Ghost-HGBlock (1) The module is six times, each of which contains 3 HG units, denoted as Ghost-HGBlock (1) -6 process; the input channel of the second stage is 96, the output channel is 512, and the module obtained in the second stage is recorded as Ghost-HGBlock (2) Modules; (5) Downsampling is performed through a depthwise separable convolution module to reduce the resolution of the feature map to 1 / 16 of the original image and adjust the number of channels to 512; (6) Entering the third stage, the network reuses Ghost-HGBlock (2) The module is six times, each of which contains 1 HG unit, denoted as Ghost-HGBlock (2) -6 process, repeat Ghost-HGBlock three times in a row (2) -6 process; the input channel of the third stage is 192, the output channel is 1024, and the module obtained in the third stage is recorded as Ghost-HGBlock (3) Modules; (7) Downsampling is performed through a depthwise separable convolution module to reduce the resolution of the feature map to 1 / 32 of the original image and adjust the number of channels to 1024; (8) Entering the fourth stage, the network reuses Ghost-HGBlock (3) The module is six times, each of which contains 1 HG unit, denoted as Ghost-HGBlock (3) -6 process; the input channel of the fourth stage is 384, the output channel is 2048, and the module obtained in the fourth stage is recorded as Ghost-HGBlock (4) Modules; (9) Finally, the SPPF module is added. The SPPF module can enhance the multi-scale representation of feature maps, capture contextual information of different scales, and further enhance the representation ability of feature maps, so that the model improves detection accuracy and adaptability to different scenarios.

6. The chip surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The step 5 is specifically as follows: The Inner-WIoU loss function calculation formula is as follows: Where Area-I is the area inside the intersection of the predicted box and the true box; Area-U is the area of ​​the union of the predicted box and the true box; The Inner-WIoU loss function calculation formula is as follows: L Inner-WIoU =1-Inner-WIoU (2) After replacing the CIoU function in the bounding box regression loss with the Inner-WIoU function in YOLOv8, the YOLOv8 bounding box regression loss calculation formula is as follows: L bbox =λ Inner-WIoU ·L Inner-WIoU +λ DFL ·L DFL (3) The total loss of YOLOv8 includes bounding box regression loss, classification loss and confidence loss; the classification loss and confidence loss are calculated as follows: L cls =-α(1-p t ) γ log(p t ) (4) L conf =-(ylog(p)+(1-y)log(1-p)) (5) Based on formulas (3), (4) and (5), the total loss calculation formula is as follows: L total =λ bbox ·L bbox +λ cls ·L cls +λ conf ·L conf (6)。 7. The chip surface defect detection method based on improved YOLOv8 according to claim 1, characterized in that: The step 6 is specifically as follows: The SE attention mechanism improves the model’s feature expression capabilities through three main steps: (1) The SE attention mechanism aggregates global information on the feature map of each channel through a "squeezing" operation; the feature map is compressed into a channel descriptor through global average pooling to capture the global context information of each channel; (2) Through the "excitation" operation, the SE attention mechanism weights these channel descriptors to readjust the weight of each channel, thereby highlighting important features and suppressing unimportant features. It can effectively enhance the expression of useful information in the feature map while suppressing the interference of irrelevant information, thereby improving the feature learning ability of the model; (3) Add the normalized weight obtained above to the features of each channel.

Citation Information

Cited By

  • Chip AFM detection degree regulation and control method and system, medium and product

    CN121746342A