Chip defect real-time detection method based on improved YOLOv13

By improving the YOLOv13 network structure and hard sample-guided preprocessing, the problems of background interference and small defect feature extraction in chip defect detection were solved, achieving high-precision, real-time chip defect detection and improving the robustness and generalization ability of the model.

CN120953200APending Publication Date: 2025-11-14JIANGSU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511047558.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing YOLO models suffer from issues such as sensitivity to background interference, weakening of features of small defects, insufficient accuracy in detecting small defects, and real-time detection accuracy in chip defect detection. Furthermore, difficulties in sample collection limit the model's generalization ability.

Method used

By improving the YOLOv13 network structure and introducing the CS-C3k2 and Sim-DS-C3k2 modules, combined with a hard sample-oriented preprocessing mechanism including virtual dense cluster enhancement, local super-resolution reconstruction, and adaptive contrast stretching, the defect feature extraction capability is enhanced, background interference is suppressed, and detection robustness and accuracy are improved.

Benefits of technology

While ensuring real-time performance, it significantly improves the accuracy and robustness of chip defect detection, meets the stringent requirements of industrial online inspection, and enhances the model's generalization ability under different batches and imaging conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005522165300000061
    Figure BDA0005522165300000061
  • Figure BDA0005522165300000072
    Figure BDA0005522165300000072
  • Figure BDA0005522165300000081
    Figure BDA0005522165300000081
Patent Text Reader

Abstract

The invention discloses a chip defect real-time detection method based on improved YOLOv13, and the method comprises the steps: constructing an industrial chip data set containing a defect-free sample, and carrying out the definition screening and class balance division; a difficult sample perception dynamic enhancement strategy (HSA-DA) is innovatively proposed, a difficult sample is identified based on dispersion, relative size and contrast feature space, and virtual dense cluster enhancement, local super-resolution reconstruction and adaptive contrast stretching are specifically implemented; a network structure is improved, a multi-scale compressed sensing module CS-C3k2 is introduced into a Backbone shallow layer to enhance defect feature extraction, a Sim-DSC3k2 module without parameter SimAM attention is integrated in a Neck deep layer, and a chip defect real-time detection model is constructed based on an improved YOLOv13 network structure and used for chip defect real-time detection. According to the method, 118FPS real-time detection is finally realized under a lightweight parameter quantity (2.45 M), mAP50 is improved to 0.807 (+ 4%), mAP50-95 is improved to 0.573 (+ 7.5%), and the detection robustness of sparse / tiny / low-contrast defects is remarkably optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial vision inspection technology, specifically to a real-time high-precision chip defect detection method based on an improved YOLOv13, which is applicable to automated chip surface defect detection scenarios in semiconductor manufacturing. Background Technology

[0002] In the semiconductor manufacturing industry, real-time detection of chip surface defects (such as contamination, cracks, and breakage) is a core aspect of ensuring product quality. Traditional manual inspection is inefficient and prone to missing defects, while deep learning-based inspection methods have become the mainstream solution. Among them, the YOLO series of models, with their excellent real-time performance and accuracy, are widely used in industrial defect detection scenarios.

[0003] However, existing YOLO models still have significant shortcomings in chip defect detection: (1) Sensitivity to background interference: chip surfaces have background noise such as metallic reflections and complex textures. When extracting high-resolution features, the shallow convolution operation of YOLOv13 is difficult to effectively distinguish between small defects and background interference, resulting in a large number of false detections; (2) Weakness of small defect features: chip defects are usually small targets (<32×32 pixels). In the deep feature extraction stage, the downsampling operation of YOLOv13 causes the loss of defect details, while existing attention mechanisms (such as HyperACE) focus on global high-order correlation modeling and fail to explicitly enhance the differential features of local defect regions; (3) Real-time detection accuracy: although the lightweight design of existing YOLOv13, such as DSConv, reduces the amount of computation and can meet the industrial detection benchmark of ≥30FPS, there is still a lot of room for improvement in the detection accuracy of small defects; (4) Difficulty in chip sample collection: the proportion of sparse, small and low-contrast defect samples in real production lines is insufficient, and existing datasets are difficult to cover complex scenes, resulting in limited model generalization ability.

[0004] Therefore, there is an urgent need for a new detection framework that can suppress background noise at a shallow level, enhance the difference in defect features at a deep level, and maintain lightweight characteristics to meet real-time requirements. Summary of the Invention

[0005] To address the shortcomings of existing chip defect detection technologies, such as difficulty in effectively suppressing complex background interference, insufficient feature extraction capability for small and low-contrast defects, and the sacrifice of real-time performance in pursuit of high accuracy, which fails to meet the stringent requirements of industrial online inspection, this invention proposes a real-time chip defect detection method based on an improved YOLOv13. By innovatively improving the network structure and introducing a difficult sample-oriented preprocessing mechanism, the detection accuracy and background robustness are significantly improved while ensuring real-time performance.

[0006] The technical solution adopted in this invention is as follows:

[0007] A real-time chip defect detection method based on an improved YOLOv13, comprising the following steps:

[0008] Step 1: Acquire images of the chip surface and construct a chip dataset; the images include chip surface images with and without defects;

[0009] Step 2: Defects are marked based on dispersion, relative size, and contrast, and a dynamic enhancement strategy for hard sample perception is implemented to preprocess the chip image;

[0010] Step 3: Improve the Backbone and Neck in the YOLOv13 network structure, and build a real-time chip defect detection model based on the improved YOLOv13 network structure;

[0011] Step 4: Train the real-time chip defect detection model using the preprocessed chip image, use the trained model to detect chip defects in real time, predict the location and category information of chip defects, and output the detection results on the image.

[0012] Furthermore, in step 1, the chip surface image is preprocessed, including the following operations:

[0013] Step 1.1: Use a dataset of chip surface defects from an industrial production line.

[0014] Step 1.2: Use the Brenner gradient evaluation function to calculate the square of the gray level difference between two adjacent pixels, which is expressed as the image sharpness calculation result D(f). Based on D(f), blurry images are removed.

[0015] Step 1.3: Mark the defects in all chip images;

[0016] Step 1.4: Divide the dataset into training set, validation set and test set.

[0017] Furthermore, step 2 is as follows:

[0018] Step 2.1: Preprocess the images in the dataset; the preprocessing includes, but is not limited to, normalization, random flipping, cropping, and data augmentation.

[0019] Step 2.2: Perform coordinate transformation on the defects in the image to obtain the pixel coordinates of the defect box, denoted as (x1, y1, x2, y2);

[0020] Step 2.3: Establish a defect dispersion model and obtain the defect dispersion D. ij ;

[0021] Step 2.4: Establish a relative size model of the defect and obtain the relative size S of the defect.i ;

[0022] Step 2.5: Establish a defect contrast model and obtain the defect contrast E. i

[0023] Step 2.6, according to D i S i E i The relationship with the threshold is used to mark the defect type, and the method is as follows:

[0024] If D ij >D th The defect is then labeled "discrete"; if S i <S th The defect is then marked as "minor"; if E i <E th The defect is then labeled "low contrast". If a defect is labeled with any two or more of the following terms, it is classified as a "compound difficult sample". th S is the dispersion threshold. th E is a relative size threshold. th The contrast threshold;

[0025] Step 2.7: Enhance the defects to obtain samples that are difficult to enhance;

[0026] Step 2.8: Incorporate the obtained enhanced hard samples into the original training set to construct a new training set containing the newly incorporated hard samples.

[0027] Furthermore, for "discrete" defects, virtual dense cluster enhancement is performed. N copies are generated in the original image according to the defect region size w×h. Seamless cloning or simple overlay is performed by randomly sampling the positions (cx, cy) that meet the non-overlapping conditions.

[0028] Furthermore, for "minor" defects, local super-resolution reconstruction is performed. The defect area is first upsampled by bicubic interpolation by 2x, and then the edges are improved by sharpening convolution kernel.

[0029] Furthermore, to address the "low contrast" defect, adaptive contrast stretching is performed by first applying CLAHE in the LAB space and then performing gamma correction.

[0030] Furthermore, for complex and difficult samples, in addition to adopting enhancement strategies for different types of defects, the intermediate results are obtained by sequentially applying local super-resolution reconstruction / virtual dense cluster enhancement or contrast enhancement to the defect area. The background is randomly selected from other images, and the intermediate results are pasted into random and reasonable positions.

[0031] Furthermore, the process of constructing a real-time chip defect detection model based on the improved YOLOv13 network structure is as follows:

[0032] Step 3.1: Replace the DSC3k2 module in the P2 layer of the shallow Backbone with the CS-C3k2 module. Also, modify the repeats of the CS-C3k2 module in the YAML file to 1. After collecting coarse-grained features, use the CS-C3k2 module as a low-order feature extractor. The CS-C3k2 module first applies a 1×1 convolutional layer to unify the channels, and then divides the features into two parts. One part is fed into multiple cascaded CS-C3k modules, and the other part is connected through shortcuts. Finally, the output is cascaded and fused through a 1×1 convolutional layer.

[0033] Step 3.2: Replace the DSC3k2 modules in the deep layers P17 and P21 of the Neck layer with Sim-DSC3k2 modules; after collecting fine-grained features, use the Sim-DSC3k2 module as a high-order feature extractor. The Sim-DSC3k2 module first performs pointwise operations through a 1×1 ordinary convolution, and then divides the number of channels into 0.5c and 0.5c through a Split block. The first part of 0.5c is fed into n cascaded Sim-DS-C3k modules and output, and then performs a channel concat operation with the second part of 0.5c. Finally, the output is fused with the 1×1 convolutional layer.

[0034] Furthermore, the CS-C3k module includes a compressed sensing block and a multi-scale feature enhancement block, as detailed below:

[0035] Step 3.1.1: Given the input feature map of the CS-C3k module Processing is performed through three parallel branches:

[0036] (1) First branch: First, halve the number of channels using a 1×1 convolution. Then, adaptive average pooling is used to downsample to 1×1 to obtain Finally, a 1×1 convolution is used to restore the number of channels to C.

[0037] (2) The second branch: The input feature map is processed by a 1×1 convolution to obtain...

[0038] (3) The third branch: keep the input feature map F unchanged;

[0039] The outputs of the three branches are summed element by element to obtain the output of the compressed sensing block:

[0040] F c =F+F expand +F conv

[0041] Step 3.1.2: Given the module input feature map of compressed sensing Generate feature maps at three scales:

[0042] (1) Low-resolution feature maps are downsampled to [resolution values] using bilinear interpolation. get

[0043] (2) Medium resolution feature map: Keep the original size F c,m =F c ;

[0044] (3) High-resolution feature map: obtained by upsampling to 2H×2W through bilinear interpolation.

[0045] Feature enhancement is performed on the feature map at each scale separately. For each feature map F at each scale... c,i For i∈{l,m,h}, the following operations are performed:

[0046] (a) First branch: Upsample the feature map to a spatial size of 1×1 using bilinear interpolation to obtain G. i1 Then compared with the original feature map F c,i Element-by-element multiplication yields Where σ is the Sigmoid activation function;

[0047] (b) Second branch: Obtain the feature map through global average pooling. Then compared with the original feature map F c,i Element-by-element multiplication yields

[0048] (c) Add the results of the two branches: F s,i =D i +A i ;

[0049] The enhanced feature maps at the three scales are adjusted to the same spatial size H×W, and the three adjusted feature maps are added element-wise to obtain the output of the multi-scale feature enhancement block:

[0050] F cs =Up(F s,l )+F s,m +Down(F s,h )

[0051] Finally, the input is fed into the next-level CS-C3k2 module; where Up represents upsampling and Down represents downsampling.

[0052] Furthermore, the Sim-DS-C3k module decomposes the feature channels into dual channels, and then performs discontinuous processing through n cascaded SimAM and DS-Bottleneck modules. Simultaneously, the input features enter the horizontal 1×1 convolutional branch, and finally the features of the two branches are cascaded along the channel dimension, and the feature channels are recovered using 1×1 convolutional layers. This design retains the cross-channel branch of the CSP structure, while integrating a depth-separable lightweight bottleneck.

[0053] The beneficial effects of this invention are:

[0054] (1) Suppressing background interference and improving detection robustness: By introducing the CS-C3k2 module in the shallow feature extraction stage of the Backbone layer, this invention can effectively compress and suppress complex backgrounds (such as circuit textures and noise) in chip images, enabling the network to focus on potential defect areas in the early stage, reducing the interference of background information on defect feature extraction, and improving the detection robustness of the model in complex industrial environments.

[0055] (2) Efficiently enhance defect features and improve detection accuracy: By introducing the Sim-DS-C3k2 module without parameter increase in the deep feature fusion stage of the Neck layer, this invention can utilize the local self-similarity of the image to efficiently capture and amplify the local regional differences of small, low-contrast or irregularly shaped defects, enhance the expressive ability of key defect features, and improve the detection accuracy of the model for small target defects.

[0056] (3) Overcoming the bottleneck of difficult sample detection: The HSA-DA strategy systematically solves the detection problems of three types of defects: spatial sparse defects, micro defects and low contrast defects. For spatial sparse defects, virtual dense cluster enhancement enables the model to learn the spatial correlation of defects; for micro defects, local super-resolution reconstruction improves the sub-pixel level feature recognition; for low contrast defects, adaptive contrast stretching enhances the defect-background distinction.

[0057] (4) Balancing real-time performance and high precision: Compared with other methods that improve precision by increasing network depth or the number of parameters, this invention focuses on the lightweight design and parameterless characteristics (SimAm) of the CS-C3k2 module and Sim-DS-C3k2 module when introducing them. This allows the improved YOLOv13 network to improve detection performance while reducing the number of parameters, thus meeting the stringent requirements of real-time and high-precision chip defect detection on industrial production lines.

[0058] (5) Enhance the generalization ability of the model: This invention suppresses the background and enhances the foreground features at different stages, making the features learned by the network more representative and discriminative, avoiding overfitting the background texture, and improving the model's ability to generalize and detect chip defects under different batches, types and imaging conditions. Attached Figure Description

[0059] Figure 1 This is the overall operational logic diagram of the method of the present invention.

[0060] Figure 2 This is the result of implementing HSA-DA in this invention.

[0061] Figure 3 This is a diagram of the improved YOLOv13 algorithm network structure provided by the method of this invention.

[0062] Figure 4 This is a structural diagram of the CS-C3k2 module provided by the present invention.

[0063] Figure 5 This is a structural diagram of the Sim-DS-C3k2 module provided by the present invention.

[0064] Figure 6 This is a schematic diagram comparing the performance of the algorithm before and after the improvement provided by this invention.

[0065] Figure 7 This is a schematic diagram comparing the detection results before and after the algorithm improvement provided by this invention. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.

[0067] Reference Figure 1 This invention designs a real-time chip defect detection method based on an improved YOLOv13, and the specific implementation steps are as follows:

[0068] Step 1: Construct a chip dataset by collecting images of real chips from real industrial production lines, and divide it into training, validation and test sets in a ratio of 7:1.5:1.5.

[0069] The specific process is as follows:

[0070] Step 1.1: Collect a dataset of chip surface defects from industrial production lines. In addition, collect a certain amount of defect-free chip surface data and incorporate defect-free chip images into the real chip dataset to enhance the model's ability to identify non-defect areas and suppress background noise. This dataset primarily contains real-world chip images and does not rely on synthetic defect samples for core model training.

[0071] Step 1.2: Use the Brenner gradient evaluation function to calculate the square of the gray-level difference between two adjacent pixels, which is expressed as the image sharpness calculation result, denoted as D(f). If D(f) < D th They are treated as blurry images and removed.

[0072]

[0073] Where f(x,y) represents the gray value of the pixel (x,y) corresponding to image f, and D th This is the image sharpness threshold.

[0074] Step 1.3: Use the LabelImg annotation tool to annotate all chip images, use horizontal boxes to outline the defect locations, and assign a serial number to each defect, such as Pollution as 0, Crack as 1, Broken as 2, and No-die as 3, and save it in YOLO format.

[0075] Step 1.4: Construct training, validation, and test sets according to a preset ratio and by category (7:1.5:1.5) to avoid class imbalance issues in the training, validation, and test sets during partitioning.

[0076] Step 2: Expand the dataset using data augmentation strategies, then select samples from the training set to calculate dispersion, relative size, and contrast, quantify and analyze defect features, and construct a defect feature space model composed of dispersion, relative size, and contrast. Based on the constructed defect feature space model, identify difficult samples and implement a dynamic hard sample perception augmentation strategy (HSA-DA) for these samples, such as... Figure 2 The HSA-DA results are shown.

[0077] The specific process is as follows:

[0078] Step 2.1: In addition to performing routine image size normalization, random flipping, and cropping on the selected training set images, we call the data augmentation strategy built into the YOLOv13 network, such as copy_paste, to further expand the diversity of the dataset by copying and pasting the defective regions onto different backgrounds.

[0079] Step 2.2: Iterate through all .txt files in the labels / directory of the training set and read each line cls x yw h. Convert the normalized center coordinates and width and height into pixel coordinate boxes, so the defect box is denoted as (x1,y1,x2,y2).

[0080] Step 2.3: Establish a defect dispersion model. If there are more than one defect of the same type in the same image, then aggregate them by type and calculate the Euclidean distance D between each pair of defect centers i and j. ij The average value is taken as the dispersion D of all defects in this category. avg Otherwise let D ij =0:

[0081] D ij =||C i -C j ||2

[0082] Among them, C i C j These are the coordinates of the center points i and j of the two defect regions, respectively. D ij It is the Euclidean distance between each pair of defect centers i and j, denoted as Calculate the distances between all defect centers and take their average as the dispersion of that category. N is the total number of defects in the same category.

[0083] Step 2.4: Establish a relative size model of the defect, obtained by dividing the pixel area of ​​the defect region by the total pixel area of ​​the image.

[0084]

[0085] Among them, A defect =(x2-x1)(y2-y1), a defect It is the pixel area of ​​the defect region, A image =H×W,A image y1 is the total pixel area of ​​the image, (x1, y1) is the center coordinate in pixel coordinates, (x2, y2) is the width and height coordinates in pixel coordinates, H is the height of the original image, and W is the width of the original image.

[0086] Step 2.5: Establish a defect contrast model:

[0087] E i =|μ defect -μ background |

[0088] Where, μ defect It is the average value of pixels in the defective region (the average brightness of the defective region), μ backgroundIt is the average value of the background pixels within 10 pixels surrounding the defect (the average brightness of the background area around the defect), E i This is the contrast index value of the i-th defect. Specific calculation: I(x,y) is the pixel value in the defect region, N d It represents the number of pixels in the defective area. I(x,y) Same as above, N b It is the effective number of pixels in the background area.

[0089] Step 2.6: Put all D ij S i E i After collecting them into the list, calculate D respectively. ij S i E i The corresponding threshold D th S th E th According to D ij S i E i Based on the relationship with the corresponding threshold, defects are marked; when a defect is marked by any two or more thresholds at the same time, the defect is judged as a "hard sample" and recorded in the list hard_samples.

[0090] More specifically, according to D ij S i E i The relationship between D and the threshold, and the method for marking defects is as follows: If D ij >D th The defect is then labeled "discrete"; if S i <S th The defect is then marked as "minor"; if E i <E th The defect is then marked as "low contrast". If a defect is marked with any two or more labels at the same time, the defect is determined to be a "compound difficult sample".

[0091] More specifically, the fields recorded in the list `hard_samples` also include the original image path (`image_path`); the defect bounding box coordinates (`x1`, `y1`, `x2`, `y2`); the category ID (`cls_id`); and the set of difficulty types (`types`).

[0092] D th =median({D ij}),S th =percentile s% ({S i}), E th =percentile c% ({Ei})

[0093] Among them, D th The dispersion threshold is the median dispersion of the entire training set, denoted as median({D ij}); S th As a relative size threshold, the s% quantile of the total training set size is taken, denoted as percentile. s% ({S i}); E th The contrast threshold represents the c% quantile of the contrast across the entire training set, denoted as percentile. c% ({E i}

[0094] Step 2.7: Apply the Hard Sample Awareness Dynamic Augmentation Strategy (HSA-DA) to each record in the hard_samples list, performing the augmentation operations sequentially and generating new images and label files respectively. The specific augmentation operations are as follows:

[0095] Perform virtual dense cluster enhancement on defects of type "discrete", such as... Figure 2 The virtual dense cluster enhancement results shown are achieved by generating several cloned defects (virtual dense clusters) in the original image, allowing the model to learn the spatial correlation and aggregation patterns between defects, thereby improving the detection capability of sparse targets. N copies are generated according to the defect region size w×h, and seamless cloning or simple overlay is performed by randomly sampling positions (cx, cy) that meet the non-overlapping conditions.

[0096] Step 2.7.1: Let the original drawing size be H×W, and the defect frame be (x1, y1, x2, y2). The width and height w of the defect area are then obtained.

[0097] x2-x1, h=y2-y1, cropping from the original image yields D, copying D yields N clones:

[0098] D = img[y1:y2,x1:x2]

[0099] {D1,D2,D3,...,D N}

[0100] Step 2.7.2: For each replica D i A center point (c) is randomly sampled from the effective area of ​​the image. x ,c y The rectangle must not strongly overlap with the original defect frame.

[0101]

[0102] Satisfy (x1-w<c) x<x2+w∧y1-h<c y <y2+h)

[0103] Step 2.7.3: Use OpenCV's seamlessClone to seamlessly clone the boundary at the target location using Poisson blending. If seamless cloning fails, simply overwrite the original. Additionally, generate a new defect box at each pasted location, convert it to YOLO format, and write the label.

[0104]

[0105] Among them, aug i mg represents the final output enhanced image, containing multiple virtual dense cluster defects, c y The vertical center coordinates (row coordinates, representing the height direction) of the defect clone pasting position are specified. `resize` indicates that the i-th cloned defect area is adjusted to the specified size w×h, i.e., the original width and height dimensions of the defect, and then pasted onto the enhancement map `aug`. i mg.

[0106] For defects classified as "minor", perform local super-resolution reconstruction, such as... Figure 2 The local super-resolution reconstruction results shown enhance the discriminative features of defects through regional magnification and sharpening, enabling downstream feature extractors to better capture sub-pixel details. The defect region is first upsampled by a factor of 2 using bicubic interpolation, and then the edges are enhanced using a sharpening convolutional kernel.

[0107] Step 2.7.4: Defect area Upsampling by a scaling factor s = 2 yields a size of 2h × 2w:

[0108] D ↑2 =cv2.resize(D,fx=2,fy=2,interpolation=INTER CUBIC )

[0109] Here, `cv2.resize` represents the OpenCV image scaling function, used to resize the image. `fx` and `fy` represent doubling the image's horizontal width (width × 2) and vertical height (height × 2), respectively. `interpolation` represents the interpolation algorithm parameters, used to determine how to fill in the grayscale values ​​of new pixels when scaling up or down. CUBIC This represents bicubic interpolation, which uses 16 pixels in a 4x4 neighborhood to calculate the new pixel value.

[0110] Step 2.7.5: Use convolution kernel K to perform a convolution on D. ↑2 Perform 2D convolution to obtain the sharpened D. sharp :

[0111]

[0112] Where u and v represent the horizontal (x-axis) offset and vertical (y-axis) offset of the convolution kernel, respectively, and K(u,v) is the (u,v)th element of the convolution kernel K. The kernel K is used to apply the input image D... ↑2 A weighted summation, or convolution operation, is performed in the neighborhood of this point.

[0113] For defects classified as "low contrast", perform adaptive contrast stretching, such as... Figure 2 The adaptive contrast stretching results shown improve the network's sensitivity to subtle boundary and texture changes by enhancing the grayscale difference between the foreground and background.

[0114] First, apply CLAHE in the LAB space, then perform gamma correction with γ = 0.5:

[0115] Step 2.7.6: Convert the color space of the defect region D from BGR to LAB, and apply contrast-limited adaptive histogram equalization to the luminance channel L. This operation can stretch the grayscale histogram within a local window, but it limits over-enhancement.

[0116] (L,a,b)=cv2.cvtColor(D,COLOR BGR2LAB )

[0117] L clahe =CLAHE(L;clipLimit=3.0,tileGridSize=(8,8))

[0118] Here, `cv2.cvtColor` represents the color space conversion function in OpenCV, used to convert an image from one color space to another. `cv2.cvtColor` also represents the color space conversion function. BGR2LAB This is a color space conversion identifier used to convert an image from the BGR color space to the CIE LAB color space. clahe This is the luminance channel result after CLAHE (Local Histogram Equalization) processing. Luminance details are locally enhanced. CLAHE represents contrast-limited adaptive histogram equalization, clipLimit represents the threshold for limiting contrast enhancement, and tileGridSize represents dividing the entire image into 8x8 local small grids and performing histogram equalization separately in each small grid.

[0119] Step 2.7.7: Perform a gamma transform on the merged image to enhance shadow details, and finally convert it back to BGR.

[0120] table[i] = (i / 255) γ×255,i=0,1,…,255,γ=0.5

[0121] L gamma (x,t)=table[L clahe (x,y)]

[0122] LAB out =(L gamma (a,b)

[0123] D enhanced =cv2.cvtColor(LAB out COLOR LAB2BGR )

[0124] Where table[i] represents the value at index iii in the gamma correction lookup table, used to quickly map the enhanced result of pixel brightness values, i is the pixel grayscale value index, ranging from 0 to 255, γ is the gamma transform coefficient, controlling the intensity of brightness adjustment, and L... gamma (x,y) represents the brightness channel L enhanced by CLAHE. clahe The gamma transform is applied to improve the contrast between light and dark areas of the image. clahe (x,y) represents the pixel value of the brightness channel of the image after processing by the CLAHE algorithm. out This is the LAB image after replacing the luminance channel (the luminance channel is enhanced with CLAHE+Gamma, while the color channels remain unchanged). a and b represent the two color channels in the LAB space of the original image, D... enhanced For the final enhanced image, convert it back to the BGR color space, COLOR LAB2BGR This is a flag indicating that a LAB format image should be converted to a BGR format image.

[0125] The three strategies mentioned above, focusing on spatial layout, resolution details, and grayscale contrast respectively, work together to improve the network's detection performance for sparse, small, and low-contrast targets.

[0126] Step 2.7.8: For complex and difficult samples, the above strategies are superimposed and fused to apply local super-resolution reconstruction / virtual dense cluster enhancement or contrast enhancement to the defective region in sequence to obtain intermediate results. The background is randomly selected from other images, and the intermediate results are pasted into random and reasonable positions.

[0127] Step 2.8: Incorporate the obtained enhanced hard samples into the training set along with the original training set to construct a new training set containing the newly incorporated hard samples.

[0128] Step 3: Based on, for example Figure 3The improved YOLOv13 network structure shown is used to build a real-time chip defect detection model, including improvements to the Backbone and Neck. The CS-C3k2 module is introduced into the shallow layer of the Backbone, which reduces background interference and enhances foreground defect features through multi-path design, multi-scale processing and feature enhancement mechanisms. The Sim-DSC3k2 module is introduced into the deep layer of the Neck, which improves the model's attention to important features without increasing parameters through its unique parameterless characteristics.

[0129] Referring to the attached figures, the network structure is as follows: The network structure of the improved YOLOv13 target detection algorithm includes a Backbone module, a Neck module, and a detection head; the Backbone module includes a first Conv unit, a second Conv unit, a first CS-C3k2 unit, a third Conv unit, a first DSC3k2 unit, a first DSConv unit, a first A2C2f unit, a second DSConv unit, a second A2C2f unit, and a HyperACE unit connected in sequence; the first DSC3k2 unit, the first A2C2f unit, and the second A2C2f unit are connected to the HyperACE unit; the Backbone module extracts features from the input chip image;

[0130] The Neck module consists of three processing paths. The first processing path includes a first Upsample unit, a first Concat module, a first Sim-DS-C3k2 unit, a second Upsample unit, a second Concat module, and a second Sim-DS-C3k2 unit connected in sequence. The first DSC3k2 unit and the HyperACE unit are added together and then connected to the second Concat module; the first A2C2f unit and the HyperACE unit are added together and then connected to the first Concat module.

[0131] The second processing route includes a fourth Conv unit, a third Concat unit, a first DS-C3k2 unit, a fifth Conv unit, and a fourth Concat unit connected in sequence. The second Sim-DS-C3k2 unit and the HyperACE unit are added together and then connected to the fourth Conv unit; the first Sim-DS-C3k2 unit and the HyperACE unit are added together and then connected to the third Concat unit.

[0132] The third processing route includes a fourth Concat unit and a second DS-C3k2 unit connected in sequence. The second A2C2f unit and the HyperACE unit are added together and then connected to the fourth Concat unit.

[0133] The detection head module consists of three heads. The first detect head is connected after the second Sim-DS-C3k2 unit and the HyperACE unit are added together. The third Concat unit is connected to the second detect head. The third detect head is connected after the second DS-C3k2 unit and the HyperACE unit are added together.

[0134] The specific processing procedure for this network structure is as follows:

[0135] Step 3.1: Replace the DSC3k2 module in the P2 layer of the Backbone shallow layer with the CS-C3k2 module. This aims to effectively reduce the interference of complex backgrounds (such as noise and texture) in the chip image during the coarse-grained stage (high-resolution, shallow feature extraction stage), while strengthening the foreground (defect) features through multi-level feature enhancement blocks. Additionally, modify the repeats of the CS-C3k2 module in the YAML file to 1. After collecting coarse-grained features, the CS-C3k2 module is used as a low-order feature extractor. The CS-C3k2 module first applies a 1×1 convolutional layer to unify the channels, then divides the features into two parts. One part is fed into multiple cascaded CS-C3k modules, and the other part is connected via shortcuts. Finally, the output is cascaded and fused through a 1×1 convolutional layer. For example... Figure 4 The CS-C3k module includes a compressed sensing block and a multi-scale feature enhancement block.

[0136] Step 3.1.1: Given the input feature map of the CS-C3k module Processing is performed through three parallel branches:

[0137] (1) First branch: First, halve the number of channels using a 1×1 convolution. Then, adaptive average pooling is used to downsample to 1×1 to obtain Finally, a 1×1 convolution is used to restore the number of channels to C.

[0138] (2) The second branch: The input feature map is processed by a 1×1 convolution to obtain...

[0139] (3) The third branch: keep the input feature map F unchanged;

[0140] The outputs of the three branches are summed element by element to obtain the output of the compressed sensing block:

[0141] F c =F+F expand +F conv

[0142] Step 3.1.2: Given the module input feature map of compressed sensing Generate feature maps at three scales:

[0143] (1) Low-resolution feature maps are downsampled to [resolution values] using bilinear interpolation. get

[0144] (2) Medium resolution feature map: Keep the original size F c,m =F c ;

[0145] (3) High-resolution feature map: obtained by upsampling to 2H×2W through bilinear interpolation.

[0146] Feature enhancement is performed on the feature map at each scale separately. For each feature map F at each scale... c,i For i∈{l,m,h}, the following operations are performed:

[0147] (a) First branch: Upsample the feature map to a spatial size of 1×1 using bilinear interpolation to obtain G. i1 Then compared with the original feature map F c,i Element-by-element multiplication yields Where σ is the Sigmoid activation function;

[0148] (b) Second branch: Obtain the feature map through global average pooling. Then compared with the original feature map F c,i Element-by-element multiplication yields

[0149] (c) Add the results of the two branches: F s,i =D i +A i ;

[0150] The enhanced feature maps at the three scales are adjusted to the same spatial size H×W, and the three adjusted feature maps are added element-wise to obtain the output of the multi-scale feature enhancement block:

[0151] F cs =Up(F s,l )+F s,m +Down(F s,h )

[0152] Finally, the input is fed into the next-level CS-C3k2 module; where Up indicates upsampling and Down indicates downsampling.

[0153] Step 3.2: Replace the DSC3k2 modules in the deep P17 and P21 layers of the Neck layer with Sim-DSC3k2 modules. After collecting fine-grained features, use the Sim-DSC3k2 module as a high-order feature extractor, such as... Figure 5 The Sim-DSC3k2 module first performs pointwise operations using a 1×1 ordinary convolution, then divides the number of channels into 0.5c and 0.5c using a Split block. The first part of 0.5c is fed into n cascaded Sim-DS-C3k modules and output, and then performs a channel concat operation with the second part of 0.5c. Finally, the output is fused with the 1×1 convolutional layer.

[0154] The SimAM module is a parameter-free attention mechanism that utilizes the spatial context information of the feature map itself to calculate the attention weight of each pixel in the fine-grained stage (low resolution, deep feature fusion stage), thereby adaptively strengthening important features (such as small defect regions), effectively amplifying the local regional differences of defects, and significantly reducing computational overhead. The Sim-DS-C3k module decomposes the feature channel into dual channels, and then performs discontinuous processing through n cascaded SimAM and DS-Bottleneck modules. Simultaneously, the input features enter the horizontal 1×1 convolutional branch, and finally the features of the two branches are cascaded along the channel dimension, and the feature channel is recovered using a 1×1 convolutional layer. This design retains the cross-channel branch of the CSP structure, while integrating a lightweight bottleneck with separable depth.

[0155] Its specific working principle is as follows: For the input feature map X, the optimal attention weight for each location is derived based on the saliency of each position within its local region (calculated using an energy function), without requiring additional learnable parameters. Finally, after evaluating the importance E groups of neurons, the calculated attention weights are multiplied element-wise with the original feature map X to obtain the feature map enhanced by the SimAM module. Specifically, as follows:

[0156]

[0157]

[0158] in, Here, E is the feature map after processing by the SimAM module, E is the energy function used to measure the importance of each neuron, and x is the input feature map, with E as above. ij Let be the neuron value in the i-th row and j-th column of the feature map, n be the total number of neurons in the feature map (i.e., the product of height and width), λ be the regularization term, and μ be the mean value in the height and width dimensions.

[0159] Step 4: Train the constructed chip defect real-time detection model using the constructed dataset, and use the trained model to perform real-time chip defect detection.

[0160] In step 4, the predicted chip defect location and category information includes selecting the optimal score threshold based on the confidence score of the evaluation results, removing bounding boxes with confidence scores lower than the threshold, and then processing the sorted bounding boxes based on non-maximum suppression (NMS) to retain the effective detection boxes with the highest confidence.

[0161] All model training was run on a GPU server, with an NVIDIA RTX 4060 GPU and an AMD Ryzen 9 7940HX CPU. The programming language was Python 3.11, and the torch version was 2.2.2+cu118.

[0162] Step 4.1: As Figure 6 Our performance evaluation method, including precision (P), recall (R), recognition speed (FPS), comprehensive performance evaluation metrics (mAP@0.5, mAP@75, mAP@50-95), and number of parameters, is calculated using the following formula, where r1, r2…rn are the recall values ​​corresponding to the first interpolation point of the Precision interpolation segment arranged in ascending order:

[0163] P = TP / (TP + FP)

[0164] R = TP / (TP + FN)

[0165]

[0166] The specific performance metrics are shown in the table below. Compared to the original YOLOv13 model, our method can significantly improve detection accuracy while meeting real-time requirements. Specifically, mAP50 is improved by 4%, mAP75 by 3.3%, and mAP50-95 by 7.5%, while reducing the number of parameters by 952.

[0167] Table 1 Comparison of performance results before and after improvement

[0168]

[0169]

[0170] Step 4.2: To further verify the effectiveness of our method, we compared it with other object detection methods. The recognition results are shown in Table 2. As can be seen from Table 2, our method outperforms some existing mature methods.

[0171] Table 2 Comparison with advanced models

[0172]

[0173] Step 4.3: Compare the detection results using both the original YOLOv13 and the improved YOLOv13 models. The recognition and detection results are as follows: Figure 7 As shown, both the traditional YOLOv13 and the improved network can accurately identify defects and defect4 defects. However, for small, difficult-to-identify defects (hard samples), the traditional YOLOv13 fails to identify them effectively, while our method can accurately identify small defects with a confidence level of 0.6. These detection results demonstrate that the improved model is more accurate in identifying both large and small targets, improving the precision and robustness of defect detection. It has greater practical value. The above embodiments are only used to illustrate the design concept and features of this invention, and their purpose is to enable those skilled in the art to understand the content of this invention and implement it accordingly. The scope of protection of this invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made based on the principles and design ideas disclosed in this invention are within the scope of protection of this invention.

Claims

1. A real-time chip defect detection method based on an improved YOLOv13, characterized in that, Includes the following steps: Step 1: Acquire images of the chip surface and construct a chip dataset; the images include chip surface images with and without defects; Step 2: Defects are marked based on dispersion, relative size, and contrast, and a dynamic enhancement strategy for hard sample perception is implemented to preprocess the chip image; Step 3: Improve the Backbone and Neck in the YOLOv13 network structure, and build a real-time chip defect detection model based on the improved YOLOv13 network structure; Step 4: Train the real-time chip defect detection model using the preprocessed chip image, use the trained model to detect chip defects in real time, predict the location and category information of chip defects, and output the detection results on the image.

2. The real-time chip defect detection method based on improved YOLOv13 according to claim 1, characterized in that, In step 1, the chip surface image is preprocessed, including the following operations: Step 1.1: Use a dataset of chip surface defects from an industrial production line. Step 1.2: Use the Brenner gradient evaluation function to calculate the square of the gray level difference between two adjacent pixels, which is expressed as the image sharpness calculation result D(f). Based on D(f), blurry images are removed. Step 1.3: Mark the defects in all chip images; Step 1.4: Divide the dataset into training set, validation set and test set.

3. The real-time chip defect detection method based on improved YOLOv13 according to claim 1, characterized in that, Step 2 is as follows: Step 2.1: Preprocess the images in the dataset; the preprocessing includes, but is not limited to, normalization, random flipping, cropping, and data augmentation. Step 2.2: Perform coordinate transformation on the defects in the image to obtain the pixel coordinates of the defect box, denoted as (x1, y1, x2, y2); Step 2.3: Establish a defect dispersion model and obtain the defect dispersion D. ij ; Step 2.4: Establish a relative size model of the defect and obtain the relative size S of the defect. i ; Step 2.5: Establish a defect contrast model and obtain the defect contrast E. i ; Step 2.6, according to D ij S i E i The relationship with the threshold is used to mark the defect type, and the method is as follows: If D ij >D th Then the defect is labeled "discrete"; if S i <S th The defect is then marked as "minor"; if E i <E th The defect is then labeled "low contrast". If a defect is labeled with any two or more of the following terms, it is classified as a "compound difficult sample". th S is the dispersion threshold. th E is a relative size threshold. th The contrast threshold; Step 2.7: Enhance the defects to obtain samples that are difficult to enhance; Step 2.8: Incorporate the obtained enhanced hard samples into the original training set to construct a new training set containing the newly incorporated hard samples.

4. The real-time chip defect detection method based on improved YOLOv13 according to claim 3, characterized in that, For "discrete" defects, virtual dense cluster enhancement is performed. N copies are generated in the original image according to the defect region size w×h. Seamless cloning or simple overlay is performed by randomly sampling the positions (cx, cy) that meet the non-overlapping conditions.

5. The chip defect real-time detection method based on improved YOLOv13 according to claim 3, characterized in that, For "minor" defects, local super-resolution reconstruction is performed. The defect area is first upsampled by bicubic interpolation by 2x, and then the edges are improved by sharpening convolution kernel.

6. The chip defect real-time detection method based on improved YOLOv13 according to claim 3, characterized in that, To address the "low contrast" defect, adaptive contrast stretching is performed by first applying CLAHE in the LAB space and then performing gamma correction.

7. A real-time chip defect detection method based on an improved YOLOv13 according to claim 4, 5, or 6, characterized in that, For complex and difficult samples, in addition to adopting enhancement strategies for different types of defects, local super-resolution reconstruction / virtual dense cluster enhancement or contrast enhancement are applied sequentially to the defect area to obtain intermediate results. Backgrounds are randomly selected from other images, and the intermediate results are pasted into random and reasonable positions.

8. The real-time chip defect detection method based on improved YOLOv13 according to claim 1, characterized in that, The process of building a real-time chip defect detection model based on the improved YOLOv13 network structure is as follows: Step 3.1: Replace the DSC3k2 module in the P2 layer of the shallow Backbone with the CS-C3k2 module. Also, modify the repeats of the CS-C3k2 module in the YAML file to 1. After collecting coarse-grained features, use the CS-C3k2 module as a low-order feature extractor. The CS-C3k2 module first applies a 1×1 convolutional layer to unify the channels, and then divides the features into two parts. One part is fed into multiple cascaded CS-C3k modules, and the other part is connected through shortcuts. Finally, the output is cascaded and fused through a 1×1 convolutional layer. Step 3.2: Replace the DSC3k2 modules in the deep layers P17 and P21 of the Neck layer with Sim-DSC3k2 modules; after collecting fine-grained features, use the Sim-DSC3k2 module as a high-order feature extractor. The Sim-DSC3k2 module first performs pointwise operations through a 1×1 ordinary convolution, and then divides the number of channels into 0.5c and 0.5c through a Split block. The first part of 0.5c is fed into n cascaded Sim-DS-C3k modules and output, and then performs a channel concat operation with the second part of 0.5c. Finally, the output is fused with the 1×1 convolutional layer.

9. A real-time chip defect detection method based on improved YOLOv13 according to claim 8, characterized in that, The CS-C3k module includes a compressed sensing block and a multi-scale feature enhancement block, as detailed below: Step 3.1.1: Given the input feature map of the CS-C3k module Processing is performed through three parallel branches: (1) First branch: First, halve the number of channels using a 1×1 convolution. Then, adaptive average pooling is used to downsample to 1×1 to obtain Finally, a 1×1 convolution is used to restore the number of channels to C. (2) The second branch: The input feature map is processed by a 1×1 convolution to obtain... (3) The third branch: keep the input feature map F unchanged; The outputs of the three branches are summed element by element to obtain the output of the compressed sensing block: F c =F+F expand +F conv Step 3.1.2: Given the module input feature map of compressed sensing Generate feature maps at three scales: (1) Low-resolution feature maps are downsampled to [resolution values] using bilinear interpolation. get (2) Medium resolution feature map: Keep the original size F c,m =F c ; (3) High-resolution feature map: obtained by upsampling to 2H×2W through bilinear interpolation. Feature enhancement is performed on the feature map at each scale separately. For each feature map F at each scale... c,i For i∈{l,m,h}, the following operations are performed: (a) First branch: Upsample the feature map to a spatial size of 1×1 using bilinear interpolation to obtain G. i1 Then compared with the original feature map F c,i Element-by-element multiplication yields Where σ is the Sigmoid activation function; (b) Second branch: Obtain the feature map through global average pooling. Then compared with the original feature map F c,i Element-by-element multiplication yields (c) Add the results of the two branches: F s,i =D i +A i ; The enhanced feature maps at the three scales are adjusted to the same spatial size H×W, and the three adjusted feature maps are added element-wise to obtain the output of the multi-scale feature enhancement block: F cs =Up(F s,l )+F s,m +Down(F s,h ) Finally, the input is fed into the next-level CS-C3k2 module; where Up represents upsampling and Down represents downsampling.

10. A real-time chip defect detection method based on an improved YOLOv13 according to claim 8, characterized in that, The Sim-DS-C3k module decomposes the feature channels into dual channels, and then performs discontinuous processing through n cascaded SimAM and DS-Bottleneck modules. Simultaneously, the input features enter the horizontal 1×1 convolutional branch. Finally, the features of the two branches are cascaded along the channel dimension, and the feature channels are recovered using 1×1 convolutional layers. This design retains the cross-channel branch of the CSP structure, while integrating a depth-separable lightweight bottleneck.

Citation Information

Cited By

  • Power quality disturbance identification method based on constraint network and priori knowledge

    CN121561372A

  • Improved StarNet-YOLOv13-based unmanned aerial vehicle field tobacco virus disease lightweight detection method

    CN121767895A