High-precision chip tiny defect detection method based on dynamic Retinex enhancement and multi-scale dynamic attention

By combining the lightweight MobileNetV3Lite and dynamic Retinex algorithms with a dynamic multi-scale attention network, the imaging quality and accuracy issues of chip micro-defect detection in highly reflective environments are resolved, achieving efficient micro-defect detection and improving recall rate and inference speed.

CN120672678APending Publication Date: 2025-09-19CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510729296.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional detection methods have poor imaging quality in highly reflective environments, insufficient accuracy in detecting tiny defects, and low efficiency in multi-scale feature fusion, resulting in halo artifacts and blurred edges, low recall rate, insufficient IoU, large number of parameters, and slow inference speed.

Method used

It adopts the lightweight MobileNetV3Lite structure, introduces edge-aware bilateral filtering and dynamic adaptive Retinex algorithm, and combines it with the dynamic multi-scale attention fusion network (DMA-YOLO). By dynamically adjusting the filtering strength and feature map compression ratio, it optimizes feature fusion and detection head design, enhances defect capture capabilities, and uses dynamic loss functions and data enhancement strategies.

Benefits of technology

It effectively eliminates halo artifacts and improves image edge clarity, with a recall rate of 94.7%, multi-scale feature fusion efficiency increased to 91%, the number of parameters reduced to 32.8M, an inference speed of 67FPS, and a positioning error of ≤1.5px.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672678A_ABST
    Figure CN120672678A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision chip tiny defect detection method based on dynamic Retinex enhancement and multi-scale dynamic attention, and belongs to the field of computer vision and industrial detection. According to the method, a dynamic illumination image is generated through a MobileNetV3Lite network, reflection components are optimized in combination with edge perception bilateral filtering, and halo artifacts (the proportion is smaller than 5%) of a high-reflection area are eliminated; based on an improved YOLOv8 framework, a dynamic mixed attention module (dynamic channel compression and deformable convolution) and a dynamic weighted bidirectional feature pyramid are constructed, and the recall rate of tiny defects is increased to 94.7%; a 160 * 160 high-resolution detection head is additionally arranged, and the positioning error is smaller than or equal to 1.5 px; a defect size sensitive loss function is introduced, and training gradient distribution is optimized. The method is superior to the prior art in parameter quantity (32.8 M), reasoning speed (67FPS) and multi-scale characteristic utilization rate (91%), and can be widely applied to the field of semiconductor manufacturing and precision machinery detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and industrial inspection technology, specifically to a method for detecting nanoscale defects during chip manufacturing. By combining the Dynamic Adaptive Retinex Image Enhancement Algorithm (DARetina) with the Dynamic Multiscale Attention Fusion Network (DMA-YOLO), it effectively addresses the issues of low imaging quality in highly reflective environments, insufficient micro-defect detection accuracy, and low multi-scale feature fusion efficiency. The method is widely applicable to fields such as semiconductor wafer defect detection, MEMS device microstructure analysis, and damage identification of precision mechanical parts. Background Art

[0002] In chip manufacturing, defects such as metal layer fractures and nano-scratches are typically smaller than 10 pixels in size. Limited by highly reflective surfaces (reflectivity > 80%), traditional detection methods face the following technical bottlenecks: poor imaging quality, with local overexposure accounting for up to 42%. The traditional Retinex algorithm uses a fixed Gaussian kernel, resulting in halo artifacts and blurred edges. The ability to detect small targets is insufficient. The standard algorithm has a recall rate of less than 85% for defects ≤10 pixels, and the deformable convolutional network has an IoU of only 72% for irregular defects. Multi-scale feature utilization is low, and the fixed-weight BiFPN algorithm results in an 18% increase in missed detection rates for defects below 20μm.

[0003] Retinex algorithm: The guided filter of He et al. (2013) cannot dynamically adjust parameters, resulting in halo artifacts exceeding 30%. Regarding the attention mechanism, SENet (Hu et al., 2018) lacks targeted processing for small objects and exhibits weak feature response. Regarding feature fusion networks, EfficientDet (Tan et al., 2020) is not optimized for chip defects, resulting in only 62% small object feature utilization. Summary of the Invention

[0004] The present invention aims to provide a high-precision and high-efficiency method for detecting micro defects in chips, which solves the existing problems through the following technical improvements:

[0005] 1. Eliminate halo artifacts in highly reflective areas (accounting for less than 5%) and improve image edge clarity;

[0006] 2. Improve the recall rate of defects ≤10 pixels to 94.7%;

[0007] 3. Optimized the multi-scale feature fusion efficiency to 91%, while reducing the number of model parameters to 32.8M and the inference speed to 67FPS.

[0008] The technical solution adopted by the present invention is: a lightweight MobileNetV3Lite structure is adopted, which includes 3 layers of depth-separable convolution, wherein the input layer: 3×3 convolution, step size 2, number of channels 16; the middle layer: 3×3 and 5×5 convolution kernels, number of channels 32 / 64, activation function HSwish; the output layer: 1×1 convolution generates a single-channel dynamic illumination map I dynamic The number of parameters is only 20.9K, the inference speed is 85FPS, and it replaces the traditional fixed Gaussian kernel (σ=80) to reduce parameter restrictions.

[0009] Edge-aware bilateral filtering is introduced, and the filtering strength is dynamically adjusted by the gradient threshold. The spatial kernel σr and the range kernel σs are adaptively adjusted according to the local gradient: When σr decreases from 2.0 to 0.5, σs decreases from 10.0 to 2.0. The filtering formula is:

[0010]

[0011] The defect edge is effectively preserved, and the edge gradient attenuation rate is reduced to 15%.

[0012] Based on the C2f structure of YOLOv8, it inputs a 640×640 image and outputs multi-scale feature maps: P3 (80×80, 128 channels): shallow high-resolution features; P4 (40×40, 256 channels): mid-level features; P5 (20×20, 512 channels): deep semantic features.

[0013] For dynamic channel compression: the compression ratio r (4 to 16) is adaptively adjusted according to the feature complexity. The formula is:

[0014] r = Sigmoid(GAP(F)·W+b)×r min -r max )+r max

[0015] Compared with the standard SE network (fixed r = 16), it avoids the loss of small target features.

[0016] The offset network (OffsetNet) predicts the 2×20×20 offset Δp, and the kernel generator (KernelNet) generates a 3×3 dynamic convolution kernel w to enhance the ability to capture irregular defects.

[0017] During the feature fusion process, the bidirectional fusion path includes top-down and bottom-up branches, with weight α i Generated by the global mean of the feature map:

[0018] α i =Sigmoid(W i ·GAP(Pi ))

[0019] Compared with fixed-weight BiFPN, the utilization of multi-scale features is improved by 29%.

[0020] New 160×160 detection head: Upsamples P3 features to 160×160, with 128 channels, specifically designed for detecting defects ≤10px, with a positioning error of ≤1.5px.

[0021] In the training strategy, 100,000 images (1280×960 pixels) of 2,000 wafers were collected, containing five types of defects (metal fractures, nano-scratches, etc.); random highlights (2000-3000 lux, radius 520px), particle noise (SNR=15dB), geometric transformation (±15° rotation, 0.8-1.2 scaling) and other methods were used to enhance the data.

[0022] The defect size sensitive weight formula introduced in CIoU Loss is:

[0023] λ k ∝1 / √Area(B k )

[0024] The proportion of training gradients for minor defects increased from 12% to 45%.

[0025] After output, a coordinate mapping model is constructed, and the physical coordinate calculation formula is:

[0026] X phy =(X px -640)×0.25+X0 Y phy =(Y px -480)×0.25+Y0

[0027] Visual output: TIFF format embedded with XMP metadata (defect coordinates, categories), CSV report supports MES system docking. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 : DARetina algorithm flow chart, showing the steps of dynamic lighting estimation and reflection component optimization.

[0029] Figure 2 : DMA-YOLO network architecture, marking the connection relationship between the dynamic hybrid attention module and DW-BiFPN.

[0030] Figure 3 : Training strategy flow chart, including data enhancement, dynamic loss function design and model optimization process. DETAILED DESCRIPTION

[0031] The implementation steps of the dynamic adaptive Retinex algorithm (DARetina) are as follows: the original chip image (resolution 1280×960) is resized to 640×640 pixels and normalized to the range [0,1].

[0032] Construct a dynamic lighting estimation network (MobileNetV3Lite): Input layer: 3×3 convolution kernel, stride 2, number of channels 16, activation function HSwish; Middle layer: Layer 1: depthwise separable convolution, kernel size 3×3, number of channels 32, stride 1; Layer 2: depthwise separable convolution, kernel size 5×5, number of channels 64, stride 2; Output layer: 1×1 convolution to generate a single-channel dynamic lighting map I dynamic , parameter size 20.9K, inference speed 85FPS.

[0033] According to Retinex theory, the input image is decomposed into illumination components I dynamic And the reflected component R, the formula is:

[0034]

[0035] Calculate the local gradient of the reflected component R like Determined as defect edge area; Dynamically adjust filter parameters: spatial kernel G σr : The gradient high area σr is reduced from 2.0 to 0.5, and the low gradient area is kept σr = 2.0; the range kernel S σs : In high gradient areas, σs is reduced from 10.0 to 2.0, while in low gradient areas, σs = 10.0 is maintained; filtering formula:

[0036]

[0037] Build a Dynamic Multi-Scale Attention Fusion Network (DMA-YOLO) backbone network. Input the optimized reflection map into the C2f backbone network of YOLOv8. Perform feature extraction. In the Focus module, downsample the 640×640 image to 320×320 with 32 channels. In the C2f module, the first layer has a bottleneck channel number of 64 and outputs P3 (80×80, 128 channels); the second layer has a bottleneck channel number of 128 and outputs P4 (40×40, 256 channels); and the third layer has a bottleneck channel number of 256 and outputs P5 (20×20, 512 channels).

[0038] Implement the dynamic hybrid attention module, input the feature map F (20×20×512), generate a 1×1×512 vector after global average pooling (GAP), and calculate the compression ratio r through the fully connected layer. min =4, r max =16. The formula is:

[0039] r=Sigmoid(W·GAP(F)+b)×(r max -r min )+r min

[0040] After compressing the channel by r, the channel weight is generated by Sigmoid and the original feature map is weighted.

[0041] Introducing deformable dynamic convolution. Input feature map, output 2×20×20 offset Δp; kernel generator (KernelNet): generates 3×3 dynamic convolution kernel w, convolution formula:

[0042] Δp=OffsetNet(F),w=KernelNet(F)

[0043] Dynamic Weighted Bidirectional Feature Pyramid (DWBiFPN) fusion process, top-down path: P5 (20×20) is upsampled to 40×40, fused with P4 (40×40) according to dynamic weights α1 and α2, and output is P4'; P4' is upsampled to 80×80, fused with P3 (80×80) according to weights α3 and α4, and output is P3'. Bottom-up path: P3' (80×80) is downsampled to 40×40, fused with P4' (40×40), and output is P4". P4" is downsampled to 20×20, fused with P5 (20×20), and output is P5". Weight generation formula:

[0044] α i =Sigmoid(W i ·GAP(P i ))

[0045] Design and output of a high-resolution detection head: input feature: P3' (80×80×128) upsampled to 160×160 by bilinear interpolation; convolution processing: 3×3 convolution kernel, 128 channels, activation function SiLU; Detection output: Output layer: 1×1 convolution generates detection results (160×160×(5+number of categories)), where 5 represents the bounding box coordinates and confidence level; the positioning error of small defects (≤10px) is ≤1.5 pixels.

[0046] To construct and enhance the training dataset, a Keyence VKX200K microscope was used to image 2,000 wafers at a resolution of 0.25 μm / px. The dataset contains 100,000 images, encompassing five defect categories (metal fractures, nanoscratches, particle contamination, voids, and shorts). Random highlights were applied to simulate a 2000-3000 lux light source with a spot radius of 520 px. Particle noise was added using Gaussian noise (SNR = 15 dB) and salt-and-pepper noise (5% density). Geometric transformations included random rotations (±15°), scaling (0.8-1.2x), and horizontal / vertical flips.

[0047] Design a dynamic loss function, improve CIoU Loss, and introduce defect size sensitive weight λ k , the formula is:

[0048] λ k ∝1 / √Area(B k )

[0049] The proportion of small defect gradients increased from 12% to 45%, accelerating model convergence.

[0050] Complete industrial integration and coordinate mapping, first pixel physical coordinate conversion, mechanical stage coordinates (X0, Y0), accuracy ± 0.1μm formula:

[0051] X phy =(X px -640)×0.25+X0 Y phy =(Y px -480)×0.25+Y0

[0052] Output includes TIFF images with embedded XMP metadata (defect category, physical coordinates, confidence level), and CSV report generation including defect size (μm 2 ), location, category, and detection timestamp, supporting MES system interface protocol. Complete visual output.

[0053] For model deployment, the NVIDIA Jetson AGX Xavier embedded platform and CUDA 11.4 acceleration are used. TensorRT is used to quantize the model (FP16 precision), increasing the inference speed to 67 FPS. The memory usage is optimized to 1.2 GB, meeting the real-time requirements of industrial equipment.

[0054] 20% of the 100,000 images were divided into a test set (20,000 images); the evaluation indicators were set as follows: halo artifact ratio: <5%; ≤10px defect recall rate: 94.7%; mAP@0.5: 96.5%; in the comparative experiment, the number of parameters was 32.8M, lower than YOLOv8+MSR (36.2M); the inference speed was 67FPS, better than BiFPN-YOLO+CBAM (55FPS).

Claims

1. A method for detecting micro defects in a chip, characterized in that: The following steps are involved: (a) Use MobileNetV3Lite network to generate dynamic light map; (b) Optimizing the reflection component through edge-aware bilateral filtering; (c) Construct a dynamic multi-scale attention fusion network based on the YOLOv8 framework; (d) Multi-scale feature fusion using dynamically weighted bidirectional feature pyramid; (e) Output of tiny defect location results through high-resolution detection head.

2. The method according to claim 1, characterized in that The dynamic lighting estimation network contains three layers of depth-wise separable convolution with kernel sizes of 3×3 and 5×5 and channel numbers of 32, 64, and 128 respectively.

3. The method according to claim 1, characterized in that The filtering strength of the edge-aware bilateral filter is dynamically adjusted according to the local gradient. When , the spatial kernel σr is reduced to 0.5 and the range kernel σs is reduced to 2.

0.

4. The method according to claim 1, wherein The dynamic hybrid attention module includes dynamic channel compression and deformable dynamic convolution, with a compression ratio ranging from 4 to 16.

5. The method according to claim 1, wherein The fusion weight of the dynamic weighted bidirectional feature pyramid is generated by the global mean of the feature map through a fully connected layer.

6. The method according to claim 1, characterized in that The input feature of the high-resolution detection head is 80×80, which is upsampled to 160×160 through bilinear interpolation, with 128 channels.

7. The method according to claim 1, characterized in that The loss function includes a defect size-sensitive weight, and the weight value is inversely proportional to the square root of the defect area.

8. The method according to claim 1, characterized in that The coordinate mapping model converts pixel coordinates into physical coordinates using the following formula: X phy =(X px -640)×0.25+X0Y phy =(Y px -480)×0.25+Y0 9. The method according to claim 1, characterized in that The visualization output includes TIFF images with embedded XMP metadata and CSV reports containing physical coordinates, dimensions, and confidence levels.

10. The method according to claim 1, characterized in that The training dataset contains 100,000 wafer images with a resolution of 1280×960, covering 5 types of defects, and is enhanced with random highlights, particle noise, and geometric transformation.