Chip defect weak supervision semantic segmentation method based on YOLO and diffusion model

By adopting the weakly supervised semantic segmentation method of YOLO and diffusion model in chip defect detection, the problems of scarce pixel-level labeling data and high labeling cost are solved, and efficient and accurate chip defect detection is achieved.

CN120125824APending Publication Date: 2025-06-10SHANGHAI UNIV
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510304277.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing semantic segmentation models require pixel-level labeling data in chip defect detection, resulting in high labeling costs and low efficiency, especially when defect samples are scarce and micro-nano-scale defects are complex.

Method used

The weak-supervised semantic segmentation method of chip defects based on YOLO and diffusion model is used to construct the data set through weak-supervised manual annotation, the defect location is predicted using the object detection model, and image reconstruction is carried out through the diffusion model, and pixel-level semantic segmentation is finally achieved through differential value comparison and heat map analysis.

Benefits of technology

It significantly reduces the cost and time overhead of data labeling, while maintaining high-precision semantic segmentation effect, effectively identifying small and complex defects, and improving the efficiency and accuracy of chip defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125824A_ABST
    Figure CN120125824A_ABST
Patent Text Reader

Abstract

The invention discloses a chip defect weak supervision semantic segmentation method based on YOLO and a diffusion model, and the method comprises the steps: collecting a chip defect weak supervision data set, carrying out the weak supervision manual marking, and constructing a chip defect image weak supervision data set; constructing and training a target detection model based on the chip defect image weak supervision data set; obtaining a to-be-detected chip defect image, inputting the to-be-detected chip defect image to the trained target detection model for defect position prediction, and outputting a chip defect area positioning image; inputting the chip defect area positioning image into a diffusion model for image reconstruction, and outputting a chip reconstruction image; and performing difference value comparison and thermodynamic diagram analysis on the to-be-detected chip defect image and the chip reconstruction image, and outputting a pixel-level precision chip defect semantic segmentation result. According to the method, the model is trained by using the low-precision weak supervision annotation data, so that high-precision semantic segmentation is realized, the training cost is reduced, and the practicability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor manufacturing technology, and in particular, to a weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion models. Background Art

[0002] The production process of semiconductor chips involves multiple main processes such as wafer manufacturing, lithography, etching, thin film deposition, and ion implantation. Each process step is further divided into several sub-steps. Any defects generated in these production steps may have a chain reaction on subsequent processes, resulting in a decrease in the yield of the production line and causing the electrical performance of the final chip to fail. Therefore, in order to reduce potential production risks and improve the overall yield of chip manufacturing, defect detection and related technologies run through the entire production process and are the key to ensuring the smooth progress of the entire process.

[0003] Image detection based on optical microscopy and electron microscopy technologies is the main means for semiconductor defect analysis, which realizes defect detection through high-resolution imaging combined with image processing algorithms and manual visual inspection. With the continuous progress of chip process technology and the continuous reduction of chip size, the difficulty of defect detection has also increased significantly. The various types of defects generated in more complex process flows have put forward higher requirements for the recognition ability of detection technologies. In recent years, artificial intelligence technologies have gradually been applied to the defect detection process of chip images. In particular, the chip defect semantic segmentation technology based on deep learning has shown significant advantages and has demonstrated extremely high practical value in pixel-level classification and recognition of chip defect images. Through its pixel-level precise classification ability, it can effectively identify tiny or complex defects existing in advanced processes, thereby helping detection personnel to discover potential problems in time and avoid the spread of defects to subsequent processes.

[0004] However, the current semantic segmentation model training usually adopts a fully supervised training paradigm, which requires accurate pixel-level annotation data as support. This kind of annotation often needs to accurately annotate each pixel where the image defect is located, which faces many challenges in practical applications. First, in the field of semiconductor manufacturing, due to the scarcity of defect samples, the scale of annotation data is often limited; second, the annotation of micro-nano scale defects needs to be completed pixel by pixel by professional detection personnel, and the annotation of a single image takes a long time and has a high annotation cost; in addition, the edges of complex defects (such as defects with blurred edges) are often difficult to define, and there are technical annotation difficulties. These factors jointly restrict the application of chip defect semantic segmentation technology in industrial scenarios.

[0005] In response to the above challenges, the semantic segmentation technology based on weakly supervised learning provides a feasible solution. By using rough annotations such as the approximate rectangular box of the defect location and the defect category as weakly supervised signals, this technology can train a semantic segmentation model while significantly reducing the annotation cost and maintaining a high semantic segmentation accuracy similar to that under full supervision. It effectively alleviates the problems of scarce pixel-level annotation data and high annotation difficulty, providing a more feasible technical path for the wide application of chip defect semantic segmentation technology. Summary of the Invention

[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide a weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion models.

[0007] According to one aspect of the present invention, a weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion models is provided, including the following steps:

[0008] Collect chip defect images and perform weakly supervised manual annotation on their defect locations to construct a weakly supervised dataset of chip defect images;

[0009] Based on the weakly supervised dataset of chip defect images, construct and train an object detection model;

[0010] Obtain the chip defect image to be detected, input it into the trained object detection model for defect location prediction, and output a chip defect area localization image;

[0011] Input the chip defect area localization image into the diffusion model for image reconstruction, and output a chip reconstruction image;

[0012] Compare the difference value between the chip defect image to be detected and the chip reconstruction image and perform heat map analysis, and output a chip defect semantic segmentation result image with pixel-level accuracy.

[0013] Preferably, the sources of the chip defect images include: the wafer-defect-rv1vx_dataset open-source dataset, which includes one or more defect categories such as bulk etching, coating defects, particle contamination, PIQ particle contamination, PO contamination, scratches, and SEZ burnout.

[0014] Preferably, performing weakly supervised manual annotation on the defect locations of chip defect images includes:

[0015] Use an annotation tool to perform weakly supervised manual annotation on the chip defect image;

[0016] Among them, the weakly supervised manual annotation means only annotating the bounding box and category information of the defect area, without annotating each pixel within the defect area.

[0017] Preferably, based on the weakly supervised dataset of chip defect images, a target detection model is constructed, including:

[0018] Construct a YOLOv8 target detection model, including a backbone network, a neck network, and a head network;

[0019] Input the chip defect image into the backbone network to extract its multi-scale features; the backbone network has multiple connected levels, and each level is sequentially configured with a convolutional layer, a C2f module, and an SPPF module; the input of each level is the chip defect image or the output of the previous level, and the output is the feature map corresponding to the scale of that level;

[0020] Input the multi-scale features into the neck network to obtain fused multi-scale features; the neck network includes an EFC module, a convolutional module, a C2f module, an upsampling, and a feature fusion module;

[0021] Input the fused multi-scale features into the head network to obtain the predicted defect positions; the head network includes a decoupled head architecture, which processes different-scale features through multiple convolutional layers.

[0022] Preferably, the method for obtaining the chip defect image to be detected, inputting it into the trained target detection model for defect position prediction, and outputting the chip defect area localization image includes:

[0023] Input the chip defect image to be detected into the trained target detection model to generate the predicted bounding box of the defect position and the corresponding defect category;

[0024] Screen and retain the predicted bounding box with the highest confidence through the non-maximum suppression algorithm, and remove redundant bounding boxes;

[0025] Mark the region of the bounding box with the highest confidence and its defect category text on the chip defect image to be detected.

[0026] Preferably, the diffusion model adopts the FLUX.1 model, which includes a CLIP and T5 text encoder, an MM-Single-DiT diffusion network architecture, and a VAE encoder;

[0027] Input the chip defect area localization image into the FLUX.1 model for image reconstruction, and output the chip reconstructed image, including:

[0028] Add Gaussian noise to the bounding box region of the chip defect area localization image, obtain the noisy chip defect image in the latent space through the VAE encoder, and input it into the MM-Single-DiT diffusion network architecture;

[0029] The defect category texts of the chip defect area location images are respectively input into the CLIP and T5 text encoders to obtain encoded texts, which are then input into the MM-Single-DiT diffusion network architecture;

[0030] Initialize a tensor of all zeros ids and perform Rotary Position Encoding (RoPE) to obtain an encoded tensor pe, which is input into the MM-Single-DiT diffusion network architecture;

[0031] Set the number of time steps and the guidance coefficient, and fuse the CLIP text features through an embedding layer to form a conditional vector vec, which is input into the MM-Single-DiT diffusion network architecture;

[0032] The input data undergoes preliminary feature fusion through the MM-DiT layer of the MM-Single-DiT diffusion network architecture, and then deep feature integration is performed through Single-DiT. Based on the flow matching denoising algorithm, the noisy chip defect image in the latent space is gradually denoised, and a reconstructed chip image in the latent space is generated by diffusion;

[0033] The reconstructed chip image in the latent space is decoded through a VAE decoder to obtain a reconstructed chip image.

[0034] Preferably, the flow matching denoising algorithm constructs a continuous probability path from the noisy chip defect image to the reconstructed chip image, dynamically optimizes the denoising path to match the velocity field, and improves the efficiency of generating the reconstructed image by diffusion.

[0035] Preferably, the diffusion model is the FLUX.1 model, and the parameters of the FLUX.1 model are fine-tuned using the LoRA fine-tuning method based on low-rank adaptation, including:

[0036] For the Transformer self-attention layer in the MM-Single-DiT network of the FLUX.1 diffusion model, trainable rank decomposition matrices are introduced into the four key sub-layers W 0 of the original weight matrix W q 、W k 、W v 、W o to obtain a weight increment matrix ΔW;

[0037] Construct a low-rank adapter module and decompose the weight increment matrix ΔW into the product form of two low-rank matrices;

[0038] Keep the original weight matrix W 0 frozen, and only perform gradient updates on the decomposed weight increment matrix ΔW to form a new weight matrix;

[0039] Use the new weight matrix as the parameter of the FLUX.1 diffusion model for image reconstruction.

[0040] Preferably, comparing the difference value between the to-be-detected chip defect image and the chip reconstruction image and performing heatmap analysis, and outputting a chip defect semantic segmentation result image with pixel-level accuracy, including:

[0041] Perform Gaussian blur processing on the to-be-detected chip defect image and the chip reconstruction image to smooth the noise and retain the feature information in the image;

[0042] Calculate the difference value between the two images pixel by pixel to generate a difference value matrix;

[0043] Based on the difference value matrix, apply a heatmap mapping algorithm to convert the corresponding numerical difference into a visual heatmap;

[0044] Based on the visual heatmap, explicitly characterize the spatial distribution of the defect area through chromaticity difference, and output a chip defect semantic segmentation result image with pixel-level accuracy.

[0045] Preferably, based on the difference value matrix, applying a heatmap mapping algorithm to convert the corresponding numerical difference into a visual heatmap, including:

[0046] Perform normalization processing on the difference value matrix;

[0047] Obtain the COLORMAP_JET chromatogram;

[0048] Linearly map and transform the normalized difference value matrix to the COLORMAP_JET chromatogram to generate a visual heatmap with continuous color scale transition.

[0049] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:

[0050] (1) In the chip defect weakly supervised semantic segmentation method based on YOLO and the diffusion model in the embodiments of the present invention, by constructing a collaborative framework of the object detection and the diffusion model, the weak supervision signals such as the defect area bounding box localization information and the defect type are effectively utilized, and a high-precision semantic segmentation training paradigm based on weak supervision learning is successfully established.

[0051] (2) In the chip defect weakly supervised semantic segmentation method based on YOLO and the diffusion model in the embodiments of the present invention, the annotation accuracy of the training data is relaxed from pixel-level accuracy to bounding box-level annotation, significantly reducing the economic cost and time overhead of data annotation while ensuring the model performance.

[0052] (3) The weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion model in the embodiments of the present invention realizes the quantitative evaluation of pixel-level semantic segmentation accuracy by comparing the difference values and heat map analysis of chip defect images and chip reconstruction images. When facing tiny and complex defects and multi-type defects, the model can still maintain high semantic segmentation accuracy.

[0053] In summary, the embodiments of the present invention construct a cascaded processing framework of "object detection - image reconstruction - semantic segmentation", and can achieve a semantic segmentation effect close to that of full-supervised training only by using weakly supervised annotation data with low annotation accuracy, thereby completing the semantic segmentation of chip defect images under the conditions of low cost and high accuracy, and effectively overcoming the problems of scarce pixel-level annotation data and high annotation cost encountered in existing semantic segmentation technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0055] Figure 1 Flowchart of the weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion model according to an embodiment of the present invention;

[0056] Figure 2 Schematic structural diagram of the YOLOv8 object detection model according to an embodiment of the present invention;

[0057] Figure 3 Schematic structural diagram of the FLUX.1 diffusion model according to an embodiment of the present invention;

[0058] Figure 4 Image of the chip to be tested according to an embodiment of the present invention;

[0059] Figure 5 Image of the located chip defect area according to an embodiment of the present invention;

[0060] Figure 6 Reconstructed chip image according to an embodiment of the present invention;

[0061] Figure 7 Image of the semantic segmentation result of chip defects with pixel-level accuracy according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0063] As Figure 1 The following is a flowchart of the weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion models according to an embodiment of the present invention. Please refer to Figure 1 As shown, the weakly supervised semantic segmentation method for chip defects based on YOLO and diffusion models in this embodiment can adopt the following steps:

[0064] Step 1: Collect chip defect images, perform weakly supervised manual annotation on their defect positions, and construct a weakly supervised dataset of chip defect images;

[0065] Step 2: Based on the weakly supervised dataset of chip defect images, construct and train an object detection model;

[0066] Step 3: Obtain the chip defect image to be detected, input it into the trained object detection model for defect position prediction, and output a chip defect area localization image;

[0067] Step 4: Input the chip defect area localization image into the diffusion model for image reconstruction, and output a chip reconstruction image;

[0068] Step 5: Compare the difference value between the chip defect image to be detected and the chip reconstruction image and perform heat map analysis, and output a chip defect semantic segmentation result image with pixel-level accuracy.

[0069] In the above embodiment, a weakly supervised dataset of chip defect images is constructed through weakly supervised manual annotation, and an object detection model is trained to predict a chip defect area localization image; the chip defect area localization image is input into the diffusion model for image reconstruction, and finally, through difference value comparison and heat map analysis, a chip defect semantic segmentation result image with pixel-level accuracy is output, realizing high-precision semantic segmentation only by using weakly supervised annotation data with low annotation accuracy.

[0070] In a preferred embodiment of the present invention, step S1 is implemented: collect a chip defect image dataset, perform weakly supervised manual annotation on the defect positions in the collected images, and construct a weakly supervised dataset of chip defect images.

[0071] Specifically, the source name of the collected training dataset is wafer - defect - rv1vx_dataset, which is an open dataset for semiconductor product defect detection research. It contains seven defect categories, namely bulk etching, poor coating, particle contamination, PIQ (Polyimide Quinoxaline) particle contamination, PO (Phosphorus Oxynitrid) contamination, scratches, and SEZ burnout. This dataset has a total of 1600 training samples, 400 validation samples, and 2532 test samples. For each image, the Labelme annotation tool is used to roughly label the picture data in the chip defect image dataset. That is: only the bounding boxes and category information of the defect areas are annotated, and the pixels within the defect areas are not finely annotated.

[0072] Through the implementation of S1 above, a weakly - supervised dataset of chip defect images can be constructed. This weakly - supervised annotation method helps to reduce the workload of manual annotation and can provide the basic training data required for subsequent object detection models.

[0073] Of course, in other embodiments, other datasets can also be used to achieve other types of defect detection.

[0074] Further, in a preferred embodiment, step S2 is implemented to construct and train a YOLOv8 object detection model based on the weakly - supervised dataset of chip defect images;

[0075] As Figure 2 shown, it is a schematic diagram of the structure of the YOLOv8 object detection model constructed in a preferred embodiment of the present invention.

[0076] The network structure of the YOLOv8 object detection model mainly includes a backbone network, a neck network, and a head network. The backbone network is based on the CSP - Darknet architecture and is responsible for extracting multi - scale features. The neck network is based on the improved PAN - FPN architecture and is responsible for fusing multi - scale features. The head network is based on the decoupled head architecture and is used to generate the final prediction results.

[0077] The backbone network has multiple connected levels, and each level is sequentially configured with a convolutional layer, a C2f (Cross Stage Partial Fusion with 2 convolutions) module, and an SPPF (Spatial Pyramid Pooling Fast) module; the input of each level is a chip defect image or the output of the previous level, and the output is the feature map corresponding to the scale of that level.

[0078] The neck network includes an EFC module, a convolutional module, a C2f module, an up - sampling module, and a Concat module;

[0079] After the feature P4 is upsampled, it is jointly input into the EFC module (Efficient Feature Concatenation) and the C2F module with the feature P3 for fusion, and then the P3 fused feature is generated, which will be used as the input of the head network.

[0080] After the feature P5 is upsampled, it is jointly input into the Concat module and the C2F module with the feature P4 for fusion, thus generating the P4 concatenated feature. Subsequently, the P3 fused feature undergoes a convolution operation and is further fused through the EFC module and the C2F module to obtain the P4 fused feature, which will be used as the input of the head network.

[0081] After the P4 feature map is convolved, it is concatenated with the P5 feature map, and then processed through the C2F module to finally obtain the P5 fused feature, which will be used as the input to the head of the network.

[0082] The head network includes a decoupled head architecture Detect, which processes features of different scales through multiple convolutional layers.

[0083] In the detection of chip defect images, there are complex detection scenarios and multi-scale feature distributions. These feature distributions include both larger-scale defects, such as scratches, and smaller-scale defects, such as PIQ particle contamination. If only relying on a simple feature concatenation module for multi-scale fusion, the detection ability of the object detection model will be limited when dealing with complex circuit backgrounds and tiny defect areas. Therefore, the above embodiments introduce an EFC enhanced feature fusion module to replace the original feature fusion module, which can improve the model's recognition ability for tiny defects. The EFC enhanced feature fusion module consists of a grouped feature focusing unit (GFF) and a multi-layer feature reconstruction unit (MFR). The role of the grouped feature focusing unit is to enhance the correlation between the features extracted at each level of the feature pyramid, while the multi-layer feature reconstruction unit is responsible for reorganizing and transforming the features at different levels to highlight the information features of tiny defect targets. By jointly optimizing the EFC enhanced feature fusion module, the improved model architecture significantly improves the object detection performance of YOLOv8 in complex background and tiny-scale defect detection tasks.

[0084] Furthermore, in a preferred embodiment, the method for training the YOLOv8 object detection model is specifically as follows:

[0085] S21. Divide the weakly supervised dataset of chip defect images into a training set and a validation set according to a ratio of 8:2, and perform preprocessing such as image enhancement, normalization, and size adjustment on the training set images;

[0086] S22. Input the preprocessed training set images into the CSP-Darknet backbone network for multi-scale feature extraction, which includes five levels (P1 - P5). Each level is sequentially configured with a 3×3 convolutional layer, a C2f module, and an SPPF module to output multi-scale image feature maps.

[0087] S23. Input the multi-scale image feature maps into the neck network of the PAN-FPN architecture. Through the EFC module, convolutional module, C2f module, upsampling, and feature fusion module, further process and fuse the features extracted by the backbone network to obtain the fused multi-scale image feature maps.

[0088] S24. Input the fused multi-scale image feature maps into the head network. The network adopts a decoupled head architecture and processes the feature maps of different scales through 3×3 and 1×1 convolutional layers respectively, and calculates the bounding box regression loss and classification loss in parallel.

[0089] S25. Update the model gradients based on the loss function and backpropagation, and use the validation set to evaluate the all-class average precision of the model. Select the model with the highest all-class average precision to obtain the trained YOLOv8 object detection model.

[0090] Through the implementation of the above S21 - S25, a trained YOLOv8 object detection model can be obtained. This model can effectively identify different defect categories in chip images. Due to adopting the YOLOv8 architecture and combining the CSP-Darknet backbone network, PAN-FPN neck network, and decoupled head architecture, and introducing the EFC module to enhance the model's detection ability for small targets (such as tiny chip defects), the model not only has efficient feature extraction ability but also can perform feature fusion at multiple scales, thereby enhancing the accuracy and robustness of defect detection.

[0091] In a preferred embodiment, in step S3, obtain the chip defect image to be detected, input it into the trained object detection model for defect location prediction, and output the chip defect area localization image.

[0092] Specifically, as Figure 4 is the chip defect image to be detected in an embodiment of the present invention. Please refer to Figure 4 shown. The chip defect image to be detected mainly consists of seven types of defect types, corresponding to the bulk etching, coating defect, particle contamination, PIQ particle contamination, PO contamination, scratch, and SEZ burn-out defects included in the wafer-defect-rv1vx_dataset public dataset. In S3, the method of inputting it into the trained YOLOv8 object detection model for defect location prediction is specifically as follows:

[0093] S31. Input the image of the chip to be tested into the trained YOLOv8 object detection model to generate the predicted bounding boxes of the defect positions and the corresponding defect categories.

[0094] S32. Screen and retain the predicted bounding box with the highest confidence through the non-maximum suppression algorithm, and remove redundant bounding boxes.

[0095] S33. Mark the region of the bounding box with the highest confidence and its defect category text on the image of the chip to be tested.

[0096] Through the implementation of the above S3, the localization image of the defect area of the chip defect image to be detected can be obtained. For example Figure 5 is the localization image of the chip defect area of an embodiment of the present invention. Please refer to Figure 5 As shown, the object detection model can accurately predict the defect areas in the image of the chip to be tested with rectangular boxes and assign corresponding defect category texts to each defect area.

[0097] Furthermore, in a preferred embodiment, step S4 is implemented. The localization image of the chip defect area is input into the FLUX.1 diffusion model for image reconstruction, and the reconstructed chip image is output.

[0098] Specifically, as Figure 3 is the schematic structural diagram of the FLUX.1 diffusion model of an embodiment of the present invention. Please refer to Figure 3 As shown, the structure of the FLUX.1 diffusion model mainly includes the CLIP and T5 text encoders, the MM-Single-DiT diffusion network architecture, and the variational autoencoder (VAE). Among them, the CLIP and T5 text encoders are responsible for the embedding and processing of text information. The MM-Single-DiT diffusion network architecture is used for feature fusion and image diffusion generation. The VAE is used to encode and decode the image into the latent space so as to effectively represent and reconstruct the image in the latent space.

[0099] The FLUX.1 diffusion model uses the LoRA (Low-Rank Adaptation) fine-tuning method based on low-rank adaptation to fine-tune the parameters, optimizes the parameters of the Transformer self-attention layer in the MM-Single-DiT network of the diffusion model, and introduces trainable rank decomposition matrices into the four key sub-layers (W 0 of the original weight matrix W q , W k , W v , W o ) to construct a low-rank adapter module, and realizes parameter update by decomposing the weight increment matrix ΔW into the product form of two low-rank matrices. The function expression is:

[0100] ΔW = A · B

[0101] where \(A\in R\) n×d and \(B\in R\) d×m are trainable parameter matrices, \(d\) is the rank dimension, satisfying the low-rank constraint condition of \(d\ll\min(n,m)\); keep the original pre-trained weight \(W\) 0 frozen, only perform gradient updates on the decomposed incremental matrix \(\Delta W\) to form a new weight matrix:

[0102] \(W' = W\) o +\(\Delta W\)

[0103] where \(W'\) are the fine-tuned Flux.1 model parameters, and the model uses \(W'\) parameters for image reconstruction.

[0104] During the above fine-tuning process, if the weight matrix \(W\) is directly updated 0 , a large number of trainable parameters will be introduced. However, by implementing the low-rank decomposition technique, the number of parameters can be reduced from \(d\times d\) (assuming \(W\) 0 is a \(d\times d\) matrix) to \(d\times r + r\times d\) (i.e., the number of parameters of matrices \(A\) and \(B\)), where \(r\) is much smaller than \(d\), representing the rank of the low-rank matrix. This low-rank decomposition method significantly reduces the number of parameters.

[0105] In addition, the low-rank decomposition technique improves the efficiency of the fine-tuning process, especially in the application of large-scale pre-trained models (such as FLUX.1), and it avoids the high cost brought by full-parameter fine-tuning.

[0106] Through the LoRA fine-tuning method, the number of fine-tuned parameters of the model is compressed from 12 billion to 273 million, while adapting the model to the chip image data domain, thus significantly improving the reconstruction effect of the defect area.

[0107] In S4, the method of inputting the chip defect area localization image into the FLUX.1 diffusion model for image reconstruction is specifically as follows:

[0108] S41. Add Gaussian noise to the bounding box area of the chip defect area localization image, obtain the noisy image of the chip defect in the latent space through the VAE encoder, and input it into the MM-Single-DiT diffusion network architecture;

[0109] S42. Input the defect category text of the chip defect area localization image into the CLIP and T5 text encoders respectively to obtain the encoded text, and input it into the MM-Single-DiT diffusion network architecture;

[0110] S43. Initialize a zero tensor ids and perform Rotary Position Encoding (RoPE) to obtain the encoded tensor pe and input it into the MM-Single-DiT diffusion network architecture;

[0111] S44. Set the number of time steps and the guidance coefficient, and fuse the CLIP text features through the embedding layer to form a conditional vector vec and input it into the MM-Single-DiT diffusion network architecture;

[0112] S45. The input data undergoes preliminary feature fusion through the MM-DiT layer of the MM-Single-DiT diffusion network architecture, and then deep feature integration is performed through Single-DiT. Based on the flow matching denoising algorithm, the noisy image of the chip defect in the latent space is gradually denoised, and the reconstructed image of the chip in the latent space is generated by diffusion;

[0113] S46. Pass the reconstructed image of the chip in the latent space through the VAE decoder to decode and obtain the reconstructed image of the chip.

[0114] Among them, the flow matching denoising algorithm can adopt the following steps:

[0115] First, construct a continuous probability path from the noisy image of the chip defect to the reconstructed image of the chip.

[0116] This path gradually transforms the initial Gaussian noise distribution of the noisy image of the chip defect into the target data distribution of the reconstructed image of the chip through a time-dependent vector field (flow field). Specifically, the flow field can be defined by an ordinary differential equation (ODE) so that the samples are gradually transformed from the initial Gaussian noise distribution to the target data distribution. The specific formula is as follows:

[0117]

[0118] Among them, z is the state of the sample at time t.

[0119] Then, dynamically optimize the continuous probability path to match the velocity field and improve the efficiency of diffusion-generated reconstructed images.

[0120] Dynamically optimize the denoising path by learning a time-dependent flow field v(t,z) to guide the data points to gradually evolve from the noise distribution to the initial Gaussian noise distribution along the direction of the flow field to improve the efficiency of diffusion-generated reconstructed images. The specific formula is as follows:

[0121]

[0122] Among them, v target (t,z) is the ideal flow field calculated according to the target distribution, and the denoising path (continuous probability path) is optimized by minimizing the difference between the flow field v(t,z) and the target flow field v target (t,z).

[0123] Therefore, the speed and direction of the change of image pixels are described by the velocity field, and the optimization process ensures smooth and effective denoising. The entire flow matching denoising algorithm can be expressed as:

[0124] x t = tx 1 + (1 - t)x 0

[0125] where x t represents the noisy image, x 1 represents the reconstructed image, x 0 represents the noise, and t represents the time step.

[0126] By implementing steps S41 - S46, the chip image reconstructed by the FLUX.1 diffusion model can be obtained. As Figure 6 is the chip reconstruction image of an embodiment of the present invention, please refer to Figure 6 shown, the chip image with original defective areas is repaired to the normal chip circuit state before the defects occurred. This process not only restores the visual details of the image but also can perform targeted repairs according to the category information of the defects, making the reconstructed image present a structure state almost the same as that of a normal chip.

[0127] Furthermore, in a preferred embodiment, step S5 is implemented to compare the difference value and perform heat map analysis between the defective area result map and the non - defective area result map, obtaining a defect segmentation result with pixel accuracy. In S5, the specific method for comparing the difference value and performing heat map analysis between the chip defective image to be measured and the chip reconstruction image and outputting the chip defect semantic segmentation result with pixel - level accuracy is as follows:

[0128] S51. Perform Gaussian blur processing on the chip defective image to be measured and the chip reconstruction image to smooth the noise and retain the feature information in the image;

[0129] S52. Calculate the difference value between the two images pixel - by - pixel to generate a difference value matrix;

[0130] S53. Based on the difference value matrix, apply the heat map mapping algorithm to convert the numerical difference into a visual heat map;

[0131] S54. Based on the visual heat map, explicitly represent the spatial distribution of the defective area through chromaticity difference and output the chip defect semantic segmentation result image with pixel - level accuracy.

[0132] By implementing the above step S5, the algorithm can efficiently extract the defective area from the difference between the two images, realizing the semantic segmentation of chip defects. At the same time, when facing different image data, the algorithm shows strong anti - noise ability, can effectively ignore noise interference, and ensure the accurate extraction of the defective area. As Figure 7 is the chip defect semantic segmentation result image with pixel - level accuracy of an embodiment of the present invention, please refer to Figure 7As shown, the semantic segmentation result image is generated by linearly transforming the normalized difference value matrix with the COLORMAP_JET color spectrum using the applyColormap function of the OpenCV library. The red channel (H = 0°) in the image represents the highest defect probability, and the blue channel (H = 240°) represents the lowest defect probability. The gradual change of hue from blue to red shows the change of defect probability, realizing the visualization of the defect probability density.

[0133] It should be clearly pointed out that in some specific embodiments, the object detection network can adopt object detection models such as Faster-Rcnn, YOLOv3, etc.; while the diffusion model can be SD1.5, SDXL, etc. In the present invention, the YOLOv8 and FLUX.1 models are used simultaneously, which makes the accuracy effect of object detection reach the best, and the same is true for the image reconstruction effect. The combination of the two can achieve the optimal semantic segmentation effect.

[0134] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be used in any combination without conflict.

Claims

1. A chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model, characterized in that: The following steps are involved: Collect chip defect images, perform weakly supervised manual annotation on their defect locations, and build a weakly supervised dataset of chip defect images; Based on the chip defect image weak supervision dataset, construct and train a target detection model; Obtain a defect image of the chip to be detected, input it into the trained target detection model to predict the defect position, and output a chip defect area positioning image; Inputting the chip defect area positioning image into a diffusion model for image reconstruction, and outputting a chip reconstructed image; The chip defect image to be detected is compared with the chip reconstructed image for difference value comparison and thermal map analysis, and a chip defect semantic segmentation result image with pixel-level accuracy is output.

2. According to claim 1, the chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model is characterized in that: Sources of the chip defect images include: wafer-defect-rv1vx_dataset open source dataset, which includes one or more defect categories of block etching, poor coating, particle contamination, PIQ particle contamination, PO contamination, scratches and SEZ burning.

3. According to claim 1, the chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model is characterized in that: Weakly supervised manual annotation of defect locations in chip defect images, including: Performing weakly supervised manual annotation on the chip defect image using an annotation tool; The weakly supervised manual labeling refers to labeling only the boundary box and category information of the defect area, without labeling each pixel in the defect area.

4. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 1, characterized in that: Based on the chip defect image weak supervision dataset, a target detection model is constructed, including: Build a YOLOv8 target detection model, including the backbone network, neck network, and head network; Input a chip defect image to the backbone network to extract its multi-scale features; the backbone network has multiple connected layers, each of which is sequentially configured with a convolutional layer, a C2f module, and an SPPF module; the input of each layer is a chip defect image or the output of the previous layer, and the output is a feature map of the corresponding scale of the layer; Inputting multi-scale features into the neck network to obtain fused multi-scale features; the neck network includes an EFC module, a convolution module, a C2f module, an upsampling and feature fusion module; The fused multi-scale features are input into the head network to obtain the predicted defect position; the head network includes a decoupled head architecture, and processes different scale features correspondingly through multiple convolutional layers.

5. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 1, characterized in that: The method of obtaining a chip defect image to be detected, inputting the image into the trained target detection model to predict the defect position, and outputting a chip defect area positioning image includes: Inputting the defect image of the chip to be detected into the trained target detection model to generate a predicted bounding box of the defect location and a corresponding defect category; The non-maximum suppression algorithm is used to filter and retain the predicted bounding boxes with the highest confidence, and to remove redundant bounding boxes. Mark the bounding box area with the highest confidence and its defect category text on the defect image of the chip to be detected.

6. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 1, characterized in that: The diffusion model adopts the FLUX.1 model, which includes CLIP and T5 text encoders, MM-Single-DiT diffusion network architecture, and VAE encoder; Inputting the chip defect area positioning image into the FLUX.1 model for image reconstruction, and outputting the chip reconstructed image, including: Adding Gaussian noise to the bounding box area of ​​the chip defect area positioning image, obtaining the chip defect noisy image in the latent space through the VAE encoder, and inputting it into the MM-Single-DiT diffusion network architecture; Input the defect category text of the chip defect area positioning image into CLIP and T5 text encoders respectively to obtain encoded text, and input it into the MM-Single-DiT diffusion network architecture; Initialize the all-zero tensor ids and perform rotation position encoding RoPE to obtain the encoded tensor pe and input it into the MM-Single-DiT diffusion network architecture; Set the time step and guidance coefficient, fuse the CLIP text features through the embedding layer to form a conditional vector vec and input it into the MM-Single-DiT diffusion network architecture; The input data is subjected to preliminary feature fusion through the MM-DiT layer of the MM-Single-DiT diffusion network architecture, and then deep feature integration is performed through Single-DiT. Based on flow matching denoising, the chip defect and noisy image in the latent space is gradually denoised, and a chip reconstruction image in the latent space is diffusely generated; The chip reconstructed image of the latent space is passed through a VAE decoder to obtain a chip reconstructed image.

7. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 6, characterized in that: The stream matching denoising comprises: Constructing a continuous probability path from the chip defect noisy image to the chip reconstructed image; The continuous probability path is dynamically optimized to match the velocity field and improve the efficiency of diffusion-generated reconstructed images.

8. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 1, characterized in that: The diffusion model is a FLUX.1 model, and the parameters of the FLUX.1 model are fine-tuned using a LoRA fine-tuning method based on low-rank adaptation, including: For the Transformer self-attention layer in the MM-Single-DiT network of the FLUX.1 diffusion model, in the four key sublayers W of the original weight matrix W0 q , W k , W v , W o Introduce a trainable rank decomposition matrix to obtain the weight increment matrix ΔW; Constructing a low-rank adapter module and decomposing the weight increment matrix ΔW into a product form of two low-rank matrices; Keep the original weight matrix W0 frozen, and only perform gradient update on the decomposed weight increment matrix ΔW to form a new weight matrix; The new weight matrix is ​​used as the parameters of the FLUX.1 diffusion model for image reconstruction.

9. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 1, characterized in that: The step of comparing the difference value of the chip defect image to be detected with the chip reconstructed image and performing a heat map analysis to output a chip defect semantic segmentation result image with pixel-level accuracy includes: Performing Gaussian blur processing on the chip defect image to be detected and the chip reconstructed image to smooth noise and retain feature information in the image; Calculate the difference between the two images pixel by pixel to generate a difference matrix; Based on the difference value matrix, a heat map mapping algorithm is applied to convert the corresponding numerical differences into a visual heat map; Based on the visualized heat map, the spatial distribution of the defect area is explicitly characterized by chromaticity difference, and the chip defect semantic segmentation result image with pixel-level accuracy is output.

10. The chip defect weakly supervised semantic segmentation method based on YOLO and diffusion model according to claim 9, characterized in that: Based on the difference value matrix, applying a heat map mapping algorithm to convert the corresponding numerical differences into a visual heat map includes: Normalizing the difference value matrix; Get COLORMAP_JET chromatogram; The normalized difference value matrix is ​​linearly mapped to the COLORMAP_JET color spectrum to generate a visualized heat map with continuous color scale transition.

Citation Information

Cited By

  • Acoustic-optical imaging semantic fusion detection method for underwater defects of hydraulic structure

    CN120656051A

  • Power transmission line insulator defect detection method and system based on machine vision

    CN120782732A

  • Chemical design and safety risk analysis integrated platform

    CN121032223A

  • Chemical engineering design and safety risk analysis integrated platform device

    CN121032223B

  • Road defect detection method based on dynamic prototype learning and weak supervision semantic segmentation

    CN121032912A