Unified multi-task computational lithography modeling and defect detection method assisted by large visual models

Through a unified multi-task lithography modeling method based on Transformer and DETR, the problems of long time consumption of lithography simulation tools and poor generalization ability of detection methods are solved, and efficient and automated lithography modeling and defect detection are achieved, which is suitable for advanced node processes.

CN120411101BActive Publication Date: 2025-09-09ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510912193.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-09
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Existing lithography simulation tools are time-consuming, rely on manual experience, and have poor generalization capabilities under multiple process conditions. Traditional detection methods make it difficult to uniformly detect lithography defects across stages.

Method used

A two-stage image generation model based on the Transformer architecture and the DETR defect detection model are used, combined with process parameters, to build a unified end-to-end lithography modeling and defect detection framework, achieving high-fidelity prediction from layout to mask and contour, and multi-stage defect identification.

Benefits of technology

It significantly improves the efficiency and accuracy of lithography simulation, has efficient and automated multi-process adaptability, reduces model deployment costs, and is suitable for lithography modeling and detection under advanced node processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411101B_ABST
    Figure CN120411101B_ABST
Patent Text Reader

Abstract

The present invention discloses a unified multi-task computational lithography modeling and defect detection method assisted by a large visual model. The method not only has high-precision modeling capabilities and defect recognition performance, but also significantly reduces model deployment costs and has advantages such as high automation, unified modeling, and engineering scalability. The method includes the following steps: constructing a dual-task dataset for image generation and defect detection; building a two-stage image generation model based on a unified Transformer architecture that integrates process conditions to generate mask and contour images from layout images; building a defect detection model based on DETR for layout, mask, and contour images to perform DRC, MRC, and LRC detection; and building an end-to-end lithography large model framework that supports inputting a layout and simultaneously outputting mask, contour, and DRC / MRC / LRC detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of semiconductor manufacturing technology, and in particular relates to a unified multi-task computational lithography modeling and defect detection method assisted by a large visual model. Background Art

[0002] As semiconductor manufacturing processes continue to advance and transistor feature sizes continue to shrink, photolithography processes face increasing challenges in manufacturing yield and pattern manufacturability. Low-fidelity structures (i.e., "hotspots") in wafer patterns have become a key factor impacting yield. To improve the manufacturability of layouts during actual photolithography, there is an urgent need to develop efficient, accurate, and multi-process adaptable process prediction and hotspot detection methods.

[0003] Currently, traditional lithography simulation tools have established high-fidelity modeling processes based on optical physics and photoresist reactions, with certain configurability and simulation capabilities, and can provide accurate predictions for layouts under different lighting and process conditions. However, the following significant issues still exist:

[0004] (1) The simulation process is complex and time-consuming: Running a complete layout evaluation on a high-performance computing cluster usually takes several days, and the simulation efficiency drops sharply with the design complexity;

[0005] (2) Process parameter adjustment is highly dependent on manual experience: Under different lithography conditions (such as exposure dose, focal length, photoresist model, etc.), engineers need to manually adjust the simulation script and parameter configuration, which is a cumbersome process with poor flexibility.

[0006] Although recent studies have introduced deep learning-based models to improve simulation efficiency, these methods are limited by the expressive capabilities of architectures such as CNN (convolutional neural networks) and GAN (generative adversarial networks). These methods perform poorly when dealing with the nonlinear coupling relationships between complex light sources, mask patterns, and imaging images. They also have poor generalization capabilities in real-world design scenarios with multi-process and multi-source lighting, making them difficult to apply stably.

[0007] In addition to process prediction, hotspot detection itself has numerous shortcomings. Hotspots manifest differently at the layout, mask, and contour stages. Defects within clips (such as edge positioning errors and necking) are often limited to subtle widths or edge offsets, making them difficult to effectively identify using traditional methods. Furthermore, most existing detection models are trained only on images from a single stage and lack the ability to uniformly detect across these stages, limiting their applicability and deployment efficiency in real-world process flows. Summary of the Invention

[0008] To overcome the shortcomings of the existing technology, the present invention aims to provide a technical solution for a unified multi-task computational lithography modeling and defect detection method assisted by a large visual model. The specific contents of this technical solution are as follows:

[0009] The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model comprises the following steps:

[0010] Step 1: Use Calibre (lithography process simulation and physical design verification software developed by Synopsys) lithography simulation tools to generate layout, mask, and contour images under multiple process conditions. This builds a dual-task dataset for image generation and defect detection, providing a data foundation for subsequent multi-task learning.

[0011] Step 2: Based on the data constructed in step 1, a two-stage image generation model integrating process conditions is constructed based on a unified Transformer (deep learning model based on self-attention mechanism) architecture to generate mask and contour images from layout images.

[0012] Step 3: Based on the real defect detection dataset constructed in Step 1, a defect detection model based on the DETR (Detection Transformer Network) is constructed for the three types of images: layout, mask, and contour. This model performs DRC (Design Rule Check), MRC (Manufacturing Rule Check), and LRC (Lithography Rule Check) inspections. This step is independent of the image generation process in Step 2 and directly utilizes real process data for defect detection modeling, ensuring the accuracy and engineering adaptability of the detection model.

[0013] Step 4. Based on the two-stage image generation model and defect detection model, a unified end-to-end lithography modeling and defect detection framework is constructed. This framework takes the layout image collected in step 1 as input and first generates mask and contour images through the two-stage image generation model in step 2. Then, the generated image is combined with the real image and the DETR detection module in step 3 is introduced to achieve comprehensive detection and localization of multiple types of defects such as DRC, MRC, and LRC.

[0014] Furthermore, in step 1, four process parameter dimensions are introduced in the simulation process of constructing the image generation subset, namely, light source type, photoresist threshold, focal offset and exposure dose, among which the light source type includes annular light source, circular light source and composite light source; the photoresist threshold selects the dose thresholds of three mainstream photoresist materials to cover high, medium and low exposure response curve characteristics; the focal offset is set to two typical values ​​of 0 nm and 50 nm; the exposure dose is set to two nominal doses of 1.0× and 1.2× to simulate process fluctuation conditions; a variety of process combinations are constructed through the four process parameter dimensions, and for each combination condition, layout-mask-contour triplet image data are simulated and generated.

[0015] Furthermore, in step 1, the defect detection subset is constructed by combining the rule checker with the automated annotation script, and the annotation is performed using the standard target detection format, and the annotation information is stored in the JSON format.

[0016] Furthermore, in step 2, the two-stage image generation model includes an input construction module, a main input encoding module, an auxiliary conditional encoding module, and a cross-modal fusion module. Specifically:

[0017] In the input construction module, the model input includes the main input layout image and auxiliary input process parameter information;

[0018] In the main input encoding module, the layout image input first passes through a patch embedding module to extract local structure, and then the backbone structure Swin Transformer (sliding window visual transformer) performs layer-by-layer feature extraction;

[0019] The auxiliary conditional encoding module receives five-channel inputs, including a light source image, a mask image, and three process parameter maps. It first performs preliminary fusion through a layer of 3×3 convolution to enhance semantic connectivity between channels. It then extracts high-level semantic information through three sets of down-block structures. The final output is dimensionally aligned through a layer of convolution, allowing it to be input into the cross-modal fusion module along with the layout encoding features.

[0020] In the cross-modal fusion module, the auxiliary input encoding result is used as the query, and the layout features are used as the key and value. The attention mechanism is used to achieve semantic alignment and response relationship modeling.

[0021] Furthermore, in step 2, the two-stage image generation model adopts a two-stage joint optimization strategy, and is jointly driven by multiple loss functions during the training process. The objective function is as follows:

[0022] ,

[0023] Among them, L mask and L contour They are the loss functions of the mask image and the contour image, respectively, and are constructed as follows:

[0024] ,

[0025] ,

[0026] Among them, P mask With P contour Represent the mask and contour prediction map generated by the model, G mask With G contour The corresponding true label map; BCE (Binary Cross Entropy, binary cross entropy loss) is used to measure the difference between the predicted value and the true value of each pixel; Dice (Dice Loss, dice coefficient loss) emphasizes the overlap between the predicted area and the true area, which is suitable for structural modeling in unbalanced data; and Triplet Loss (triplet loss) 、 The discriminative ability of the model is enhanced by constraining similar samples to be closer and heterogeneous samples to be farther apart in the feature space.

[0027] Furthermore, in step 2, a cross-process comparison optimization mechanism is introduced into the two-stage image generation model, and Triplet Loss process-aware supervision is introduced into the loss function. During the training process, for the same input layout, different process parameters p are used for simulation to obtain multiple generated image results. For images with the same process conditions and images with different process conditions, triplets are constructed. The Triplet Loss expression is as follows:

[0028] ,

[0029] Where x is the input layout; p, p+, and p− represent the current process parameters, the positive sample process (same parameters), and the negative sample process (different parameters), respectively; G(x, p) represents the feature representation extracted by the model under the conditions of input image x and process parameter p; is the minimum distinguishing distance between positive and negative samples; the symbol Represents the square of the Euclidean distance (L2 norm), which is used to measure the similarity between feature vectors. The smaller the value, the more similar they are.

[0030] Furthermore, in step 3, the specific contents of building a defect detection model based on DETR are as follows:

[0031] Step 3-1, feature extraction and Transformer encoder construction: Input the defect detection task dataset constructed in step 1 into the convolutional neural network feature extraction backbone network to obtain a dimension of The image feature map is then flattened into a sequence representation and fed into the standard Transformer encoder module; where H and W are the height and width of the original image, respectively, and p is the patch size, indicating that the image is divided into patches, d is the feature dimension after each patch is encoded; the entire feature map is represented as tokens, each token is a d-dimensional vector;

[0032] Step 3-2, Transformer decoder and target representation generation mechanism: Introducing the object query vector q∈R N×d As the input of the Transformer decoder, the decoder interacts with the encoded features and outputs a prediction result of a candidate object for each query, including the defect bounding box and class probability, thus forming the final object set;

[0033] Step 3-3, target matching and loss function design: Introduce the Hungarian matching algorithm and calculate the DETR total loss function;

[0034] Step 3-4, three-stage defect detection model deployment: Build three DETR-based defect detection models for layout, mask, and contour images, respectively, to identify layout design violations (DRC), mask defects (MRC), and resist imaging anomalies (LRC). Each model is trained based on an independent training dataset and outputs a bounding box of the defect target.

[0035] Furthermore, in step 4, the end-to-end lithography model uses a unified Transformer backbone as a shared encoder. The input is the layout image and related process condition information. After the encoder outputs common semantic features, the multi-branch decoder performs the image generation task and the hotspot detection task respectively.

[0036] ① For image generation tasks: Use the image generation branch to perform high-precision modeling on the layout image and gradually output the mask graphic x mask With imaging profile x contour ,Among them, mask generation depends on layout and process condition input, and contour generation introduces mask image as auxiliary information on this basis;

[0037] ② For the hotspot detection task: Based on the shared feature representation, the structural defects in the layout, mask, and contour images are identified through the DRC, MRC, and LRC detection modules respectively, and the target set Y is output. layout 、Y mask 、Y contour ,Each item includes the bounding box location of the defect, forming a full-process structural violation detection capability.

[0038] Furthermore, in step 4, to achieve dual-path collaborative training, the following joint loss function is designed:

[0039] ,

[0040] in, Represents the joint loss function of the image generation task, including the mask and contour stages; Denotes the DETR category and position loss term of defect detection at each stage; It is a weight adjustment factor between tasks, used to balance the training priority of structure reconstruction and defect recognition.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] (1) This paper proposes a two-stage image generation network based on the Vision Transformer architecture, which uses a cross-attention mechanism guided by process conditions to achieve high-fidelity prediction from layout to mask and contour. In the first stage, a mask image is generated based on the input layout and process parameters (such as dose and focus). In the second stage, a contour image is further generated based on the generated mask, thereby significantly improving the structural continuity and geometric consistency of the contour image. At the same time, in order to enhance the model's application ability under different process conditions and enhance process sensitivity, this paper introduces a process-aware contrastive learning method. By shortening the distance between positive sample pairs (generated images and real images under the same process conditions) and extending the distance between negative sample pairs (images under different process conditions), the model is guided to explicitly model the coupling relationship between layout and process conditions, thereby improving the discriminability and stability of image generation under different process perturbations.

[0043] (2) The present invention constructs a structural defect detection model based on DETR for three types of images: layout, mask, and contour, to achieve unified recognition of DRC, MRC, and LRC multi-stage process violations; it does not require candidate boxes and post-processing processes, and has efficient modeling and end-to-end reasoning capabilities.

[0044] (3) The present invention integrates image generation and structure detection functions into a unified architecture, realizes feature sharing through a shared Transformer backbone network, fuses the two tasks of image generation and target detection, and realizes an integrated processing flow of layout → mask / contour + DRC / MRC / LRC under the conditions of satisfying process rule constraints and graphic connectivity. It not only has high-precision modeling capabilities and defect recognition performance, but also significantly reduces the model deployment cost. It has the advantages of high automation, unified modeling and engineering scalability, and is suitable for the integrated application of lithography modeling and detection systems under advanced node processes. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Flow chart of the method of the present invention;

[0046] Figure 2 The present invention is a lithography generator based on Vision Transformer;

[0047] Figure 3 It is a hotspot detector based on the DETR architecture of the present invention;

[0048] Figure 4 This is a schematic diagram of the framework structure of the end-to-end integrated lithography large model of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0050] like Figure 1 As shown, the unified multi-task computational lithography modeling and defect detection method assisted by a large visual model of the present invention includes the following steps:

[0051] (I) Step 1: Use the Calibre simulation tool to generate layout, mask, and contour images under multiple process conditions, and construct a dual-task dataset of image generation and defect detection to provide a data foundation for subsequent multi-task learning.

[0052] To achieve unified multi-task lithography modeling and defect detection, this paper first constructs a training dataset covering a variety of process conditions and defect types to support the joint training of large models. This dataset consists of two major components: an image generation subset and a defect detection subset, serving the image generation task and the multi-stage defect recognition task, respectively.

[0053] (1) Image generation subset construction

[0054] The present invention uses the commercial photolithography simulation software Calibre to perform simulation modeling on a standard integrated circuit layout image (layout) under multiple process conditions, and generates the corresponding mask pattern (mask) and resist contour (contour).

[0055] To improve the model's ability to generalize to actual manufacturing process variations, four process parameter dimensions (i.e., light source type, photoresist threshold, focus offset, and exposure dose) are introduced into the simulation process. The light source type (Illumination Type) includes three typical industrial configurations: annular light source (Annular), circular light source (Coh. Disk), and composite light source (Annular + Center Disk). The photoresist threshold (Resist Threshold) selects the dose thresholds of three mainstream photoresist materials to cover the high, medium, and low exposure response curve characteristics. The focus offset (Defocus) is set to two typical values: 0 nm (ideal focus state) and 50 nm (common offset range). The exposure dose (Dose) is set to two nominal doses of 1.0× and 1.2× to simulate process fluctuation conditions.

[0056] The four process parameter dimensions above constitute 3 × 3 × 2 × 2 = 36 process combinations. For each combination, 8,000 layout–mask–contour image triplets were simulated and generated, ultimately generating a total of 36 (combinations) × 8,000 = 288,000 sets of image data.

[0057] The image generation subset provides sufficient sample support for the two-stage image generation network proposed in this invention, and has good structural expression and simulation accuracy in different process scenarios.

[0058] (2) Construction of defect detection subset

[0059] In order to support multi-stage defect recognition tasks, the present invention further performs defect annotation on the layout, mask and contour images respectively, covering three types of typical lithography violation problems: DRC, MRC and LRC.

[0060] Defect samples were generated using a combination of a rule checker and automated annotation scripts. These samples were annotated using a standard object detection format, and the annotation information was stored in JSON format. The following image statistics were collected for each defect type: 50,000 layout images with DRC hotspots, 72,000 mask images with MRC hotspots, 144,000 contour images with LRC hotspots, and 100,000 defect-free sample images (at any stage).

[0061] (2) Step 2: Based on the data constructed in step 1, a two-stage image generation model that integrates process conditions is constructed based on a unified Transformer architecture to generate mask and contour images from the layout image.

[0062] The two-stage image generation model integrates layout structure information and lithography process parameter prompt information to complete mask prediction and contour simulation tasks in a unified model. It has end-to-end joint modeling capabilities and multi-process controllability. Its technical process is as follows Figure 2 shown.

[0063] The two-stage image generation model includes an input construction module, a main input encoding module, an auxiliary conditional encoding module, and a cross-modal fusion module.

[0064] (1) Input building blocks

[0065] The model inputs include a primary input layout image and auxiliary input process parameter information. Specifically, the layout image is 512×512 in size and is represented as a single-channel binary image. The auxiliary inputs include a light source image (source) and three process parameters (exposure dose, defocus, and photoresist threshold). The source is a two-dimensional image generated under actual lighting conditions (original size is 256×256). The remaining three parameters are given in scalar form, expanded to a constant image, and then concatenated with the light source image in the channel dimension to form a complete auxiliary input tensor. In the second stage of image generation, the mask image output from the first stage is added as an additional channel to the auxiliary input, forming a five-channel input tensor consisting of the source image, the mask image, and the three process parameter maps, which serves as the input for the subsequent process perception path.

[0066] (2) Main input encoding module

[0067] The layout image input first passes through a Patch Embedding module to extract local structure, followed by layer-by-layer feature extraction by the Swin Transformer backbone. Swin ViT utilizes sliding window partitioning and Shifted Window Attention to achieve local modeling and global information flow. Its attention mechanism is implemented using the following formula:

[0068] ,

[0069] Among them, Q, K, V represent the query, key and value vectors obtained from the layout feature map respectively; B is the relative position bias term; K Tis the transpose of the key vector K; d is the dimension of each head in the attention mechanism, used for scaling to prevent excessive inner product size. The encoder models the non-local dependencies between structures in the layout image and outputs a semantic feature tensor of the graph structure, providing a structural foundation for subsequent image generation.

[0070] (3) Auxiliary conditional coding module

[0071] The auxiliary conditional encoding module receives five-channel inputs, including a light source image, a mask image, and three process parameter maps. It first performs a preliminary fusion through a layer of 3×3 convolution to enhance semantic connections between channels. It then extracts high-level semantic information through three sets of down-block structures. Each down-block includes convolution, normalization, and activation operations, as follows:

[0072] ,

[0073] This formula describes the processing flow of the feature extraction module in a typical convolutional neural network. i Represents the input feature map of the i-th layer, which can be an intermediate representation of the image or the output of the previous layer; Conv3×3 is a 3×3 convolution operation used to extract local spatial structure information; BN stands for batch normalization, which is used to standardize the convolution results to improve training stability; ReLU is a nonlinear activation function used to enhance the expressive power of the network, defined as max(0,x); the final output F i+1 This output serves as the feature map output by the current module and is used for processing in subsequent layers. This module combination (Conv + BN + ReLU) is a common basic unit in deep learning models, effectively capturing the lighting pattern in the light source image, the geometric boundaries in the mask image, and the control characteristics of the process parameters on imaging behavior. The final output undergoes a layer of convolution for dimensionality matching, allowing it to be input into the fusion module along with the layout encoding features.

[0074] (4) Cross-modal fusion module (Cross Attention)

[0075] To achieve efficient integration of graphical structure and process conditions, this paper adopts a Cross Attention fusion structure. In this module, the auxiliary input encoding result is used as the query, and the layout features are used as the key and value. The attention mechanism is used to achieve semantic alignment and response relationship modeling. The core calculation is:

[0076] ,

[0077] Where Q = F cond W Q、K=F layout W K 、V=F layout W V , respectively represent the auxiliary condition feature F cond Mapped to query (Query), the image structure feature F layout Mapped to key (Key) and value (Value); W Q 、W K 、W V They are the linear transformation weight matrices corresponding to query, key, and value respectively; d is the dimension of the key vector; and finally through Achieve alignment and fusion, and output the fused feature F fused This mechanism helps capture the semantic relationship between structural features and conditions, thereby improving the expressiveness of the model.

[0078] (5) Joint optimization loss function design

[0079] The two-stage image generation model of the present invention adopts a two-stage joint optimization strategy, and is jointly driven by multiple loss functions during the training process. The objective function is as follows:

[0080] ,

[0081] Among them, L mask and L contour They are the loss functions of the mask image and the contour image, respectively, and are constructed as follows:

[0082] ,

[0083] ,

[0084] Among them, P mask With P contour Represent the mask and contour prediction map generated by the model, G mask With G contour The corresponding true label map. Binary Cross Entropy (BCE) is used to measure the difference between the predicted value and the true value of each pixel; Dice Loss emphasizes the overlap between the predicted area and the true area, and is suitable for structural modeling in imbalanced data; and Triplet Loss 、 By constraining similar samples to be closer and heterogeneous samples to be farther apart in the feature space, the model's discriminative ability is enhanced. Overall, this design combines pixel-level accuracy, regional structure matching, and feature semantic differentiation, improving the model's ability to consistently and accurately model mask and contour structures in images.

[0085] (6) Cross-process parameter perception modeling mechanism

[0086] To address the inability of traditional image generation models to distinguish fine-grained differences in images under different process conditions (such as exposure dose and focus offset), this paper further introduces a cross-process comparison optimization mechanism and incorporates TripletLoss process-aware supervision into the loss function to enhance the model's ability to discern process variations. The core idea is to simulate the same input layout using different process parameters (p) during training, generating multiple generated images. For images with the same process conditions (called "positive samples") and images with different process conditions (called "negative samples"), a triplet (Anchor, Positive, Negative) is constructed. The optimization goal is to minimize the distance between the Anchor and the Positive, while maximizing the distance between the Anchor and the Negative, thereby forming a clear boundary for characterizing process features.

[0087] The triplet loss expression is as follows:

[0088] ,

[0089] Where x is the input layout, p, p+, and p− represent the current process parameters, the positive sample process (same parameters), and the negative sample process (different parameters), respectively. G(x, p) represents the feature representation extracted by the model under the conditions of the input image x and the process parameter p. is the minimum distinguishing distance between positive and negative samples, symbol Represents the square of the Euclidean distance (L2 norm), used to measure the similarity between feature vectors; smaller values ​​represent greater similarity. This process-aware optimization mechanism applies comparative constraints to images of the same layout under different process parameter conditions, guiding the model to explicitly learn the correspondence between layout and process parameters during training. This enables the model to not only generate images consistent with the current process conditions but also distinguish subtle changes in images under different process conditions, thereby improving its adaptability and resolution to process perturbations such as focus and exposure, effectively enhancing the model's generalization performance in multi-process environments.

[0090] (III) Step 3: Based on the real defect detection dataset constructed in Step 1, a DETR-based defect detection model is constructed for layout, mask, and contour images, and DRC, MRC, and LRC detection are performed. This step is independent of the image generation process in Step 2 and directly utilizes real process data for defect detection modeling, ensuring the detection model has good accuracy and engineering adaptability.

[0091] like Figure 3 As shown in the figure, the specific workflow of the DETR-based defect detection model is as follows:

[0092] (1) Feature extraction and Transformer encoder construction

[0093] First, the input image (which can be a layout, mask or contour image) is fed into the convolutional neural network backbone to extract the underlying semantic features, resulting in a shape of Image feature map; where H and W are the height and width of the original image, respectively, and p is the patch size, indicating that the image is divided into patches, d is the feature dimension after each patch is encoded; the entire feature map is represented as tokens, each of which is a d-dimensional vector. These serialized image features are then fed into a standard Transformer encoder for modeling. The Transformer encoder, consisting of a multi-layered, multi-head self-attention mechanism superimposed on a feedforward neural network, effectively captures long-range dependencies and contextual semantic information between different image locations, thereby enhancing the model's ability to express defect structures and perceptual robustness.

[0094] (2) Transformer decoder and target representation generation mechanism

[0095] Introduce a set of learnable object query vectors q∈R N×d As input to the Transformer decoder, the decoder interacts with the encoded features and outputs a prediction result for a candidate object for each query, including the defect bounding box and class probability, thus forming the final object set:

[0096] ,

[0097] in, is the category distribution prediction, Represents the normalized bounding box coordinates, including the center point position (x, y) and width and height (w, h).

[0098] (3) Target matching and loss function design

[0099] In order to achieve a one-to-one match between the prediction and the true label, the Hungarian matching algorithm is introduced. The matching cost function is defined as follows:

[0100] ,

[0101] Among them, y jrepresents the label of the jth true target, is the prediction result that matches it; the first is the classification loss, Indicates the prediction result Position pair categories The predicted probability of is used to measure the confidence of the classification; the second term is the bounding box regression loss, Measure the j-th ground-truth bounding box b j and predicted bounding box The position difference between them includes L1 distance and generalized IoU (GIoU); λ box is a weight hyperparameter used to adjust the relative importance of classification loss and box regression loss in the matching cost.

[0102] The final DETR total loss function is as follows:

[0103] ,

[0104] in, Cross entropy classification loss measures the difference between the predicted category and the true category; represents the L1 distance between bounding box coordinates; is the generalized IoU loss, which more comprehensively evaluates the degree of bounding box overlap; λ1 and λ2 are hyperparameters that adjust the weights of different loss terms. This total loss function jointly optimizes classification accuracy and bounding box localization accuracy during training, thereby improving overall object detection performance.

[0105] (4) Three-stage defect detection model deployment plan

[0106] This paper constructs three DETR-based defect detection models for layout, mask, and contour images, respectively, to identify design layout violations (DRCs), mask defects (MRCs), and resist imaging anomalies (LRCs). Each model is trained on an independent training dataset and outputs a bounding box representing the defect.

[0107] (IV) Step 4: Based on the two-stage image generation model and defect detection model, a unified end-to-end lithography modeling and defect detection framework is constructed. This framework takes the layout image acquired in Step 1 as input and first generates mask and contour images using the two-stage image generation model in Step 2. Subsequently, the generated image is combined with the real image and the DETR detection module in Step 3 is introduced to achieve comprehensive detection and localization of multiple types of defects, including DRC, MRC, and LRC. In other words, this step combines the image generation in Step 2 and the defect detection in Step 3, and then jointly optimizes them to achieve the end-to-end goal.

[0108] By integrating image generation and defect detection tasks into the same network architecture, we can build an end-to-end intelligent system with dual capabilities of image generation and structural analysis, and achieve integrated joint modeling from layout input to mask and contour prediction, and then to DRC / MRC / LRC defect recognition result output.

[0109] like Figure 4 As shown in the figure, the end-to-end lithography model uses a unified Transformer backbone as a shared encoder. The input is the layout image and related process condition information (including the source map and process parameters). After the encoder outputs common semantic features, it performs the following two main tasks through a multi-branch decoder:

[0110] (1) Image generation task

[0111] Use the image generation branch to perform high-precision modeling on the layout image and gradually output the mask graphic x mask With imaging profile x contour Among them, mask generation depends on layout and process condition input; contour generation introduces mask image as auxiliary information on this basis to enhance geometric consistency modeling.

[0112] (2) Hotspot detection task

[0113] Based on the shared feature representation, the structural defects in the layout, mask and contour images are identified through the DRC, MRC and LRC detection modules respectively, and the target set Y is output. layout 、Y mask 、Y contour ,Each item includes the bounding box location of the defect, forming a full-process structural violation detection capability.

[0114] To achieve the above dual-path collaborative training, the present invention designs the following joint loss function:

[0115] ,

[0116] in, represents the joint loss function of the image generation task (including mask and contour stages), represents the DETR category and position loss terms of defect detection at each stage, and α is the weight adjustment factor between tasks, which is used to balance the training priority of structure reconstruction and defect recognition.

[0117] This paper utilizes a unified Transformer backbone to achieve feature sharing, integrating image generation and object detection tasks. While satisfying process rule constraints and graph connectivity, it implements an integrated workflow from layout to mask / contour + DRC / MRC / LRC. This system not only delivers high-precision modeling and defect recognition performance, but also significantly reduces model deployment costs. Its advantages include high automation, unified modeling, and engineering scalability, making it suitable for integrated lithography modeling and detection systems at advanced process nodes.

[0118] Example

[0119] This embodiment uses the industry-standard lithography modeling software Calibre in a real manufacturing simulation environment to systematically simulate the integrated circuit layout under various process combination conditions, and constructs an integrated image dataset that can be used for image generation and defect detection tasks.

[0120] During the simulation preparation phase, the input is a large-scale layout drawing with a size of 1900 μm × 1500 μm. The present invention sets the simulation process through a supporting recipe file, defining optical model parameters and key process variables such as exposure dose, focus offset, photoresist threshold, and light source type. It also integrates design inspection rules for DRC, MRC, and LRC to ensure that the output data has a basis for defect annotation.

[0121] After completing the simulation on the complete layout, the layout layer, mask layer, and resist contour layer are segmented into 6000 nm × 6000 nm clips. These clips are then saved as 512×512 pixel images to facilitate subsequent large-scale model learning. The simulation log file (.log) records the entire execution process, including the number of parallel threads used (96 CPU cores), the total task duration (29 minutes), and the execution status of DRC and LRC checks, providing detailed resource and process traceability for the dataset construction process. Finally, the simulation output image is combined with the rule-based checking module for automatic defect detection and annotation. The output format is standard JSON, containing the bounding box coordinates and corresponding defect type of each defect target, supporting subsequent DETR model training.

[0122] The above process not only ensures the coverage of image data in terms of spatial resolution and process diversity, but also guarantees the accuracy and consistency of defect annotation, providing a reliable, real, and scalable high-quality data foundation for the unified modeling framework proposed in this invention.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A unified multi-task computational lithography modeling and defect detection method assisted by a large visual model, characterized by: The following steps are involved: Step 1: Use the Calibre simulation tool to generate layout, mask, and contour images under multiple process conditions, and build a dual-task dataset of image generation and defect detection to provide a data foundation for subsequent multi-task learning; Step 2: Based on the data constructed in step 1, a two-stage image generation model integrating process conditions is constructed based on a unified Transformer architecture to generate mask and contour images from the layout image. Step 3: Based on the real defect detection dataset constructed in step 1, a DETR-based defect detection model is constructed for the three types of images: layout, mask, and contour, and DRC, MRC, and LRC detection are performed; Step 4. Based on the two-stage image generation model and defect detection model, a unified end-to-end lithography modeling and defect detection framework is constructed. This framework takes the layout image acquired in Step 1 as input and first generates mask and contour images through the two-stage image generation model in Step 2. Then, the generated image is combined with the real image and the DETR detection module in Step 3 is introduced to achieve comprehensive detection and localization of multiple types of defects such as DRC, MRC, and LRC.

2. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1, characterized in that: In step 1, four process parameter dimensions are introduced in the simulation process of constructing the image generation subset, namely, light source type, photoresist threshold, focal offset and exposure dose. Among them, the light source type includes annular light source, circular light source and composite light source; the photoresist threshold selects the dose thresholds of three mainstream photoresist materials to cover high, medium and low exposure response curve characteristics; the focal offset is set to two typical values ​​of 0 nm and 50 nm; the exposure dose is set to two nominal doses of 1.0× and 1.2× to simulate process fluctuation conditions; a variety of process combinations are constructed through the four process parameter dimensions, and for each combination condition, layout-mask-contour triplet image data are simulated and generated.

3. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1, characterized in that: In step 1, the defect detection subset is generated by combining the rule checker with the automated annotation script, and the annotation is performed using the standard target detection format. The annotation information is stored in JSON format.

4. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1, characterized in that: In step 2, the two-stage image generation model includes an input construction module, a main input encoding module, an auxiliary conditional encoding module, and a cross-modal fusion module. Specifically: In the input construction module, the model input includes the main input layout image and auxiliary input process parameter information; In the main input encoding module, the layout image input first passes through a Patch Embedding module for extracting local structures, and then the backbone structure Swin Transformer performs layer-by-layer feature extraction; The auxiliary conditional encoding module receives five-channel inputs, including a light source image, a mask image, and three process parameter maps. It first performs preliminary fusion through a layer of 3×3 convolution to enhance semantic connectivity between channels. It then extracts high-level semantic information through three sets of DownBlock structures. The final output is dimensionally aligned through a layer of convolution, allowing it to be input into the cross-modal fusion module along with the layout encoding features. In the cross-modal fusion module, the auxiliary input encoding result is used as the query, and the layout features are used as the key and value. The attention mechanism is used to achieve semantic alignment and response relationship modeling.

5. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1 or 4, characterized in that: In step 2, the two-stage image generation model adopts a two-stage joint optimization strategy. During the training process, multiple loss functions are jointly driven. The objective function is as follows: , Among them, L mask and L contour They are the loss functions of the mask image and the contour image, respectively, and are constructed as follows: , , Among them, P mask With P contour Represent the mask and contour prediction map generated by the model, G mask With G contour The corresponding true label map; BCE is used to measure the difference between the predicted value and the true value of each pixel; Dice loss emphasizes the overlap between the predicted area and the true area, which is suitable for structural modeling in unbalanced data; and triple loss and The discriminative ability of the model is enhanced by constraining similar samples to be closer and heterogeneous samples to be farther apart in the feature space.

6. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1 or 4, characterized in that: In step 2, a cross-process comparison optimization mechanism is introduced into the two-stage image generation model, and Triplet Loss process-aware supervision is introduced into the loss function. During the training process, for the same input layout, simulations are performed using different process parameters p to obtain multiple generated image results. Triplet loss is constructed for images with the same process conditions and images with different process conditions. The expression for Triplet Loss is as follows: , Where x is the input layout; p, p+, and p− represent the current process parameters, positive sample process, and negative sample process, respectively; G(x, p) represents the feature representation extracted by the model under the conditions of input image x and process parameter p; is the minimum distinguishing distance between positive and negative samples; the symbol Represents the square of the L2 norm of the Euclidean distance, which is used to measure the similarity between feature vectors.

7. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1, characterized in that: In step 3, the specific contents of building a defect detection model based on DETR are as follows: Step 3-1, feature extraction and Transformer encoder construction: Input the defect detection task dataset constructed in step 1 into the convolutional neural network feature extraction backbone network to obtain a dimension of The image feature map is then flattened into a sequence representation and fed into the standard Transformer encoder module; where H and W are the height and width of the original image, respectively, and p is the patch size, indicating that the image is divided into patches, d is the feature dimension after each patch is encoded; the entire feature map is represented as tokens, each token is a d-dimensional vector; Step 3-2, Transformer decoder and target representation generation mechanism: Introducing the object query vector q∈R N×d As the input of the Transformer decoder, the decoder interacts with the encoded features and outputs a prediction result of a candidate object for each query, including the defect bounding box and class probability, thus forming the final object set; Step 3-3, target matching and loss function design: Introduce the Hungarian matching algorithm and calculate the DETR total loss function; Step 3-4, three-stage defect detection model deployment: Build three DETR-based defect detection models for layout, mask, and contour images, respectively, to identify layout design violations (DRC), mask defects (MRC), and resist imaging anomalies (LRC). Each model is trained based on an independent training dataset and outputs a bounding box of the defect target.

8. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 1, characterized in that: In step 4, the end-to-end lithography model uses a unified Transformer backbone as a shared encoder. The input is the layout image and related process condition information. After the encoder outputs common semantic features, it is used by a multi-branch decoder to perform the image generation task and the hotspot detection task respectively. ① For image generation tasks: Use the image generation branch to perform high-precision modeling on the layout image and gradually output the mask graphic x mask With imaging profile x contour ,Among them, mask generation depends on layout and process condition input, and contour generation introduces mask image as auxiliary information on this basis; ② For the hotspot detection task: Based on the shared feature representation, the structural defects in the layout, mask, and contour images are identified through the DRC, MRC, and LRC detection modules respectively, and the target set Y is output. layout 、Y mask 、Y contour ,Each item includes the bounding box location of the defect, forming a full-process structural violation detection capability.

9. The unified multi-task computational lithography modeling and defect detection method assisted by a large visual model according to claim 8, characterized in that: In step 4, to achieve dual-path collaborative training, the following joint loss function is designed: , in, Represents the joint loss function of the image generation task, including the mask and contour stages; Denotes the DETR category and position loss term of defect detection at each stage; It is a weight adjustment factor between tasks, used to balance the training priority of structure reconstruction and defect recognition.

Citation Information

Patent Citations

  • Industrial product defect detection system based on multi-task learning

    CN111179250A

  • Casting surface defect detection method and system based on improved DETR

    CN117314837A