Mask optimization method and system based on UC-ILT framework

Through the mask optimization method of the UC-ILT framework, combined with the encoder and multi-scale feature fusion device, the accuracy problem of mask optimization under complex layout is solved, and the generation of high-quality masks and photolithography is improved.

CN120491395APending Publication Date: 2025-08-15ANHUI UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510596348.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing mask optimization methods cannot generate high-quality masks under complex layouts, and traditional sampling methods have caused noise and blur problems, affecting the quality of lithographic imaging.

Method used

The mask optimization method based on the UC-ILT framework is adopted, and cross-channel cross-attention fusion and high-resolution feature map fusion are performed through an encoder and multi-scale feature fusion device combined with an attention module to generate the optimized mask image.

Benefits of technology

It improves the accuracy and robustness of mask optimization, reduces noise interference, enhances the accuracy and detailed expression of image reconstruction, and is suitable for complex lithography mask optimization tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491395A_ABST
    Figure CN120491395A_ABST
Patent Text Reader

Abstract

The invention discloses a UC-ILT framework-based mask optimization method and system, and the method comprises the steps: inputting a target design layout into a pre-trained mask optimization model, the mask optimization model comprises an encoder, a multi-scale feature fusion device and a decoder, the output of the encoder is connected with the decoder through the multi-scale feature fusion device, and the output of the encoder is connected with the decoder through the multi-scale feature fusion device; the encoder comprises a plurality of convolution blocks, and the adjacent convolution blocks are connected through an attention module. Performing local feature extraction on the input target layout by using an encoder to obtain multi-scale local features; performing cross-channel cross attention fusion on the multi-scale local features by using a multi-scale feature fusion device to obtain a fusion feature map; fusing the fused feature map with the high-resolution feature map recovered by the decoder to obtain a mask image; according to the method, the problems of noise and fuzziness caused by a traditional up-sampling method can be avoided, and mask optimization is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of integrated circuit technology, and in particular to a mask optimization method and system based on a UC-ILT framework. Background Art

[0002] Integrated circuit manufacturing photolithography typically uses a 193nm wavelength light source. However, in subwavelength lithography, where the wavelength is much larger than the integrated circuit feature size, light diffraction can cause image distortion, a phenomenon known as the optical proximity effect. To improve image quality, resolution enhancement technologies (such as optical proximity correction (OPC)) are widely used to effectively mitigate this diffraction-induced distortion.

[0003] Existing mask optimization methods mainly include rule-based OPC, model-based OPC, and learning-based OPC, among which: (1) Rule-based OPC originated from the early manual adjustment of mask pattern design rules. Its core is to achieve correction of different structures by establishing a correction rule table containing a variety of simple and complex graphics. Although this method is more efficient in simple layouts, as the complexity of integrated circuit design increases, the size of the rule table grows exponentially, and the difficulty of matching the design graphics with the rule table also increases. (2) Model-based OPC establishes an OPC model through an accurate mathematical model, calculates the graphics generated on the silicon wafer, and uses an optimization algorithm to adjust the edge position of the mask graphics to achieve high-fidelity graphics transfer; however, due to the high complexity of the mask graphics generated by this method and the low optimization efficiency, subsequent researchers have proposed improvement strategies to reduce complexity. (3) The learning-based OPC uses a plate diagram and a corresponding optimized mask to train a neural network model through deep learning. The mask is obtained by using the trained model. It has a faster generation speed and good printability and robustness. It has been widely studied. For example, the patent application document with publication number CN115222829A proposes to learn its OCP correction process through the AutoEncoder neural network model to generate a mask image. The patent application document with publication number CN117058362A proposes a mask optimization method based on a semantic segmentation network, using the backbone network SegNet for mask optimization; however, the neural network architecture used in these schemes is fuzzy. For complex layouts, high-quality masks cannot be generated. In addition, the traditional upsampling method for image reconstruction will bring about problems such as noise and blur. For example, the paper "Research on Lithography Mask Optimization Technology Based on Deep Learning, Master's Thesis by Tang Fuxin" proposes an end-to-end mask optimization framework, TransU-ILT, based on an improved TransUnet. This framework combines the advantages of CNN in extracting local details with the Transformer in learning global features, effectively extracting deep features of the target layout. However, while this solution also improves on U-Net, it is essentially supervised learning and requires the use of optimized masks to guide model training. However, in actual industrial processes, the optimal mask shape for any target layout is not known a priori. In this case, because the optimized mask is approximated by the traditional ILT algorithm rather than the true optimal solution, the network may be guided towards inaccurate optimization directions during the learning process. The accumulation of such deviations will cause the network to gradually deviate from the true optimal solution space, thus affecting the model's adaptability and generalization performance for different target designs. Furthermore, the Transformer architecture used in this solution is only used in the intermediate range from 64×64×512 to 32×32×768, not a multi-scale feature fusion architecture. Summary of the Invention

[0004] The technical problem to be solved by the present invention is how to make mask optimization more accurate.

[0005] The present invention solves the above technical problems through the following technical means:

[0006] A mask optimization method based on the UC-ILT framework is proposed, which is characterized by including:

[0007] Input the target design layout into a pre-trained mask optimization model, which includes an encoder, a multi-scale feature fuser, and a decoder. The output of the encoder is connected to the decoder via the multi-scale feature fuser. The encoder includes several convolution blocks, and adjacent convolution blocks are connected via an attention module.

[0008] The encoder is used to extract local features of the input target map to obtain multi-scale local features;

[0009] Use the multi-scale feature fuser to perform cross-channel attention fusion on the multi-scale local features to obtain a fused feature map;

[0010] The fused feature map is fused with the high-resolution feature map restored by the decoder to obtain a mask image.

[0011] Furthermore, each layer of the decoder exchanges multi-scale information with the encoder through a multi-scale feature fuser.

[0012] Furthermore, the encoder and the decoder adopt the encoder and decoder in the U-Net framework.

[0013] Furthermore, the attention module includes a channel attention mechanism and a spatial attention mechanism. The input feature map of the channel attention mechanism is multiplied point by point with its output feature map as the input feature map of the spatial attention mechanism, and the output feature map of the spatial attention mechanism is multiplied point by point with its input feature map as the output of the attention module.

[0014] Furthermore, the decoder includes a plurality of upsampling layers or transposed convolution layers, and adjacent upsampling layers or adjacent transposed convolution layers are connected via a pixel reassembly layer.

[0015] Furthermore, the multi-scale feature fuser includes four channel cross-fusion Transformer modules connected in sequence, each of the channel cross-fusion Transformer modules includes a multi-head cross-attention mechanism, the input of the multi-head cross-attention mechanism is connected to several first normalization layers, the multi-scale local features output by the encoder are respectively used as inputs of several first normalization layers, the output of the multi-head cross-attention mechanism is connected to several second normalization layers, the output of each second normalization layer is connected to a multi-layer perceptron, and the input of the second normalization layer is element-wise added to the output of its corresponding multi-layer perceptron.

[0016] Furthermore, before inputting the target layout into the pre-trained mask optimization model, the method further includes:

[0017] Collect open source layout files to build a dataset, and divide the dataset into training dataset and test dataset;

[0018] The mask optimization model is trained using a training data set, and during the training process, the difference between the target design layout and the wafer image after simulated lithography is used as an optimization target;

[0019] The trained mask optimization model is tested using a test data set to obtain a detection result of mask optimization.

[0020] Furthermore, the performance evaluation indicators set during the testing process include the printability of the mask, edge placement error, process variation band, and inference time of the mask optimization model.

[0021] In addition, the present invention also proposes a mask optimization system based on the UC-ILT framework, comprising:

[0022] The target design layout acquisition module is used to acquire the target design layout and input it into the mask optimization module. The mask optimization module deploys a pre-trained mask optimization model. The mask optimization model includes an encoder, a multi-scale feature fuser, and a decoder. The output of the encoder is connected to the decoder via the multi-scale feature fuser. The encoder includes several convolution blocks, and adjacent convolution blocks are connected via an attention module.

[0023] The encoder is used to extract local features of the input target layout and obtain multi-scale local features;

[0024] Multi-scale feature fuser, used to perform cross-channel attention fusion of multi-scale local features to obtain a fused feature map;

[0025] The decoder is used to fuse the fused feature map with the high-resolution feature map restored by the decoder to obtain a mask image.

[0026] In addition, the present invention also proposes a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the mask optimization method based on the UC-ILT framework as described above is implemented.

[0027] The advantages of the present invention are:

[0028] (1) Since traditional learning-based mask optimization methods usually use U-Net as the main framework, the jump connection of U-Net itself directly copies shallow features to deep layers. For complex layouts, U-Net cannot generate high-quality masks. The present invention combines U-Net and a multi-scale feature extractor, replacing the original U-Net method of directly copying shallow features to deep features with features fused by a multi-scale feature fuser, effectively fusing shallow information and making the network features contain more information, which is conducive to the decoder gradually generating better quality masks. In addition, an attention module is introduced between adjacent convolutional blocks of the encoder, so that the network can adaptively focus on the most informative part of the image, effectively remove redundant features, and reduce noise interference, thereby ensuring the accuracy of mask optimization.

[0029] (2) The attention module combines channel attention and spatial attention mechanisms. First, it weights the features of each channel to ensure that the network focuses on the channel that contributes most to the task. Then, through the spatial attention mechanism, it focuses on important spatial regions in the image, further improving the model's ability to identify key features. This mechanism not only improves image reconstruction accuracy, but also makes the model's decision-making process more transparent and enhances the network's interpretability. In addition, the attention module has a low computational overhead and can improve model performance while ensuring high efficiency. It is particularly effective for photolithography mask optimization tasks and can significantly improve the accuracy and robustness of mask optimization.

[0030] (3) The multi-scale feature extractor adopts a channel cross-fusion Transformer module, which combines the multi-scale local features extracted by the encoder with the Transformer's ability to process global context information. It can efficiently fuse multi-scale features and generate optimized masks to cope with complex lithography mask optimization tasks.

[0031] (4) By using a pixel reorganization layer in the decoder, the spatial resolution of the image is enhanced by rearranging the pixels in the feature map, thereby improving the detail expression of the image. This avoids the noise and blurring problems caused by traditional upsampling methods, ensures more accurate restoration of mask details, and improves the printability of the mask. This method improves the quality of image reconstruction without increasing the amount of computation, making it particularly suitable for the high-precision requirements of mask optimization tasks.

[0032] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 1 is a flowchart of a mask optimization method based on the UC-ILT framework proposed in one embodiment of the present invention;

[0034] Figure 2 1 is a network structure diagram of a mask optimization model of a UC-ILT framework in one embodiment of the present invention;

[0035] Figure 3 is a structural diagram of an attention module in one embodiment of the present invention;

[0036] Figure 4 is a structural diagram of a pixel recombination layer in one embodiment of the present invention;

[0037] Figure 5 is a structural diagram of a multi-scale feature extractor in one embodiment of the present invention;

[0038] Figure 6 It is a structural diagram of a mask optimization system based on the UC-ILT framework proposed in one embodiment of the present invention. DETAILED DESCRIPTION

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0040] like Figure 1 and Figure 2 As shown, an embodiment of the present invention proposes a mask optimization method based on the UC-ILT framework, the method comprising the following steps:

[0041] S10, inputting the target design layout into a pre-trained mask optimization model, wherein the mask optimization model includes an encoder, a multi-scale feature fuser, and a decoder, wherein the output of the encoder is connected to the decoder via the multi-scale feature fuser, and the encoder includes a plurality of convolution blocks, and adjacent convolution blocks are connected via an attention module;

[0042] S20, using the encoder to extract local features of the input target layout to obtain multi-scale local features;

[0043] S30, using a multi-scale feature fuser to perform cross-channel cross-attention fusion on the multi-scale local features to obtain a fused feature map;

[0044] S40: Fusing the fused feature map with the high-resolution feature map restored by the decoder to obtain a mask image.

[0045] Specifically, the encoder and decoder use the encoder and decoder in the U-Net framework, replacing traditional skip connections with a multi-scale feature fuser. U-Net can effectively distinguish which pixels are foreground and which pixels are background; the multi-scale feature fuser can fuse information at different scales, giving the model better feature extraction and semantic understanding capabilities; each layer of the decoder exchanges multi-scale information with the encoder through the multi-scale feature fuser. The specific process of mask optimization is as follows: the input image data first passes through the U-Net encoder, which includes a series of convolution block operations for feature extraction. As the depth of the encoder increases, the resolution of the image will gradually decrease, but the number of channels will increase. The goal of the encoder is to compress the input image into a set of highly expressive multi-scale local feature maps, which facilitates the subsequent decoder to partially restore the details of the image. By performing a cross-channel cross-attention mechanism on the multi-scale local feature maps extracted by the encoder, low-level detail features and high-level abstract features are effectively combined. In this way, the network can more accurately understand the global structure of the image and retain important detail information in the image. Through multi-scale feature fusion, U-Net can take into account both global information and local details in the process of feature extraction and image reconstruction, thereby generating more refined and accurate masks or output images.

[0046] In the decoder, each layer exchanges information with the encoder via skip connections. This ensures that the decoder can leverage the low-level spatial information retained by the encoder and scale the image up through inverse convolution to recover detailed image information. This allows the decoder to convert compact feature maps into high-resolution ones, avoiding detail loss during image reconstruction. The fused feature map is then fused with the high-resolution feature map recovered by the decoder to create a mask image. These operations enable the decoder to gradually generate higher-quality masks.

[0047] Furthermore, by introducing an attention module between adjacent convolutional blocks of the encoder, the network can adaptively focus on the most informative parts of the image, effectively remove redundant features, and reduce noise interference, thereby ensuring the accuracy of mask optimization.

[0048] As a further preferred technical solution, Figure 3As shown, the attention module includes a channel attention mechanism and a spatial attention mechanism. The input feature map of the channel attention mechanism is multiplied point by point with its output feature map as the input feature map of the spatial attention mechanism, and the output feature map of the spatial attention mechanism is multiplied point by point with its input feature map as the output of the attention module.

[0049] Specifically, the attention module CBAM includes a channel attention module CAM and a spatial attention module SAM. By dynamically adjusting the importance of channel and spatial features, the network can focus on the most representative features, further improving the accuracy of mask optimization.

[0050] The attention module adaptively weights local and global features in the image, effectively highlighting the most informative areas and suppressing redundant or irrelevant features, thereby enhancing feature extraction accuracy. The convolutional block attention mechanism combines channel attention and spatial attention mechanisms. It first weights the features of each channel to ensure that the network focuses on the channels that contribute most to the task. Then, through the spatial attention mechanism, it focuses on important spatial regions in the image, further improving the model's ability to identify key features. This mechanism not only improves image reconstruction accuracy but also makes the model's decision-making process more transparent, enhancing the network's interpretability. Furthermore, the convolutional fast attention mechanism has a low computational overhead, improving model performance while maintaining high efficiency. It is particularly effective for photolithography mask optimization tasks, significantly improving the accuracy and robustness of mask optimization.

[0051] As a further preferred technical solution, Figure 4 As shown, the decoder includes several upsampling layers or transposed convolution layers, and adjacent upsampling layers or adjacent transposed convolution layers are connected via a pixel reassembly layer.

[0052] It should be noted that this embodiment enhances the spatial resolution of the image by rearranging the pixels in the feature map, thereby improving the image's detail representation and the accuracy of image reconstruction, thereby enhancing the quality of mask generation. Compared with traditional upsampling methods, the pixel reorganization layer can more accurately restore image details and avoid common noise and blur issues. This method improves the quality of image reconstruction without increasing the amount of computation. It is particularly suitable for the high-precision requirements of mask optimization tasks, making the optimized mask more accurate during the lithography process.

[0053] It should be noted that due to the particularity of the mask optimization task, the training process is often unstable. Therefore, this embodiment introduces the convolutional block attention mechanism and the pixel reconstruction layer to stabilize the training process.

[0054] As a further preferred technical solution, Figure 5As shown, the multi-scale feature fuser includes four channel cross-fusion Transformer modules connected in sequence, each of the channel cross-fusion Transformer modules includes a multi-head cross-attention mechanism, the input of the multi-head cross-attention mechanism is connected to several first normalization layers, the multi-scale local features output by the encoder are respectively used as the input of several first normalization layers, the output of the multi-head cross-attention mechanism is connected to several second normalization layers, the output of each second normalization layer is connected to a multi-layer perceptron, and the input of the second normalization layer is element-wise added to the output of its corresponding multi-layer perceptron.

[0055] It should be noted that the multi-scale feature fuser is composed of several channel-cross-fusion Transformer modules. Its function is to perform cross-scale and cross-channel global interaction and semantic alignment on the multi-scale features output by the encoder. By fusing local and global information through multi-head cross-attention, it improves the decoder's perception of fine-grained structure and overall context, thereby significantly reducing the semantic gap between the encoder and decoder and enhancing the subsequent reconstruction quality. The use of residual connections to directly add feature maps helps alleviate the gradient vanishing problem, accelerate network training, and facilitate deeper network structures.

[0056] Therefore, this embodiment adopts a channel cross-fusion Transformer module to combine the local features extracted by CNN with the Transformer's ability to process global context information. It can efficiently fuse multi-scale features and generate optimized masks to cope with complex lithography mask optimization tasks.

[0057] As a further preferred technical solution, the convolution block in the encoder includes a convolution layer, a pooling layer, and a batch normalization layer connected in sequence. The convolution layer and the pooling layer are composed of multiple convolution layers and pooling layers, which are used to extract local features of the input image. In each layer of convolution operation, the spatial resolution of the image gradually decreases, but the number of channels increases to capture deeper feature information. The batch normalization layer is used in the network, which helps to accelerate the convergence speed of the network and reduce training time.

[0058] As a further preferred technical solution, before step S10: inputting the target layout into the pre-trained mask optimization model, the method further includes the following steps:

[0059] Collect open source layout files to build a dataset, and divide the dataset into training dataset and test dataset;

[0060] The mask optimization model is trained using a training data set, and during the training process, the difference between the target design layout and the wafer image after simulated lithography is used as an optimization target;

[0061] The trained mask optimization model is tested using a test data set to obtain a detection result of mask optimization.

[0062] It should be noted that the wafer image after simulated lithography is a wafer image (that is, chip imaging) obtained by lithography simulation of the mask, which is used for comparison with the target design layout.

[0063] As a further preferred technical solution, the performance evaluation indicators set in the test process of this embodiment include the printability of the mask, edge placement error, process variation band and inference time of the mask optimization model; the defined performance indicators are used to comprehensively evaluate the performance of the generated mask under different process conditions and the optimization effect of the model.

[0064] Furthermore, printability (L2 error) refers to the difference between the wafer image obtained after lithography simulation using the optimized mask and the original layout. This difference is quantified by calculating the Euclidean distance between the two, thereby evaluating the printability of the mask in the actual lithography process. Robustness (PVBand) refers to the stability of the mask in the face of process variations during the lithography process. This is achieved by simulating lithography on the mask under different process conditions (such as minimum and maximum process deviations) and calculating the difference between the resulting wafer image and the target layout image. Edge placement error (EPE) is used to evaluate the geometric deviation between the mask pattern and the target layout. Process variation band (PVBand) is used to measure the stability of the mask under different process variation conditions, ultimately providing a quantitative evaluation of the model's performance. The inference time of the mask optimization model refers to the total computational time required when performing lithography simulation using the optimized mask. This time includes all computational steps in the mask detection and simulation processes and primarily reflects the model's operational efficiency and processing speed during the optimization process.

[0065] Specifically, in practical applications, the mask optimization model is trained first, and then the mask optimization of the actual design layout is performed:

[0066] The training process includes:

[0067] Step S1: Collect 4875 open source layout files as the total dataset, and divide the total dataset into a training dataset and a test dataset.

[0068] Step S2: Input the training data set obtained in step S1 into the mask optimization network for learning, and select a suitable algorithm as the network loss function during the learning process of the mask optimization network.

[0069] Step S3: Put the test data set obtained in step S1 into the network learned in step S2 to obtain an optimized mask, perform testing and obtain the detection result.

[0070] The process of mask optimization using the UC-ILT framework is as follows:

[0071] Step S1': Input a 512×512 image to the encoder part of the network.

[0072] Step S2': Feed the features of different scales in the encoder into the multi-scale feature fusion module.

[0073] Step S3': input the feature map after the network encoder into the decoder part, and fuse the feature map after multi-scale feature fusion with the feature map of the corresponding size of the decoder.

[0074] Step S4 ′: outputting an optimized mask image of the same size as the input.

[0075] Specifically, to validate the superiority of the proposed network model, this example used the widely used ICCAD2013 competition dataset, consisting of ten layouts. Results show that this experiment compared the methods used by MOSAIC, GAN-OPC, Neural-ILT, and A2-ILT on the ICCAD2013 competition dataset, as shown in Tables 1 and 2. These results demonstrate that this example achieves lower printability (L2 error), robustness (PVBand), edge placement error (EPE), and inference time across all datasets.

[0076] Table 1 Comparison of different mask optimization methods on some evaluation indicators

[0077]

[0078] Table 2 Comparison with different mask optimization methods on other evaluation indicators

[0079]

[0080] In addition, if Figure 6 As shown, another embodiment of the present invention further proposes a mask optimization system based on the UC-ILT framework, the system comprising:

[0081] The target design layout acquisition module 10 is used to acquire the target design layout and input it into the mask optimization module 20. The mask optimization module 20 deploys a pre-trained mask optimization model. The mask optimization model includes an encoder, a multi-scale feature fuser, and a decoder. The output of the encoder is connected to the decoder via the multi-scale feature fuser. The encoder includes several convolution blocks, and adjacent convolution blocks are connected via an attention module.

[0082] The encoder is used to extract local features of the input target layout and obtain multi-scale local features;

[0083] Multi-scale feature fuser, used to perform cross-channel attention fusion of multi-scale local features to obtain a fused feature map;

[0084] The decoder is used to fuse the fused feature map with the high-resolution feature map restored by the decoder to obtain a mask image.

[0085] It should be noted that other embodiments or specific implementation methods of the mask optimization model in the mask optimization system based on the UC-ILT framework of the present invention can refer to the above-mentioned method embodiments and will not be repeated here.

[0086] In addition, another embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the mask optimization method based on the UC-ILT framework as described in the above embodiment is implemented.

[0087] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0088] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0089] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0090] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0091] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A mask optimization method based on the UC-ILT framework, characterized in that: include: Input the target design layout into a pre-trained mask optimization model, which includes an encoder, a multi-scale feature fuser, and a decoder. The output of the encoder is connected to the decoder via the multi-scale feature fuser. The encoder includes several convolution blocks, and adjacent convolution blocks are connected via an attention module. The encoder is used to extract local features of the input target map to obtain multi-scale local features; Use the multi-scale feature fuser to perform cross-channel attention fusion on the multi-scale local features to obtain a fused feature map; The fused feature map is fused with the high-resolution feature map restored by the decoder to obtain a mask image.

2. The mask optimization method based on the UC-ILT framework according to claim 1, wherein: Each layer of the decoder exchanges multi-scale information with the encoder through a multi-scale feature fuser.

3. The mask optimization method based on the UC-ILT framework according to claim 1, wherein: The encoder and the decoder adopt the encoder and decoder in the U-Net framework.

4. The mask optimization method based on the UC-ILT framework according to claim 1, wherein: The attention module includes a channel attention mechanism and a spatial attention mechanism. The input feature map of the channel attention mechanism is multiplied point by point with its output feature map as the input feature map of the spatial attention mechanism, and the output feature map of the spatial attention mechanism is multiplied point by point with its input feature map as the output of the attention module.

5. The mask optimization method based on the UC-ILT framework according to claim 1, wherein: The decoder includes a plurality of upsampling layers or transposed convolution layers, and adjacent upsampling layers or adjacent transposed convolution layers are connected via a pixel reassembly layer.

6. The mask optimization method based on the UC-ILT framework according to claim 1, wherein: The multi-scale feature fuser includes four channel cross-fusion Transformer modules connected in sequence, each of the channel cross-fusion Transformer modules includes a multi-head cross-attention mechanism, the input of the multi-head cross-attention mechanism is connected to several first normalization layers, the multi-scale local features output by the encoder are respectively used as the input of several first normalization layers, the output of the multi-head cross-attention mechanism is connected to several second normalization layers, the output of each second normalization layer is connected to a multi-layer perceptron, and the input of the second normalization layer and the output of its corresponding multi-layer perceptron are added element by element.

7. The mask optimization method based on the UC-ILT framework according to claim 1, wherein: Before inputting the target layout into the pre-trained mask optimization model, the method further includes: Collect open source layout files to build a dataset, and divide the dataset into training dataset and test dataset; The mask optimization model is trained using a training data set, and during the training process, the difference between the target design layout and the wafer image after simulated lithography is used as an optimization target; The trained mask optimization model is tested using a test data set to obtain a detection result of mask optimization.

8. The mask optimization method based on the UC-ILT framework according to claim 6, wherein: The performance evaluation metrics set during the test include mask printability, edge placement error, process variation band, and inference time of the mask optimization model.

9. A mask optimization system based on the UC-ILT framework, characterized in that: include: The target design layout acquisition module is used to acquire the target design layout and input it into the mask optimization module. The mask optimization module deploys a pre-trained mask optimization model. The mask optimization model includes an encoder, a multi-scale feature fuser, and a decoder. The output of the encoder is connected to the decoder via the multi-scale feature fuser. The encoder includes several convolution blocks, and adjacent convolution blocks are connected via an attention module. The encoder is used to extract local features of the input target layout and obtain multi-scale local features; Multi-scale feature fuser, used to perform cross-channel attention fusion of multi-scale local features to obtain a fused feature map; The decoder is used to fuse the fused feature map with the high-resolution feature map restored by the decoder to obtain a mask image.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the mask optimization method based on the UC-ILT framework according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Mask optimization method and device based on neural network model

    CN115222829A

  • Mask optimization method based on semantic segmentation network

    CN117058362A