Photoetching mask optimization method and system based on automatic encoder network
Through the lithographic mask optimization method based on the automatic encoder network, the convolutional attention mechanism and multi-scale feature extraction layer are used to optimize the network structure, which solves the problem of large time overhead of the existing mask optimization method and the small feature redundancy in the generated mask, and achieves efficient and robust mask optimization effect.
Patent Information
- Application Number
- CN202510603884.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing mask optimization methods have large time overhead and the generated mask may have tiny feature redundancy, affecting robustness and manufacturability.
The lithography mask optimization method based on the automatic encoder network is adopted. By constructing a data set containing the target layout layout and the real reference mask, the mask optimization network is trained, and the prediction mask corresponding to the target layout layout is generated, and the network structure is optimized using the convolutional attention mechanism and multi-scale feature extraction layer.
It significantly improves training efficiency, reduces the time overhead during network operation, and generates masks closer to the optimal solution, improving printability and robustness.
Smart Images

Figure CN120215201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of integrated circuits, and particularly to a lithography mask optimization method and system based on an autoencoder network. Background Art
[0002] In the current era of the rapid development of very large scale integration (VLSI), the integrated circuit manufacturing process is moving towards increasing complexity and fineness. Among them, lithography technology, as the core driving force for the development of the integrated circuit manufacturing process and one of the most complex technologies, plays a crucial role. Lithography is a key means to accurately transfer the integrated circuit design pattern on the mask plate to the silicon wafer with the help of a lithography machine. Its specific process includes steps such as photoresist coating, exposure, development, and etching. In the integrated circuit layout design, the patterns are rich and diverse, including both dense patterns and sparse patterns. In particular, the design shapes of logic devices are more complex, and the lithography exposure process windows of different types of patterns often vary. In addition, when the photomask pattern design contains patterns of different lengths arranged side by side, the ends of the longer patterns will be affected by the optical proximity effect (OPE) and show the phenomenon of "end bulging" due to the lack of shielding by adjacent patterns, which has an adverse impact on the electrical characteristics and production yield of integrated circuits.
[0003] In the design for manufacturability (DFM) process of integrated circuits, lithography is a key link in converting the design layout into an actual physical structure. The mask, as an important carrier for transmitting light and printing an image on the wafer during the lithography process, has strict requirements for the proximity between the wafer image and the target image. However, as the critical dimensions of integrated circuits continue to shrink, the wavelength of light in the lithography process is much larger than the critical dimensions, resulting in inevitable diffraction effects when the light source passes through the mask. This causes problems such as line-end shortage and corner rounding after the unoptimized mask is printed, thus seriously affecting the electrical performance and production yield of integrated circuits. Therefore, optimizing the mask has become an inevitable trend. Currently, mask optimization methods mainly include optical proximity correction and sub-resolution assist features. However, most existing mask optimization methods generally have the problem of large time overhead, and the generated mask may have redundant small features, which will affect the robustness and manufacturability of the mask.
[0004] To further improve the mask optimization effect, in the case of continuous reduction of feature sizes, it is necessary to introduce forward sub-resolution assist features (SRAF) or reverse sub-resolution assist features (SRIF) in the mask optimization design to reduce the process deviation caused by the difference in pattern density and improve the uniformity of depth of focus and process window. The size of the added assist features is smaller than the imaging resolution of the lithography system. Although they do not form exposure patterns themselves, they will affect the light intensity distribution of the surrounding mask patterns. The SRAF technology introduced since the 90nm node has become a key assist means for mask optimization in the manufacture of integrated circuits at the 40nm and smaller nodes. Currently, the generation methods of SRAF mainly include rule-based and model-based. The rule-based method relies on a large amount of experimental experience accumulation to set the placement rules of SRAF patterns. Once the process changes, the rule table needs to be re-established; while the model-based method also requires experience to determine the placement quantity and initial position of SRAF patterns. Subsequently, it is necessary to continuously adjust parameters and judge through mask rule checking. The process is complex and cumbersome, with poor adaptability to new processes, difficult to quickly generate effective assist features, and low efficiency.
[0005] Therefore, current lithography technology faces many challenges in integrated circuit manufacturing, especially in mask optimization, and there are some deficiencies in existing methods. For example, most existing mask optimization methods generally have the problem of large time overhead, and the generated masks may have redundant small features, which will affect the robustness and manufacturability of the masks. For this reason, in the publicly available materials and literature, it is proposed to train a neural network model using deep learning. By using the trained model to obtain the mask, it has a faster generation speed, and both printability and robustness are better. For example, in the patent application document with the publication number CN117058362A, a mask optimization method based on a semantic segmentation network is proposed, and the backbone network SegNet is used for mask optimization. However, the neural network architecture adopted in the solution is fuzzy, and for complex layouts, high-quality masks cannot be generated. Moreover, the traditional method of using upsampling to implement image reconstruction will bring problems such as noise and blurring. Another example is the literature "Research on Lithography Mask Optimization Technology Based on Deep Learning, Tang Fuxin, Master's Thesis", which proposes an end-to-end mask optimization framework TransU-ILT based on the improved TransUnet. This solution is essentially supervised learning and requires the use of optimized masks to guide model training. However, in the actual industrial process, for any target layout, the corresponding optimal mask shape is not a priori known. In this case, since the optimized mask is approximately obtained through the traditional ILT algorithm rather than the true optimal solution, the network may be guided to an inaccurate optimization direction during the learning process, and the accumulation of this deviation will cause the network to gradually deviate from the true optimal solution space, thus affecting the adaptability and generalization performance of the model under different target designs. Summary of the Invention
[0006] The present invention aims to solve the problem of large training overhead of existing methods.
[0007] The present invention solves the above technical problems through the following technical means:
[0008] A lithography mask optimization method based on an autoencoder network is proposed, and the method includes:
[0009] Construct a data set including the target layout pattern and the true reference mask;
[0010] Use the data set to train the mask optimization network, generate a predicted mask corresponding to the target layout pattern, and calculate the loss between the predicted mask and the true reference mask to form a loss function for mask optimization;
[0011] Use the trained mask optimization network to obtain the optimized mask image;
[0012] Among them, the mask optimization network includes an encoder and a decoder connected in sequence. The encoder includes a first encoding module to a fifth encoding module connected in sequence, and the decoder includes a first decoding module to a fifth decoding module connected in sequence. Convolutional attention blocks are provided in both the second encoding module and the fourth encoding module, and multi-scale feature extraction layers are provided in both the third encoding module and the fifth encoding module.
[0013] Further, the first encoding module includes 1 convolutional layer.
[0014] Further, both the second encoding module and the fourth encoding module include a convolutional layer Conv_1, a convolutional layer Conv_2, a convolutional layer Conv_3, a convolutional attention block SEB lock, and a pooling layer Max Pool_1. The input features are respectively used as the inputs of the convolutional layer Conv_1, the convolutional layer Conv_2, and the convolutional attention block SEB lock. The outputs of the convolutional layer Conv_1, the convolutional layer Conv_2, and the convolutional attention block SEB lock are added element-wise and then used as the input of the convolutional layer Conv_3. The output of the convolutional layer Conv_3 is connected to the pooling layer Max Pool_1.
[0015] Further, both the third encoding module and the fifth encoding module include a convolutional layer Conv_4, a convolutional layer Conv_5, a convolutional layer Conv_6, a multi-scale feature extraction layer MSFE, and a pooling layer Max Pool_2. The input features are respectively used as the inputs of the convolutional layer Conv_4 and the multi-scale feature extraction layer MSFE. The output of the convolutional layer Conv_4 is connected to the input of the convolutional layer Conv_5. The output of the convolutional layer Conv_5 and the output of the multi-scale feature extraction layer MSFE are added element-wise and then used as the input of the convolutional layer Conv_6. The output of the convolutional layer Conv_6 is connected to the pooling layer Max Pool_2.
[0016] Further, the first decoding module includes an upsampling layer.
[0017] Further, the second decoding module to the fifth decoding module all include a convolutional layer Conv_7, a convolutional layer Conv_8, a normalization layer, and an upsampling layer connected in sequence.
[0018] Further, there is a skip connection between the encoder and the decoder.
[0019] Further, the loss function includes a mean squared loss function and a process variation band loss function. The mean squared loss function is used to detect the difference between the wafer image after lithography processing of the predicted mask and the true target layout; the process variation band loss function is used to measure the robustness of the optimized mask.
[0020] Further, the training of the mask optimization network using the data set includes:
[0021] Dividing the data set into a training data set and a test data set;
[0022] Training the mask optimization network using the training data set, calculating the corresponding gradient using a loss function, and reversely optimizing the network parameters according to the gradient descent method;
[0023] Testing the mask optimization network after learning using the test data set.
[0024] In addition, the present invention also proposes a lithography mask optimization system based on an autoencoder network, including:
[0025] A data set construction module for constructing a data set including a target layout layout and a true reference mask;
[0026] A training module for training the mask optimization network using the data set, generating a predicted mask corresponding to the target layout layout, and calculating the loss between the predicted mask and the true reference mask to form a loss function for mask optimization;
[0027] A mask optimization module for obtaining an optimized mask image using the trained mask optimization network;
[0028] Wherein, the mask optimization network includes an encoder and a decoder connected in sequence. The encoder includes a first coding module to a fifth coding module connected in sequence, and the decoder includes a first decoding module to a fifth decoding module connected in sequence. Convolutional attention blocks are provided in both the second coding module and the fourth coding module, and multi-scale feature extraction layers are provided in both the third coding module and the fifth coding module.
[0029] The advantages of the present invention are:
[0030] (1) By introducing convolutional attention mechanism and multi-scale feature extraction in the encoder, the convolutional attention mechanism can guide the model to pay more attention to the important features in the target layout, while multi-scale feature extraction enables the model to widely and deeply extract feature information in different dimensions. Thanks to the synergistic effect of these two methods, the network can not only effectively extract richer features, but also optimize the overall structure and reduce the scale. This not only speeds up the convergence speed of the network, significantly improves the training efficiency, but also reduces the time overhead during the operation of the network, providing strong support for the efficient operation of the model.
[0031] (2) The present invention reconstructs the autoencoder model. In the feature extraction stage of downsampling the model, a convolutional attention module and a multi-scale feature extraction module are used. The former can automatically learn the important regions and detailed features in the image, focus more attention on the key information, suppress irrelevant or secondary information, thereby enhancing the feature representation. The latter can process data at different scales, so as to capture more comprehensive feature information. Features at different scales have different sensitivities to data changes and noises. Fusing multi-scale features can make the model more robust when facing various complex situations, reducing the performance degradation caused by factors such as scale changes. The combined use of the two can give full play to their respective advantages, enabling the model to focus on both key details and the overall structure during learning, thus significantly improving the model performance and generalization ability.
[0032] (3) The present invention designs a network loss function for the mask optimization task. This loss function combines the traditional MSE with PVB. The MSE loss function focuses on the Euclidean distance between the coarsened mask and the true mask, while the PVB loss function focuses on the robustness of the predicted mask. This function fully considers the uncertainty brought by process variations, evaluates the advantages and disadvantages of mask design by integrating various factors, and combines the two loss functions to make the mask generated by the network closer to the optimal solution.
[0033] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic flowchart of a lithography mask optimization method based on an autoencoder network proposed in an embodiment of the present invention;
[0035] Figure 2 is a schematic structural diagram of a mask optimization network in an embodiment of the present invention;
[0036] Figure 3 is a mask optimization flowchart in an embodiment of the present invention;
[0037] Figure 4 is a comparison diagram of mask optimization visualization results in an embodiment of the present invention;
[0038] Figure 5 is a performance comparison diagram between the mask optimization network of the present invention and other existing technologies in an embodiment of the present invention;
[0039] Figure 6 is a schematic structural diagram of a lithography mask optimization system based on an autoencoder network proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] As Figure 1 and Figure 2 shown, an embodiment of the present invention proposes a lithography mask optimization method based on an autoencoder network, and the method includes the following steps:
[0042] S10. Construct a data set including a target layout layout and a true reference mask;
[0043] S20. Use the data set to train a mask optimization network, generate a predicted mask corresponding to the target layout layout, and calculate the loss between the predicted mask and the true reference mask to form a loss function for mask optimization;
[0044] Among them, the mask optimization network includes an encoder and a decoder connected in sequence. The encoder includes a first encoding module to a fifth encoding module connected in sequence. The decoder includes a first decoding module to a fifth decoding module connected in sequence. Convolutional attention blocks are provided in both the second encoding module and the fourth encoding module, and multi-scale feature extraction layers are provided in both the third encoding module and the fifth encoding module;
[0045] S30. Use the trained mask optimization network to obtain an optimized mask image;
[0046] It should be noted that the obtained optimized mask image is subjected to a lithography simulation operation to generate a corresponding wafer image.
[0047] It should be noted that in this embodiment, the backbone network in the previous machine learning-based framework is replaced with a self-designed encoder network. By introducing a convolutional attention mechanism and multi-scale feature extraction in the encoder, the convolutional attention mechanism can guide the model to pay more attention to the important features in the target layout, while multi-scale feature extraction enables the model to widely and deeply extract feature information in different dimensions. Thanks to the synergistic effect of these two methods, the network can not only effectively extract richer features, but also optimize the overall structure and reduce the scale, which not only speeds up the convergence speed of the network, significantly improves the training efficiency, but also reduces the time overhead during the operation of the network, providing strong support for the efficient operation of the model.
[0048] As a further preferred technical solution, the first encoding module includes 1 convolutional layer.
[0049] Both the second encoding module and the fourth encoding module include a convolutional layer Conv_1, a convolutional layer Conv_2, a convolutional layer Conv_3, a convolutional attention block SEB lock, and a pooling layer Max Pool_1. The input features are respectively used as the inputs of the convolutional layer Conv_1, the convolutional layer Conv_2, and the convolutional attention block SEB lock. The outputs of the convolutional layer Conv_1, the convolutional layer Conv_2, and the convolutional attention block SEB lock are added element-wise and then used as the input of the convolutional layer Conv_3. The output of the convolutional layer Conv_3 is connected to the pooling layer Max Pool_1.
[0050] It should be noted that in the second encoding module and the fourth encoding module designed in this embodiment, the convolutional layer Conv_1 has the ability to integrate information across channels, while the convolutional layer Conv_2 has a large receptive field and can capture local spatial features. The two cooperate with each other to extract features from different scales and angles, greatly enriching the feature expression, enabling the model to take into account both local details and overall structural information.
[0051] Moreover, this module introduces a multi-scale convolutional attention mechanism, extracts multi-scale context information by means of a multi-branch convolutional structure, replaces the global calculation in the traditional self-attention mechanism, and effectively reduces the computational complexity. When generating attention weights, channel-wise convolution is adopted, abandoning complex matrix multiplication, significantly reducing the number of parameters while maintaining the global perception ability. Subsequently, the different features extracted are fused to integrate the complementary information therein, further reducing feature redundancy. Finally, by performing secondary fusion on the features, the feature dimension is reduced, the amount of calculation and the number of parameters are reduced, the computational efficiency is improved, and at the same time, the occurrence of overfitting is effectively prevented.
[0052] Both the third encoding module and the fifth encoding module include a convolutional layer Conv_4, a convolutional layer Conv_5, a convolutional layer Conv_6, a multi-scale feature extraction layer MSFE, and a pooling layer Max Pool_2. The input features are respectively used as the inputs of the convolutional layer Conv_4 and the multi-scale feature extraction layer MSFE. The output of the convolutional layer Conv_4 is connected to the input of the convolutional layer Conv_5. The output of the convolutional layer Conv_5 and the output of the multi-scale feature extraction layer MSFE are added element-wise and then used as the input of the convolutional layer Conv_6. The output of the convolutional layer Conv_6 is connected to the pooling layer Max Pool_2.
[0053] In the construction of these two encoding modules, namely the third encoding module and the fifth encoding module, this embodiment first uses two consecutive 3×3 convolutions (Conv_4, Conv_5) to perform feature extraction tasks. This design takes advantage of the property that two 3×3 convolutions are equivalent to the receptive field of a 5×5 convolution. Without sacrificing the feature extraction range, it significantly reduces the number of parameters, effectively reducing the computational complexity and the risk of overfitting. In addition, the non-linear activation function paired with the convolutional layer endows the model with stronger non-linear expression ability through multiple superimposed non-linear transformations, enabling it to extract more complex and abstract feature representations, thereby enhancing the overall performance.
[0054] To further improve the model's ability to handle complex target layouts, a multi-scale feature extraction module (MSFE layer) is introduced into the module. This module can comprehensively capture multi-scale feature information from details to the whole in the image by performing convolutional operations at different scales. In the specific implementation, the MSFE layer was originally planned to use large kernel convolutions of 5×5 and 7×7. However, to balance performance and computational overhead, this solution uses a stack of 3×3 convolutions for equivalent replacement. This replacement strategy not only ensures that the receptive field does not decrease but also effectively reduces the computational complexity, ultimately achieving a significant improvement in model performance.
[0055] It should be noted that the autoencoder constructed in this embodiment by repeatedly using SEConv and MSFEBlock has many advantages. From the perspective of feature extraction, the multi-scale feature extraction module (MSFE) in the MSFEBlock can capture the detailed and structural information of the image at different scales through different-scale convolutional kernels (such as 3×3Conv_4, 3×3Conv_5, etc.), effectively extracting features from the fine details to the overall layout of the target, enriching the feature dimension, and enhancing the expression ability for complex target layouts.
[0056] In terms of the attention mechanism, the multi-scale convolutional attention module in the SEConv module can assign weights to the extracted features. It enables the model to focus on key features, suppress unimportant information, makes the model pay more attention to the regions crucial for the image reconstruction task, and improves the utilization efficiency of features.
[0057] In addition, by alternating the use of the two, a cycle of feature extraction and attention optimization is formed during the encoding process. The encoding module continuously extracts and filters features, which helps to better compress the input information, so that the desired pattern output can be generated more accurately during decoding, enhancing the model's understanding and processing ability of data, and improving the overall performance and robustness of the model.
[0058] As a further preferred technical solution, the first decoding module includes an upsampling layer.
[0059] The second to fifth decoding modules all include a convolutional layer Conv_7, a convolutional layer Conv_8, a normalization layer, and an upsampling layer connected in sequence.
[0060] It should be noted that in order to gradually restore the image to its original resolution in this embodiment, an upsampling layer is added to each decoding module to progressively increase the image resolution. The transposed convolutional layer in the constructed decoder dynamically learns the mapping relationship from the low-resolution feature map to the high-resolution output through a parameterized convolutional kernel, rather than relying on a fixed upsampling interpolation algorithm. This learnability enables it to adaptively restore the detailed information.
[0061] In this embodiment, the maximum pooling index is replaced with a skip connection. Each decoding module combines the skip connection to introduce the multi-scale features of the encoder (such as the local details extracted by SEConv and the global context captured by MSFE), enabling the decoder to have all-round features from local to global, enriching the feature representation, and contributing to improving the overall performance of the model. At the same time, the multi-scale features introduced by the skip connection provide more valuable information for the decoder. During the learning process, the decoder can capture the key feature information faster, thereby improving the model learning speed, accelerating the model convergence speed, and enhancing the training efficiency.
[0062] Furthermore, there is a skip connection between the encoder and the decoder.
[0063] Specifically, the output of the first encoding module and the output of the fourth decoding module are added element-wise and then used as the input of the fifth decoding module. The output of the second encoding module and the output of the third decoding module are added element-wise and then used as the input of the fourth decoding module. The output of the third encoding module and the output of the second decoding module are added element-wise and then used as the input of the third decoding module. The output of the fourth encoding module and the output of the first decoding module are added element-wise and then used as the input of the second decoding module. The output of the fifth encoding module is connected to the input of the first decoding module.
[0064] It should be noted that in this embodiment, the target layout layout is input into the autoencoder network, and the network mainly generates the corresponding mask through convolutional operations. Subsequently, the generated mask is compared with the real mask to calculate the loss. The loss function here consists of two parts: the mean squared error loss function MSE and the process variation band loss function PVB. Among them, the MSE loss function is used to calculate the mean squared error (MSE) between the mask generated by the autoencoder network and the real mask to measure the similarity between the two. The PVB loss function focuses on the robustness of the predicted mask. This function fully considers the uncertainty brought by process variations and comprehensively evaluates the quality of the mask design from multiple aspects.
[0065] As a further preferred technical solution, such asFigure 3 As shown in Figure 3 , step S20: training the mask optimization network using the data set specifically includes the following steps:
[0066] S21. Divide the data set into a training data set and a test data set;
[0067] S22. Train the mask optimization network using the training data set, calculate the corresponding gradient using the loss function, and reverse-optimize the network parameters according to the gradient descent method;
[0068] S23. Test the learned mask optimization network using the test data set.
[0069] Specifically, in the training stage, 4,875 pairs of data sets in the ICCAD2013 competition are selected to construct a training set, and these data are combined in pairs to form training pairs. Through a dedicated data set preparation function, the data containing the target layout and the true mask are read into the algorithm. Among them, the target layout is used as the input and is fed into the first convolutional block of the encoder. After four modules of convolutional calculations, a feature map corresponding to the target layout is generated. Subsequently, this feature map is fed into the first upsampling module of the decoder, and after four deconvolutional blocks of calculations, a mask corresponding to the target layout is generated. Then, the generated mask and the given reference true mask are used for loss calculation. The loss function uses MSE and PVB as the combined loss to calculate the loss value between the mask and the true mask, thereby constituting the loss function for mask optimization.
[0070] It should be noted that in the model training stage, according to the composition of the loss function, first calculate the loss value between the mask generated by the autoencoder network and the true mask. Based on this loss value, further calculate the gradient. Then, the model updates the weights of the network model along the direction of the loss function gradient descent. This process is continuously iterated until the model converges, thus completing the training work of the model, and finally achieving the convergence of the network and saving the corresponding network parameters. In this way, the autoencoder network is continuously optimized to make the generated mask closer to the true mask, improving the accuracy and reliability of mask generation.
[0071] After the network iteration is completed in this embodiment, the function in the network that saves the network parameters is used to save the weights in the network as a weight file for reading the network parameters when used.
[0072] In a single test process, first, the network needs to be trained. It is necessary to prepare the target layout data set in advance, reasonably divide the data set into a training set and a test set, and then input the training set into the network for training. Inside the network model, a series of modeling and calculation operations will be carried out, and after multiple iterations until the loss function reaches the convergence state.
[0073] When the loss function converges, the network parameters of the converged network need to be saved. These parameters record the important information learned by the network during the training process. After that, input a single test layout that needs to be optimized, and the network will perform single-sample optimization for the test layout. After the network is processed, the optimized mask corresponding to the test layout will be output. Finally, save the corresponding mask as the output result of the network, which means that the mask optimization operation is completed.
[0074] From the above-mentioned overall process of a single mask, it can be seen that the lithography mask optimization method based on the automatic encoder network of this embodiment can perform mask optimization tasks for both single target layout and batch target layout. When the training is completed, the target layout can achieve the goal of the optimal mask directly in the network.
[0075] from Figure 4 It can be seen that compared with the wafer image directly generated based on the original layout, the wafer image generated by the existing model has improved the printability of the mask to a certain extent. However, due to the influence of physical effects such as diffraction, the generated wafer image still has edge mismatches and even bridging (connection) defects. For example, the mask generated by the traditional model has a large gap in the wafer image. Figure 4 There are obvious connection defects in (a), (b) and (c). The mask generated by the network proposed in this embodiment is as follows Figure 4 In (d), the wafer image obtained after lithography simulation is not only closer to the target layout in terms of edge matching, but also significantly improves the printability of the mask, greatly reduces the connection phenomenon, effectively improves the image generation quality, and further verifies its superiority.
[0076] according to Figure 5 The comparative analysis of the experimental results shows that this embodiment shows significant advantages in key performance indicators. In terms of accuracy indicators, the proposed method is superior to the comparative method in both squared L2 error (L2) and process variation band (PVB): compared with the traditional ILT-based method, this embodiment reduces the squared L2 error by 27.4% and PVB by 18.7%; compared with the Neural-ILT-based learning method, the squared L2 error is further optimized by 14.9% and PVB is reduced by 23.5%; compared with the A2-ILT method, this embodiment can still achieve an 8.3% reduction in squared L2 error and a 3.2% improvement in PVB. In terms of computational efficiency, the turnaround time (TAT) of this embodiment is particularly outstanding: its TAT value is only 0.0012 times that of the traditional ILT method, achieving a 35.8-fold acceleration ratio compared to Neural-ILT, and also achieving a 16-fold speed increase compared to A2-ILT. This result shows that the method proposed in this embodiment significantly improves the printability and robustness of the mask.
[0077] In addition, compared with the two technical solutions cited in the background art, this embodiment achieves a significant breakthrough in technical effects. From Figure 5 the data, the L2 error and PVB of this solution are 31,941 and 41,382 respectively. Compared with the data of 37,412 and 49,316 in Table 3 of the literature "Research on Lithography Mask Optimization Technology Based on Deep Learning" and the data of 37,216 and 51,511 in the patent application document with the publication number CN117058362A, Figure 3 this embodiment greatly improves the printability and robustness of the mask and can generate a printing mask with better quality.
[0078] In addition, as Figure 6 shown, another embodiment of the present invention further proposes a lithography mask optimization system based on an autoencoder network, and the system includes:
[0079] a dataset construction module 10 for constructing a dataset including a target layout layout and a true reference mask;
[0080] a training module 20 for training a mask optimization network using the dataset, generating a predicted mask corresponding to the target layout layout, and calculating a loss between the predicted mask and the true reference mask to form a loss function for mask optimization;
[0081] a mask optimization module 30 for obtaining an optimized mask image using the trained mask optimization network;
[0082] wherein, the mask optimization network includes an encoder and a decoder connected in sequence, the encoder includes a first encoding module to a fifth encoding module connected in sequence, the decoder includes a first decoding module to a fifth decoding module connected in sequence, and convolutional attention blocks are arranged in both the second encoding module and the fourth encoding module, and multi-scale feature extraction layers are arranged in both the third encoding module and the fifth encoding module.
[0083] It should be noted that the mask optimization network and the loss function adopted in the lithography mask optimization system based on the autoencoder network of the present invention can refer to the above method embodiments and will not be elaborated here.
[0084] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0085] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0086] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0087] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0088] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A lithography mask optimization method based on an autoencoder network, characterized in that: include: Construct a dataset containing the target layout and the real reference mask; The mask optimization network is trained using the dataset to generate a predicted mask corresponding to the target layout, and the loss function of the mask optimization is constructed by calculating the loss between the predicted mask and the real reference mask. Using the trained mask optimization network to obtain an optimized mask image; Among them, the mask optimization network includes an encoder and a decoder connected in sequence, the encoder includes a first encoding module to a fifth encoding module connected in sequence, and the decoder includes a first decoding module to a fifth decoding module connected in sequence, wherein the second encoding module and the fourth encoding module are both provided with a convolutional attention block, and the third encoding module and the fifth encoding module are both provided with a multi-scale feature extraction layer.
2. The method for optimizing lithography masks based on an autoencoder network according to claim 1, characterized in that: The first encoding module includes one convolutional layer.
3. The method for optimizing lithography masks based on an autoencoder network according to claim 1, characterized in that: The second encoding module and the fourth encoding module both include a convolution layer Conv_1, a convolution layer Conv_2, a convolution layer Conv_3, a convolution attention block SEB lock and a pooling layer Max Pool_1. The input features are respectively used as the input of the convolution layer Conv_1, the convolution layer Conv_2 and the convolution attention block SEB lock. The outputs of the convolution layer Conv_1, the convolution layer Conv_2 and the convolution attention block SEB lock are added element by element as the input of the convolution layer Conv_3. The output of the convolution layer Conv_3 is connected to the pooling layer Max Pool_1.
4. The method for optimizing lithography masks based on an autoencoder network according to claim 1, characterized in that: The third encoding module and the fifth encoding module both include a convolution layer Conv_4, a convolution layer Conv_5, a convolution layer Conv_6, a multi-scale feature extraction layer MSFE and a pooling layer Max Pool_2. The input features are respectively used as inputs of the convolution layer Conv_4 and the multi-scale feature extraction layer MSFE. The output of the convolution layer Conv_4 is connected to the input of the convolution layer Conv_5. The output of the convolution layer Conv_5 and the output of the multi-scale feature extraction layer MSFE are added element by element as the input of the convolution layer Conv_6. The output of the convolution layer Conv_6 is connected to the pooling layer Max Pool_2.
5. The method for optimizing photolithography mask based on an autoencoder network according to claim 1, characterized in that: The first decoding module includes an upsampling layer.
6. The method for optimizing photolithography mask based on an autoencoder network according to claim 1, characterized in that: The second to fifth decoding modules each include a convolution layer Conv_7, a convolution layer Conv_8, a normalization layer and an upsampling layer connected in sequence.
7. The photolithography mask optimization method based on the autoencoder network according to claim 1, characterized in that: The encoder and the decoder are skip-connected.
8. The method for optimizing photolithography mask based on an autoencoder network according to claim 1, characterized in that: The loss function includes an average square loss function and a process variation band loss function. The average square loss function is used to detect the difference between the wafer image of the predicted mask after lithography processing and the actual target layout; the process variation band loss function is used to measure the robustness of the optimized mask.
9. The method for optimizing photolithography mask based on an autoencoder network according to claim 1, characterized in that: The method of using the data set to train the mask optimization network includes: Divide the dataset into training dataset and test dataset; The mask optimization network is trained using the training data set, the corresponding gradient is calculated using the loss function, and the network parameters are reversely optimized according to the gradient descent method; The learned mask optimization network is tested using the test dataset.
10. A lithography mask optimization system based on an autoencoder network, characterized in that: include: A data set construction module, used to construct a data set including a target layout and a true reference mask; A training module is used to train the mask optimization network using the data set, generate a predicted mask corresponding to the target layout, and perform loss calculation between the predicted mask and the real reference mask to form a loss function for mask optimization; A mask optimization module, used to obtain an optimized mask image using a trained mask optimization network; Among them, the mask optimization network includes an encoder and a decoder connected in sequence, the encoder includes a first encoding module to a fifth encoding module connected in sequence, and the decoder includes a first decoding module to a fifth decoding module connected in sequence, wherein the second encoding module and the fourth encoding module are both provided with a convolutional attention block, and the third encoding module and the fifth encoding module are both provided with a multi-scale feature extraction layer.
Citation Information
Patent Citations
Mask optimization method based on semantic segmentation network
CN117058362A