Method and system for multi channel mask prediction
A machine learning model generates mask patterns as separate channels for main features and SRAFs, addressing the merging issue in conventional methods by penalizing distance violations, enhancing accuracy and efficiency in semiconductor manufacturing.
Patent Information
- Application Number
- PCT/EP2024/083855
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-03
AI Technical Summary
Conventional methods for generating mask patterns in semiconductor manufacturing often result in undesirable merging of main features and sub-resolution assist features (SRAFs), leading to defective patterns and requiring time-consuming post-processing to isolate features, which complicates the process and increases computational resources.
A machine learning model is trained to generate mask patterns as separate channels for main features and SRAFs, using a cost function to penalize violations of specified distances between them, allowing for independent processing and reducing the need for post-processing.
This approach prevents undesirable merging of features, simplifies post-processing, and reduces computational and time requirements by isolating different types of features, thereby improving the accuracy and efficiency of mask pattern generation.
Smart Images

Figure EP2024083855_03072025_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEM FOR MULTI CHANNEL MASK PREDICTIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority of US application 63 / 614,803 which was filed on 26 December 2023 and which is incorporated herein in its entirety by reference.FIELD
[0002] The embodiments provided herein relate to semiconductor manufacturing, and more particularly to mask pattern design through computational lithography.BACKGROUND
[0003] A lithographic apparatus is a machine that applies a desired pattern onto a target portion of a substrate. The lithographic apparatus can be used, for example, in the manufacture of integrated circuits (ICs). For example, an IC chip in a smart phone can be as small as a person’s thumbnail, and may include over 2 billion transistors. Making an IC is a complex and time-consuming process, with circuit components in different layers and including hundreds of individual steps. Errors in even one step have the potential to result in problems with the final IC and can cause device failure. High process yield and high wafer throughput can be impacted by the presence of defects.BRIEF SUMMARY
[0004] In some embodiments, the techniques described herein relate to a method of generating a mask pattern using a machine learning model, the method including: providing a target pattern to a machine learning (ML) model; and executing the ML model to generate an output having multiple channels of information associated with a mask pattern corresponding to the target pattern, wherein the multiple channels include: a first channel that is configured to generate first mask information corresponding to main features to be printed on a substrate, and a second channel that is configured to generate second mask information corresponding to sub-resolution assist features (SRAFs), and wherein the ML model is trained by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the main features and an SRAF of the SRAFs that is adjacent to the main feature.
[0005] In some embodiments, the techniques described herein relate to a method of generating a mask pattern using a machine learning model, the method including: providing a training dataset to a machine learning (ML) model, wherein the training dataset includes a target pattern and a set of ground truth including (a) mask information corresponding to main features of a mask pattern corresponding to the target pattern, and (b) mask information corresponding to sub-resolution assist features (SRAFs) of the mask pattern; and training the ML model with the training dataset to generate multiple channels of information associated with the mask pattern, wherein the multiple channels include: a first channel thatis configured to generate predicted main features to be printed on a substrate, and a second channel that is configured to generate predicted SRAFs, and wherein the ML model is trained by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the predicted main features and an SRAF of the predicted SRAFs that is adjacent to the main feature.
[0006] In some embodiments, there is provided a non-transitory computer readable medium having instructions that, when executed by a computer, cause the computer to execute a method of any of the above embodiments.
[0007] In some embodiments, there is provided an apparatus includes a memory storing a set of instructions and a processor configured to execute the set of instructions to cause the apparatus to perform a method of any of the above embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Embodiments will now be described, by way of example only, with reference to the accompanying drawings in which:
[0009] Figure 1 illustrates a block diagram of various subsystems of a lithographic projection apparatus, according to an embodiment.
[0010] Figure 2 is a schematic diagram of a lithographic projection apparatus, according to an embodiment.
[0011] Figure 3 illustrates an exemplary flow chart for simulating lithography in a lithographic projection apparatus, according to an embodiment.
[0012] Figure 4 is a block diagram for generating a mask pattern as multiple channels of information, consistent with various embodiments.
[0013] Figure 5 is a block diagram for verifying a violation of a specified distance to be maintained between a main feature and an SRAF in a predicted output of the mask prediction model, consistent with various embodiments.
[0014] Figure 6 is a block diagram illustrating training of a mask prediction model to generate a mask pattern as multiple channels of information, consistent with various embodiments.
[0015] Figure 7 illustrates ground truth generation for training a mask prediction model to generate a mask pattern as multiple channels of information, consistent with various embodiments.
[0016] Figures 8A and 8B illustrate computing a cost function to penalize a violation of a specified relationship between multiple channels of information output by a mask prediction model, consistent with various embodiments.
[0017] Figure 9 is a flow diagram of a method for generating a mask pattern as multiple channels of information, consistent with various embodiments.
[0018] Figure 10 is a block diagram that illustrates a computer system which can assist in implementing various methods and systems disclosed herein.
[0019] Embodiments will now be described in detail with reference to the drawings, which are provided as illustrative examples so as to enable those skilled in the art to practice the embodiments. Notably, the figures and examples below are not meant to limit the scope to a single embodiment, but other embodiments are possible by way of interchange of some or all of the described or illustrated elements. Wherever convenient, the same reference numbers will be used throughout the drawings to refer to same or like parts. Where certain elements of these embodiments can be partially or fully implemented using known components, only those portions of such known components that are necessary for an understanding of the embodiments will be described, and detailed descriptions of other portions of such known components will be omitted so as not to obscure the description of the embodiments. In the present specification, an embodiment showing a singular component should not be considered limiting; rather, the scope is intended to encompass other embodiments including a plurality of the same component, and vice-versa, unless explicitly stated otherwise herein. Moreover, applicants do not intend for any term in the specification or claims to be ascribed an uncommon or special meaning unless explicitly set forth as such. Further, the scope encompasses present and future known equivalents to the components referred to herein by way of illustration.DETAILED DESCRIPTION
[0020] A lithographic apparatus is a machine that applies a designed pattern onto a target portion of a substrate. This process of transferring the designed pattern to the substrate is called a patterning process. The patterning process can include a patterning step to transfer a pattern from a patterning device (such as a mask) to the substrate. A mask pattern typically includes several types of features. For example, a mask pattern includes “main” features of a target pattern, which are features to be printed on a substrate, and sub-resolution assist features (SRAFs), which are features that do not print on the substrate but aids in printing of the main features. Various methods are used to generate a mask pattern. For example, a machine learning (ML) model is used to generate a mask pattern for a given target pattern. Typically, the ML model is trained to predict a mask image having both the main features and the SRAFs, that is, all the features of the mask pattern are predicted in a single image (or other format) in which all the features are combined. Such conventional methods have drawbacks or limitations. For example, the main features and the SRAFs may be undesirably merged in the predicted mask pattern which may result in a pattern printed on the substrate being defective. Additionally, some post processing (e.g., modification or further adjustment of a particular type of feature) may have to be performed on a particular type of feature of the predicted mask pattern, and such post processing may require the particular type of feature to be isolated from the other features for performing the post processing, and the isolation process may not only be time consuming or compute-intensive but also challenging. Further, performing post processing may be problematic if the different types of features have merged. These and other drawbacks exist.
[0021] Disclosed are embodiments for predicting a mask pattern as multiple channels of information, where each channel generates different types of features of the mask pattern. In some embodiments, a mask prediction model may be trained to predict main features of a mask pattern as a first channel of information, and predict the SRAFs of the mask pattern as a separate second channel of information. Information from each of the channels may be accessed and then processed independently. In some embodiments, each channel may output the information as an image (e.g., image of the main features and SRAFs). Various cost terms may be used in a cost function for training the mask prediction model. In some embodiments, a cost term may be defined to control a prescribed relationship between the channels of information. In some embodiments, a cost term may be defined to maintain a specified distance between a main feature and an SRAF and any violation of that distance may be penalized. Another cost term that is indicative of data fidelity of each channel may be defined. For example, a cost term may be defined to compute a difference between predicted data (e.g., predicted main features) and ground truth channel data (e.g., ground truth main features). Another cost term that is indicative of data fidelity of aggregated output of some or all of the channels may be defined. For example, a cost term may be defined to compute a difference between aggregated data (e.g., aggregated mask pattern generated by aggregating the main features and the SRAFs from the respective channels) and ground truth data (e.g., ground truth mask pattern). The mask prediction model may be iteratively trained to optimize (e.g., reduce or minimize) the cost function.
[0022] While the above paragraphs describe predicting a mask pattern as two channels of information, where a first channel has the main features and a second channel has the SRAFs, in other embodiments, the mask pattern may be generated as more than two channels of information. For example, the mask prediction model may be trained to output a first channel having information regarding the main features, a second channel having information regarding SRAFs of a first size and a third channel having information regarding SRAFs of a second size, etc. While the above describes generating main features and SRAFs of a mask pattern as separate channels of information, the features of a mask pattern may be categorized or classified to different types or sets of features (e.g., according to a user-defined classification) and each type or set of features may be generated as a separate channel of information. In some embodiments, the main features may be classified into different types of features based on their geometry, size, density in the target pattern, or some other parameter, and each type of main feature may be generated as a separate channel of information. For example, groups of main features with a first density may be classified as a first feature type and groups of main features with a second density may be categorized as a second feature type and so on. In some embodiment, SRAFs may be classified into different types of features based on their geometry, size, distance from main feature, their impact on process window conditions of a lithography process, or other parameters. For example, SRAFs with a first minimum distance from the main features may be classified as a first feature type and SRAFs with a second minimum distance from the main features may be categorized as a second feature type and so on.
[0023] In some embodiments, training a mask prediction model to generate a mask pattern as separate channels of information by using a cost function to maintain relationships between channels of information has various benefits. For example, by using a cost function to penalize any overlap, or violation of a specified distance, between the main features and the SRAFs, undesirable merging of the main features with the SRAFs may be prevented. Additionally, by generating different types of features of the mask pattern as different channels of information, the post processing of any type of feature in the predicted mask pattern may be simplified as the different types of features are isolated, thereby reducing the time and computing resource that may otherwise have been consumed in isolating the features for performing the post processing method. For example, the conventional prediction method, which predicts the SRAFs and main features as combined information (e.g., as a single image), poses a challenge of identifying what feature is an SRAF or what feature is main feature. In contrast, since the disclosed embodiments generates different types of features as separate channels of information and each feature type may be accessed individually and independently, the challenge in identifying what feature is what type is eliminated. Further, by generating different types of features of the mask pattern as different channels of information, the problem of undesirable merging of different types of features, which prevents the application of any post processing method, is eliminated.
[0024] In the present disclosure, although specific reference may be made to the manufacture of ICs, it should be explicitly understood that the description herein has many other possible applications. For example, it may be employed in the manufacture of integrated optical systems, guidance and detection patterns for magnetic domain memories, liquid crystal display panels, thin film magnetic heads, etc. The skilled artisan will appreciate that, in the context of such alternative applications, any use of the terms “reticle”, “wafer” or “die” in this text should be considered as interchangeable with the more general terms “mask”, “substrate” and “target portion”, respectively.
[0025] In the present document, the terms “radiation” and “beam” are used to encompass all types of electromagnetic radiation, including ultraviolet radiation (e.g., with a wavelength of 365, 248, 193, 157 or 126 nm) and EUV (extreme ultra-violet radiation, e.g., having a wavelength in the range of about 5- 100 nm). In the present document, the term “radiation source” or “source” is used to encompass all types of sources of radiation, including laser sources, incandescent sources, etc. which may include treatment of the radiation between the radiation source and the target or other parts of the optics, including filtering, collimating, focusing, etc.
[0026] A patterning device can comprise, or can form, one or more design layouts. The design layout can be generated utilizing CAD (computer-aided design) programs. This process is often referred to as EDA (electronic design automation). Most CAD programs follow a set of predetermined design rules in order to create functional design layouts / patterning devices. These rules are set based processing and design limitations. For example, design rules define the space tolerance between devices (such as gates, capacitors, etc.) or interconnect lines, to ensure that the devices or lines do not interact with one another in an undesirable way. One or more of the design rule limitations may be referred to as a “criticaldimension” (CD). A critical dimension of a device can be defined as the smallest width of a line or hole, or the smallest space between two lines or two holes. Thus, the CD regulates the overall size and density of the designed device. One of the goals in device fabrication is to faithfully reproduce the original design intent on the substrate (via the patterning device).
[0027] The term “mask” or “patterning device” as employed in this text may be broadly interpreted as referring to a generic patterning device that can be used to endow an incoming radiation beam with a patterned cross-section, corresponding to a pattern that is to be created in a target portion of the substrate. The term “light valve” can also be used in this context. Besides the classic mask (transmissive or reflective; binary, phase-shifting, hybrid, etc.), examples of other such patterning devices include a programmable mirror array. An example of such a device is a matrix-addressable surface having a viscoelastic control layer and a reflective surface. The basic principle behind such an apparatus is that (for example) addressed areas of the reflective surface reflect incident radiation as diffracted radiation, whereas unaddressed areas reflect incident radiation as undiffracted radiation. Using an appropriate filter, the said undiffracted radiation can be filtered out of the reflected beam, leaving only the diffracted radiation behind; in this manner, the beam becomes patterned according to the addressing pattern of the matrix-addressable surface. The required matrix addressing can be performed using suitable electronic means. Examples of other such patterning devices also include a programmable LCD array. An example of such a construction is given in U.S. Patent No. 5,229,872, which is incorporated herein by reference.
[0028] The term “projection optics” as used herein should be broadly interpreted as encompassing various types of optical systems, including refractive optics, reflective optics, apertures and catadioptric optics, for example. The term “projection optics” may also include components operating according to any of these design types for directing, shaping or controlling the projection beam of radiation, collectively or singularly. The term “projection optics” may include any optical component in the lithographic projection apparatus, no matter where the optical component is located on an optical path of the lithographic projection apparatus. Projection optics may include optical components for shaping, adjusting and / or projecting radiation from the source before the radiation passes the patterning device, and / or optical components for shaping, adjusting and / or projecting the radiation after the radiation passes the patterning device. The projection optics generally exclude the source and the patterning device.
[0029] Figure 1 illustrates a block diagram of various subsystems of a lithographic projection apparatus 10A, according to an embodiment. Major components are a radiation source 12A, which may be a deepultraviolet excimer laser source or other type of source including an extreme ultra violet (EUV) source (the lithographic projection apparatus itself need not have the radiation source), illumination optics which, e.g., define the partial coherence (denoted as sigma) and which may include optics 14A, 16Aa and 16Ab that shape radiation from the source 12A; a patterning device (or mask) 18A; and transmission optics 16Ac that project an image of the patterning device pattern onto a substrate plane 22A.
[0030] A pupil 20A can be included with transmission optics 16Ac. In some embodiments, there canbe one or more pupils before and / or after mask 18 A. As described in further detail herein, pupil 20A can provide patterning of the light that ultimately reaches substrate plane 22A. An adjustable filter or aperture at the pupil plane of the projection optics may restrict the range of beam angles that impinge on the substrate plane 22A, where the largest possible angle defines the numerical aperture of the projection optics NA= n sin(0max), wherein n is the refractive index of the media between the substrate and the last element of the projection optics, and ©max is the largest angle of the beam exiting from the projection optics that can still impinge on the substrate plane 22A.
[0031] In a lithographic projection apparatus, a source provides illumination (i.e., radiation) to a patterning device and projection optics direct and shape the illumination, via the patterning device, onto a substrate. This is not to disclaim that the source does not itself provide patterning, directing, or shaping to the radiation or that patterning, directing, or shaping does not occur between the source and the projection optics. The projection optics may include at least some of the components 14A, 16Aa, 16Ab and 16Ac. An aerial image (Al) is the radiation intensity distribution at substrate level. A resist model can be used to calculate the resist image from the aerial image, an example of which can be found in U.S. Patent Application Publication No. US 2009-0157360, the disclosure of which is hereby incorporated by reference in its entirety. The resist model is related to properties of the resist layer (e.g., effects of chemical processes which occur during exposure, post-exposure bake (PEB) and development). Optical properties of the lithographic projection apparatus (e.g., properties of the illumination, the patterning device and the projection optics) dictate the aerial image and can be defined in an optical model. Since the patterning device used in the lithographic projection apparatus can be changed, it is desirable to separate the optical properties of the patterning device from the optical properties of the rest of the lithographic projection apparatus including at least the source and the projection optics. Details of techniques and models used to transform a design layout into various lithographic images (e.g., an aerial image, a resist image, etc.), apply optical proximity correction (OPC) using those techniques and models and evaluate performance (e.g., in terms of process window) are described in U.S. Patent Application Publication Nos. US 2008-0301620, 2007-0050749, 2007- 0031745, 2008-0309897, 2010-0162197, and 2010-0180251, the disclosure of each which is hereby incorporated by reference in its entirety.
[0032] One aspect of understanding a lithographic process is understanding the interaction of the radiation and the patterning device. The electromagnetic field of the radiation after the radiation passes the patterning device may be determined from the electromagnetic field of the radiation before the radiation reaches the patterning device and a function that characterizes the interaction. This function may be referred to as the mask transmission function (which can be used to describe the interaction by a transmissive patterning device and / or a reflective patterning device).
[0033] The mask transmission function may have a variety of different forms. One form is binary. A binary mask transmission function has either of two values (e.g., zero and a positive constant) at any given location on the patterning device. A mask transmission function in the binary form may bereferred to as a binary mask. Another form is continuous. Namely, the modulus of the transmittance (or reflectance) of the patterning device is a continuous function of the location on the patterning device. The phase of the transmittance (or reflectance) may also be a continuous function of the location on the patterning device. A mask transmission function in the continuous form may be referred to as a continuous tone mask or a continuous transmission mask (CTM). For example, the CTM may be represented as a pixelated image, where each pixel may be assigned a value between 0 and 1 (e.g., 0.1, 0.2, 0.3, etc.) instead of binary value of either 0 or 1. In an embodiment, CTM may be a pixelated gray scale image, where each pixel has values (e.g., within a range [-255, 255], normalized values within a range [0, 1] or [-1, 1] or other appropriate ranges).
[0034] The thin-mask approximation, also called the Kirchhoff boundary condition, is widely used to simplify the determination of the interaction of the radiation and the patterning device. The thin-mask approximation assumes that the thickness of the structures on the patterning device is very small compared with the wavelength and that the widths of the structures on the mask are very large compared with the wavelength. Therefore, the thin-mask approximation assumes the electromagnetic field after the patterning device is the multiplication of the incident electromagnetic field with the mask transmission function. However, as lithographic processes use radiation of shorter and shorter wavelengths, and the structures on the patterning device become smaller and smaller, the assumption of the thin-mask approximation can break down. For example, interaction of the radiation with the structures (e.g., edges between the top surface and a sidewall) because of their finite thicknesses (“mask 3D effect” or “M3D”) may become significant. Encompassing this scattering in the mask transmission function may enable the mask transmission function to better capture the interaction of the radiation with the patterning device. A mask transmission function under the thin-mask approximation may be referred to as a thin-mask transmission function. A mask transmission function encompassing M3D may be referred to as a M3D mask transmission function.
[0035] Figure 2 schematically depicts an exemplary lithographic projection apparatus whose illumination source could be optimized utilizing the methods described herein. The apparatus comprises:- an illumination system IL, to condition a beam B of radiation. In this particular case, the illumination system also comprises a radiation source SO;- a first object table (e.g., mask table, patterning device table or reticle stage) MT provided with a patterning device holder to hold a patterning device MA (e.g., a reticle), and connected to a first positioner to accurately position the patterning device with respect to item PS;- a second object table (substrate table or wafer stage) WT provided with a substrate holder to hold a substrate W (e.g., a resist-coated silicon wafer), and connected to a second positioner to accurately position the substrate with respect to item PS;- a projection system (“lens”) PS (e.g., a refractive, catoptric or catadioptric optical system) to image an irradiated portion of the patterning device MA onto a target portion C (e.g., comprising one or more dies) of the substrate W.
[0036] As depicted herein, the apparatus is of a transmissive type (i.e., has a transmissive mask). However, in general, it may also be of a reflective type, for example (with a reflective mask). Alternatively, the apparatus may employ another kind of patterning device as an alternative to the use of a classic mask; examples include a programmable mirror array or LCD matrix.
[0037] The source SO (e.g., a mercury lamp or excimer laser) produces a beam of radiation. This beam is fed into an illumination system (illuminator) IL, either directly or after having traversed conditioning means, such as a beam expander Ex, for example. The illuminator IL may comprise adjusting means AD for setting the outer or inner radial extent (commonly referred to as o-outer and o-inner, respectively) of the intensity distribution in the beam. In addition, it will generally comprise various other components, such as an integrator IN and a condenser CO. In this way, the beam B impinging on the patterning device MA has a desired uniformity and intensity distribution in its cross-section.
[0038] It should be noted with regard to Figure 2 that the source SO may be within the housing of the lithographic projection apparatus (as is often the case when the source SO is a mercury lamp, for example), but that it may also be remote from the lithographic projection apparatus, the radiation beam that it produces being led into the apparatus (e.g., with the aid of suitable directing mirrors); this latter scenario is often the case when the source SO is an excimer laser (e.g., based on KrF, ArF or F2 lasing).
[0039] The beam B subsequently intercepts the patterning device MA, which is held on a patterning device table MT. Having traversed the patterning device MA, the beam B passes through the lens PS, which focuses the beam B onto a target portion C of the substrate W. With the aid of the second positioning means (and interferometric measuring means IF), the substrate table WT can be moved accurately, e.g., so as to position different target portions C in the path of the beam B. Similarly, the first positioning means can be used to accurately position the patterning device MA with respect to the path of the beam B, e.g., after mechanical retrieval of the patterning device MA from a patterning device library, or during a scan. In general, movement of the object tables MT, WT will be realized with the aid of a long-stroke module (coarse positioning) and a short-stroke module (fine positioning), which are not explicitly depicted in Figure 2. However, in the case of a wafer stepper (as opposed to a step- and-scan tool) the patterning device table MT may just be connected to a short stroke actuator, or may be fixed.
[0040] The depicted tool can be used in two different modes:- In step mode, the patterning device table MT is kept essentially stationary, and an entire patterning device image is projected in one go (i.e., a single “flash”) onto a target portion C. The substrate table WT is then shifted in the x or y directions so that a different target portion C can be irradiated by the beam B;- In scan mode, essentially the same scenario applies, except that a given target portion C is not exposed in a single “flash”. Instead, the patterning device table MT is movable in a given direction (the so-called “scan direction”, e.g., the y direction) with a speed v, so that the projection beam B is caused to scan over a patterning device image; concurrently, the substrate table WT is simultaneously movedin the same or opposite direction at a speed V = Mv, in which M is the magnification of the lens PS (typically, M = 1 / 4 or 1 / 5). In this manner, a relatively large target portion C can be exposed, without having to compromise on resolution.
[0041] Figure 3 illustrates an exemplary flow chart for simulating lithography in a lithographic projection apparatus, according to an embodiment. As will be appreciated, the models may represent a different patterning process and need not comprise all the models described below. A source model 300 represents optical characteristics (including radiation intensity distribution, bandwidth and / or phase distribution) of the illumination of a patterning device. The source model 300 can represent the optical characteristics of the illumination that include, but not limited to, numerical aperture settings, illumination sigma (o) settings as well as any particular illumination shape (e.g., off-axis radiation shape such as annular, quadrupole, dipole, etc.), where o (or sigma) is outer radial extent of the illuminator.
[0042] A projection optics model 310 represents optical characteristics (including changes to the radiation intensity distribution and / or the phase distribution caused by the projection optics) of the projection optics. The projection optics model 310 can represent the optical characteristics of the projection optics, including aberration, distortion, one or more refractive indexes, one or more physical sizes, one or more physical dimensions, etc.
[0043] The patterning device / design layout model module 320 captures how the design features are laid out in the pattern of the patterning device and may include a representation of detailed physical properties of the patterning device, as described, for example, in U.S. Patent No. 7,587,704, which is incorporated by reference in its entirety. In an embodiment, the patterning device / design layout model module 320 represents optical characteristics (including changes to the radiation intensity distribution and / or the phase distribution caused by a given design layout) of a design layout (e.g., a device design layout corresponding to a feature of an integrated circuit, a memory, an electronic device, etc.), which is the representation of an arrangement of features on or formed by the patterning device. Since the patterning device used in the lithographic projection apparatus can be changed, it is desirable to separate the optical properties of the patterning device from the optical properties of the rest of the lithographic projection apparatus including at least the illumination and the projection optics. The objective of the simulation is often to accurately predict, for example, edge placements and CDs, which can then be compared against the device design. The device design is generally defined as the pre-OPC patterning device layout, and will be provided in a standardized digital file format such as GDSII or OASIS.
[0044] An aerial image 330 can be simulated from the source model 300, the projection optics model 310 and the patterning device / design layout model module 320. An aerial image (Al) is the radiation intensity distribution at substrate level. Optical properties of the lithographic projection apparatus (e.g., properties of the illumination, the patterning device, and the projection optics) dictate the aerial image.
[0045] A resist layer on a substrate is exposed by the aerial image and the aerial image is transferred to the resist layer as a latent “resist image” (RI) therein. The resist image (RI) can be defined as a spatial distribution of solubility of the resist in the resist layer. A resist image 350 can be simulated from theaerial image 330 using a resist model 340. The resist model can be used to calculate the resist image from the aerial image, an example of which can be found in U.S. Patent Application No. 8,200,468, the disclosure of which is hereby incorporated by reference in its entirety. The resist model 340 typically describes the effects of chemical processes which occur during resist exposure, post exposure bake (PEB) and development, in order to predict, for example, contours of resist features formed on the substrate and so it typically related only to such properties of the resist layer (e.g., effects of chemical processes which occur during exposure, post-exposure bake and development). In an embodiment, the optical properties of the resist layer, e.g., refractive index, film thickness, propagation, and polarization effects — may be captured as part of the projection optics model 310.
[0046] So, in general, the connection between the optical and the resist model is a simulated aerial image intensity within the resist layer, which arises from the projection of radiation onto the substrate, refraction at the resist interface and multiple reflections in the resist film stack. The radiation intensity distribution (aerial image intensity) is turned into a latent “resist image” by absorption of incident energy, which is further modified by diffusion processes and various loading effects. Efficient simulation methods that are fast enough for full-chip applications approximate the realistic 3- dimensional intensity distribution in the resist stack by a 3-dimensional aerial (and resist) image.
[0047] In an embodiment, the resist image 350 can be used as an input to a post-pattern transfer process model module 360. The post-pattern transfer process model module 360 defines performance of one or more post-resist development processes (e.g., etch, development, etc.).
[0048] Simulation of the patterning process can, for example, predict contours, CDs, edge placement (e.g., edge placement error), etc. in the resist and / or etched image. Thus, the objective of the simulation is to accurately predict, for example, edge placement, and / or aerial image intensity slope, and / or CD, etc. of the printed pattern. These values can be compared against an intended design to, e.g., correct the patterning process, identify where a defect is predicted to occur, etc. The intended design is generally defined as a pre-OPC design layout which can be provided in a standardized digital file format such as GDSII or OASIS or other file format.
[0049] Thus, the model formulation describes most, if not all, of the known physics and chemistry of the overall process, and each of the model parameters desirably corresponds to a distinct physical or chemical effect. The model formulation thus sets an upper bound on how well the model can be used to simulate the overall manufacturing process.
[0050] The following paragraphs describe a system and a method for predicting a mask pattern as multiple channels of information, where each channel generates different types of features of the mask pattern. For example, a mask prediction model may be trained to predict main features of a mask pattern as a first channel of information and the SRAFs of the mask pattern as a second channel of information. The mask prediction model may be trained using a cost function that may be used to penalize any violation of a specified relationship to be maintained between features across the channels. For example, the mask prediction model may be trained using a cost function that may be used to penalize any overlap,or a violation of a specified distance (e.g., minimum distance) to be maintained between, a predicted main feature and SRAF.
[0051] Figure 4 is a block diagram for generating a mask pattern as multiple channels of information, consistent with various embodiments. A target pattern 405 is input to a mask prediction model 450 that is trained to generate a mask pattern as multiple channels of information. The mask prediction model 450 generates an output having multiple channels of information. For example, the mask prediction model 450 generates a first channel of information 408 including mask information corresponding to main features 410 of a mask pattern, and a second channel of information 409 including mask information corresponding to SRAFs 415 of the mask pattern. In some embodiments, the predicted main features 410 are post-optical proximity correction (OPC) features. The predicted SRAFs, while they are not printed on the substrate, aid in the printing of the main features. The predicted mask pattern information in each of the channels may be output in various formats. In some embodiments, the mask information generated in each of the channels may be output as an image. For example, the first channel 408 outputs the main features 410 as an image of the main features 410 and the second channel 409 outputs the SRAFs 415 as an image of the SRAFs 415.
[0052] The two channels of information may be aggregated to obtain a mask pattern 420 corresponding to the target pattern 405. The aggregation may be performed in any of a known number of ways. In some embodiments, the aggregation may include aggregating pixel values of the images output by the individual channels (e.g., by aggregating pixel values of the image of the main features 410 and the image of the SRAFs 415). In some embodiments, the aggregation may include obtaining an average of the pixel values of the corresponding locations in the images output by the channels. For example, the pixel value of a location (x,y) in an aggregated image is an average of the pixel values in the location (x,y) of the images output by the channels. In some embodiments, the aggregation may include obtaining a maximum of the pixel values of the corresponding locations in the images output by the channels (e.g., assuming a presence of feature has a higher pixel value than an absence of a feature). For example, the pixel value of a location (x,y) in an aggregated image is a max(pixel value (x,y) of each image). The different channels of information may be output as separate outputs of the mask prediction model 450 or as different channels within a single output of the mask prediction model 450. For example, the first channel 408 containing the main features 410 and the second channel 409 containing the SRAFs 415 may be output as two separate outputs of the mask prediction model 450, or as two separate channels within a single output of the mask prediction model 450. Regardless of how the channels of information are output, each channel of information may be accessed individually and independently. The mask prediction model may be implemented using any of a number of ML models (e.g., convolutional neural network (CNN), deep CNN (DCNN), U-NET, etc.).
[0053] In some embodiments, the mask prediction model 450 is trained to maintain a specified relationship between any two channels of information. For example, the mask prediction model 450 is trained to maintain a specified distance between a main feature generated in the first channel 408 andan SRAF generated in the second channel 409. A cost function may be used to determine if there is an overlap, or a violation of the specified distance, between the main feature and the SRAF and such an overlap or violation may be penalized. While the details of computing the cost function and training of the mask prediction model 450 using the cost function are described with respect to Figures 6 and 8A- 8B below, Figure 5 shows checking the predicted output (e.g., main features 410 and the SRAF 415) to verify if there is any violation of the specified distance (e.g., minimum distance) between a main feature and an SRAF in the predicted output.
[0054] While the above paragraphs describe predicting a mask pattern as two channels of information, where a first channel has the main features of the mask pattern and a second channel has the SRAFs of the mask pattern, in other embodiments, the mask pattern may be generated as more than two channels of information. For example, the mask prediction model 450 may be trained to output a first channel having information regarding the main features, a second channel having information regarding SRAFs of a first size and a third channel having information regarding SRAFs of a second size, etc. The features of a mask pattern may be categorized or classified to different types or sets of features (e.g., according to a user-defined classification) and each type or set of features may be generated as a separate channel of information.
[0055] Figure 5 is a block diagram for verifying a violation of a specified distance to be maintained between a main feature and an SRAF in a predicted output of the mask prediction model, consistent with various embodiments. If the distance 535 between a main feature 531 and an SRAF 532 that is adjacent to the main feature 531 in the predicted output (e.g., main features 410 and the SRAF 415) is less than a specified distance to be maintained between the two types of features (e.g., minimum distance between a main feature and an SRAF), then the specified distance is violated. Such a violation may lead to unfavorable consequences (e.g., resulting in a mask pattern design that is difficult to manufacture, or a pattern printed on the substrate using the main features 410 and the SRAFs 415 may be defective). The image 521 shows an enlarged version of a portion 421 of the mask pattern 420 containing the main feature 531 and the SRAF 532. While Figure 5 shows an aggregated mask pattern 420 being used to determine the distance violation for the sake of convenience, note that the distance violation may be determined using the two separate channels of information (e.g., the image of main features 410 and the image of SRAFs 415) without having to generate the aggregated mask pattern 420. The main feature image 541 shows an enlarged version of a portion (e.g., indicated by a rectangular dotted box) of the main features 410 containing the main feature 531. The SRAF image 542 shows an enlarged version of a portion (e.g., indicated by a rectangular dotted box) of the SRAFs 415 containing the SRAF 532.
[0056] In some embodiments, the features across the channels are positioned relative to each other. That is, the position of a feature in the output of the channel may be similar to the position of the feature in an aggregated output of all the channels. For example, if the channels generate an output in the form of an image, the position of a main feature in the image output by a first channel and the position of anSRAF in the image output by a second channel may be similar to their respective positions in an image of a predicted mask pattern generated by aggregating images output by all the channels. The positions of the features may be encoded in the pixel locations and therefore, there may be one to one correspondence on the position between the channels.
[0057] The distance violation between the main feature 531 and the SRAF 532 in the predicted output may be determined in a number of ways. In some embodiments, if there are no common pixels that are non-zero in a region around the features defined by a specified distance d, then it may be considered that there is no distance violation between the two features. For example, a first set of pixel values in a first region 533a proximate the main feature 531 may be obtained from the main feature image 541 generated in the first channel 408. The size of the first region 533a may be defined by the specified distance to be maintained between the two types of features. A second set of pixel values in a second region 533b corresponding to the first region 533a may be obtained from the SRAF image 542 generated in the second channel 409. If there are no non-zero pixel values in either the first region 533a or the second region 533b, then there may be no distance violation between the main feature 531 and the SRAF 532. If one or more pixel values from at least one of the first set of pixel values or the second set of pixel values are non-zero, then it may be considered that there is a distance violation between the main feature 531 and the SRAF 532. In some embodiments, if a distance violation is detected, the mask prediction model 450 may be retrained (e.g., trained with additional training datasets to improve the accuracy of prediction).
[0058] In some embodiments, a cost term may be used to quantify the distance violation between the main feature 531 and the SRAF 532. The cost term may determine an overlap by: (a) convolving a feature from each of the channels with a kernel (e.g., using one-dimensional or two-dimensional kernel) whose size is defined by the specified distance (e.g., minimum distance) to be maintained between the features from two or more channels to obtain convolved features, and (b) determining if there is any overlap between the convolved features from two or more channels (e.g., based on a product of the convolved features across the channels; if the product includes any non-zero values, then there may be an overlap between the features.) For example, the cost term, c, may be represented as:... Eq. (l) where KSRAFand KMFare averaging kernels for SRAF and main feature, respectively, and whose size is determined based on a specified distance to be maintained between the main feature and the SRAF, MMFis a signal representing pixel values of the main feature image in a first region where the distance violation is to be detected, MSRAFis a signal representing pixel values in the SRAF image in a regioncorresponding to the first region, ® is an element wise multiplication, * represents a convolution operation, and | • |Frepresents the Frobenius norm.
[0059] In some embodiments, the cost term may also be represented as:... Eq. (2) where | • ^ representsnorm.
[0060] In some embodiments, the greater the value of the cost term, the greater is the distance violation. Additional details with respect to the cost term is discussed at least with reference to Figures 6 and 8A and 8B. While Figure 5 shows verifying the distance violation in a single region near the main feature 531 and SRAF 532 (e.g., regions 533a and 533b, respectively) the distance violation may be checked for other regions proximate the main feature 531 and SRAF 532 (e.g., any region between the main feature 531 and the SRAF 532). Further, the distance violation verification may be performed for some or all of the main feature-SRAF pairs.
[0061] Figure 6 is a block diagram illustrating training of a mask prediction model to generate a mask pattern as multiple channels of information, consistent with various embodiments. The mask prediction model 450 is trained using a training dataset 602 that includes a target pattern 605 and a set of ground truth data including (a) a mask pattern 615 corresponding to the target pattern 605, (b) main features 610a of the mask pattern 615, and (c) SRAFs 610b of the mask pattern 615. The ground truth main features 610a and the SRAFs 610b may be generated in a number of ways. Figure 7 illustrates ground truth generation for training a mask prediction model to generate a mask pattern as multiple channels of information. In some embodiments, the ground truth main features 712 and SRAFs 714 of a mask pattern 710 may be generated by decomposing (e.g., using known methods) a ground truth mask pattern 710 corresponding to a target pattern 705 into one or more constituent features (e.g., main features 712 and SRAFs 714) of the mask pattern 710. The decomposing method may be configured to obtain any types of features (e.g., based on user-classification of features). The mask pattern 710 may be generated in a number of ways (e.g., via ME model prediction, inverse lithography method, or other such processes). In some embodiments, the mask pattern 710 is in the form of a continuous transmission mask (CTM). The target pattern 705 may be obtained in a number of ways (e.g., as an image from GDS file). In some embodiments, the target pattern 605 is similar to the target pattern 705, the mask pattern 615 is similar to the mask pattern 710, the main features 610a is similar to the main features 712, and the SRAFs 610b is similar to the SRAFs 714.
[0062] After receiving the training dataset 602, the mask prediction model 450 generates a predicted set of main features 625a and a predicted set of SRAFs 625b as two separate channels of information. For example, the first channel 621 is configured to generate the predicted set of main features 625a andthe second channel 622 is configured to generate a predicted set of SRAFs 625b. The different channels of information may be output as separate outputs of the mask prediction model 450, or as different channels within a single output of the mask prediction model 450. For example, the first channel 621 containing the predicted set of main features 625a and the second channel 622 containing the predicted set of SRAFs 625b may be output as two separate outputs of the mask prediction model 450, or as two separate channels within a single output. In the case where the two channels are output as a single output of the mask prediction model 450, the information from the two channels may be concatenated or otherwise combined using a known function, which may then be used to retrieve or extract the information of each of the channels from the output. Regardless of how the channels of information are output, each channel of information may be accessed separately.
[0063] The training process may continue with aggregating the information from multiple channels to generate an aggregated output. For example, an aggregator 650 may aggregate the predicted set of main features 625a with the predicted set of SRAFs 625b to generate an aggregated or a predicted mask pattern 625c. The aggregator 650 may aggregate the multiple channels of information using any of a number of known aggregation methods (e.g., as described at least with reference to Figure 4 above). For example, in a case where the predicted set of main features 625a and the predicted set of SRAFs 625b are generated as images, the aggregator 650 may generate the predicted mask pattern 625c by aggregating the corresponding pixel values from each of the images.
[0064] The mask prediction model 450 may be trained using a cost function that has one or more cost terms. In some embodiments, a first cost term that is indicative of data fidelity of each channel may be defined. The first cost term may be defined to compute a difference between predicted data (e.g., predicted main features) and ground truth channel data (e.g., ground truth main features), and the mask prediction model 450 may be trained to optimize (e.g., reduce or minimize) the first cost term. For example, the first cost term may be indicative of a difference between the ground truth main features 610a and the predicted set of main features 625a. The first cost term may also be indicative of a difference between the ground truth SRAFs 610b and the predicted set of SRAFs 625b.
[0065] In some embodiments, a second cost term that is indicative of data fidelity of an aggregated output of some or all of the channels may be defined. The second cost term may be defined to compute a difference between aggregated data (e.g., aggregated mask pattern generated by aggregating the main features and the SRAFs from the respective channels) and ground truth data corresponding to the aggregated output (e.g., ground truth mask pattern), and the mask prediction model 450 may be trained to optimize the second cost term. For example, the second cost term may be defined to compute a difference between the predicted mask pattern 625c and the ground truth mask pattern 615.
[0066] In some embodiments, a third cost term may be defined to maintain a relationship between the channels of information. For example, a third cost term may be defined to maintain a specified distance between a main feature and an SRAF of the predicted output and any violation of that distance may bepenalized. Figures 8A and 8B describe the details regarding computing the third cost term that is defined to maintain a specified distance between a main feature and an SRAF of the predicted output.
[0067] Figures 8A and 8B illustrate computing a cost term to penalize a violation of a specified relationship between multiple channels of information output by a mask prediction model, consistent with various embodiments. In some embodiments, the main feature image 808 shows an enlarged version of a portion of an image of the predicted set of main features 625a containing the main feature831. The SRAF image 828 shows an enlarged version of a portion of an image of the predicted set of SRAFs 625b containing the SRAF 832. The mask pattern image 838 shows an illustration of how a portion of the mask pattern 625c may look like when the images of the predicted main feature 831 and the predicted SRAF 832 are superimposed with distance 834 between the main feature 831 and the SRAF 832. As described at least with reference to Figure 5 above, the distance violation between the main feature 831 and the SRAF 832 may be determined in a number of ways. In some embodiments, if there are no common pixels that are non-zero in a region around the features defined by a specified distance d, then it may be considered that there is no distance violation between the two features. For example, a first set of pixel values in a first region 833a proximate the main feature 831 may be obtained from the main feature image 808. The size of the first region 833a may be defined by the specified distance to be maintained between the two types of features. A second set of pixel values in a second region 833b corresponding to the first region 833a may be obtained from the SRAF image 828. If there are no pixel values that are non-zero in either the first region 833a or the second region 833b, then there may be no distance violation between the main feature 831 and the SRAF 832. If one or more pixel values from at least one of the first set of pixel values or the second set of pixel values are non-zero, then it may be considered that there is a distance violation between the main feature 831 and the SRAF832.
[0068] The third cost term, such as the cost term of Eqs. (1) or (2) may be used to quantify the distance violation between the predicted main feature 831 and the SRAF 832, and the distance violation may be reduced by optimizing the third cost term. The third cost term may be computed as follows. A onedimensional signal, such as a first signal 802, MMF, having a first set of pixel values of a cross section of the main feature image 808 is obtained. The first peak 804 is indicative of non-zero pixel values, which is indicative of a presence of the main feature 831 in the corresponding location of the image. Another one-dimensional signal, such as a second signal 822, MSRAF, having a second set of pixel values of a cross section of the SRAF image 828 is obtained. A second peak 824 is indicative of nonzero pixel values, which is indicative of a presence of the SRAF 832 in the corresponding location of the image. The distance between two peaks 804 and 824 (e.g., between ending of the peak 804 and beginning of the peak 824) may be equal to or greater than the specified distance to be maintained between the main feature 831 and the SRAF 832 for avoiding any distance violation (e.g., undesired merging of the main feature and the SRAF). A kernel 852 having a size 854 defined based on thespecified distance to be maintained between the main feature 831 and the SRAF 832 may be used in computing the third cost term. For example, KSRAFand KMFare averaging kernels for the SRAF 832 and the main feature 831, respectively, and KSRAF= KMF. A first convolution output 842, KMF* MMF^, may be obtained by convolving the first signal 802 with the kernel 852. A second convolution output 844, (KSRAF * MSRAF^, may be obtained by convolving the second signal 822 with the kernel 852. The output 862 of the third cost term may be obtained based on the first and second convolution outputs 842 and 844, respectively, for example, as an element wise multiplication the first and second convolution outputs 842 and 844, respectively. If there is a violation of the specified distance, the distance violation is indicated by a peak 864, which is indicative of a presence of non-zero pixels either in the first region 833a of the main feature image 808 or the second region 833b of the SRAF image 828 causing the features to overlap or merge or be closer than they are defined to be. The location of the peak 864 may indicate the location of the violation in the image. The distance violation can be reduced by optimizing the third cost term.
[0069] In some embodiments, different shapes of kernels may be used to compute the third cost term. For example, as illustrated in Figure 8B, a convex shaped kernel 873 may be used to determine the third cost term. In some embodiments, the shape of the kernel may be chosen based on the amount of distance to be maintained between a predicted main feature and a predicted SRAF. Figure 8B also shows the first convolution output 843,and the second convolution output 853, KSRAF* MSRAF), the output 863 of the third cost term obtained using the kernel 873. The peak 865 of the output 863 is indicative of a distance violation causing the main feature and the SRAF to overlap.
[0070] While the above paragraphs describe convolution using a one-dimensional kernel, in some embodiments, two-dimensional kernels may also be used to convolve the image.
[0071] Referring to the training of the mask prediction model 450 in Figure 6, after one or more cost terms (e.g., the first, second or third cost terms) of the cost function are computed, the training may be continued by adjusting the parameters of the mask prediction model 450 (e.g., weights or biases) to optimize the cost function, and the mask prediction model 450 is executed / trained again with the training dataset (e.g., same target pattern and same set of ground truth, or other target patterns and corresponding set of ground truth data). The training process may be continued until a specified training condition is satisfied. For example, the training process may be continued for a specified number of iterations. In another example, the training process may be continued until the cost function is optimized. Once the training condition is satisfied, the mask prediction model 450 is considered to be trained, and the mask prediction model 450 may be deployed to generate a mask pattern as multiple channels of information for any new target pattern (e.g., unseen target pattern or a target pattern that the mask prediction model 450 is neither trained on, nor executed on), as illustrated in Figure 4 above.
[0072] In some embodiments, the mask prediction model 450 may be trained without any ground truth data. For example, a partially trained mask prediction model 450 (e.g., trained with a specified numberof datasets) may be trained without any ground truth data (e.g., without ground truth main features, SRAFs or entire mask pattern). In such a training, the accuracy or performance of the mask prediction model 450 in maintaining a specified relationship between predicted channels of information may be improved without the need for the generation of ground truth, thereby reducing the time and computing resources required for generating the ground truth. For example, since the determination of the third cost term (e.g., cost term of Eqs. (1) or (2)) does not depend on the ground truth data (e.g., ground truth main features, SRAFs or entire mask pattern), the mask prediction model 450 may be trained using a cost function that includes the third cost term but not the first and second cost terms. For example, in the training stage, when a target pattern is input to a partially trained mask prediction model 450, the mask prediction model 450 generates mask information corresponding to main features in a first channel and mask information corresponding to SRAFs in a second channel. Since no ground truth is input for the main features and the SRAFs, the cost function may be configured to ignore the first and second cost terms but compute the third cost term, which, as described above, determines a specified distance between the predicted main feature and SRAF. Any violation of the specified distance may be penalized, and the mask prediction model 450 may be iteratively trained (e.g., without the ground truth data) until the cost function is optimized (e.g., third cost term is optimized). In this way, the performance of the mask prediction model 450 in maintaining a specified relationship between predicted channels of information may be improved without the need for the generation of ground truth.
[0073] Figure 9 is a flow diagram of a method for generating a mask pattern as multiple channels of information, consistent with various embodiments. The method of Figure 9 is described at least with reference to Figure 4 above.
[0074] At process P902, a target pattern is provided as an input to a mask prediction model (e.g., an ML model) that is trained to generate a mask pattern as multiple channels of information. For example, a target pattern 405 is input to a mask prediction model 450.
[0075] At process P904, the mask prediction model is executed to generate an output having multiple channels of information associated with a mask pattern corresponding to the target pattern. Different channels are configured to output different types of feature of a mask pattern. The multiple channels include: (a) a first channel that is configured to generate first mask information corresponding to main features to be printed on a substrate, and (b) a second channel that is configured to generate second mask information corresponding to SRAFs. For example, the mask prediction model 450 generates a first channel of information 408 including mask information corresponding to main features 410 of a mask pattern, and a second channel of information 409 including mask information corresponding to SRAFs 415 of the mask pattern. The mask information in each of the channels may be output in various formats. In some embodiments, the mask information generated in each of the channels may be output as an image. For example, the first channel 408 outputs the main features 410 as an image of the main features 410 and the second channel 409 outputs the SRAFs 415 as an image of the SRAFs 415.
[0076] The two channels of information may be aggregated to obtain a mask pattern 420 corresponding to the target pattern 405. In some embodiments, the mask pattern 420 may be generated by aggregating pixel values of the images output by the individual channels (e.g., by aggregating pixel values of the image of the main features 410 and the image of the SRAFs 415). The mask prediction model may be implemented using any of a number of ML models (e.g., CNN, DCNN, U-NET, etc.).
[0077] In some embodiments, the mask prediction model 450 is trained to maintain a specified relationship between any two channels of information. For example, the mask prediction model 450 is trained to maintain a specified distance between a main feature generated in the first channel 408 and an SRAF generated in the second channel 409. A cost function may be used in training the mask prediction model 450 to penalize any overlap, or a violation of the specified distance, between the main feature and the SRAF. In the inference stage, the cost function may also be used for checking the predicted output (e.g., main features 410 and the SRAF 415) to verify if there is any overlap, or violation of the specified distance, between a main feature and an SRAF in the predicted output.
[0078] Figure 10 is a block diagram that illustrates a computer system 100 which can assist in implementing various methods and systems disclosed herein. The computer system 100 may be used to implement any of the entities, components, modules, or services depicted in the examples of the figures (and any other entities, components, modules, or services described in this specification). The computer system 100 may be programmed to execute computer program instructions to perform functions, methods, flows, or services (e.g., of any of the entities, components, or modules) described herein. The computer system 100 may be programmed to execute computer program instructions by at least one of software, hardware, or firmware.
[0079] Computer system 100 includes a bus 102 or other communication mechanism for communicating information, and a processor 104 (or multiple processors 104 and 105) coupled with bus 102 for processing information. Computer system 100 also includes a main memory 106, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 102 for storing information and instructions to be executed by processor 104. Main memory 106 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 104. Computer system 100 further includes a read only memory (ROM) 108 or other static storage device coupled to bus 102 for storing static information and instructions for processor 104. A storage device 110, such as a magnetic disk or optical disk, is provided and coupled to bus 102 for storing information and instructions.
[0080] Computer system 100 may be coupled via bus 102 to a display 112, such as a cathode ray tube (CRT) or flat panel or touch panel display for displaying information to a computer user. An input device 114, including alphanumeric and other keys, is coupled to bus 102 for communicating information and command selections to processor 104. Another type of user input device is cursor control 116, such as a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 104 and for controlling cursor movement on display112. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane. A touch panel (screen) display may also be used as an input device.
[0081] According to one embodiment, portions of one or more methods described herein may be performed by computer system 100 in response to processor 104 executing one or more sequences of one or more instructions contained in main memory 106. Such instructions may be read into main memory 106 from another computer-readable medium, such as storage device 110. Execution of the sequences of instructions contained in main memory 106 causes processor 104 to perform the process steps described herein. One or more processors in a multi-processing arrangement may also be employed to execute the sequences of instructions contained in main memory 106. In an alternative embodiment, hard-wired circuitry may be used in place of or in combination with software instructions. Thus, the description herein is not limited to any specific combination of hardware circuitry and software.
[0082] The term “computer-readable medium” as used herein refers to any medium that participates in providing instructions to processor 104 for execution. Such a medium may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 110. Volatile media include dynamic memory, such as main memory 106. Transmission media include coaxial cables, copper wire and fiber optics, including the wires that comprise bus 102. Transmission media can also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media include, for example, a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave as described hereinafter, or any other medium from which a computer can read.
[0083] Various forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to processor 104 for execution. For example, the instructions may initially be borne on a magnetic disk of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 100 can receive the data on the telephone line and use an infrared transmitter to convert the data to an infrared signal. An infrared detector coupled to bus 102 can receive the data carried in the infrared signal and place the data on bus 102. Bus 102 carries the data to main memory 106, from which processor 104 retrieves and executes the instructions. The instructions received by main memory 106 may optionally be stored on storage device 110 either before or after execution by processor 104.
[0084] Computer system 100 also preferably includes a communication interface 118 coupled to bus 102. Communication interface 118 provides a two-way data communication coupling to a network link 120 that is connected to a local network 122. For example, communication interface 118 may be anintegrated services digital network (ISDN) card or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 118 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 118 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0085] Network link 120 typically provides data communication through one or more networks to other data devices. For example, network link 120 may provide a connection through local network 122 to a host computer 124 or to data equipment operated by an Internet Service Provider (ISP) 126. ISP 126 in turn provides data communication services through the worldwide packet data communication network, now commonly referred to as the “Internet” 128. Local network 122 and Internet 128 both use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 120 and through communication interface 118, which carry the digital data to and from computer system 100, are exemplary forms of carrier waves transporting the information.
[0086] Computer system 100 can send messages and receive data, including program code, through the network(s), network link 120, and communication interface 118. In the Internet example, a server 130 might transmit a requested code for an application program through Internet 128, ISP 126, local network 122 and communication interface 118. One such downloaded application may provide for the illumination optimization of the embodiment, for example. The received code may be executed by processor 104 as it is received, or stored in storage device 110, or other non-volatile storage for later execution. In this manner, computer system 100 may obtain application code in the form of a carrier wave.
[0087] While the concepts disclosed herein may be used for imaging on a substrate such as a silicon wafer, it shall be understood that the disclosed concepts may be used with any type of lithographic imaging systems, e.g., those used for imaging on substrates other than silicon wafers.
[0088] The terms “optimizing” and “optimization” as used herein refers to or means adjusting a patterning apparatus (e.g., a lithography apparatus), a patterning process, etc. such that results and / or processes have more desirable characteristics, such as higher accuracy of projection of a design pattern on a substrate, a larger process window, etc. Thus, the term “optimizing” and “optimization” as used herein refers to or means a process that identifies one or more values for one or more parameters that provide an improvement, e.g., a local optimum, in at least one relevant metric, compared to an initial set of one or more values for those one or more parameters. "Optimum" and other related terms should be construed accordingly. In an embodiment, optimization steps can be applied iteratively to provide further improvements in one or more metrics.
[0089] Embodiments of the present disclosure can be further described by the following clauses.1. A method of generating a mask pattern using a machine learning model, the method comprising:providing a target pattern to a machine learning (ML) model; and executing the ML model to generate an output having multiple channels of information associated with a mask pattern corresponding to the target pattern, wherein the multiple channels include: a first channel that is configured to generate first mask information corresponding to main features to be printed on a substrate, and a second channel that is configured to generate second mask information corresponding to subresolution assist features (SRAFs), and wherein the ML model is trained by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the main features and an SRAF of the SRAFs that is adjacent to the main feature.2. The method of clause 1 , wherein the first mask information and the second mask information are images of the corresponding features.3. The method of clause 1, further comprising: determining a distance between a first main feature of the main features and a first SRAF of the SRAFs that is adjacent to the first main feature; and determining whether there is an overlap between the first main feature and the first SRAF based on the distance.4. The method of clause 1 , further comprising: training the ML model by: obtaining a first ground truth image of a set of main features and a second ground truth image of a set of SRAFs, determining, in the first ground truth image, a first set of pixel values in a first region proximate a first main feature of the set of main features, wherein the first region is defined by the specified distance, determining, in the second ground truth image, a second set of pixel values in a second region proximate a first SRAF of the set of SRAFs, wherein the second region corresponds to the first region, and determining the violation based on a determination of at least one of the first set of pixel values or the second set of pixel values being non-zero.5. The method of clause 4, wherein determining the violation includes: computing the cost function, wherein computing the cost function includes: defining a size of a kernel based on the specified distance, convolving each of the first main feature and the first SRAF with the kernel to obtain a first convolved feature and a second convolved feature, respectively, and determining the violation based on a product of the first convolved feature and the second convolved feature.6. The method of clause 1 , further comprising training the ML model by using a cost function that penalizes a violation of a specified relationship between two or more channels of the multiple channels.7. The method of clause 1 further comprising: obtaining a training dataset having a first target pattern and a first ground truth mask pattern,obtaining from the first ground truth mask pattern (a) a first image having a set of main features of the first ground truth mask pattern, and (b) a second image having a set of SRAFs of the first ground truth mask pattern, and training the ML model, with the first target pattern and using the first image and the second image as ground truth for a first channel and second channel of the multiple channels, respectively, to generate a predicted set of main features as the first channel and a predicted set of SRAFs as the second channel.8. The method of clause 7, further comprising training the ML model by using a cost function that penalizes a difference between the set of main features and the predicted set of main features, and a difference between the set of SRAFs and the predicted set of SRAFs.9. The method of clause 7, further comprising aggregating the first image having the set of main features with the second image having the set of SRAFs to generate an aggregated mask pattern.10. The method of clause 9, further comprising training the ML model by using a cost function that penalizes a difference between the aggregated mask pattern and the first ground truth mask pattern.11. The method of clause 7, further comprising training the ML model by using a cost function that includes multiple cost terms, wherein the multiple cost terms include:(i) a first cost term that penalizes a difference between (a) the set of main features and the predicted set of main features and (b) the set of SRAFs and the predicted set of SRAFs,(ii) a second cost term that penalizes a difference between an aggregated mask pattern and the first ground truth mask pattern, and(iii) a third cost term that penalizes a violation of a specified relationship between two or more channels of the multiple channels.12. The method of clause 1, further comprising: obtaining a training dataset having a second target pattern, and training the ML model with the second target pattern to generate a second predicted set of main features and a second predicted set of SRAFs as the first and second channels of the multiple channels, respectively.13. The method of clause 12, further comprising training the ML model by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the second predicted set of main features and an SRAF of the second predicted set of SRAFs that is adjacent to the main feature.14. A method of generating a mask pattern using a machine learning model, the method comprising: providing a training dataset to a machine learning (ML) model, wherein the training dataset includes a target pattern and a set of ground truth including (a) mask information corresponding to main features of a mask pattern corresponding to the target pattern, and (b) mask information corresponding to subresolution assist features (SRAFs) of the mask pattern; and training the ML model with the training dataset to generate multiple channels of information associated with the mask pattern, wherein the multiple channels include:a first channel that is configured to generate predicted main features to be printed on a substrate, and a second channel that is configured to generate predicted SRAFs, and wherein the ML model is trained by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the predicted main features and an SRAF of the predicted SRAFs that is adjacent to the main feature.15. The method of clause 14, further comprising training the ML model by: obtaining a first image of the predicted main features and a second image of the predicted SRAFs, determining, in the first image, a first set of pixel values in a first region proximate the main feature, wherein the first region is defined by the specified distance, determining, in the second image, a second set of pixel values in a second region proximate the SRAF, wherein the second region corresponds to the first region, and determining the violation based on a determination of at least one of the first set of pixel values or the second set of pixel values being non-zero.16. The method of clause 15, wherein determining the violation includes: computing the cost function, wherein computing the cost function includes: defining a size of a kernel based on the specified distance, convolving each of the main feature and the SRAF with the kernel to obtain a first convolved feature and a second convolved feature, respectively, and determining the violation based on a product of the first convolved feature and the second convolved feature.17. The method of clause 14, further comprising training the ML model by using a cost function that penalizes a violation of a specified relationship between two or more channels of the multiple channels.18. The method of clause 14, wherein the training includes: obtaining the mask pattern, and obtaining from the mask pattern (a) a first image having the main features of the mask pattern, and (b) a second image having the SRAFs of the mask pattern.19. The method of clause 14, further comprising training the ML model by using a cost function that penalizes a difference between the main features and the predicted main features, and a difference between the SRAFs and the predicted SRAFs.20. The method of clause 14 further comprising: aggregating a first image having the predicted main features with a second image having the predicted SRAFs to generate an aggregated mask pattern.21. The method of clause 20, further comprising training the ML model by using a cost function that penalizes the difference between the aggregated mask pattern and the mask pattern.22. The method of clause 14, further comprising training the ML model by using a cost function that includes multiple cost terms, wherein the multiple cost terms include:(i) a first cost term that penalizes a difference between (a) the main features and the predicted main features and (b) the SRAFs and the predicted SRAFs,(ii) a second cost term that penalizes a difference between an aggregated mask pattern and the mask pattern, and(iii) a third cost term that penalizes a violation of the specified distance to be maintained between the main feature and the SRAF of the predicted SRAFs.23. The method of clause 14, further comprising: obtaining a training dataset having a second target pattern, and training the ML model with the second target pattern to generate a second predicted set of main features and a second predicted set of SRAFs as the first and second channels of the multiple channels, respectively.24. The method of clause 23, further comprising training the ML model by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the second predicted set of main features and an SRAF of the second predicted set of SRAFs that is adjacent to the main feature.25. The method of clause 14, further comprising: providing a first target pattern to the ML model; and executing the ML model to generate an output having the multiple channels of information associated with a first mask pattern corresponding to the first target pattern, wherein the first channel includes a predicted set of main features to be printed on the substrate, and the second channel includes a predicted set of SRAFs.26. The method of clause 25, further comprising: determining a distance between a first main feature of the predicted set of main features and a first SRAF of the predicted set of SRAFs that is adjacent to the first main feature; and determining whether there is an overlap between the first main feature and the first SRAF based on the distance.27. An apparatus comprising: a memory storing a set of instructions; and a processor configured to execute the set of instructions to cause the apparatus to perform a method of any of the preceding clauses.28. A non-transitory computer-readable medium having instructions recorded thereon, the instructions when executed by a computer implementing the method of any of clauses 1-26.
[0090] Aspects of the invention can be implemented in any convenient form. For example, an embodiment may be implemented by one or more appropriate computer programs which may be carried on an appropriate carrier medium which may be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communications signal). Embodiments of the invention may be implemented using suitable apparatus which may specifically take the form of a programmable computer running acomputer program arranged to implement a method as described herein. Thus, embodiments of the disclosure may be implemented in hardware, firmware, software, or any combination thereof. Embodiments of the disclosure may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine -readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium may include read only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory devices; electrical, optical, acoustical, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.), and others. Further, firmware, software, routines, instructions may be described herein as performing certain actions. However, it should be appreciated that such descriptions are merely for convenience and that such actions in fact result from computing devices, processors, controllers, or other devices executing the firmware, software, routines, instructions, etc.
[0091] In block diagrams, illustrated components are depicted as discrete functional blocks, but embodiments are not limited to systems in which the functionality described herein is organized as illustrated. The functionality provided by each of the components may be provided by software or hardware modules that are differently organized than is presently depicted, for example such software or hardware may be intermingled, conjoined, replicated, broken up, distributed (e.g., within a data center or geographically), or otherwise differently organized. The functionality described herein may be provided by one or more processors of one or more computers executing code stored on a tangible, non- transitory, machine-readable medium. In some cases, third party content delivery networks may host some or all of the information conveyed over networks, in which case, to the extent information (e.g., content) is said to be supplied or otherwise provided, the information may be provided by sending instructions to retrieve that information from a content delivery network.
[0092] Unless specifically stated otherwise, as apparent from the discussion, it is appreciated that throughout this specification discussions utilizing terms such as “processing,” “computing,” “calculating,” “determining” or the like refer to actions or processes of a specific apparatus, such as a special purpose computer or a similar special purpose electronic processing / computing device.
[0093] The reader should appreciate that the present application describes several inventions. Rather than separating those inventions into multiple isolated patent applications, these inventions have been grouped into a single document because their related subject matter lends itself to economies in the application process. But the distinct advantages and aspects of such inventions should not be conflated. In some cases, embodiments address all of the deficiencies noted herein, but it should be understood that the inventions are independently useful, and some embodiments address only a subset of such problems or offer other, unmentioned benefits that will be apparent to those of skill in the art reviewing the present disclosure. Due to cost constraints, some inventions disclosed herein may not be presently claimed and may be claimed in later filings, such as continuation applications or by amending the present claims. Similarly, due to space constraints, neither the Abstract nor the Summary sections ofthe present document should be taken as containing a comprehensive listing of all such inventions or all aspects of such inventions.
[0094] It should be understood that the description and the drawings are not intended to limit the present disclosure to the particular form disclosed, but to the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the inventions as defined by the appended claims.
[0095] Modifications and alternative embodiments of various aspects of the inventions will be apparent to those skilled in the art in view of this description. Accordingly, this description and the drawings are to be construed as illustrative only and are for the purpose of teaching those skilled in the art the general manner of carrying out the inventions. It is to be understood that the forms of the inventions shown and described herein are to be taken as examples of embodiments. Elements and materials may be substituted for those illustrated and described herein, parts and processes may be reversed or omitted, certain features may be utilized independently, and embodiments or features of embodiments may be combined, all as would be apparent to one skilled in the art after having the benefit of this description. Changes may be made in the elements described herein without departing from the spirit and scope of the invention as described in the following claims. Headings used herein are for organizational purposes only and are not meant to be used to limit the scope of the description.
[0096] As used herein, unless specifically stated otherwise, the term “or” encompasses all possible combinations, except where infeasible. For example, if it is stated that a component includes A or B, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component includes A, B, or C, then, unless specifically stated otherwise or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C. Expressions such as “at least one of’ do not necessarily modify an entirety of a following list and do not necessarily modify each member of the list, such that “at least one of A, B, and C” should be understood as including only one of A, only one of B, only one of C, or any combination of A, B, and C. The phrase “one of A and B” or “any one of A and B” shall be interpreted in the broadest sense to include one of A, or one of B.
[0097] The descriptions herein are intended to be illustrative, not limiting. Thus, it will be apparent to one skilled in the art that modifications may be made as described without departing from the scope of the claims set out below.
Claims
CLAIMS1. A method of generating a mask pattern using a machine learning model, the method comprising: providing a target pattern to a machine learning (ML) model; and executing the ML model to generate an output having multiple channels of information associated with a mask pattern corresponding to the target pattern, wherein the multiple channels include: a first channel that is configured to generate first mask information corresponding to main features to be printed on a substrate, and a second channel that is configured to generate second mask information corresponding to sub-resolution assist features (SRAFs), and wherein the ML model is trained by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the main features and an SRAF of the SRAFs that is adjacent to the main feature.
2. The method of claim 1 , wherein the first mask information and the second mask information are images of the corresponding features.
3. The method of claim 1, further comprising: determining a distance between a first main feature of the main features and a first SRAF of the SRAFs that is adjacent to the first main feature; and determining whether there is an overlap between the first main feature and the first SRAF based on the distance.
4. The method of claim 1, further comprising: training the ML model by: obtaining a first ground truth image of a set of main features and a second ground truth image of a set of SRAFs, determining, in the first ground truth image, a first set of pixel values in a first region proximate a first main feature of the set of main features, wherein the first region is defined by the specified distance, determining, in the second ground truth image, a second set of pixel values in a second region proximate a first SRAF of the set of SRAFs, wherein the second region corresponds to the first region, and determining the violation based on a determination of at least one of the first set of pixel values or the second set of pixel values being non-zero.
5. The method of claim 4, wherein determining the violation includes: computing the cost function, wherein computing the cost function includes: defining a size of a kernel based on the specified distance, convolving each of the first main feature and the first SRAF with the kernel to obtain a first convolved feature and a second convolved feature, respectively, and determining the violation based on a product of the first convolved feature and the second convolved feature.
6. The method of claim 1 , further comprising training the ML model by using a cost function that penalizes a violation of a specified relationship between two or more channels of the multiple channels.
7. The method of claim 1 further comprising: obtaining a training dataset having a first target pattern and a first ground truth mask pattern, obtaining from the first ground truth mask pattern (a) a first image having a set of main features of the first ground truth mask pattern, and (b) a second image having a set of SRAFs of the first ground truth mask pattern, and training the ML model, with the first target pattern and using the first image and the second image as ground truth for a first channel and second channel of the multiple channels, respectively, to generate a predicted set of main features as the first channel and a predicted set of SRAFs as the second channel.
8. The method of claim 7, further comprising training the ML model by using a cost function that penalizes a difference between the set of main features and the predicted set of main features, and a difference between the set of SRAFs and the predicted set of SRAFs.
9. The method of claim 7, further comprising aggregating the first image having the set of main features with the second image having the set of SRAFs to generate an aggregated mask pattern.
10. The method of claim 9, further comprising training the ML model by using a cost function that penalizes a difference between the aggregated mask pattern and the first ground truth mask pattern.
11. The method of claim 7, further comprising training the ML model by using a cost function that includes multiple cost terms, wherein the multiple cost terms include:(i) a first cost term that penalizes a difference between (a) the set of main features and the predicted set of main features and (b) the set of SRAFs and the predicted set of SRAFs,(ii) a second cost term that penalizes a difference between an aggregated mask pattern and the first ground truth mask pattern, and(iii) a third cost term that penalizes a violation of a specified relationship between two or more channels of the multiple channels.
12. The method of claim 1, further comprising: obtaining a training dataset having a second target pattern, and training the ML model with the second target pattern to generate a second predicted set of main features and a second predicted set of SRAFs as the first and second channels of the multiple channels, respectively, and wherein the method further comprises training the ML model by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the second predicted set of main features and an SRAF of the second predicted set of SRAFs that is adjacent to the main feature.
13. A method of generating a mask pattern using a machine learning model, the method comprising: providing a training dataset to a machine learning (ML) model, wherein the training dataset includes a target pattern and a set of ground truth including (a) mask information corresponding to main features of a mask pattern corresponding to the target pattern, and (b) mask information corresponding to sub-resolution assist features (SRAFs) of the mask pattern; and training the ML model with the training dataset to generate multiple channels of information associated with the mask pattern, wherein the multiple channels include: a first channel that is configured to generate predicted main features to be printed on a substrate, and a second channel that is configured to generate predicted SRAFs, and wherein the ML model is trained by using a cost function configured to penalize a violation of a specified distance to be maintained between a main feature of the predicted main features and an SRAF of the predicted SRAFs that is adjacent to the main feature.
14. The method of claim 13, further comprising training the ML model by: obtaining a first image of the predicted main features and a second image of the predicted SRAFs, determining, in the first image, a first set of pixel values in a first region proximate the main feature, wherein the first region is defined by the specified distance, determining, in the second image, a second set of pixel values in a second region proximate the SRAF, wherein the second region corresponds to the first region, anddetermining the violation based on a determination of at least one of the first set of pixel values or the second set of pixel values being non-zero.
15. The method of claim 13, further comprising training the ML model by using a cost function that penalizes a violation of a specified relationship between two or more channels of the multiple channels.
Citation Information
Patent Citations
System and method for creating a focus-exposure model of a lithography process
US20070031745A1
Method for identifying and using process window signature patterns for lithography process control
US20070050749A1
System and method for model-based sub-resolution assist feature generation
US20080301620A1
Multivariable solver for optical proximity correction
US20080309897A1
Methods and system for lithography process window simulation
US20090157360A1