Method of predicting parameter of interest in semiconductor manufacturing process
By using prediction submodules in the lithography equipment to process abnormalities in the measurement data, the problem of difficulty in interpreting the measurement data and determining the optimal process correction in the prior art is solved, and the accuracy of accurate prediction of key parameters of the lithography equipment manufacturing process is achieved and the accuracy of pattern reproduction is improved.
Patent Information
- Application Number
- CN202380074112.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-09-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to determine the optimal process correction when interpreting the measurement data in a lithography device, resulting in possible errors in the measurement data of subsequent exposure wafers.
A method is adopted to obtain measurement data related to the parameter of interest, and to process non-systematic and systematic anomalies using the first and second prediction submodules, generate anomaly prediction data, and combine it with the non-analysis prediction data to obtain predictions of the parameter of interest.
By separating systematic and non-systematic abnormalities, we can accurately predict key parameters in the manufacturing process of lithography equipment, reduce errors in measurement data, and improve the accuracy of pattern reproduction.
Smart Images

Figure CN120077331A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to European Application No. 22202732.8, filed on October 20, 2022, the entire content of which is incorporated herein by reference. Technical field
[0003] The present invention relates to semiconductor manufacturing processes, and more particularly to methods for inspection or metrology in semiconductor manufacturing processes. Background art
[0004] A lithographic apparatus is a machine configured to apply a desired pattern onto a substrate. A lithographic apparatus can be used, for example, in the manufacture of integrated circuits (ICs). A lithographic apparatus can project a pattern (also often referred to as a "design layout" or "design") present on a patterning device (e.g., a mask) onto a layer of radiation - sensitive material (resist) provided on a substrate (e.g., a wafer).
[0005] In order to project a pattern onto a substrate, a lithographic apparatus can use electromagnetic radiation. The wavelength of this radiation determines the minimum size of the features that can be formed on the substrate. Typical wavelengths currently in use are 365 nm (i - line), 248 nm, 193 nm, and 13.5 nm. Compared to a lithographic apparatus using radiation having a wavelength of, for example, 193 nm, a lithographic apparatus using extreme ultraviolet (EUV) radiation having a wavelength in the range of 4 nm to 20 nm (e.g., 6.7 nm or 13.5 nm) can be used to form smaller features on a substrate.
[0006] Low k 1 Lithography can be used to process features smaller than the classical resolution limit of a lithographic apparatus. In such a process, the resolution formula can be expressed as CD = k 1 ×λ / NA, where λ is the wavelength of the radiation employed, NA is the numerical aperture of the projection optics in the lithographic apparatus, CD is the "critical dimension" (usually the smallest feature size printed, but in this case the half - pitch), and k 1 is an empirical resolution factor. Typically, k 1The smaller it is, the more difficult it is to reproduce on a substrate a pattern similar to the shape and dimensions planned by the circuit designer in order to achieve specific electrical functionality and performance. To overcome these difficulties, complex fine-tuning steps can be applied to the lithographic projection apparatus and / or the design layout. These steps include, for example but not limited to, optimization of the NA, customized illumination schemes, use of a phase-shifting patterning device, various optimizations of the design layout such as optical proximity correction (OPC, sometimes also referred to as "optical and process correction") in the design layout, or other methods generally defined as "resolution enhancement techniques" (RET). Alternatively, a strict control loop for controlling the stability of the lithographic apparatus can be used to improve the reproduction of the pattern at low k1.
[0007] These strict control loops are typically based on measurement data obtained using metrology tools that measure characteristics of the applied pattern or a metrology target representative of the applied pattern. Typically, the metrology tools are based on optical measurements of the position and / or dimensions of the pattern and / or target. It is essentially assumed that these optical measurements represent the quality of the integrated circuit manufacturing process.
[0008] Process corrections for the IC manufacturing process can be determined from the measurement data of previously exposed wafers (the terms wafer and substrate are used interchangeably and / or synonymously throughout this disclosure) in order to minimize any errors in the measurement data for subsequently exposed wafers. However, it can sometimes be difficult to interpret the measurement data, i.e., the measurement data does not always represent the optimal correction. SUMMARY OF THE INVENTION
[0009] An object of the inventors is to solve the mentioned disadvantages of the prior art.
[0010] In a first aspect of the invention, there is provided a method for predicting a parameter of interest in a manufacturing process for manufacturing an integrated circuit, the method comprising: obtaining measurement data related to the parameter of interest; applying a first prediction sub-module to the measurement data to obtain non-anomalous prediction data; detecting an anomaly in the measurement data (e.g., using an anomaly detection module); classifying the anomaly into a systematic anomaly and a non-systematic anomaly; using a first prediction strategy for the non-systematic anomaly to obtain first anomaly prediction data; using a second prediction strategy for the systematic anomaly to obtain second anomaly prediction data; wherein the first prediction strategy is different from the second prediction strategy; and combining the first anomaly prediction data and / or the second anomaly prediction data with the non-anomalous prediction data to obtain a prediction of the parameter of interest.
[0011] There is also disclosed a computer program and various devices operable to perform the method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Embodiments of the present invention will now be described, by way of example only, with reference to the accompanying schematic diagrams, in which:
[0013] Figure 1 A schematic overview of a lithographic apparatus is depicted;
[0014] Figure 2 A schematic overview of a lithographic cell is depicted;
[0015] Figure 3 A schematic representation of overall lithography is depicted, which represents the cooperation between three key technologies for optimizing semiconductor manufacturing; and
[0016] Figure 4 a and Figure 4 b schematically illustrate two known feedback control methods applied to a manufacturing facility.
[0017] Figure 5 is a simplified schematic flow chart of a part of an IC manufacturing method according to a known method;
[0018] Figure 6 is a graph of the control parameter value PV (or the measured parameter value depending on the control parameter) against time t for each of the following: (a) a positive transition / event and exponentially weighted moving average (EWMA) time-domain filtering method; (b) a positive transition / event and neural network (NN)-based time-domain filtering method; (c) a negative transition / event and EWMA time-domain filtering method; and (d) a negative transition / event and NN-based time-domain filtering method;
[0019] Figure 7 is a measured signal trace against time showing different types of anomalies and drifts;
[0020] Figure 8 is a flow chart of a correction determination method according to an embodiment;
[0021] Figure 9 is a diagram of an anomaly detection architecture according to an embodiment; and
[0022] Figure 10 is a visualization of a continuous learning strategy according to an embodiment. Detailed Description
[0023] In this document, the terms "radiation" and "beam" are used to encompass all types of electromagnetic radiation, including ultraviolet radiation (e.g., having a wavelength of 365 nm, 248 nm, 193 nm, 157 nm, or 126 nm) and EUV (extreme ultraviolet radiation, e.g., having a wavelength in the range of about 5 nm to 100 nm).
[0024] As used herein, the terms "reticle", "mask", or "patterning device" can be broadly interpreted as referring to a general patterning device that can be used to impart a patterned cross-section to an incident radiation beam, the patterned cross-section corresponding to a pattern to be created in a target portion of a substrate. In such a context, the term "light valve" may also be used. Examples of such other patterning devices include, in addition to classical masks (transmission or reflection; binary, phase-shifting, hybrid, etc.):
[0025] - programmable mirror arrays. More information about such mirror arrays is given in U.S. Patent Nos. 5,296,891 and 5,523,193, which are incorporated herein by reference.
[0026] - programmable LCD arrays. An example of such a configuration is given in U.S. Patent No. 5,229,872, which is incorporated herein by reference.
[0027] Figure 1 The lithographic apparatus LA is schematically depicted. The lithographic apparatus LA includes: an illumination system (also referred to as an illuminator) IL configured to condition a radiation beam B (e.g., UV radiation, DUV radiation, or EUV radiation); a support structure (e.g., a mask table) MT configured to support a patterning device (e.g., a mask) MA and connected to a first positioner PM configured to accurately position the patterning device MA according to certain parameters; a substrate table (e.g., a wafer table) WT configured to hold a substrate (e.g., a wafer coated with resist) W and connected to a second positioner PW configured to accurately position the substrate according to certain parameters; and a projection system (e.g., a refractive projection lens system) PS configured to project the pattern imparted to the radiation beam B by the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W.
[0028] In operation, the illuminator IL receives the radiation beam from a radiation source SO, e.g., via a beam delivery system BD. The illumination system IL may include various types of optical components for guiding, shaping, or controlling the radiation, such as refractive, reflective, magnetic, electromagnetic, electrostatic, or other types of optical components, or any combination thereof. The illuminator IL may be used to condition the radiation beam B to have a desired spatial and angular intensity distribution in the cross-section of the radiation beam at the plane of the patterning device MA.
[0029] As used herein, the term "projection system" PS should be interpreted broadly to cover any type of projection system suitable for the exposure radiation used or for other factors such as the use of an immersion liquid or the use of a vacuum, including refractive, reflective, catadioptric, anamorphic, magnetic, electromagnetic, and / or electrostatic optical systems or any combination thereof. Any use herein of the term "projection lens" may be considered synonymous with the more general term "projection system" PS.
[0030] A lithographic apparatus may belong to a type in which at least a portion of the substrate is covered by a liquid having a relatively high refractive index (e.g., water) in order to fill the space between the projection system and the substrate - this is also known as immersion lithography. More information on immersion techniques is given in U.S. Patent No. 6,952,253, which is incorporated herein by reference.
[0031] A lithographic apparatus LA may also belong to a type having two (dual-platform) or more substrate tables WT and, for example, two or more support structures MT (not shown). In these "multi-platform" machines, additional tables / structures may be used in parallel, or preparatory steps may be carried out on one or more tables while one or more other tables are used to expose the design layout of the patterning device MA to the substrate W.
[0032] In operation, a radiation beam B is incident on a patterning device (e.g., a mask MA) held on a support structure (e.g., a mask table MT) and is patterned by the patterning device MA. After traversing the mask MA, the radiation beam B passes through a projection system PS which focuses the beam onto a target portion C of the substrate W. By means of a second positioner PW and a position sensor IF (e.g., an interferometric device, a linear encoder, a 2D encoder, or a capacitive sensor), the substrate table WT can be accurately moved, for example, so as to position different target portions C in the path of the radiation beam B. Similarly, a first positioner PM and possibly another position sensor (which is not explicitly depicted in Figure 1 can be used to accurately position the mask MA relative to the path of the radiation beam B. Mask alignment marks M1, M2 and substrate alignment marks P1, P2 may be used to align the mask MA and the substrate W. Although the illustrated substrate alignment marks occupy dedicated target portions, the substrate alignment marks may be located in the space between the target portions (these substrate alignment marks are called scribe alignment marks).
[0033] In Figure 2As shown, the lithographic apparatus LA can form part of a lithographic cell LC (sometimes also referred to as a lithographic cell or (lithographic) cluster), which often also includes equipment for performing pre-exposure and post-exposure processes on the substrate W. Conventionally, such equipment includes a spin coater SC for depositing a resist layer, a developer DE for developing the exposed resist, a chill plate CH for adjusting the temperature of the substrate W (e.g., for adjusting the solvent in the resist layer), and a bake plate BK. A substrate transfer device or robot RO picks up the substrate W from the input / output ports I / O1, I / O2, moves the substrate between different process equipment, and transfers the substrate W to the feed table LB of the lithographic apparatus LA. The devices in the lithographic cell, often collectively referred to as the track or coat and develop system, are typically under the control of a track or coat and develop system control unit TCU, which itself may be controlled by a management control system SCS, which may also control the lithographic apparatus LA via, for example, a lithography control unit LACU.
[0034] To correctly and consistently expose the substrate W exposed by the lithographic apparatus LA, it is desirable to inspect the substrate to measure properties of the patterned structures, such as overlay errors between subsequent layers, line thickness, critical dimension (CD), etc. For this purpose, an inspection tool (not shown) may be included in the lithographic cell LC. If an error is detected, the exposure of subsequent substrates or other processing steps to be performed on the substrate W can be adjusted, especially in cases where inspections are carried out before other substrates W in the same batch or lot are still to be exposed or processed.
[0035] An inspection device, which may also be referred to as a metrology device, is used to determine the properties of the substrate W and, in particular, how the properties of different substrates W vary or how the properties associated with different layers of the same substrate W vary from layer to layer. The inspection device is alternatively configured to identify defects on the substrate W and may be, for example, part of the lithographic cell LC, or may be integrated into the lithographic apparatus LA, or may even be a separate device. The inspection device can measure properties on a latent image (the image in the resist layer after exposure), or a semi-latent image (the image in the resist layer after a post-exposure bake step PEB), or a developed resist image (where the exposed or unexposed portions of the resist have been removed), or even properties on an etched image (after a pattern transfer step such as etching).
[0036] Typically, the patterning process in the lithographic apparatus LA is one of the most important steps in the process, which requires a high degree of accuracy in the sizing and placement of the structures on the substrate W. To ensure such high accuracy, three systems can be combined in a so-called "integrated" control environment, as Figure 3is schematically depicted. One of these systems is a lithographic apparatus LA, which is (virtually) connected to a metrology tool MT (second system) and to a computer system CL (third system). The key to such an "integrated" environment is to optimize the collaboration between these three systems to enhance the overall process window and to provide a tight control loop, thereby ensuring that the patterning performed by the lithographic apparatus LA remains within the process window. The process window defines a range of process parameters (e.g., dose, focus, overlay) within which a particular manufacturing process yields a defined result (e.g., a functional semiconductor device) - typically allowing process parameter variations in the lithography process or patterning process within the defined result.
[0037] The computer system CL can use (parts of) the design layout to be patterned to predict which resolution enhancement techniques to use and perform computational lithography simulations and calculations to determine which mask layout and lithographic apparatus settings achieve the maximum overall process window for the patterning process (depicted by the double white arrows in Figure 3 the first scale SC1). Typically, the resolution enhancement techniques are arranged to match the patterning capabilities of the lithographic apparatus LA. The computer system CL can also be used to detect where within the process window the lithographic apparatus LA is currently operating (e.g., using input from the metrology tool MT) in order to predict whether there might be defects, for example due to sub-optimal processing (depicted by the arrow pointing to "0" in Figure 3 the second scale SC2).
[0038] The metrology tool MT can provide input to the computer system CL to enable accurate simulations and predictions, and can provide feedback to the lithographic apparatus LA to identify possible drifts, for example in the calibration state of the lithographic apparatus LA (depicted by the multiple arrows in Figure 3 the third scale SC3).
[0039] The lithographic apparatus LA is configured to reproduce a pattern accurately on a substrate. The position and size of the features applied need to be within certain tolerances. The position error can occur due to overlay error (often referred to as "overlay"). Overlay is the error in placing a first feature during a first exposure relative to placing a second feature during a second exposure. The lithographic apparatus minimizes the overlay error by accurately aligning each wafer with a reference before patterning. This is done by using alignment sensors to measure the position of alignment marks on the substrate. More information about the alignment process can be found in U.S. Patent Application Publication No. US20100214550, which is incorporated herein by reference. For example, pattern size calibration (e.g., CD) errors can occur when the substrate is not correctly positioned relative to the focal plane of the lithographic apparatus. These focus position errors can be associated with the non-planarity of the substrate surface. The lithographic apparatus minimizes these focus position errors by using a leveling sensor to measure the substrate surface topography before patterning. A substrate height correction is applied during subsequent patterning to ensure correct imaging (focusing) from the patterning device onto the substrate. More information about the leveling sensor system can be found in U.S. Patent Application Publication No. US20070085991, which is incorporated herein by reference.
[0040] In addition to the lithographic apparatus LA and the metrology apparatus MT, other processing equipment can also be used during IC production. An etch station (not shown) processes the substrate after the pattern has been exposed into the resist. The etch station transfers the pattern from the resist into one or more layers below the resist layer. Typically, the etch is based on the application of a plasma medium. Local etch characteristics can be controlled, for example, by using temperature control of the substrate or by using a voltage control loop to direct the plasma medium. More information about etch control can be found in International Patent Application Publication No. WO2011081645 and U.S. Patent Application Publication No. US 20060016561, which are incorporated herein by reference.
[0041] During the manufacture of an IC, it is very important that the process conditions for processing a substrate using processing equipment such as a lithography apparatus or an etching station remain stable so that the properties of the features remain within certain control limits. The stability of the process is particularly important for the features of the functional parts of the IC, i.e., the product features. To ensure stable processing, process control capabilities need to be in place. Process control involves monitoring of process data and implementation of means for process correction, such as controlling the processing equipment based on the characteristics of the process data. Process control can be based on periodic measurements made by a metrology device MT, often referred to as "advanced process control" (also alternatively referred to as APC). More information about APC can be found in U.S. Patent Application Publication No. US20120008127, which is incorporated herein by reference. Typical APC implementation involves periodic measurements of metrology features on a substrate to monitor and correct for drifts associated with one or more processing devices. The metrology features reflect the response of the process variations to the product features.
[0042] In US20120008127, a lithography apparatus is calibrated with reference to a primary reference substrate. Using equipment that does not need to be the same as the calibrated lithography apparatus, device-specific characteristic identifications of the primary reference substrate are obtained. Using the same settings, device-specific characteristic identifications of a secondary reference substrate are then obtained. The device-specific characteristic identifications of the secondary reference substrate are subtracted from the device-specific characteristic identifications of the primary reference substrate to obtain and store device-independent characteristic identifications of the secondary reference substrate. The secondary reference substrate and the stored device-independent characteristic identifications are then used together as a reference for calibration of the lithography apparatus to be calibrated, instead of the primary reference substrate. The initial setup of a cluster of lithography tools can be performed with less use of expensive primary reference substrates and less interruption of normal production. The initial setup can be integrated with the ongoing monitoring and recalibration of the equipment.
[0043] The term characteristic identification can refer to the main (systematic) contributing factors ("latent factors") of the measured signal, and in particular to contributing factors related to performance effects on the wafer or related to previous processing steps. Such characteristic identifications can refer to substrate (grid) patterns (e.g., from alignment, horizontal metrology, overlay, focus, CD), field patterns (e.g., from in-field alignment, horizontal metrology, overlay, focus, CD), substrate partition patterns (e.g., the outermost radius of wafer measurement), or even patterns in scanner measurements related to wafer exposure (e.g., inter-batch heating signatures from mask alignment measurements, temperature / pressure / servo profiles, etc.). Characteristic identifications can be included in a set of characteristic identifications and can be encoded uniformly or non-uniformly therein.
[0044] Thus, the APC recognizes correctable variations in performance parameters such as overlay, and applies a set of corrections to a lot (batch) of wafers. When determining these corrections, corrections from previous lots are considered to avoid over-correcting for noise in the measurements. For sufficient smoothing of the current correction with the previous correction, the history of the corrections considered can match the context of the current lot. In this regard, "context" encompasses any parameter that identifies variants present within the same overall industrial process. Layer ID, layer type, product ID, product type, reticle ID, etc. are all context parameters that can produce different signatures in the final performance. In addition to individual scanners that can be used in high-volume manufacturing (HVM) facilities, individual tools for each of the other steps involved in coating, etching, and semiconductor manufacturing can also vary between different lots or between different wafers. Each of these tools can impose a specific error "signature" on the product. Outside of the field of semiconductor manufacturing, similar situations can occur in any industrial process.
[0045] To ensure accurate feedback control suitable for a particular context, product units of different lots (batches) can be treated as separate "threads" in the APC algorithm. Context data can be used to assign each product unit to the correct thread. In cases where a manufacturing fab typically produces large volumes of only a few types of products through the same process steps, the number of different contexts can be relatively small, and the number of product units in each thread will be sufficient to smooth the noise. All lots with a common context can be assigned to their own thread to optimize the feedback correction and the final performance. In cases where a chip foundry produces many different types of products in very small production runs, the context may change more frequently, and the number of lots with exactly the same context data may be quite small. Assigning lots to different APC "threads" using only context data can then result in a large number of threads, each with a small number of lots. The complexity of the feedback control increases, and the ability to improve the performance of low-volume products decreases. Combining different lots into the same thread without sufficient consideration of the different contexts of the lots will also result in a loss of performance.
[0046] Figure 4 (a) Schematically illustrates the operation of one type of control method implemented by the APC system 250. Historical performance data PDAT is received from the memory 252, which has been obtained by the metrology device 240 or other means from the wafer 220 that has been processed by the lithography apparatus 200 and the associated equipment of the lithography cell. The feedback controller 300 analyzes the performance parameters represented in the performance data of the most recent lot, and calculates a process correction PC that is fed to the lithography apparatus 200. These process corrections are added to wafer-specific corrections derived from the alignment sensors and other sensors of the lithography apparatus to obtain a combined correction for processing each new lot.
[0047] Figure 4 (b) Schematically illustrates the operation of another type of control method implemented by the known APC system 250. It can be seen that the general form of the feedback control method is the same as that Figure 4 shown in (a), but in this example, the context data related to the historical wafers and the context data CTX related to the current wafer are used to provide a more selective use of the performance data PDAT. In particular, while in the earlier example, the performance data of all historical wafers were combined in a single stream 302 and the modified method, the context data from the memory 256 is used to assign the performance data of each historical batch to one of several threads 304. These threads are effectively processed by the feedback controller 300 in parallel feedback loops, resulting in a plurality of process corrections 306, each process correction based on the historical performance data of the wafers in one of the threads 304. Then, when a new batch is received for processing, its individual context data CTX can be used to select which of the threads provides the appropriate context data 306 to the current wafer.
[0048] There are various alternative methods for time processing and / or filtering of (feedback) control data (e.g., overlay feature identification or EPE feature identification). These methods include using low-pass time-domain filtering methods or moving average processing methods; e.g., weighted moving average or exponentially weighted moving average EWMA. Other methods include machine learning models, such as neural networks (NN). For example, advanced NN filtering methods can be trained to learn an appropriate response to the time behavior based on historical control parameter data and, for example, provide (feedback) correction predictions for the next batch in the APC control loop.
[0049] The disadvantages of these existing methods are as follows: The time processor (e.g., NN or EWMA filter) "learns" only based on the behavior of the control parameters, which can be determined from the metrology data of the exposed structure. The measured parameter values in the metrology data (e.g., overlay, edge placement error, critical dimension, focus) change as the behavior of these control parameters changes over time (i.e., different outputs are measured for the same control input). The control parameters can be any input parameters of scanners or other tools used in IC manufacturing (e.g., etch chambers, deposition chambers, bonding tools, etc.) that control the manufacturing process (e.g., exposure process, etch process, deposition process, bonding process, etc.). Therefore, the control of the output process, and more particularly, the formation, configuration, and / or positioning of the exposed and / or etched structure, depends on these control parameters. Thus, these control parameters can be controlled to correct or compensate for any measured errors in the metrology data, correct future wafers / batches in a feedback loop, or correct the current wafer / batch as a feed-forward correction. Note that "learning" in this context includes averaging in the moving average example, as the moving average output effectively "learns" (in a loose sense) to respond to the input data by averaging the data (its output response changes over time based on the previous few inputs).
[0050] The APC control loop described above has the following main tasks: monitoring the drift in the metrology data indicating the drift of the behavior of the control parameters; and determining the appropriate correction to the control parameters to address this drift and maintain the measured metrology parameter values within the specifications (i.e., within a certain acceptable tolerance or "process window") within which the IC device can be expected to function with a good probability.
[0051] However, not all drifts in the metrology data should be corrected (or followed); rather, only the drifts in the actual parameters of the product features caused by the drift of the control parameter behavior (true drift or systematic change) should be corrected (or followed). Other sources that may cause drifts in the metrology data, such as metrology tool drift or metrology target defects that are not replicated in the product structure (e.g., feature identification introduced by overlay target deformation) (due to their large size, metrology targets can behave differently from the product structure when being imaged and / or measured). These drifts are not "real", i.e., they do not actually indicate a drift in the exposure (or other processing) process that affects the quality of the product structure. Of course, a metrology tool that has drifted and measures inaccurately, resulting in out-of-spec metrology parameter values, does not mean that the product on the wafer is out of spec; therefore, such metrology tool drift should be ignored by the APC loop. In addition, alignment mark deformation may cause a drift that can be captured by the APC loop, and such drift should not be followed.
[0052] In addition to drift (e.g., relatively stable), the measurement data can also indicate "jumps" or "steps" in the process, such as a sudden relatively large change in the measured parameter value indicating a relatively sudden change in the control behavior. Similar to drift, these jumps can indicate something that should be followed and corrected (systematic anomalies or interference events), or alternatively something that should be ignored (e.g., non-systematic / instantaneous anomalies or interference events). A specific example of an anomaly indicating something that should be followed is when the calibration state of the scanner changes. For example, this can manifest as a jump in the magnification. Such a state change should be incorporated into the updated feedback control because the change is permanent. In contrast, the scanner lens can experience a lens "hiccup", i.e., a non-reproducible problem of uncertain cause with the lens, or an instantaneous lens anomaly, which can also manifest as a jump in the magnification. However, such a non-reproducible problem of uncertain cause with the lens is a one-time deviation and should not be followed by the feedback control because the non-reproducible problem of uncertain cause will not exist in the next batch.
[0053] The APC controller cannot know which drifts and / or jumps it should act on and which should be ignored based only on the measurement data because these different types of drifts and jumps are indistinguishable within the measurement data; i.e., a lens hiccup or a non-reproducible problem of uncertain cause with the lens, and a calibration state jump will look the same in the measurement data. More particularly, a neural network used as a time-domain filter can be trained to learn how to respond to certain trends or events (e.g., drifts and / or jumps). However, without any knowledge or understanding of the underlying causes of these effects, a neural-network-based control system cannot determine to respond appropriately (e.g., follow, ignore, or partially follow / partially ignore (e.g., according to appropriate non-binary weighting)).
[0054] Figure 5is a flowchart of an IC manufacturing process (e.g., part of) related to exposure, metrology, and APC control for processes for a large number of wafer lots. Examples are abbreviated and may include etch steps, deposition steps, wafer bonding steps, etc. depending on the particular process. Time t is shown as progressing from left to right. Exposure EXP N-1 of lot N-1 is performed and then measured MET N-1. The modeled step MOD N-1 is performed to fit the model to the metrology data, e.g., such that the metrology data can be described more effectively. Within the APC controller, the filtered feature signature FP is determined (i.e., based on the modeled metrology data and the feature signature FP N-2 from at least the previous lot, and possibly additional data (e.g., from other lots) in the case where NN is used as a time domain filter). Based on this feature signature, the process correction PC N for the next lot (lot N) is determined. Although the metrology, modeling, feature signature, and correction determination steps are shown as occurring simultaneously, they of course cannot literally occur simultaneously, but only effectively occur simultaneously within the context of the flow shown. After this, the exposure EXP N of lot N is performed using the determined correction PC N. However, during this exposure, a disturbance event DE occurs, which may, for example, manifest as a jump in the metrology data MET N. The remainder of the flow is the same as for lot N-1, but the gray arrows indicate that this flow and the modeled data MOD N, feature signature, and correction PC N+1 for the next lot will be affected by the disturbance event DE.
[0055] Figure 6 The figure shows how control parameters can follow data according to a control strategy, which depends on whether the jump or disturbance event is a real event (systematic) or a false / one-time event (non-systematic). Each graph is a graph of a control parameter value PV or a metrology parameter value that depends on the control parameter (e.g., overlay) versus time (or lot). In each graph, each point up to lot N represents the value of that lot. The point for lot N (represented as a white circle) indicates the jump ( Figure 6 (a) and Figure 6 (b) a positive jump and Figure 6 (c) and Figure 6 (d) a negative jump). In addition, each lot is also represented by two points: a first point (black), which represents what is expected to be seen in the case where the jump is real; and a second point (gray), which represents what is expected to be seen in the case where the jump is a false / one-time event. The line represents the control signal correction, which can be determined by the APC loop based on the metrology points. Similarly, after lot N, there are two lines, a black line following the black dot and a gray line following the gray dot. The distance between the control signal correction (line) and the corresponding control parameter point indicates the control performance; the closer the line is to the corresponding point, the better the correction and control performance.
[0056] Figure 6 (a) shows a positive jump and EWMA-based control. The EWMA is slow in following the true jump. The non-reproducible cause of uncertainty (in this one example) actually helps with control because it reduces the control latency (the grey line for batch N+1 is closer to the parameter value than it would be without the jump occurring). Figure 6 (b) shows a positive jump and NN-based control. The NN more preferably follows the correction and preferably follows the true jump, assuming it has been trained to do so. It also misinterprets the non-reproducible cause of uncertainty as a true jump, meaning that in the case of a non-reproducible cause of uncertainty, batch N+1 can preferably be exposed out of specification. This can also be the other way around depending on how the NN is trained; i.e., if it is trained to ignore jumps and assume it is a non-reproducible cause of uncertainty or somewhere in between (e.g., when it has been trained on both and can respond with an intermediate correction). Figure 6 (c) and Figure 6 (d) respectively show Figure 6 (a) and Figure 6 (b)'s equivalent curve graphs, with negative jumps. Comparing Figure 6 (c) with Figure 6 (a), it is now clear that the negative non-reproducible cause of uncertainty (false jump) significantly impairs control of a large number of future batches in this example.
[0057] During the lithography cycle, metrology tools are used to measure the parameter of interest (e.g., overlay, but it can be another parameter indicating imaging performance) for each exposed batch, resulting in a large stream of multivariate time series data (i.e., the metrology data described above Figure 5 and below Figure 8 can include such multivariate time series data). This allows for the incorporation of a feedback loop into the control strategy, where the temporal relationships within the measurements of previous exposed batches can be analyzed by a controller to predict the parameter of interest (e.g., overlay) for future batches. Based on this, one or more control parameters of the manufacturing tools used in the integrated circuit (IC) manufacturing process can be corrected and optimized to generate a correction. The correction can be for scanner parameters (i.e., the control parameters of the lithography exposure tool), but it can be for other tools such as an etcher. The correction optimization can even include co-optimization of two or more of these tools. The correction optimization aims to determine a correction that improves the performance of the parameter of interest (e.g., minimizes the overlay compared to the predicted overlay) relative to the predicted performance of the parameter of interest.
[0058] In an overlay scenario, exposure batches are typically consecutive, and overlay parameters are continuously estimated, resulting in a large amount of multivariate time series data. To minimize overlay during the lithography cycle, a feedback loop can be incorporated into the control strategy, such as according to the aforementioned APC control loop. As part of the feedback control, mathematical relationships within the estimated overlay parameters of previous exposure batches can be analyzed to predict the parameters of future batches. The semiconductor process is highly dynamic and non-stationary, which can cause the overlay to follow complex spatio-temporal patterns. The controller aims to compensate for this by modifying the recipe. In particular, before exposing the next batch, the settings of the scanner can be adjusted to predict the overlay using correction optimization based on a set of predicted overlay parameters predicted by the controller. Typically, this can be done by calculating the feature signature for correcting the overlay of the next batch, or if it can be determined that the prediction may be related to an operating condition within the scanner (such as heating), the information can be directly fed back to the scanner to adjust these operating conditions within the scanner.
[0059] Accordingly, the goal of the controller can be to predict a set of overlay parameters for all future wafers based on the available history at the time point before exposure. This can include a multivariate time series forecasting task, where the goal is to prevent future errors by predicting the behavior of the machine.
[0060] Additional constraints can be imposed through practical applications; for example, the algorithm should consider the current and previous states to predict the overlay parameters before the batch corresponding to the exposure, in order to adjust the machine settings in a timely manner, thus limiting the allowed computation time (e.g., imposing a time constraint on performing these predictions based on the machine settings). The system continuously observes each data record as it arrives, and any processing or learning can be automated without manual intervention or parameter fine-tuning.
[0061] In real-world time series data from complex machinery, there are situations where the behavior of the machine unexpectedly changes based on usage or external factors. A sudden change in an overlay parameter can be caused by sources of systematic variation that often lead to a new definition of "normal" for the system that can be compensated for by appropriate control adjustments. The causes include multiple manual mechanism or control setting changes, external phenomena (e.g., environmental conditions), automatic calibration, auxiliary information (e.g., test runs), or any additional factors not accounted for by the controller. In addition to the mechanisms of the exposure system, other non-systematic sources can vary significantly enough to cause random high-frequency fluctuations, thereby reducing overlay accuracy. These situations can occur at any time during the process, resulting in the period in the streaming data being essentially unpredictable and thus rendering forecasting methods ineffective, even in the near future. A controller that cannot adapt to unpredictable behavior can have serious consequences for the prediction of overlay parameters, causing the feedback loop to make incorrect adjustments, which can ultimately lead to an increase in overlay and failures in the final device.
[0062] To predict overlay parameters for overlay control (e.g., APC), it is known to use a controller based on exponentially weighted moving average (EWMA). EWMA requires very little computational effort and storage of measurement data. This low-pass filter has been shown to smooth out small variations in the data stream. However, by definition, EWMA is a reactive control scheme that lags in time and is thus unable to keep up with sudden but permanent systematic changes in the time series. Additionally, EWMA responds to non-systematic high frequencies, which can have serious consequences for the prediction of overlay parameters in subsequent batches. EWMA can be designed to respond quickly to these rapid changes, but will then be very unstable and lose accuracy elsewhere. Furthermore, EWMA does not capture the spatio-temporal correlation structure and other dependencies among multiple overlay parameter time series by addressing the forecasting problem in a multi-variate setting.
[0063] The goal of the concepts disclosed herein is to develop a predictive controller that improves the prediction accuracy on which control adjustments are based and thereby reduces the parameter error of interest (e.g., reduces overlay). However, as has been described, there are situations where the behavior of a lithography machine unexpectedly changes based on usage or external factors. These situations can result in short-term or long-term sudden changes in the streaming data being essentially unpredictable. Unpredictable behavior can have serious consequences for the controller's prediction of future batches, causing the feedback loop to make incorrect control adjustments, which can ultimately lead to an increase in overlay or even failures in the final device. Given the unpredictable behavior of the lithography machine, it is proposed to enable the controller to automatically adapt to systematic long-term changes while preventing it from responding to non-systematic transient changes that do not carry any useful information.
[0064] An adaptive controller is proposed, which includes two main modules: a prediction module and an anomaly detection module. The prediction module predicts a future set (e.g., the next set) of one or more parameters of interest (e.g., one or more overlay parameters) at each step of the lithography cycle. The anomaly detection module can be configured to detect anomalies or interference events in the streaming data. An automated system can be provided, which combines two networks by differentiating between at least two types of anomalies identified by the anomaly detection module and changes the prediction strategy depending on the specific type.
[0065] Thus, a method for predicting parameters of interest in a manufacturing process for manufacturing integrated circuits, the method comprising: obtaining measurement data related to the parameters of interest; applying a first prediction sub-module to the measurement data to obtain non-anomaly prediction data; detecting anomalies in the measurement data (e.g., using an anomaly detection module); classifying the anomalies into systematic anomalies and non-systematic anomalies; using a first prediction strategy for the non-systematic anomalies to obtain first anomaly prediction data; using a second prediction strategy for the systematic anomalies to obtain second anomaly prediction data; wherein the first prediction strategy is different from the second prediction strategy; and combining the first anomaly prediction data and / or the second anomaly prediction data with the non-anomaly prediction data to obtain a prediction of the parameters of interest.
[0066] The measurement data may include at least current measurement data (e.g., from the current processed batch) and historical measurement data (e.g., from previous processed batches and / or as simulated). For example, the measurement data may include current measurement values of the parameters of interest and historical measurement values of the parameters of interest. The measurement data may include simulated or artificial data and / or real (i.e., actually measured) measurement data. In the case of using artificial data, there will be a dependence on what actually occurs before training. Reference may be made to US20210405544 A1, which is incorporated herein by reference. The measurement data may include pre-exposure measurement data (e.g., level measurement data, heating data, alignment data) and / or post-exposure measurement data (e.g., overlay, CD, focus or another measured parameter of interest). The measurement data may include context data (such as exposure time, changes in device design, production gap data, status changes in non-scanner tools such as etchers / etching chambers). Any of these types of measurement data may be included in the measurement data.
[0067] The step of detecting anomalies in the measurement data and classifying the anomalies may include detecting and classifying the anomalies in at least the current measurement data.
[0068] The prediction can be used, for example, to determine corrections for a lithography system or other tools in an IC manufacturing system (such as a scanner, an etcher, or other tools used in IC manufacturing) using a correction module (e.g., to determine correction feature identifiers), or can be used directly by the scanner or other manufacturing device to modify the operating conditions of the scanner / other manufacturing device.
[0069] The anomaly detection module can be configured to trigger an alarm when anomalies or unpredictable behavior in the (streamed) measurement data are detected. An automated system that detects these anomalies and classifies them as either systematic or non-systematic enables the prediction module to automatically adapt to sudden permanent changes and avoid responding to sudden transient changes that do not carry any useful information.
[0070] A continuous learning strategy can be used to train and optimize the module while the lithography cycle is running. The continuous learning strategy incrementally updates the controller system to account for any gradual changes in the streamed data distribution. For example, at each iteration, the anomaly detection module receives new observations from the metrology tool and appends the new observations to the previous observations it has received. Based on the entire (currently available) multivariate time series, the anomaly detection module uses a reconstruction-based unsupervised anomaly detection method to detect unpredictable behavior in the streamed data. The output of the anomaly detection module can be a set of subsequences identified as irregular or anomalous. Next, the anomaly detection module can distinguish the nature of each anomalous subsequence and classify it as either systematic or non-systematic.
[0071] All subsequences classified as non-systematic can be applied to a first prediction strategy (e.g., using a first prediction sub-module). The first prediction strategy can initially include processing the measurement data corresponding to the non-systematic anomaly; for example, to smooth the measurement data subsequence by applying a filter (e.g., a smoothing filter). The first prediction sub-module can include a prediction model or a prediction network to determine a prediction based on this smoothed measurement data at the next training step. In this way, the prediction module is prevented from responding to sudden transient changes.
[0072] For a systematic anomaly (e.g., a long-term change detected in real time), the first prediction strategy can be automatically replaced by a second prediction strategy (e.g., using a second prediction sub-module). The second prediction strategy can include applying a low-pass time-domain filter or a moving average processing method; e.g., a weighted moving average or an exponentially weighted moving average EWMA. Such algorithms can specifically utilize the data arriving after an anomaly has occurred, which allows the controller to immediately adjust the system's new definition of normal. While adopting the newly initialized EWMA, the first prediction sub-module (prediction network) can be automatically fine-tuned using the incoming data and can replace the EWMA after a set number of new batches of data have been received.
[0073] In an embodiment, the EWMA can include a short-sequence EWMA (ssEWMA). Such ssEWMA is described in R. Good and K Chamness: “Small-Sample Controller State Estimation: Initializing the EWMA Filter”, AEC / APC Europe, April 20, 2007, which is incorporated herein by reference.
[0074] It should be understood that the application of a low-pass time-domain filter or a moving average processing method (such as EWMA) is only an example of the second prediction strategy. The second prediction strategy can include applying any suitable predictor network, such as a neural network or a Bayesian predictor. In an embodiment, a machine learning model or a neural network can be trained to select an appropriate second prediction strategy based on the detected systematic anomaly.
[0075] More specifically, if the current observation is identified as a systematic anomaly, the prediction of the prediction network is not passed to the correction module for the corresponding time series. Instead, a new EWMA algorithm can be initialized to predict a (e.g., predefined) number (t EWMA ) of time points of the time series that include the anomaly, where these t EWMA time periods are the t EWMA immediately following the anomaly. The reason for this is that, given the sharp shift in the input data, the first few predictions generated by the prediction network after a systematic anomaly may be unstable for the time series with the anomaly. This phenomenon is called catastrophic forgetting; the network tends to completely and suddenly forget the previously learned information when learning new information.
[0076] The newly initialized EWMA uses only data that arrives after a systematic anomaly. Therefore, the controller does not require time to adjust. For illustrative purposes, if the first input to the new EWMA is the first observation after a sudden change, its prediction for the next observation is exactly equal to the input. The number of time periods t predicted using the EMWA can be determined empirically EWMA (e.g., it can be between 10 and 1000, between 10 and 500, between 10 and 400, between 10 and 300, between 20 and 300, or between 30 and 300).
[0077] In an embodiment, t EWMA can be set by the user to a preferred value, or the user can use the recommended settings based on the observations. Additionally, using relative weights, the transition from the EWMA to the predictor can be gradual, which increases the influence of the prediction-based parameter and decreases the influence of the EWMA-based parameter over a period of time.
[0078] The prediction network can be fine-tuned using the newly arrived data by retraining the prediction network before the anomaly (e.g., multiple times) after detecting a systematic anomaly, such that the prediction has time to remain stable using a continuous learning strategy.
[0079] Only all non-anomaly subsequences are passed to the prediction model / prediction network to determine non-anomaly prediction data related to the non-anomaly measurement data.
[0080] Figure 7 The figure illustrates three different types of irregularities that can be defined as follows: a sudden instantaneous change (non-systematic anomaly) NSA in the streaming data distribution, a long-term sudden shift (systematic anomaly) SA, and a gradual drift GD. The goal of the anomaly detection module is to alert for two sudden changes, NSA and SA. A continuous learning strategy can be employed such that the prediction network can adapt to the gradual change GD; therefore, it is proposed that the anomaly detection module does not alert for the gradual shift of the distribution. Although these categories do not cover all anomaly behaviors, these sudden anomaly types often occur in overlapping control and cause the greatest reduction in the performance of the current EWMA-controller based method. However, the types of anomalies can vary to serve different tasks in predictive maintenance.
[0081] Non-systematic anomalies can include (problematic) random high frequencies that the prediction network should ignore. Non-systematic anomalies can be defined as sudden instantaneous changes. These types do not have a systematic cause and are assumed to occur individually in the parameter time series. Non-systematic anomalies do not include periodic shifts. Batches corresponding to the non-systematic anomalies may cause failures in the final device and require rework. These anomalies are unforeseeable and the goal is not to predict the anomalies because, by definition, this is impossible. Since EWMA control schemes are reactive, they lag in time, and the prediction after a non-systematic anomaly will be affected. This deviation or difference may cause error correction and ultimately cause increased overlap.
[0082] Systematic anomalies can be defined as a long-term sudden shift in the signal average. Often, the systematic anomalies exist for a very long time or do not recover to their previous state at all. Such anomalies do not necessarily indicate machine problems and have related causes that the controller should adapt to. The causes can include external phenomena (e.g., environmental conditions), auxiliary information (e.g., test runs), or any additional factors that the controller did not take into account, including mechanism or control setting changes. The systematic anomalies occur only occasionally and have a high probability of occurring simultaneously in multiple parameter time series. For example, calibration often causes unpredictable shifts in multiple overlapping model parameters simultaneously. For example, many calibrations will actually cause predictable shifts that can be coupled to the neural network for control.
[0083] Anomalies are usually caused by unknown effects, indicating that the anomaly network may not have all the information it should have. As an example, if the exposure time is not considered as an input to the anomaly detector, long gaps in batch production will often cause anomalies. Therefore, significant anomalies may trigger the anomaly detection module to request more information from the user.
[0084] At each iteration of the lithography cycle, the anomaly detection module receives vector-valued observations (e.g., new elements of each parameter time series of interest in the set of time series). The output of the anomaly detection module is a set of anomaly subsequences for each time series in the set. Considering the rarity of systematic anomalies and that they occur simultaneously in multiple time series by definition, the type of anomaly can be distinguished based on the following domain rule: If multiple elements of a single observation at a particular time point are identified as anomalies, those elements are identified as systematic anomalies; otherwise they are identified as non-systematic anomalies. This is one way to distinguish anomalies, and other ways to distinguish anomalies are possible, including methods for identifying additional types of anomalies.
[0085] Figure 8is a flowchart depicting such an arrangement. According to the memory MEM, previous or historical measurement data HMET from historical batches / wafers is available. Current measurement data CMET can be obtained from a measurement tool MTL that measures the current batch / wafer. This information is fed into an anomaly detection module ADM that can include a machine learning model or neural network to identify anomalies that may pose problems for the predictor. The output from the anomaly detection module ADM is an anomaly sequence ANO, which is partitioned into systematic anomalies SYS and non-systematic anomalies NSYS.
[0086] One or more pre-established rules RL can be applied to the systematic anomalies SYS. Optionally, the information from the current measurement data can be used to update the rules when it arrives. The pre-established rules RL can be set based on known domain knowledge. Once the network has enough data, the network can determine its own rules and further optimize performance. By way of a specific example, these rules RL can include short computations to determine whether the latest observation is included in a systematic anomaly subsequence. If so, the index of the subsequence in which the latest observation occurs can be saved, and EWMA is used for the time series corresponding to these indexes. The data DAS after a systematic anomaly includes the data for the period after the last systematic anomaly that occurred with the maximum number of time instances t EWMA If EWMA is being used, it can output a univariate one-step ahead prediction for each time series where an anomaly occurred. These values replace their corresponding values in the multivariate one-step ahead prediction generated by the prediction network PN.
[0087] Systematic anomalies are used in predictors that are more robust (compared to the prediction network PN) when processing or transporting these anomalies. As already described, a simple low-cost implementation can include EWMA triggered by the anomaly detection module ADM. Alternatively, this first prediction sub-module can also include any network or model designed to transport systematic anomalies, including another machine learning model or neural network.
[0088] The non-systematic anomalies NSYS can be smoothed using a low-pass filter FIL, combined with a continuous data stream (e.g., the smoothed data can replace this stream during the time period that includes the non-systematic anomaly), and fed into a prediction module or prediction network PN. Such a prediction network can include a machine learning model or neural network.
[0089] The outputs from the EWMA and the prediction network PN can be combined to produce a prediction PD for the next batch of overlapping control. Depending on the scenario, the output can come entirely from the EWMA, entirely from the prediction network PN, or can be a combination of both, such as a weighted combination with weights automatically determined or selected by the user. For example, the weighting can vary over time to achieve a gradual transition from one strategy to another (e.g., from EWMA to the prediction network). The calibration optimization CO step can use the prediction to determine the calibration CR.
[0090] It can be appreciated that the anomaly detection module ADM can receive all available history and reclassify all available data points as anomalous or non-anomalous (before differentiating anomaly types) at each time point. The anomaly detection module ADM always receives the ground truth rather than a smoothed version. Thus, the anomaly detection module ADM can be able to identify a subsequence of a time series as anomalous at one time point and reclassify the subsequence as non-anomalous at another time point, or vice versa. To illustrate, if, later in the loop, an observation previously identified as anomalous becomes part of the time pattern, the disclosed arrangement can automatically adapt.
[0091] In an embodiment, the anomaly detection module ADM can include a reconstruction-based method such as implemented by an adversarial machine learning model; the adversarial machine learning model is, for example, a network based on a generative adversarial network (GAN). In particular, the anomaly detection module ADM can include a network based on a cycle generative adversarial network or CycleGAN that acts as a reconstruction-based method for anomaly detection. A particular implementation can include a Wasserstein GAN with gradient penalty for cycle consistency (CycleGAN-WP), such as described in "TadGAN: Time series anomaly detection using generative adversarial networks" by Alexander Geiger et al. in the 2020 IEEE International Conference on Big Data (Big Data). IEEE. 2020, pp. 33 - 43, which is incorporated herein by reference.
[0092] These reconstruction-based networks have advantages over traditional fixed-outlier-based methods (e.g., the n-σ filter), which predict that a signal is an anomaly if it exceeds a predefined threshold. The main problem with such outlier-based methods is that they can lose accuracy when the data is dynamic. Reconstruction-based methods have the following advantages: mapping the observations to a lower-dimensional space and then decoding or reconstructing the encoded points, assuming that anomalous observations cannot be reconstructed like regular samples because information is lost during encoding.
[0093] Figure 9 An exemplary anomaly detection module ADM architecture based on CycleGAN-WP is shown. The path of reconstruction is depicted by the dashed line. The dash represents the adversarial Wasserstein loss, and the dotted line represents the forward cycle-consistency loss.
[0094] Formally, given a multivariate time series X = [x <1> , ..., x <t>< / t> , the goal of unsupervised reconstruction-based time series anomaly detection is to identify the set of anomalous subsequences for each time series i , where is a contiguous time series that appears to deviate from the expected behavior, and i is the index of the parameter time series of interest. The data received by the anomaly detection module ADM from the measurement tool can be divided into multiple subsequences (training samples), e.g., using a sliding window with a pre-fixed window size and step size (e.g., step size of one). Here represents a time series window of length m T ending at time step t, and is the total number of subsequences. The model learns two mapping functions G ENC : X → Z and F DEC : Z → X by fully exploiting adversarial learning techniques. The two mapping functions can be regarded as generators. Here, X represents the input data domain and Z represents the latent domain, where the random vector z can be sampled from a standard multivariate normal distribution to represent white noise, e.g., .
[0095] Using the two generators, the input time series can be reconstructed as follows: , where represents the reconstructed time series window. In such an exemplary arrangement, the two generators are not two separate encoder-decoder networks, such as in a standard CycleGAN. Instead, the combined generators form an encoder-decoder architecture. In other words, both the encoder and the decoder are regarded as separate generators. The generator is used as a single encoder that maps the input time series into a latent space, and the generator It is a single decoder that transforms the latent vector into a time series reconstruction. A bidirectional long short-term memory network (LSTM) can be used as the base model of the generator to capture the temporal correlation of the time series distribution.
[0096] The full high-level objective differs from the full objective of CycleGAN in two aspects:
[0097] (1) CycleGAN applies the adversarial loss from the original GAN to both the generator and their associated discriminators. Instead, the described network employs two adversarial WGAN critics C X and C Z , where Wasserstein loss is used as adversarial loss. Commentator C X Prompt the generator Transform z into an output that is indistinguishable from a true time series sequence from domain X, while the critic C Z Evaluate the performance of the mapping into the latent space. Similar to the generator, a bidirectional LSTM can be used for two critics. The Wasserstein loss is less affected by gradient vanishing and mode collapse. In addition, the critic network returns a score of how real or fake the input sequence is. Therefore, the trained critic C X Can be used directly as an anomaly measure.
[0098] (2) The network only uses the forward consistency loss: for the original input time series from domain X, the reconstruction cycle must be able to transform x t Bring it back to the original time series, that is Without the adversarial loss of positive cycle consistency, it is uncertain whether the learned function can transform each specific input into the expected output. Adding a positive cycle consistency loss prevents and However, the generator does not have to satisfy the backward consistency loss, i.e. , since we are not interested in reconstructing the exact latent vector. In addition, considering that the goal is anomaly detection, the L2 norm (Euclidean norm) can be used instead of the L1 norm used in CycleGAN because it imposes a larger penalty when the observation distance is farther, thereby emphasizing the impact of outliers. The generator can be trained using an adapted cycle consistency loss between the original and reconstructed samples.
[0099] Assuming the step size of the sliding window method is one, each time point in the time series belongs to multiple input time windows. Obviously, each element of the time series has multiple reconstruction values. In order to calculate a single reconstruction value for each element, we can take the median of the multiple reconstruction values, which corresponds to a specific time point in the time series. The result is a reconstructed time series of each time series in the set. To compute the true univariate time series Instead of reconstructing The deviation between them is called the reconstruction error r(x i ), dynamic time warping (DTW) can be applied since it is robust to temporal shifts by allowing the time axis to warp. After training the network with a cycle consistency loss, the critic C can be optimized X To distinguish whether the given sample is from the training data or the generative model F with high sensitivity DEC Therefore, the commentator C X Viewed as a network that captures the distribution of the input data. Reviewer C X The output score is analogous to the confidence of the model on how sure it is that the signal it is given is real. The obtained score is therefore relevant for distinguishing anomalous sequences from normal ones, since anomalies, by definition, do not follow expected temporal patterns. Informally, if the critic returns a low score (false) given real training sample inputs, the training sample does not follow the expected behavior and can therefore be assumed to be anomalous.
[0100] For example, Calculate the critic score for each univariate training sample of C. X The critic scores for each subsequence are returned, resulting in multiple critic scores for each time point per time series, which is similar to the reconstruction error. Kernel density estimation (KDE) with a Gaussian kernel can be applied to each set of critic scores corresponding to a particular time point with the smoothed value set equal to the maximum value. The result can be a per time series Smoothed critic scores at each time point.
[0101] The two scores can be combined to obtain the final anomaly score for each time step. However, it is not straightforward to combine the critic score and the reconstruction error to obtain the anomaly score. All non-outliers will have similarly high critic scores and low reconstruction errors, while outliers will have unusually low critic scores and high reconstruction errors. Therefore, the reconstruction error and reviewer ratings By calculating their corresponding absolute z scores (Z r and Z c)(to be normalized). Given that there are no extreme values with very high reviewer scores and very low reconstruction errors by definition, large z-scores indicate high anomaly scores. The scores can be combined into a single value α(x) (anomaly score) for each time point by taking the pointwise product between two vectors α(x) = Z r (x) ⊙ Z c (x).
[0102] To determine whether the calculated anomaly scores constitute anomalies, local adaptive thresholding techniques can be applied to classify each point as an anomaly or normal. First, the anomaly scores can be passed through a smoothed moving average algorithm to obtain smoothed anomaly scores. Then, a sliding window method can be applied to the smoothed error sequence, where for each window, the current static threshold is set to a number of standard deviations (e.g., 4) from the window mean. Intuitively, thresholds for computational time windows rather than the entire sequence help identify anomalies that are normal relative to the entire data stream but are anomalies only in the context of the data around them. Finally, points with smoothed anomaly scores higher than the local threshold can be classified as anomalies, resulting in a set of anomaly subsequences for each time series: , where .
[0103] An intelligent continuous learner (CL) can be employed in the concepts disclosed herein, with the aim that the hyperparameter optimization required for the prediction network does not slow down the overlay control. Hyperparameter optimization and the model given a set of hyperparameters can both take longer than the allowed computational time, e.g., as defined by the separation between two consecutive batch exposures. Thus, a continuous learning strategy can be used to obtain the optimal model while the lithography cycle is operating and automatically update the model to change gradually.
[0104] CL can employ Bayesian optimization and a "gap and stride" strategy to continuously learn while the predictor network is operating. The Bayesian optimization can be used to create surrogate models so that the predictor model does not have to be recreated from scratch each time. The stride determines the number of time points between two consecutive training steps. The gap determines the number of time points between the start and end of hyperparameter optimization.
[0105] The learning task includes streaming data arriving in a continuous mode, and the underlying data distribution is constantly changing due to the non-stationary nature of the real-world lithography environment. Both the anomaly detection module network and the prediction network should be trained while the lithography cycle is running. For the automatic hyperparameter optimization of the prediction network, Bayesian optimization (BO) is proposed, which enables the application of incremental model learning, where the model is not recreated from scratch each time, but is automatically updated using the new information received. Thus, the network can reduce the cost associated with training the model and allow for a more gradual adaptation to data changes, which is particularly useful in a streaming setting. By selecting the next input value based on input values that have performed well in the past, BO limits the cost of evaluating the objective function. In short, Bayesian optimization finds the value that minimizes the objective function by creating a probability model (surrogate function) based on past evaluation results of the objective. The surrogate that maps input values to loss probabilities is optimized more cheaply than the objective. By applying a criterion (e.g., expected improvement) to the surrogate, the next input value to be evaluated is selected. After sufficient evaluation, the surrogate function will resemble the objective function. In the current case, the objective function can be the validation error of the prediction network given a set of hyperparameters.
[0106] Consider using the current history to start a new hyperparameter optimization. The model that existed before the start of the optimization continues to make the predictions while the optimization is running. The current optimization cannot take into account new real observations that arrive while it is running. Once the optimization terminates due to a time limit, the model is retrained using the optimally obtained hyperparameter combination for the current history (both the training set and the validation set), and the updated model takes over until the next retraining step. Immediately after the optimization is completed, a new optimization continues, including the data that arrived during the previous optimization. This allows the optimization to run for a longer time and consider more hyperparameter combinations.
[0107] The anomaly detection module may have a pre-fixed set of hyperparameters (e.g., based on expert knowledge and / or training on artificial data), and may be retrained after a set number of new observations have arrived to effectively update the model using the new information. Two different possible strategies can be defined by introducing the following two parameters: stride and gap. The stride determines the number of time points between two consecutive training steps. The gap determines the number of time points between the start and end of hyperparameter optimization. In other words, the gap determines the allowed computation time for each hyperparameter optimization defined as the number of time points. It is expected that the two variables have different effects on the model performance. There is a trade-off: if the gap is chosen to be large, the hyperparameter optimization can evaluate more combinations and will theoretically result in a lower validation error, but the history used for hyperparameter optimization is far from the prediction range of the model based on the resulting hyperparameter set. For a fixed set of hyperparameters, retraining after a large number of new observations have entered essentially uses the new information to update the model. Intuitively, a smaller stride will lead to better predictions because the model is up-to-date. The optimal values of the stride and gap can be determined empirically.
[0108] Figure 10 is a visualization of the continuous learning strategy according to an embodiment. The step of initializing Bayesian optimization Int BO can be performed while predicting the EWMA. This Bayesian optimization can be terminated Term BO after a period (dashed line) to obtain a first set of hyperparameters a 1 . The first set of hyperparameters a 1 and the predicted next stride time step can be used to train a first model TR M 1 . Meanwhile, the Bayesian optimization can continue Cont BO where the previous optimization ended (including the time points reached during the previous optimization). The same set of hyperparameters a 1 can be used to retrain the model M 1 TR M 2 to obtain a second model M 2 after the number of strides (dotted line) of observations have been received. This Bayesian optimization can be terminated Term BO’ after a second gap to obtain a second set of hyperparameters a 2 . The new hyperparameters a 2 can be used to retrain the model M 2 TR M 3 to obtain the model M 3 . These steps can be repeated as needed. It should be noted that the stride is set to half the size of the gap.
[0109] The output from the methods disclosed herein can be used in traditional APC control (such as creating new feature signatures (spatial distributions) for the next batch), and can also be directly fed back to the scanner as a method for improving scanner settings (feedback control based on the scanner). For example, if certain parameters have a direct correlation with the scanner settings, predictions for these parameters can be used to directly change the scanner settings, rather than correcting them via the APC loop.
[0110] It should be understood that the measurement data used in the methods described herein can include synthetic measurement data (alternatively or in combination with non-synthetic measurement data measured from one or more physical wafers), for example, obtained via computational lithography techniques that simulate one or more steps of a semiconductor manufacturing process.
[0111] Other embodiments of the present invention are disclosed in the following numbered list of aspects:
[0112] 1. A method for predicting a parameter of interest in a manufacturing process for manufacturing an integrated circuit, the method comprising: obtaining measurement data related to the parameter of interest; applying a first prediction sub-module to the measurement data to obtain non-anomalous prediction data; detecting anomalies in the measurement data; classifying the anomalies into systematic anomalies and non-systematic anomalies; using a first prediction strategy for the non-systematic anomalies to obtain first anomaly prediction data; using a second prediction strategy for the systematic anomalies to obtain second anomaly prediction data; wherein the first prediction strategy is different from the second prediction strategy; and combining the first anomaly prediction data and / or the second anomaly prediction data with the non-anomalous prediction data to obtain a prediction of the parameter of interest.
[0113] 2. The method according to aspect 1, wherein the first prediction sub-module comprises a machine learning model or neural network trained to predict the parameter of interest based on the measurement data.
[0114] 3. The method according to aspect 1 or 2, wherein the measurement data comprises current measurement data and historical measurement data.
[0115] 4. The method according to aspect 3, comprising performing the detection step on both the current measurement data and the historical measurement data before any processing of the measurement data according to the first prediction strategy.
[0116] 5. The method according to any of the preceding aspects, wherein the steps of detecting anomalies in the measurement data and classifying the anomalies comprise detecting and classifying the anomalies included in at least the current measurement data.
[0117] 6. The method according to any of the preceding aspects, wherein the first prediction strategy includes processing the measurement data corresponding to the non-systematic anomaly to obtain processed measurement data in which the non-systematic anomaly is smoothed, reduced or removed.
[0118] 7. The method according to aspect 6, including: applying a smoothing filter smoothly to the measurement data corresponding to the non-systematic anomaly to obtain the processed measurement data.
[0119] 8. The method according to aspect 6 or 7, wherein the first prediction strategy includes inputting the processed measurement data into the first prediction sub-module in place of the measurement data obtained in the obtaining step to obtain the first anomaly prediction data.
[0120] 9. The method according to any of the preceding aspects, wherein the second prediction strategy includes applying a second prediction sub-module only to a corresponding subset of the measurement data immediately following each of the systematic anomalies.
[0121] 10. The method according to aspect 9, wherein each of the subsets is related to a set duration following each of the systematic anomalies.
[0122] 11. The method according to aspect 10, wherein the second prediction strategy is applied only within the set duration such that the second prediction sub-module is used in place of the first prediction sub-module within the set duration.
[0123] 12. The method according to aspect 10 or 11, wherein the number of cycles of the set duration is determined empirically.
[0124] 13. The method according to any one of aspects 9 to 12, wherein the second prediction sub-module includes a low-pass time-domain filtering module and / or a prediction sub-module based on moving average.
[0125] 14. The method according to aspect 13, wherein the second prediction sub-module includes a prediction sub-module based on exponentially weighted moving average.
[0126] 15. The method according to aspect 13, wherein the second prediction sub-module includes a neural network trained to predict a parameter of interest based on the subset of the measurement data.
[0127] 16. The method according to any of the preceding aspects, including re-training the first prediction sub-module using newly arrived measurement data while applying the second prediction strategy.
[0128] 17. The method according to any of the foregoing aspects includes identifying a systematic anomaly as a sudden change in the measurement data that occurs simultaneously in multiple time series of the measurement data.
[0129] 18. The method according to any of the foregoing aspects includes identifying a non-systematic anomaly as a sudden change in the measurement data that does not occur simultaneously in multiple time series of the measurement data.
[0130] 19. The method according to any of the foregoing aspects includes: identifying a detected anomaly where multiple elements of a single observation within the measurement data at a specific time point are anomalous as a systematic anomaly; and identifying all other detected anomalies as non-systematic anomalies.
[0131] 20. The method according to any of the foregoing aspects, wherein the step of detecting anomalies in the measurement data is performed using an adversarial machine learning model.
[0132] 21. The method according to aspect 20, wherein the adversarial machine learning model includes a machine learning model based on a generative adversarial network.
[0133] 22. The method according to aspect 20, wherein the adversarial machine learning model includes a machine learning model based on a recurrent generative adversarial network.
[0134] 23. The method according to aspect 20, wherein the adversarial machine learning model includes a recurrent-consistent Wasserstein generative adversarial network with gradient penalty.
[0135] 24. The method according to any one of aspects 20 to 23, wherein the adversarial machine learning model includes: a first mapping function that can operate to map from a measurement data domain to a latent domain; and a second mapping function that can operate to map from the latent domain to the measurement data domain; a first critic that can operate to prompt the first mapping function to convert a noise representation into data indistinguishable from the measurement data; and a second critic that can operate to evaluate the performance of mapping to the latent domain.
[0136] 25. The method according to aspect 24, wherein anomalies are detected based on a critic score associated with the first critic and / or a reconstruction error score of the adversarial machine learning model.
[0137] 26. The method according to any one of aspects 20 to 25 includes continuously training the adversarial machine learning model via automatic hyperparameter optimization of the adversarial machine learning model.
[0138] 27. The method according to any of the foregoing aspects includes continuously training the first prediction sub-module via automatic hyperparameter optimization of the first prediction sub-module.
[0139] 28. The method according to aspect 26 or 27, wherein the hyperparameter optimization includes Bayesian optimization.
[0140] 29. The method according to any one of aspects 26 to 28, wherein the hyperparameter optimization employs a gap and stride strategy, wherein the stride determines the number of time points between two consecutive training steps, and the gap determines the number of time points between the start and end of the hyperparameter optimization.
[0141] 30. The method according to any one of aspects 26 to 29, wherein the automatic hyperparameter optimization includes routinely performing the following iterations: hyperparameter optimization, which is performed based on all the acquired measurement data to obtain an optimized set of hyperparameters; and model retraining, which is performed based on the latest optimized set of hyperparameters.
[0142] 31. The method according to any of the foregoing aspects, wherein the measurement data includes one or more of the following: simulated measurement data; pre-exposure measurement data; post-exposure measurement data; and / or context data.
[0143] 32. The method according to any of the foregoing aspects includes determining a correction to the manufacturing process from the prediction of the parameter of interest.
[0144] 33. The method according to aspect 32, wherein the determining the correction includes determining a spatial distribution of corrections for subsequent substrates or substrate batches.
[0145] 34. The method according to aspect 32 or 33, wherein the determining the correction includes determining a correction to the settings of the equipment used in the manufacturing process.
[0146] 35. The method according to any one of aspects 32 to 34 includes manufacturing additional integrated circuits using the correction.
[0147] 36. The method according to any of the foregoing aspects includes routinely measuring a substrate to obtain the measurement data.
[0148] 37. A computer program comprising program instructions operable to perform the method according to any one of aspects 1 to 35 when run on a suitable device.
[0149] 38. A non-transitory computer program carrier comprising the computer program according to aspect 37.
[0150] 39. A processing system includes a processor and a storage device including a computer program according to aspect 37.
[0151] 40. A lithographic apparatus arrangement includes: a lithographic exposure apparatus; and a processing system according to aspect 39.
[0152] 41. A lithography cell includes: a lithographic apparatus arrangement according to aspect 40; and a metrology device, the metrology device including a processing system according to aspect 39 and further operable to perform the method according to aspect 36.
[0153] Although specific reference may be made herein to the use of a lithographic apparatus in the manufacture of ICs, it should be understood that the lithographic apparatus described herein may have other applications. Other possible applications include the manufacture of integrated optical systems, the guidance and detection of magnetic domain memories, flat panel displays, liquid crystal displays (LCDs), thin film magnetic heads, and the like.
[0154] Although embodiments of the invention may be specifically referred to herein in the context of a lithographic apparatus, embodiments of the invention may be used in other apparatuses. Embodiments of the invention may form part of a mask inspection apparatus, a metrology apparatus, or any apparatus for measuring or processing an object such as a wafer (or other substrate) or a mask (or other patterning device). Such apparatuses may generally be referred to as lithographic tools. Such lithographic tools may operate under vacuum conditions or ambient (non-vacuum) conditions.
[0155] Although specific reference has been made above to the use of embodiments of the invention in the context of optical lithography, it should be appreciated that, where the context allows, the invention is not limited to optical lithography and may be used in other applications, such as imprint lithography.
[0156] Although specific embodiments of the invention have been described above, it should be understood that the invention may be practiced in other ways different from those described. The above description is intended to be illustrative rather than restrictive. Accordingly, those skilled in the art will appreciate that the invention described may be modified without departing from the scope of the claims set forth below.
Claims
1. A method for predicting a parameter of interest in a manufacturing process for manufacturing an integrated circuit, the method comprises: obtaining measurement data related to the parameter of interest; applying a first prediction sub-module to the measurement data to obtain non-anomaly prediction data; detecting anomalies in the measurement data; classifying the anomalies into systematic anomalies and non-systematic anomalies; using a first prediction strategy for the non-systematic anomalies to obtain first anomaly prediction data; using a second prediction strategy for the systematic anomalies to obtain second anomaly prediction data; wherein the first prediction strategy is different from the second prediction strategy; and combining the first anomaly prediction data and / or the second anomaly prediction data with the non-anomaly prediction data to obtain a prediction of the parameter of interest.
2. The method according to claim 1, wherein, the first prediction sub-module includes a machine learning model or a neural network trained to predict the parameter of interest based on the measurement data.
3. The method according to claim 1, wherein, the measurement data includes current measurement data and historical measurement data.
4. The method according to claim 3, comprising performing the detection step on both the current measurement data and the historical measurement data before any processing of the measurement data according to the first prediction strategy.
5. The method according to claim 1, wherein, the first prediction strategy includes processing the measurement data corresponding to the non-systematic anomalies to obtain processed measurement data in which the non-systematic anomalies are smoothed, reduced or removed.
6. The method according to claim 5, comprises: applying a smoothing filter to the measurement data corresponding to the non-systematic anomalies to obtain the processed measurement data.
7. The method according to claim 5, wherein, the first prediction strategy includes inputting the processed measurement data into the first prediction sub-module in place of the measurement data obtained in the obtaining step to obtain the first anomaly prediction data.
8. The method according to claim 1, wherein, the second prediction strategy includes applying a second prediction sub-module only to a corresponding subset of the measurement data immediately following each of the systematic anomalies.
9. The method according to claim 8, wherein, each subset is related to a set duration following each of the systematic anomalies.
10. The method according to claim 9, wherein, the second prediction strategy is applied only within the set duration, such that the second prediction sub-module is used in place of the first prediction sub-module within the set duration.
11. The method according to claim 8, wherein, the second prediction sub-module includes a low-pass time-domain filtering module and / or a prediction sub-module based on moving average.
12. The method according to claim 1, comprises: identifying a systematic anomaly as a sudden change in the measurement data that occurs simultaneously in multiple time series of the measurement data; and identifying non-systematic anomalies as sudden changes in the measurement data that do not occur simultaneously in multiple time series of the measurement data.
13. The method according to claim 1, comprising: identifying the detected anomalies where multiple elements of a single observation within the measurement data at a particular time point are anomalies as systematic anomalies; and identifying all other detected anomalies as non-systematic anomalies.
14. The method according to claim 1, wherein, an adversarial machine learning model such as, for example, a machine learning model based on a generative adversarial network is used to perform the step of detecting anomalies in the measurement data.
15. A computer program comprising program instructions operable to perform the method according to any one of the preceding claims when run on a suitable device.
Citation Information
Patent Citations
Semiconductor etching apparatus
US20060016561A1
Density-aware dynamic leveling in scanning exposure systems
US20070085991A1
Alignment System and Alignment Marks for Use Therewith
US20100214550A1
Method Of Calibrating A Lithographic Apparatus, Device Manufacturing Method and Associated Data Processing Apparatus and Computer Program Product
US20120008127A1
Method for obtaining training data for training a model of a semiconductor manufacturing process
US20210405544A1
Cited By
Self-adaptive optimization and collaborative decision-making system for spin-coating process of large-diameter substrate
CN121613736A
A large-diameter substrate spin coating process adaptive optimization and collaborative decision system
CN121613736B