Laser maintenance using tree-based machine learning models incorporating target-based functions
A decision tree-based maintenance process using a hybrid cost function optimizes maintenance timing for light sources in semiconductor photolithography, addressing the challenge of inefficient maintenance scheduling and improving productivity and reliability.
Patent Information
- Application Number
- PCT/IB2024/061732
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-11-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing light sources used in semiconductor photolithography face challenges in determining the optimal time for maintenance, leading to either too frequent or too late maintenance, which affects productivity and product quality.
A decision tree-based maintenance process using a hybrid cost function that combines increasing orderliness within subsets and closeness to target parameters to classify light sources or modules as needing maintenance, allowing for timely and efficient maintenance decisions.
This approach optimizes maintenance timing, reducing downtime and improving productivity by predicting maintenance needs accurately, thus enhancing the reliability and efficiency of light sources in photolithographic processes.
Smart Images

Figure IB2024061732_03072025_PF_FP_ABST
Abstract
Description
LASER MAINTENANCE USING TREE-BASED MACHINE LEARNING MODELSINCORPORATING TARGET-BASED COST FUNCTIONSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to US Application No. 63 / 615,063, filed December 27, 2023, titled LASER MAINTENANCE USING TREE-BASED MACHINE LEARNING MODELS INCORPORATING TARGET-BASED COST FUNCTIONS, which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] The disclosed subject matter relates to maintenance of light sources such as those used for integrated circuit photolithographic manufacturing processes including deep ultraviolet (DUV) light sources.BACKGROUND
[0003] Light that is used for semiconductor photolithography, which can be in the form of laser radiation, is typically supplied by a system referred to herein as a light source. These light sources produce radiation as a series of pulses at specified repetition rates, for example, in the range of about 500 Hz to about 6 kHz. Additionally, such light sources conventionally have expected useful lifetimes measured in terms of the number of pulses they are able to produce before requiring repair or replacement or major maintenance, typically expressed as billions of pulses.
[0004] One system for generating light such as laser radiation at frequencies useful for semiconductor photolithography (such as at deep-ultraviolet (DUV) wavelengths) involves use of a master oscillator power amplifier (MOPA) dual-gas-discharge-chamber configuration. This configuration has two chambers, a master oscillator chamber (MO chamber) and a power amplifier chamber (PA chamber). These chambers and many other system components can be regarded as modules, and the light source overall can be regarded as an ensemble of modules. Each module in general has a lifetime that is shorter than the lifetime of the overall light source. Thus, over the course of the lifetime of the light source, the health of the light source and the health of individual modules can be evaluated to determine whether specific modules should be repaired or replaced, and modules or the light source as a whole can be repaired or replaced in accordance with such evaluation.SUMMARY
[0005] In some general aspects, a process for maintaining light sources including laser systems or modules of laser systems includes obtaining one or more features of the light sources, the one or more features representing one or more operational aspects of the light sources stored overtime; generating a decision tree, the tree having one or more levels, for classifying the light sources or modules thereof,based on the one or more features, as either (1) no technical issue or (2) maintenance needed, the generating using the one or more features of the light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function including a combination of (1) a first cost function minimized with increasing order (or orderliness) within resulting subsets and (2) one or more second cost functions minimized with increasing closeness to respective target parameters; for a given light source, applying the decision tree to current or recent features of the given light source not within the one or more features previously obtained to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed; and performing maintenance on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed.
[0006] Implementations can include one or more of the following. The features can include one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltage. The hybrid cost function can be minimized over a split of a single feature. The hybrid cost function can be minimized over splits of two or more features. The features can consist of electrode voltage and the first cost function can be minimized over a split of electrode voltage. The first cost function can be minimized with increasing information gain or decreasing impurity of resulting subsets. The first cost function can be minimized with increasing information gain.
[0007] The second cost function can have a number of lost pulses of the light source upon a maintenance needed classification as a target parameter. The second cost function can have a number of lost pulses of the light source upon a maintenance needed classification as the only target parameter of the second cost function. The hybrid cost function can be a weighted combination of the first cost function and the second cost function, without a relative ranking. The hybrid cost function can be a weighted combination of a rank value of the first cost function and a rank value of the second cost function.
[0008] The first cost function for a given split of a parent set into one or more child sets can be given by the entropy of the parent set minus the weighted average entropy of the child sets. The first cost function for a given split of a parent set into one or more child sets can thus be given byV lslIG = £(d) - ) ^-E(s) |d|
[0009] where E(d) is the entropy of the parent set d, E(s) is the entropy of a respective child set s, |x| is the cardinality of a set x, and E(y) is entropy of a set y given by
[0010] wherein C is the number of classes of a set y and , is the probability of class i (in other words, where p, is the probability of randomly picking an element of class i from set y (i.e. the proportion of the scty made up of class i). Alternatively, the first cost function for a given split of a parent set into one or more child sets can be given by a weighted sum of each child set of
[0011] where p, denotes the probability of a respective element of a child set being randomly classified correctly relying only on the distribution of classes in the child set. Other representations of a decrease in impurity or an increase in orderliness and / or of information gain can also be used.
[0012] The second cost function for the split can be given by\TV- v\
[0013] where Tvis a target of a variable v and v is average of the variable v for the split.
[0014] The second cost function for the split can be given by\TLP- LP\
[0015] where TLP is a target number of lost pulses and LP is the average number of lost pulses for the split. TLP can be in the range of 1 to 4 billion or in the range of 1 to 2 billion.
[0016] The hybrid cost function can be given by
[0017] where w, are n weighting parameters, is the summation of the n weighting parameters, Rank(O) is the relative rank of a given split, among all splits, of the order (or orderliness) of a given split, and Rank | Tv — | is the relative rank of the given split of the average of the variablerelative to its target.
[0018] The hybrid cost function can be given by
[0019] wherein H’ / is a first weighting value, w2is a second weighting value, Rank(IG) is the relative rank among all splits of the information gain of the given split, and Rank(LP) is the relative rank among all splits of the average lost pulses of the given split.
[0020] In additional aspects, a process for generating a decision tree for classifying states of light sources including laser systems or modules of laser systems as having no technical issue or requiring maintenance can include obtaining features of the light sources, the features representing operational aspects of the light sources stored over time; generating a decision tree based on the stored features for classifying the light sources or modules thereof as having either (1) no technical issue or (2) maintenance needed, the tree having one or more levels, the generating using the features of the light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function including a combination of (1) a first cost function minimized with increasing order within resulting subsets and (2) one or more second cost functions minimized with closeness to target parameters.
[0021] Implementations can include one or more of the following. The first cost function can be minimized with increasing information gain or decreasing impurity of resulting subsets, the second cost function can include a target parameter of lost pulses within the range of 1 to 4 billion lost pulses, and the hybrid cost function can be a weighted combination of the first cost function and the second cost function without relative ranking, or of a rank value of the first cost function and a rank value of the second cost function.
[0022] In still more aspects, a process for classifying states of light sources including laser systems or modules of laser systems can include acquiring current features of a given light source; and applying a decision tree to the current features of the given light source to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed, the decision tree having one or more levels, wherein the decision tree is a decision tree generated using a hybrid cost function including a combination of (1) a first cost function minimized with increasing information gain or decreasing impurity and (2) one or more second cost functions minimized with closeness to target parameters.
[0023] Implementations can include one or more of the following. The first cost function can be minimized with increasing information gain or decreasing impurity of resulting subsets over splits of one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltage, the second cost function can include a target parameter of lost pulses within the range of 1 to 4 billion lost pulses, and the hybrid cost function can be a weighted combination of the first cost function and the second cost function or of a rank value of the first cost function and a rank value of the second cost function. The process can further include performing maintenance on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed.
[0024] In another aspect, a light source for use with a photolithography apparatus is provided, the light source including first and second modules and a sensing and evaluation apparatus in communication with the first and second modules and / or additional parts of the light source, wherein the sensing and evaluation apparatus is configured to receive and store features of the first and second modules and / or the light source relating to performance of the first and second modules and / or the light source and to apply a decision tree to the features to classify the light source, the first module, and / or the second module as having either (1) no technical issue or (2) maintenance needed The decision tree has one or more levels and is a decision tree generated using a hybrid cost function including a combination of (1) a first cost function minimized with increasing orderliness of resulting subsets and (2) one or more second cost functions minimized with closeness to respective target parameters.
[0025] Implementations can include one or more of the following. The first cost function can be a cost function minimized with increasing information gain or decreasing impurity of resulting subsets. The second cost function can have a number of lost pulses of the light source, upon a maintenance needed classification, as a target parameter. The second cost function can have an amount of lost productive time, of the light source upon a maintenance needed classification, as a target parameter.
[0026] In still another aspect, a processing module configured for processing historical data of multiple light sources used in photolithography is provided, the processing module including data storage configured to receive and store data including performance data of the multiple light sources and / or of modules thereof and one or more processors configured to generate, based on the stored data, a decision tree classifying the light sources or modules thereof as having either (1) no technical issue or (2) maintenance needed. The tree has one or more levels, and the generating uses the stored data to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function. The hybrid cost function includes a combination of (1) a first cost function minimized with increasing order within resulting subsets and (2) one or more respective second cost functions minimized with closeness to respective target parameters.
[0027] Implementations can include one or more of the following. The first cost function can be a cost function minimized with increasing information gain or decreasing impurity of resulting subsets. The second cost function can have a number of lost pulses of the light source, upon a maintenance needed classification, as a target parameter. The second cost function can have an amount of lost productive time, of the light source upon a maintenance needed classification, as a target parameter.
[0028] In yet another aspect, a sensing and evaluation apparatus for use with respective one or more light sources used in photolithography is provided, the apparatus including (a) data storage configured to receive and store respective data including performance data of the respective one or more light sources and / or respective modules thereof and (b) one or more processors configured to apply a decision-tree machine learning model to the respective data to classify the respective one or more light sources and / or one or more respective modules thereof as having either (1) no technical issue or(2) maintenance needed, wherein the decision-tree machine learning model comprises one or more levels generated by using historical data including performance data from multiple light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function including a combination of (1) a first cost function minimized with increasing order within resulting subsets and (2) one or more second cost functions minimized with closeness to target parameters.
[0029] Implementations can include one or more of the following. The first cost function can be a cost function minimized with increasing information gain or decreasing impurity of resulting subsets. The second cost function can have a number of lost pulses of the light source, upon a maintenance needed classification, as a target parameter. The second cost function can have an amount of lost productive time, of the light source upon a maintenance needed classification, as a target parameter.
[0030] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.DRAWING DESCRIPTION
[0031] FIG. 1 is a schematic diagram of aspects of a light source.
[0032] FIG. 2 is a schematic diagram of aspects of a lithography exposure apparatus that can be used with a light source such as the light source of FIG. 1.
[0033] FIG. 3 is a flow chart of steps of a process of maintaining a light source.
[0034] FIG. 3 A is a diagram of an example decision tree structure.
[0035] FIG. 4 is a flow chart of a process of generating a decision tree that can be used in the process of FIG. 3.
[0036] FIG. 5 is a schematic diagram of one or more light sources together with a sensing and evaluation apparatus.
[0037] FIG. 6 is a schematic diagram of multiple light sources together with a processing module.DETAILED DESCRIPTION
[0038] Referring to FIG. 1, a light source 100 in the form of a deep UV (DUV) light source 100 is shown. The light source 100 can be in the form of a dual stage pulsed light source that produces as a light beam a pulsed amplified light beam 105. The light source 100 can include a solid state or gas discharge master oscillator (MO) system 160, a power amplification (PA) system such as a power ring amplifier (PRA) system 165, relay optics 170, and an optical output subsystem 175.
[0039] The MO system 160 can include, for example, an MO chamber module 161, in which electrical discharges between electrodes (not shown) can cause lasing gas discharges in a lasing gas to create an inverted population of high energy molecules or dimers or exciplexes. Gases employed can include argon, krypton, xenon, or mixtures thereof, typically combined with fluorine or chlorine toform the excited states. Radiation emitted by the inverted population can produce relatively broad band radiation which is then center-wavelength-selected and line-narrowed to a relatively very narrow bandwidth in a line narrowing module (‘LNM’) 162. The MO system 160 can also include an MO output coupler (MO OC) 164, which can include a partially reflective mirror (not shown). The partially reflective mirror in the MO OC 164 can form, with a reflective grating (not shown) in the LNM 162, an oscillator cavity. The MO system 160 can also include a line-center analysis module (LAM) 163. The LAM 180 can include, for example, an etalon spectrometer for fine wavelength measurement and a coarser resolution grating spectrometer.
[0040] The relay optics 170 can include an MO wavefront engineering box (WEB) 171 that serves to redirect the output of the MO system 160 toward the PA system 165. The relay optics 160 can also include, for example, beam expansion capability such as, for example, a multi-prism beam expander (not shown), and coherence-busting capability, for example, in the form of an optical delay path (not shown).
[0041] The PA system 165 includes a PRA chamber module 166, for receiving an output light beam from the MO system 160 and propagating it as a light beam within the PA system 165. The PA system 165 can further include output coupling optics incorporated into a PRA WEB 167. The light beam within the PA system 165 can be redirected back through a gain medium in the chamber module 166 by way of a beam reverser 168. The PRA WEB 167 can incorporate a partially reflective input / output coupler (not shown) and a maximally reflective mirror for the nominal operating wavelength (which can be at around 193 nm, for example, for an ArF system) and one or more prisms. The PA system 165 operates to optically amplify the output light beam from the MO system 160.
[0042] The optical output subsystem 175 can include a bandwidth analysis module (BAM) 176 adj acent to the output of the PA system 165. The BAM 176 can receive the output light beam of pulses from the PA system 165 and pick off or split off a portion of the light beam for metrology purposes, such as to measure the output bandwidth and pulse energy, for example. The output light beam of pulses then passes through an optical pulse stretcher module (OPuS) 177 and an output combined autoshutter metrology module (CASMM) 178. The CASMM 178 can include a pulse energy meter. One purpose of the OPuS 177 can be to convert a single output pulse into a pulse train. Secondary pulses created from the original single output pulse can be delayed with respect to each other. By distributing the original laser pulse energy into a train of secondary pulses, the effective pulse length of the light beam can be expanded and at the same time the peak pulse intensity reduced.
[0043] The light source 100 is made up of modules. Each of the components (such as the MO chamber 161, the LNM 162, the MO WEB 171, the PRA chamber 166, the PRA WEB 167, the OPuS 177, the BAM 176) of the light source 100 are modules. The overall availability of the light source 100 is the direct result of the respective availabilities of these individual modules making up the light source 100. In other words, the light source 100 cannot be available unless all of these modulesmaking up the light source 100 are available. A sensing and evaluation apparatus 120 can monitor these modules so that they can be adjusted, refreshed, or replaced, generally before they fail, in order to maintain the operation of the light source 100 and optimize and improve productivity of an associated photolithography apparatus 210 (FIG. 2, described below). The sensing and evaluation apparatus 120 can record and store overtime features of the light source and / or of the modules in the form of information relating to performance and other properties of the modules and / or of the light source as a whole. The sensing and evaluation apparatus can evaluate the condition or status of the light source and / or one or more modules of the light source and provide maintenance alerts that can be used to perform automated and / or manual maintenance tasks for one or more specific modules. These maintenance tasks can include replacement tasks when needed.
[0044] Referring to FIG. 2, the amplified light beam 105 (FIG. 1) can be a light beam 205 used by a photolithography exposure apparatus 210 to pattern features on a substrate or wafer 211. The wafer 211 is placed on a wafer table 212 structured to hold the wafer 211 and connected to a positioner configured to position the wafer 211 accurately in accordance with certain parameters. The light beam 205 can have a wavelength in the deep ultraviolet (DUV) range, which can include wavelengths from, for example, about 100 nanometers (nm) to about 400 nm. The light source 100 can be an excimer light source. Thus, for example, the gain medium can include argon fluoride (ArF), krypton fluoride (KrF), or xenon chloride (XeCl). If the gain medium includes argon fluoride, then the wavelength of the amplified light beam 205 is typically about 193 nm. If the gain medium includes krypton fluoride, then the wavelength of the amplified light beam 205 is typically about 248 nm. The size of the microelectronic features patterned on the wafer 211 depends on the wavelength of the light beam 205, with a lower wavelength resulting in a smaller minimum feature size. When the wavelength of the light beam 205 is about 248 nm or about 193 nm, the minimum size of the microelectronic features can be, for example, 50 nm or less. The bandwidth of the light beam 205 can be the actual, instantaneous bandwidth of its optical spectrum (or emission spectrum), including information on how the optical energy of the light beam 205 is distributed over different wavelengths.
[0045] The photolithography exposure apparatus 210 includes an optical arrangement having, for example, one or more condenser lenses, a mask, and an objective arrangement. The mask is movable along one or more directions, such as along an optical axis of the light beam 205 or in a plane that is perpendicular to the optical axis. The objective arrangement includes a projection lens and enables an image transfer to occur from the mask to the photoresist on the wafer 211. The photolithography exposure apparatus 210 also includes an illumination system that adjusts the range of angles for the light beam 205 impinging on the mask. The illumination system also homogenizes (makes uniform) the intensity distribution of the light beam 205 across the mask.
[0046] The photolithography exposure apparatus 210 can also include, among other features, a lithography controller 213 that controls how patterns are irradiated at or on the wafer 211. The lithography controller 213 can include a memory that stores information such as process recipes. Aprocess program or recipe determines the length of the exposure on the wafer 211, the mask used, and other factors that affect the exposure. The lithography controller 213, if desired, can communicate with a sensing and evaluation apparatus of an associated light source, such as the sensing and evaluation apparatus 120 of the light source 100 of FIG. 1. Information related to the performance of the light source 100, and / or of the light source 100 together with the lithography apparatus 210, can also be recorded and stored by the sensing and evaluation apparatus 120.
[0047] The quality of the features produced on the wafer 211 by the photolithography exposure apparatus 210 depends directly upon the quality and reliability of the light pulses from the light source 100. Pulses having lower than desired power can result in underexposure of an area of the wafer 211. Missing pulses can similarly result in underexposure. Shifts in wavelength or bandwidth distribution can result in shifts in image position at the wafer 211, with resulting alterations in patterns produced at the wafer 211.
[0048] Information or data from various sources within a light source such as light source 100, and optionally information from an associated lithography apparatus such as lithography apparatus 210, can be used to assess the need for maintenance or replacement of the entire light source 100 or of a module within the light source 100 such as the MO chamber module 161 or in the PRA chamber module 166, for example. The information used can include current, real-time or near real-time information, as well as information stored over some preceding time period. Information gathered over time can be used, or example, to assess trends, central measures, variability, and / or other statistical measures of performance.
[0049] ft is desirable to perform maintenance such as replacement of a module at a suitable time. Too early maintenance is too frequent maintenance and reduces the fraction of the time that the light source is available for productive operation. Too late maintenance can potentially result in improperly processed product — product that may in some cases have to be scrapped, resulting in significant losses from high production costs in preceding production steps. Too late maintenance can also result in maintenance being performed without any time flexibility, that is, at inefficient or otherwise inopportune times.
[0050] Machine learning can be applied to determine a suitable time for maintenance to be performed. Data from various parts of the light source, additionally from an associated lithography apparatus, if desired, can be used by a machine learning model to determine a reasonably ideal time for maintenance of a light source module. Replacement of the MO chamber module 161 and / or the PRA chamber module 166 is required periodically, and thus a benefit can be obtained by performing module replacement close to but before a module fails, for optimizing “up-time” and productivity for a user of the light source.
[0051] FIG. 3 is a flow chart of steps of a process 330 for maintaining a light source, such as a light source similar to light source 100 of FIG. 1.
[0052] As shown in FIG. 3, the process 330 includes obtaining one or more features representing one or more operational aspects of multiple light sources over time (331). The features can include recorded performance data or monitoring data from multiple light sources of the same type of similar types. Features can include one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltages of the light sources, as well as other data, such as the model and / or type of light source, the software and / or firmware version of the light source, the age and other status aspects of the light source and / or various light source modules and / or components, all recorded overtime. The features can be recorded at a fixed interval, for example, once per day. Features at each fixed interval can then be labeled or grouped as representing either (1) a light source or light source module having no technical issue, or (2) a light source or light source module having maintenance required, such as maintenance required (or predicted to be required) within a given time period, such as 14 days, or within a given number of pulses produced, such as a given number of pulses in the range of 1 to 4 billion pulses. This labeling or grouping can be accomplished by labeling, as data indicating maintenance is required, the recorded performance or monitoring data occurring within the given time period, or within the given number of pulses, from a removal from service of the corresponding light source or module.
[0053] The machine learning task then becomes to “learn” from the stored and labeled features to reasonably accurately recognize or predict when maintenance is required for a given light source or module, given the current and / or recent features (the performance or monitoring data or the like, which is unlabeled) of the given light source. In the process 330 of FIG. 3, this is accomplished by generating a decision tree (332).
[0054] FIG. 3 A is a diagram of a decision tree 380. With reference to FIG. 3 A, the term “tree” or “decision tree” as understood in machine learning can be described as a process structure in which one ormore levels of testing, such as levels 382a-382c of the tree 380 of FIG. 3A, are applied to classify a given instance / set of data.
[0055] The tree processing structure can be thought of as having similarities to a classification tree such as is sometimes used in life sciences to identify or classify an organism, for example. “Nodes” or “splits” 384a, 384b, 384c correspond to forks in the tree structure and represent tests in the form of expressions to be evaluated to determine which “branch” of the tree to follow, for the given instance, to the next level of the tree, if any. The initial node 384a may be referred to as the root node, with additional nodes or splits 384b, 384c referred to as interior nodes. “Leaves” or “leaf nodes” 386 are not splits, but rather represent the final outcomes or categories produced by application of the tree 380. In a variation, below the first level 382a, not every level of the tree needs to be fully populated with interior nodes. In another variation, a decision tree can have just one level. Such a decision tree can sometimes be referred to as a “trunk” or “decision trunk.”
[0056] As further shown in item 332 of FIG. 3, the decision tree is generated by using a feature of the one or more features of the light sources to determine a split or splits for each of the one or morelevels of the tree by minimizing a hybrid cost function. The hybrid cost function comprises a combination of (1) a first cost function minimized with increasing order (or orderliness) within resulting subsets and (2) one or more second cost functions minimized with increasing closeness to respective target parameters (332). In other words, “training data” such as a set of historical data from light sources of the desired type is used to generate (create or “train”) a decision tree, generating one or more levels each having one or more splits. In principle, this proceeds by choosing, for each node generated, a feature within the training or historical data and a value or expression based on the feature, on which to “split” or divide the universe of instances of training or historical data present at that node. A given node or split (if binary, as is typical) thus effectively produces two subsets of the total data present at the node. Three-way and higher number splits can also be used if desired, producing three or more subsets. In the case of a tree with multiple levels, the subsets may be further split, based on other features, in lower levels of the tree.
[0057] The first cost function measures the “cost” of increasing “disorder” of increasingly less ideal division between (1) instances of feature sets representing modules that need or are soon to need maintenance (as reflected by the historical data) (“maintenance needed”) and (2) instances of features sets of modules that do not need and are not soon to need maintenance (“no technical issue”). Minimizing the first cost function thus finds a preferred split, i.e., an optimized division into sets of “maintenance needed” and “no technical issue” for the given training or historical data at the given node. The second cost function measures closeness to one or more respective target parameters unrelated (or less related) to distinguishing between “maintenance needed” and “no technical issue.” Minimizing the second cost function allows the decision tree to optimize for such target parameters. The target parameters can relate to economic considerations of the timing of maintenance actions. Minimizing the hybrid cost function including both the first and second cost functions allows generation of a compromise decision tree, with a split or splits that result both in relatively low disorder in separating “maintenance needed” states from “no technical issue” states and in better target parameter performance, or better economics of maintenance decisions or maintenance timing, than without the hybrid cost function.
[0058] As further shown in FIG. 3, the decision tree is then applied to current or recent features of the given light source (features that are current or near-current and not in the set used to generate the tree) to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed (333). Then maintenance is performed on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed (334).
[0059] As mentioned, the one or more target parameters of the one or more second cost functions (332 of FIG. 3) can include factors that are related to performance metrics not directly tied to the predictive power of the split regarding whether maintenance is needed. Such factors can include performance metrics that have economic impacts in or on the operation or use of the light source. Forexample, the number of “lost pulses” or alternatively “lost days” of good operation, known retrospectively from the stored, previously collected features of the light sources, represents a direct economic impact relating to availability or “uptime” of a given light source. A target of a specific number of lost pulses, or of a specific number of lost days can thus be used as a target of a second cost function and combined with the first cost function to bias the resulting hybrid cost function toward limiting the amount of lost production time and / or capability. For instance, a target time or number of days, such as a number of days in the range of from 3 to 21, such as 3, 5, 7, 14, 21, or some other value of days or time before failure (when maintenance is absolutely required) can help bias the split or splits in the decision tree away from excessively large amounts of lost good operation time, while simultaneously biasing the split or splits in the decision tree away from calling for maintenance too close in time to actual failure (or even after). Similarly, a target of a number of lost pulses in the range of 1 to 4 billion lost pulses, such as 1, 2, 3, or 4 billion lost pulses, can help bias the decision tree away from too large numbers of lost pulses, as well as away from calling for maintenance too close to failure. In other words, the hybrid cost function, by using a lost pulses or lost days target or both, can decrease the impact of lost pulses and / or lost days that would result from a split based only on relative order (orderliness) of a split with respect to maintenance requirements, as well as increase the margin of safety between signaling maintenance is needed and actual failure. Two or more second cost functions can be used, if desired, with a corresponding number of targets. Both days and lost pulses could be used as targets for two second cost functions, for example. Other second cost functions are not limited to but can include, for example, (1) cost functions that bias toward the triggering of “maintenance needed” at economically advantageous times such as at scheduled shut-downs or other planned otherwise economically practical breaks in production, such as when maintenance technicians are available, (2) cost functions that bias toward the triggering of “maintenance needed” for multiple modules or other components together to avoid or reduce separate, individual module- or componentspecific downtime intervals, and (3) cost functions that bias away from the simultaneous or overlapping triggering of “maintenance needed” for particular pairs or combinations of modules or components, if any, that are desirably replaced or serviced and brought back up individually rather than simultaneously. Such pairs or combinations of modules or components can include, for example, those that are desirably serviced and returned to good performance individually in order to help preserve light source system stability by limiting the number of simultaneous interacting changes.
[0060] FIG. 4 is a flow chart showing an implementation of a process 440 for generating a decision tree, such as generating the decision tree of box 332 of FIG. 3. As seen in FIG. 4, the process 440 can include setting a level counter to a first level of the tree (441), then selecting a variable from the stored features (i.e., from the training data set) on which to select splits for evaluation (442). Multiple variables can be evaluated to find a best variable and a best split on that variable, or the variable for the current level can be user-directed, such as by a subject matter expert, if desired (442). For a given variable at a given level, the values of the variable are ordered and the midpoint between each pair inthe order is taken as a split to be evaluated (443). For each split to be evaluated, the first cost function and the one or more second cost functions and the hybrid cost function in the form of a weighted sum of the first and the one or more second cost functions are evaluated (444) until all splits previously found (443) have been evaluated (445). When all splits have been evaluated (445), if the variable selection for the current level of the tree was directed (446, “yes” branch), then the split (on the directed variable) with the lowest cost function is set as the split for the current level (447). If no more levels are to be evaluated (448, “no” branch) (such as if a preset maximum level has been reached, or such as if a single level or “trunk” model is being used, having only a root node), the process 400 ends. If more levels are to be evaluated (448, “yes” branch), the process 400 sets the next level for which a split is to be determined (451) and returns to select a next candidate variable, or next directed variable if a directed variable is to be used (442), for which to evaluate potential splits. The process 400 then repeats as described above.
[0061] If, however, for a first or subsequent level, a variable selection was not directed (446, “no” branch), and if there are more variables to be evaluated at the current level (449, “yes” branch), then the process 400 returns to select a next candidate variable (442) and evaluate all splits of that next candidate variable (443, 444, and 445). This loop repeats until there are no more variables to be evaluated at the current level (449, “no” branch), then the variable and split of the variable with the lowest hybrid cost function over all candidate variables at the current level is set as the split for the current level (450). If more levels are to be evaluated (448, “yes” branch), the process 400 loops as described above, first setting a next level (451). Also as described above, if no more levels are to be evaluated (448, “no” branch), the process 400 ends.
[0062] The hybrid cost function with one or more second cost functions can be given by
[0063] where w, are n weighting parameters, £ is the summation of the n weighting parameters, Rank(O) is the relative rank of a given split, among all splits, of the order (or orderliness) of a given split, and Rank | Tvt— | is the relative rank for the given split of the distance of the average vLof the variable relative to its target Try. As an alternative, the n weighting parameters can be applied directly to the orderliness O and the target metrics | Tv; — iy | through | Tvn— vn| of the given split, without the relative rank order being used.
[0064] In one implementation having a single second cost function, the hybrid cost function can be given by
[0065] wherein H’ / is a first weighting parameter, w2is a second weighting parameter, Rank(lG) is the relative rank among all splits of an information gain of the given split, and Rank(LP) is the relative rank among all splits of the average lost pulses of the given split. As in the case of multiple second cost functions, as an alternative, if desired, weighting parameters wi, w2can be applied directly to the information gain IG and the lost pulses metric LP without first ranking the respective values of IG and LP.
[0066] The order or relative order (or orderliness) of a split can be given or represented, for example, by increasing information gain or decreasing impurity of resulting subsets, with the relevant cost function increasing with increasing disorder.
[0067] Information gain can be given by
[0068] where E(d) is the entropy of the parent set d, E(s) is the entropy of a respective child set s, |x| is the cardinality of a set x, and E(y) is entropy of a sety, given by
[0069] wherein C is the number of classes (categories) and p, is the probability of class i (in other words, where p, is the probability of randomly picking an element of class i from sety (i.e. the proportion of the set y made up of class / ).
[0070] If impurity, in the form of Gini impurity (or Gini index) is used, impurity can be given, for the split of a parent set into one or more child sets, by weighting the following expression over the child sets: l - S?=i -(Pi)2(5)
[0071] where p, denotes the probability of a respective element of a child set being randomly classified correctly based only on the distnbution of classes in the child set, and n is the number of elements in the child set.
[0072] FIG. 5 is a schematic diagram of one or more light sources 500a, 500b, and 500c together with a sensing and evaluation apparatus 520, and each light source can include respective first and second modules 560a-560c and 565a-565c. Note that, as shown and discussed above with reference to FIG. 1, a light source 100 for use with a photolithography apparatus can include first and second modules in the form of MO 160 and PRA165 as well as multiple additional modules and features, and a sensing and evaluation apparatus 120 in communication with the first and second modules and / oradditional parts of the light source. As shown in FIG. 5, if desired, one sensing and evaluation apparatus 520 can provide sensing and evaluation for multiple light sources. In either case, the sensing and evaluation apparatus (120 or 520) is configured to receive and store features of the first and second modules (160 and 165 or 560a-c and 565a-c) and / orthe light source (100 or 500a-500c) relating to performance of the first and second modules and / or the light source, and to apply a decision tree to the features to classify the light source, the first module, and / or the second module as having either (1) no technical issue or (2) maintenance needed, as explained above with respect to FIGS. 3 an 3A. As previously described, the decision tree has one or more levels and is a decision tree generated using a hybrid cost function including a combination of (1) a first cost function minimized with increasing orderliness of resulting subsets and (2) one or more second cost functions minimized with closeness to respective target parameters.
[0073] The sensing and evaluation apparatus 520 for use with respective one or more light sources 500a-500c can include data storage DS configured to receive and store respective data including performance data of the respective one or more light sources 500-a-500c and / or respective modules 560a-560c, 565a-565c thereof. The sensing and evaluation apparatus 520 can also include one or more processors Pl, P2, and P3 configured to apply the decision-tree machine learning model to the respective data to classify the respective one or more light sources 500a-500c and / or one or more respective modules 560a-560c, 565a-565c as having either (1) no technical issue or (2) maintenance needed. Note that data including performance data includes more than just performance data, and can be coextensive with the term “features.” As noted above, “features” as used herein can include one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltages of the light sources, as well as other performance data and other data not as directly related to performance, such as the model and / or type of light source, the software and / or firmware version of the light source, the age and other status aspects of the light source and / or various light source modules and / or components.
[0074] FIG. 6 is a schematic diagram of multiple light sources 600a, 600b, 600c, 600d, 600e and 600f in communication with a processing module 625. The communication need not be hard wired and need not be simultaneous or in parallel, and can also be indirect. In the example of FIG. 6, data D from multiple light sources 600a-600f (and more, if available) is provided to and received by data storage DS of the processing module 625. The data storage is configured to receive and store data including performance data of the multiple light sources 200a-600f and / or of modules thereof (not shown). The processing module also includes one or more processors Pl, P2, P3, and P4 configured to generate, based on the stored data, a decision tree classifying the light sources or modules thereof as having either (1) no technical issue or (2) maintenance needed. As above, the tree has one or more levels, the generating uses the stored data to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function comprising a combination of (1)a first cost function minimized with increasing order within resulting subsets and (2) one or more respective second cost functions minimized with closeness to respective target parameters.
[0075] Aspects and implementations of the present disclosure can be further described using the following clauses:1. A process for maintaining light sources including laser systems or modules of laser systems, the process including: obtaining one or more features of the light sources, the one or more features representing one or more operational aspects of the light sources stored overtime; generating a decision tree, the tree having one or more levels, for classifying the light sources or modules thereof, based on the one or more features, as either (1) no technical issue or (2) maintenance needed, the generating using the one or more features of the light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function including a combination of (1) a first cost function minimized with increasing orderliness within resulting subsets and (2) one or more second cost functions minimized with increasing closeness to respective target parameters; for a given light source, applying the decision tree to current or recent features of the given light source not within the one or more features previously obtained to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed; and performing maintenance on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed.2. The process of clause 1, wherein the features include one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltage.3. The process of clause 2, wherein the hybrid cost function is minimized over a split of a single feature.4. The process of clause 2, wherein the hybrid cost function is minimized over splits of two or more features.5. The process of clause 1, wherein the features consist of electrode voltage and the first cost function is minimized over a split of electrode voltage.6. The process of clause 1, wherein the first cost function is minimized with increasing information gain or decreasing impurity of resulting subsets.7. The process of clause 1, wherein the first cost function is minimized with increasing information gain.8. The process of clause 1, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as a target parameter.9. The process of clause 1, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as the only target parameter of the second cost function.10. The process of clause 1, wherein the hybrid cost function is a weighted combination of the first cost function and the second cost function.11. The process of clause 1, wherein the hybrid cost function is a weighted combination of a rank value of the first cost function and a rank value of the second cost function.12. The process of clause 1, wherein the first cost function for a given split of a parent set into one or more child sets is given by the entropy of the parent set minus the weighted average entropy of the child sets.13. The process of clause 1, wherein the first cost function for a given split of a parent set into one or more child sets is given bywhere E(d) is the entropy of the parent set d, E(s) is the entropy of a respective child set 5, | x | is the cardinality of a set r, and E(y) is entropy of a set y given bywherein C is the number of classes and p, is the probability of class i.14. The process of clause 1, wherein the first cost function for a given split of a parent set into one or more child sets is given a weighted sum of each child set ofwhere p> denotes the probability of a respective element of a child set being randomly classified correctly relying only on the distribution of classes in the child set.15. The process of clause 13, wherein the second cost function for the split is given by\TV- v\ where T, is a target of a variable v and v is average of the variable v for the split.16. The process of clause 13, wherein the second cost function for the split is given by\TLP- LP\ where TLPis a target number of lost pulses and LP is the average number of lost pulses for the split.17. The process of clause 16, wherein TLPis in the range of 1 to 4 billion.18. The process of clause 16, wherein TLPis in the range of 1 to 2 billion.19. The process of clause 1, wherein the hybrid cost function is given bywherein w, are n weighting parameters, £ wtis the summation of the n weighting parameters, Rank(O) is the relative rank of a given split, among all splits, of the order (or orderliness) of a given split, and Rank | Tvt — | is the relative rank of the given split of the average Vj of the variable vtrelative to its target T vL.20. The process of clause 1, wherein the hybrid cost function is given bywherein wj is a first weighting value, w2is a second weighting value, Rank(IG) is the relative rank among all splits of the information gain of the given split, and Rank(LP) is the relative rank among all splits of the average lost pulses of the given split.21. A process for generating a decision tree for classifying states of light sources including laser systems or modules of laser systems as having no technical issue or requiring maintenance, the process including: obtaining features of the light sources, the features representing operational aspects of the light sources stored over time; generating a decision tree based on the stored features for classifying the light sources or modules thereof as having either (1) no technical issue or (2) maintenance needed, the tree having one or more levels, the generating using the features of the light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function including a combination of (1) a first cost function minimized with increasing order within resulting subsets and (2) one or more second cost functions minimized with closeness to target parameters.22. The process of clause 21, wherein the first cost function is minimized with increasing information gain or decreasing impurity of resulting subsets, the second cost function includes a target parameter of lost pulses within the range of 1 to 4 billion lost pulses, and the hybrid cost function is a weighted combination of the first cost function and the second cost function or of a rank value of the first cost function and a rank value of the second cost function.23. A process for classifying states of light sources including laser systems or modules of laser systems, the process including: acquiring current features of a given light source; and applying a decision tree to the current features of the given light source to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed, the decision tree having one or more levels, wherein the decision tree is a decision tree generated using a hybrid cost function including a combination of (1) a first cost function minimized with increasing information gain or decreasing impurity and (2) one or more second cost functions minimized with closeness to target parameters.24. The process of clause 23, wherein the first cost function is minimized with increasing information gain or decreasing impurity of resulting subsets over splits of one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltage, the second cost function includes a target parameter of lost pulses within the range of 1 to 4 billion lost pulses, and the hybrid cost function is a weighted combination of the first cost function and the second cost function or of a rank value of the first cost function and a rank value of the second cost function.25. The process of clause 24, further including performing maintenance on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed26. A light source for use with a photolithography apparatus, the light source including: first and second modules; and a sensing and evaluation apparatus in communication with the first and second modules and / or additional parts of the light source; wherein the sensing and evaluation apparatus is configured to receive and store features of the first and second modules and / or the light source relating to performance of the first and second modules and / or the light source, and to apply a decision tree to the features to classify the light source, the first module, and / or the second module as having either (1) no technical issue or (2) maintenance needed, wherein the decision tree has one or more levels and is a decision tree generated using a hybrid cost function including a combination of(1) a first cost function minimized with increasing orderliness of resulting subsets and (2) one or more second cost functions minimized with closeness to respective target parameters.27. The light source of clause 26, wherein the first cost function is a cost function minimized with increasing information gain or decreasing impurity of resulting subsets.28. The light source of clause 26, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as a target parameter.29. The light source of clause 26, wherein the second cost function has an amount of lost productive time, of the light source upon a maintenance needed classification, as a target parameter.30. A processing module configured for processing historical data of multiple light sources used in photolithography, the processing module including: data storage configured to receive and store data including performance data of the multiple light sources and / or of modules thereof; one or more processors configured to generate, based on the stored data, a decision tree classifying the light sources or modules thereof as having either (1) no technical issue or (2) maintenance needed, the tree having one or more levels, the generating using the stored data to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function including a combination of (1) a first cost function minimized with increasing order within resulting subsets and(2) one or more respective second cost functions minimized with closeness to respective target parameters.31. The processing module of clause 30, wherein the first cost function is a cost function minimized with increasing information gain or decreasing impurity of resulting subsets.32. The processing module of clause 30, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as a target parameter.33. The processing module of clause 30, wherein the second cost function has an amount of lost productive time, of the light source upon a maintenance needed classification, as a target parameter.34. A sensing and evaluation apparatus for use with respective one or more light sources used in photolithography, the apparatus including: data storage configured to receive and store respective data including performance data of the respective one or more light sources and / or respective modules thereof; and one or more processors configured to apply a decision-tree machine learning model to the respective data to classify the respective one or more light sources and / or one or more respectivemodules thereof as having either (1) no technical issue or (2) maintenance needed; wherein the decision-tree machine learning model comprises one or more levels generated by using historical data including performance data from multiple light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function including a combination of (1) a first cost function minimized with increasing order within resulting subsets and (2) one or more second cost functions minimized with closeness to target parameters.35. The sensing and evaluation apparatus of clause 34, wherein the first cost function is a cost function minimized with increasing information gain or decreasing impurity of resulting subsets.36. The sensing and evaluation apparatus of clause 34, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as a target parameter.37. The sensing and evaluation apparatus of clause 34, wherein the second cost function has an amount of lost productive time, of the light source upon a maintenance needed classification, as a target parameter.
[0076] The above-described aspects and implementations and other implementations are within the scope of the following claims.
Claims
CLAIMS1. A process for maintaining light sources including laser systems or modules of laser systems, the process comprising: obtaining one or more features of the light sources, the one or more features representing one or more operational aspects of the light sources stored over time; generating a decision tree, the decision tree having one or more levels, for classifying the light sources or modules thereof, based on the one or more features, as either (1) no technical issue or (2) maintenance needed, the generating using the one or more features of the light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function comprising a combination of (1) a first cost function minimized with increasing order (or orderliness) within resulting subsets and (2) one or more second cost functions minimized with increasing closeness to respective target parameters; for a given light source, applying the decision tree to current or recent features of the given light source not within the one or more features previously obtained to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed; and performing maintenance on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed.
2. The process of claim 1, wherein the one or more features include one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltage.
3. The process of claim 1, wherein the hybrid cost function is minimized over a split of a single feature.
4. The process of claim 1 wherein the hybrid cost function is minimized over splits of two or more features.
5. The process of claim 1, wherein the first cost function is minimized with increasing information gain or decreasing impurity of resulting subsets.
6. The process of claim 1, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as a target parameter.
7. The process of claim 1, wherein the second cost function has a number of lost pulses of the light source upon a maintenance needed classification as the only target parameter of the second cost function.
8. The process of claim 1, wherein the hybrid cost function is a weighted combination of the first cost function and the second cost function.
9. The process of claim 1, wherein the hybrid cost function is a weighted combination of a rank value of the first cost function and a rank value of the second cost function.
10. The process of claim 1, wherein the first cost function for a given split of a parent set into one or more child sets is given by an entropy of the parent set minus a weighted average entropy of the child sets.
11. The process of claim 1, wherein the first cost function for a given split of a parent set into one or more child sets is given bywhere E(d) is the entropy of the parent set d, E(s) is the entropy of a respective child set s, |x| is the cardinality of a set x, and E(y) is entropy of a set y given bywherein C is the number of classes and p> is the probability of class i.
12. The process of claim 1, wherein the first cost function for a given split of a parent set into one or more child sets is given by a weighted sum of each child set ofwhere p, denotes the probability of a respective element of a child set being randomly classified correctly relying only on a distribution of classes in the child set.
13. The process of claim 11, wherein the second cost function for the split is given by\TV- v\ where Tvis a target of a variable v and v is an average of the variable v for the split.
14. The process of claim 11, wherein the second cost function for the split is given by\TLP- LP\ where TLp is a target number of lost pulses and LP is an average number of lost pulses for the split.
15. The process of claim 14, wherein TIP is in the range of 1 to 4 billion.
16. The process of claim 1, wherein the hybrid cost function is given bywherein w, are n weighting parameters, Xwi is the summation of the n weighting parameters, Rank(O) is the relative rank of a given split, among all splits, of the order (or orderliness) of a given split, and Rank | Tvi— v is the relative rank of the given split of the average of the variable vLrelative to its target Tvi.
17. The process of claim 1, wherein the hybrid cost function is given bywherein wj is a first weighting value, w2is a second weighting value, Rank(IG) is the relative rank among all splits of the information gain of the given split, and Rank(LP) is the relative rank among all splits of the average lost pulses of the given split.
18. A process for generating a decision tree for classifying states of light sources including laser systems or modules of laser systems as having no technical issue or requiring maintenance, the process comprising: obtaining features of the light sources, the features representing operational aspects of the light sources stored over time; generating a decision tree based on the stored features for classifying the light sources or modules thereof as having either (1) no technical issue or (2) maintenance needed, the decision tree having one or more levels, the generating using the features of the light sources to determine a split or splits for each of the one or more levels by minimizing a hybrid cost function, the hybrid cost function comprising a combination of (1) a first cost function minimized with increasing order within resulting subsets and (2) one or more second cost functions minimized with closeness to target parameters.
19. The process of claim 18, wherein the first cost function is minimized with increasing information gain or decreasing impurity of resulting subsets, the second cost function includes a target parameter of lost pulses within the range of 1 to 4 billion lost pulses, and the hybrid cost function is a weighted combination of the first cost function and the second cost function or of a rank value of the first cost function and a rank value of the second cost function.
20. A process for classifying states of light sources including laser systems or modules of laser systems, the process comprising:acquiring current features of a given light source; and applying a decision tree to the current features of the given light source to classify the given light source or a module thereof as having either (1) no technical issue or (2) maintenance needed, the decision tree having one or more levels, wherein the decision tree is a decision tree generated using a hybrid cost function comprising a combination of (1) a first cost function minimized with increasing information gain or decreasing impurity and (2) one or more second cost functions minimized with closeness to target parameters.
21. The process of claim 20, wherein the first cost function is minimized with increasing information gain or decreasing impurity of resulting subsets over splits of one or more of bandwidth, wavelength, wavelength variation, power, and electrode voltage, the second cost function includes a target parameter of lost pulses within the range of 1 to 4 billion lost pulses, and the hybrid cost function is a weighted combination of the first cost function and the second cost function or of a rank value of the first cost function and a rank value of the second cost function.
22. The process of claim 21, further comprising performing maintenance on the given light source or the module thereof when the given light source or the module thereof has been classified as maintenance needed.
Citation Information
Patent Citations
Energy consumption reduction of gas discharge chamber blower
CN116636100A
Maintenance of module of light source for semiconductor lithography
CN116997864A
Maintenance of modules for light sources in semiconductor photolithography
WO2023121798A1