Method for marking time series data related to one or more machines

By segmenting and graphically encode the time series data of the lithography device, combined with the semi-supervised learning method, the problem of large amount of data and multiple noise in the health status monitoring of the lithography device hardware components is solved, and more accurate health status estimation and prediction are achieved.

CN120303671APending Publication Date: 2025-07-11ASML NETHERLANDS BV
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202380082945.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-06
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When monitoring and diagnosing the health status of hardware components in lithography equipment, the prior art has problems such as large amount of data, high noise and difficult to effectively classify, resulting in inconsistent predictions and errors.

Method used

The time series data is segmented to identify similar patterns, define the graph structure to encode similarity, and use semi-supervised learning methods to classify and indicate unlabeled patterns, and use similarity and domain knowledge in the graph structure to perform health status estimation.

Benefits of technology

It improves the accuracy and efficiency of monitoring the health status of the lithography equipment hardware components, reduces dependence on labeled data, reduces the impact of noise, and achieves more stable prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303671A_ABST
    Figure CN120303671A_ABST
Patent Text Reader

Abstract

A method for marking time series data associated with one or more machines is disclosed. The method comprises the following steps: acquiring the time sequence data; segmenting the time series data to obtain a plurality of patterns grouped according to pattern similarity; marking a subset of the plurality of patterns to obtain a marked subset of patterns, the remaining patterns of the plurality of patterns including five unmarked patterns; defining a graph structure throughout the patterns, the graph structure describing similarities between the patterns; and classifying and / or marking the unmarked pattern using the graph structure and the marked subset of patterns to obtain a marked pattern.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This application claims priority to European Application No. 22211052.0, filed on December 02, 2022, the entire content of which is incorporated herein by reference. Technical field

[0002] The present invention relates to methods and apparatus, such as may be used in the manufacture of devices by lithographic techniques, and to methods of manufacturing devices using lithographic techniques. The present invention more particularly relates to fault detection for such devices. Background art

[0003] A lithographic apparatus is a machine that applies a desired pattern onto a substrate, usually onto a target portion of the substrate. A lithographic apparatus can be used, for example, in the manufacture of integrated circuits (ICs). In that case, a patterning device, alternatively referred to as a mask or a reticle, can be used to generate a circuit pattern to be formed on a single layer of the IC. This pattern can be transferred onto a target portion (e.g., including a part of a die, a single die, or multiple dies) on the substrate (e.g., a silicon wafer). The transfer of the pattern is typically via imaging onto a layer of radiation-sensitive material (resist) provided on the substrate. Usually, a single substrate will include a network of adjacent target portions that are successively patterned. These target portions are typically referred to as “fields”.

[0004] In the manufacture of complex devices, typically many lithographic patterning steps are performed, whereby functional features are formed in successive layers on a substrate. As such, a key aspect of the performance of a lithographic apparatus is the ability to place the applied pattern relative to features placed in a previous layer (by the same apparatus or a different lithographic apparatus) properly and accurately. For this purpose, the substrate is provided with one or more sets of alignment marks. Each mark is a structure whose position can later be measured using a position sensor, typically an optical position sensor. A lithographic apparatus includes one or more alignment sensors by which the position of the marks on the substrate can be accurately measured. Different types of marks and different types of alignment sensors are known from different manufacturers and different products of the same manufacturer.

[0005] In other applications, metrology sensors are used to measure exposed structures (in the resist and / or after etching) on a substrate. A specialized inspection tool in the form of a scatterometer, in which a radiation beam is directed onto a target located on the surface of the substrate and the properties of the scattered or reflected beam are measured. Examples of known scatterometers include angular-resolved scatterometers of the type described in US2006033921A1 and US2010201963A1. In addition to measuring feature shapes by reconstruction, such devices can also be used to measure diffraction-based overlay, as described in the published patent application US2006066855A1. Diffraction-based overlay metrology using dark-field imaging of diffraction orders enables overlay measurements of smaller targets. Examples of dark-field imaging metrology can be found in international patent applications WO 2009 / 078708 and WO 2009 / 106279, which are hereby incorporated by reference in their entirety. Further developments of such techniques are described in the published patent applications US20110027704A, US20110043791A, US2011102753A1, US20120044470A, US20120123581A, US20130258310A, US20130271740A and WO2013178422A1. These targets can be smaller than the illumination spot and can be surrounded by product structures on the wafer. Composite grating targets can be used to measure multiple gratings in one image. The content of all these applications is also incorporated herein by reference.

[0006] Hardware components in a machine, such as a lithography apparatus or other equipment used in the manufacture of integrated circuits (ICs), may deteriorate over time. Therefore, it is necessary to monitor the health status of these hardware components in order to prevent unscheduled downtime and / or non-functional ICs. These devices are very complex, including many modules and a very large number of sensors, and thus provide a very large amount of data (e.g., such as time series data). Analyzing such a large amount of data to estimate the health status is difficult. Due to this, supervised machine learning techniques are sometimes used to analyze the data and estimate the health status (or more generally, the device status).

[0007] There is a desire to improve such a method for estimating the system state or health status based on supervised machine learning. Summary of the Invention

[0008] The present invention provides and discloses, in a first aspect, a method for labeling time series data associated with one or more machines, the method comprising: obtaining the time series data; segmenting the time series data to obtain a plurality of patterns grouped according to pattern similarity; labeling a subset of the plurality of patterns to obtain a labeled pattern subset, the remaining patterns in the plurality of patterns including unlabeled patterns; defining a graph structure over the patterns, the graph structure describing the similarity between the patterns; and using the graph structure and the labeled pattern subset to classify and / or label the unlabeled patterns to obtain labeled patterns.

[0009] There is also disclosed a computer program operable to perform the method according to the first aspect.

[0010] The above and other aspects of the invention will be understood from consideration of the examples described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0012] Figure 1 A lithographic apparatus is depicted;

[0013] Figure 2 Schematically illustrates Figure 1 the measurement and exposure processes in the apparatus of;

[0014] Figure 3 is a graph of sensor signal versus time, showing different patterns, each pattern representing a corresponding equipment state or health state;

[0015] Figure 4 is a flow chart of a prior art method for determining the equipment state or health state of a machine;

[0016] Figure 5 is a flow chart of a method for determining the equipment state or health state of a machine according to an embodiment;

[0017] Figure 6 (a) is a flow chart of step 525 of FIG. (5) according to an embodiment, and Figure 6 (b) is an exemplary similarity graph according to an embodiment;

[0018] Figure 7 is a flow chart of rule generation according to an embodiment; and

[0019] Figure 8 is a flow chart of generating labeled data for training other machine learning algorithms according to an embodiment. DETAILED DESCRIPTION

[0020] Before describing embodiments of the present invention in detail, it is beneficial to present an exemplary environment that can be used to implement embodiments of the present invention.

[0021] Figure 1 A lithographic apparatus LA is schematically depicted. The apparatus includes: an illumination system (illuminator) IL configured to condition a radiation beam B (e.g., UV radiation or DUV radiation); a patterning device support or support structure (e.g., a mask table) MT configured to support a patterning device (e.g., a mask) MA and connected to a first positioner PM configured to accurately position the patterning device according to certain parameters; two substrate tables (e.g., wafer tables) WTa and WTb, each configured to hold a substrate (e.g., a wafer coated with resist) W and each connected to a second positioner PW configured to accurately position the substrate according to certain parameters; and a projection system (e.g., a refractive projection lens system) PS configured to project a pattern imparted to the radiation beam B by the patterning device MA onto a target portion C (e.g., including one or more dies) of the substrate W. A reference frame RF connects the various components and serves as a reference for setting and measuring the positions of the patterning device and the substrate, and the positions of features on the patterning device and the substrate.

[0022] The illumination system may include various types of optical components for guiding, shaping, or controlling the radiation, such as refractive, reflective, magnetic, electromagnetic, electrostatic, or other types of optical components, or any combination thereof.

[0023] The patterning device support MT holds the patterning device in a manner that depends on the orientation of the patterning device, the design of the lithographic apparatus, and other conditions (such as, for example, whether the patterning device is held in a vacuum environment). The patterning device support may use mechanical, vacuum, electrostatic, or other clamping techniques to hold the patterning device. The patterning device support MT may be, for example, a frame or table that can be fixed or movable as required. The patterning device support can ensure that the patterning device is, for example, in a desired position relative to the projection system.

[0024] The term "patterning device" as used herein should be broadly interpreted as meaning any device that can be used to impart a pattern in a cross-section of a radiation beam so as to create a pattern in a target portion of a substrate. It should be noted that, for example, if the pattern imparted to the radiation beam includes phase-shifting features or so-called assist features, the pattern may not exactly correspond to the desired pattern in the target portion of the substrate. In general, the pattern imparted to the radiation beam will correspond to a particular functional layer in a device (such as an integrated circuit) created in the target portion.

[0025] As depicted herein, the apparatus is of the transmissive type (e.g., employing a transmissive patterning device). Alternatively, the apparatus may be of the reflective type (e.g., employing a programmable mirror array of the type mentioned above, or a reflective mask). Examples of patterning devices include masks, programmable mirror arrays, and programmable LCD panels. Any use herein of the term "reticle" or "mask" may be considered synonymous with the more general term "patterning device". The term "patterning device" may also be construed to refer to a device that stores pattern information digitally for controlling such a programmable patterning device.

[0026] The term "projection system" as used herein should be broadly construed to cover any type of projection system suitable for the exposure radiation being used or for other factors such as the use of an immersion liquid or the use of a vacuum, including refractive, reflective, catadioptric, magnetic, electromagnetic, and electrostatic optical systems, or any combination thereof. Any use herein of the term "projection lens" may be considered synonymous with the more general term "projection system".

[0027] The lithographic apparatus may also be of the type in which at least a portion of the substrate is covered by a liquid having a relatively high refractive index (e.g., water) in order to fill the space between the projection system and the substrate. The immersion liquid may also be applied to other spaces in the lithographic apparatus, such as the space between the mask and the projection system. Immersion techniques are well known in the art for increasing the numerical aperture of the projection system.

[0028] In operation, the illuminator IL receives a radiation beam from the radiation source SO. For example, when the source is an excimer laser, the source and the lithographic apparatus may be separate entities. In such cases, the source is not considered to form part of the lithographic apparatus, and the radiation beam is transmitted from the source SO to the illuminator IL by means of a beam delivery system BD including, for example, suitable directing mirrors and / or beam expanders. In other cases, for example, when the source is a mercury lamp, the source may be an integral part of the lithographic apparatus. The source SO, the illuminator IL, and the beam delivery system BD (when required) may be referred to as the radiation system.

[0029] The illuminator IL may include, for example, an adjuster AD for adjusting the angular intensity distribution of the radiation beam, an integrator IN, and a condenser CO. The illuminator may be used to condition the radiation beam to have a desired uniformity and intensity distribution in the cross-section of the radiation beam.

[0030] The radiation beam B is incident on a patterning device MA which is held on a patterning device support MT, and is patterned by the patterning device. After having traversed the patterning device (e.g., mask) MA, the radiation beam B passes through a projection system PS which focuses the beam onto a target portion C of a substrate W. By means of a second positioner PW and a position sensor IF (e.g., an interferometric device, a linear encoder, a 2D encoder or a capacitive sensor), the substrate table WTa or WTb can be accurately moved, e.g., in order to position different target portions C in the path of the radiation beam B. Similarly, e.g., after a mechanical retrieval from a mask library or during a scan, a first positioner PM and another position sensor ( Figure 1 the other position sensor is not explicitly depicted in) can be used to accurately position the patterning device (e.g., mask) MA relative to the path of the radiation beam B.

[0031] Mask alignment marks M1, M2 and substrate alignment marks P1, P2 can be used to align the patterning device (e.g., mask) MA and the substrate W. Although the illustrated substrate alignment marks occupy dedicated target portions, the illustrated substrate alignment marks can be located in the space between the target portions (these labels are referred to as scribe alignment marks). Similarly, in the case of setting more than one die on the patterning device (e.g., mask) MA, the mask alignment marks can be located between the dies. Smaller alignment marks can also be included within the die among the device features, in which case it is desirable to make the identification as small as possible and without any imaging or process conditions different from adjacent features. An alignment system for detecting the alignment marks is further described below.

[0032] The depicted apparatus can be used in a variety of modes. In a scanning mode, while projecting the pattern imparted to the radiation beam onto the target portion C, the patterning device support (e.g., mask table) MT and the substrate table WT are scanned synchronously (i.e., single dynamic exposure). The speed and direction of the substrate table WT relative to the patterning device support (e.g., mask table) MT can be determined by the magnification (reduction ratio) and image inversion characteristics of the projection system PS. In the scanning mode, the maximum size of the exposure field limits the width of the target portion (in the non-scanning direction) in a single dynamic exposure, while the length of the scanning movement determines the height of the target portion (in the scanning direction). As is well known in the art, other types of lithographic apparatuses and operating modes are possible. For example, the step mode is well known. In so-called “maskless” lithography, a programmable patterning device is kept stationary but has a changing pattern, and the substrate table WT is moved or scanned.

[0033] Combinations and / or variants of the usage modes described above and / or completely different usage modes can also be employed.

[0034] The lithographic apparatus LA belongs to the so-called dual-platform type, which has two substrate tables WTa, WTb and two stations - an exposure station EXP and a measurement station MEA - between which the substrate tables can be exchanged. While exposing a substrate on one substrate table at the exposure station, another substrate can be loaded onto the other substrate table at the measurement station and various preparatory steps can be carried out. This results in a significant increase in the throughput of the apparatus. The preparatory steps can include using a level sensor LS to map the surface height profile of the substrate and using an alignment sensor AS to measure the position of alignment marks on the substrate. If a position sensor IF is unable to measure the position of the substrate table when the substrate table is at the measurement station and at the exposure station, a second position sensor can be provided to enable tracking of the position of the substrate table relative to a reference frame RF at both stations. Instead of the dual-platform arrangement shown, other arrangements are known and available. For example, other lithographic apparatuses with a substrate table and a measurement table are known. These substrate table and measurement table dock together when performing preparatory measurements and then undock when the substrate table undergoes exposure.

[0035] Figure 2 The figure illustrates the steps of exposing a target portion (e.g., a die) onto Figure 1 a substrate W in a dual-platform apparatus. The steps performed at the measurement station MEA are within the left dashed box, while the steps performed at the exposure station EXP are shown on the right. Sometimes, one of the substrate tables WTa, WTb will be at the exposure station while the other of the substrate tables WTa, WTb is at the measurement station, as described above. For the purposes of this description, it is assumed that the substrate W has been loaded into the exposure station. At step 200, a new substrate W' is loaded into the apparatus by a mechanism (not shown). Such two substrates are processed in parallel to increase the throughput of the lithographic apparatus.

[0036] First referring to the newly loaded substrate W', this substrate can be a substrate that has not been processed previously and is prepared with a new resist for the first exposure in the apparatus. However, typically, the lithographic process described will be just one step in a series of exposure and processing steps such that the substrate W' has passed through this apparatus and / or other lithographic apparatuses several times and can also undergo subsequent processes. Specifically for the problem of improving overlay performance, the task is to ensure that a new pattern is accurately applied in the correct position on a substrate that has undergone one or more cycles of patterning and processing. These processing steps gradually introduce deformations in the substrate that must be measured and corrected to achieve satisfactory overlay performance.

[0037] The previous and / or subsequent patterning steps (such as those just mentioned) can be performed in other lithography apparatuses, and even the previous and / or subsequent patterning steps can be performed in different types of lithography apparatuses. For example, some layers in a device manufacturing process that require very high parameters such as resolution and overlay can be performed in a higher-order lithography tool compared to other layers with less stringent requirements. Thus, some layers can be exposed in an immersion lithography tool, while other layers are exposed in a "dry" tool. Some layers can be exposed in a tool operating at a DUV wavelength, while other layers are exposed using EUV wavelength radiation.

[0038] At 202, alignment measurements using a substrate label P1, etc. and an image sensor (not shown) are used to measure and record the alignment of the substrate relative to the substrate table WTa / WTb. Additionally, an alignment sensor AS will be used to measure a number of alignment marks across the substrate W'. In one embodiment, these measurements are used to establish a "wafer grid" that very accurately maps the distribution of labels across the substrate, including any deformations relative to a nominal rectangular grid.

[0039] At step 204, a level sensor LS is also used to measure a map of the wafer height (Z) relative to the X-Y position. Conventionally, the height map is only used to achieve accurate focusing of the exposed pattern. The height map can additionally be used for other purposes.

[0040] When loading the substrate W', option data 206 is received. The option data 206 defines the exposure to be performed and also defines the nature of the wafer and the previously generated pattern and the pattern to be generated on the substrate W'. The measurements of the wafer position, wafer grid, and height map performed at 202 and 204 are added to these option data, such that a complete set 208 of option and measurement data can be transferred to the exposure station EXP. The measurement of the alignment data includes, for example, the X and Y positions of alignment targets formed in a fixed or nominally fixed relationship to the product pattern that is the product of the lithography process. These alignment data obtained just before exposure are used to generate an alignment model with parameters that fit the model to the data. These parameters and the alignment model will be used during the exposure operation to correct the position of the pattern applied in the current lithography step. The model interpolates the position deviation between the measured positions. Conventional alignment models may include four, five, or six parameters that together define the translation, rotation, and scaling of an "ideal" grid in different dimensions. Higher-order models using more parameters are well known.

[0041] At 210, the wafers W' and W are swapped such that the measured substrate W' becomes the substrate W and enters the exposure station EXP. At Figure 1In an exemplary apparatus, such swapping is performed by exchanging the support members WTa and WTb within the apparatus such that the substrates W, W' remain accurately held and positioned on those support members to retain the relative alignment between the substrate table and the substrates themselves. Thus, once the table has been swapped, in order to utilize the measurement information 202, 204 for the substrate W (previously W') to control the exposure step, it is necessary to determine the relative position between the projection system PS and the substrate table WTb (previously WTa). At step 212, reticle alignment is performed using the mask alignment marks M1, M2. At steps 214, 216, 218, a scanning motion and radiation pulses are applied at successive target sites across the substrate W in order to complete the exposure of a plurality of patterns.

[0042] By using the alignment data and height maps obtained at the measurement station during the execution of the exposure step, these patterns are accurately aligned relative to the desired sites and in particular relative to features previously placed on the same substrate. At step 220, the exposed substrate, now labeled W'', is unloaded from the apparatus to subject the exposed substrate to an etching or other process in accordance with the exposed patterns.

[0043] Those skilled in the art will recognize the simplified schematic of the many very detailed steps involved in an example of a real manufacturing scenario as described above. For example, there will often be separate stages of rough and fine measurement using the same or different marks rather than measuring alignment in a single pass. The rough and / or fine alignment measurement steps may be performed before or after the height measurement, or the rough and / or fine alignment measurement steps may be interleaved.

[0044] In a lithography system, an important issue that has a significant impact on the normal operating time of the system is the ability to quickly and effectively detect and / or diagnose events or trends (e.g., equipment status, health status, and / or fault events) that may indicate irregular or abnormal behavior. However, these systems are very complex, and these systems include a plurality of different modules (e.g., in particular, including a projection optics module, a wafer stage module, a reticle stage module, a reticle shading module), each of which generates a large amount of data. Due to the lack of data on fault events, complex problems involving multiple modules can be a particular challenge for diagnosis.

[0045] The performance of the hardware components in a machine (such as a lithography apparatus (scanner) used in IC manufacturing or other machines) deteriorates over time due to wear and / or aging. If the deteriorated components are not replaced, refurbished, or otherwise maintained, the machine functionality will not be able to remain within specifications, which will result in a yield loss (non-functional ICs). Therefore, maintaining the deteriorated hardware components in a machine is very important or crucial, which affects its availability and productivity.

[0046] Sensor measurement results (sensor signals) are typically used as an indication of the health state of a hardware component. Typically, these sensor measurement results include multiple signals and are thus high-dimensional. The sensor measurement results exhibit different patterns corresponding to the health state of the hardware component (i.e., the device state). For example, the health state or device state can be classified into two or more categories of interest; for example, a three-category system can classify the health state into three categories: "healthy / good", "deteriorated", and "unhealthy / bad". These categories are merely exemplary, and the number and / or their definitions of the categories can depend on the use case.

[0047] Figure 3 is a graph of sensor signals over time that is an example of the corresponding relationship between the signal behavior and the hardware state. More specifically, the graph shows the labeled sensor measurements of a machine hardware component during four different time periods distinguishable by the signal behavior. During the first time period TP1, the signal indicates a deterioration behavior DG of the component (i.e., the component is deteriorating and immediate maintenance action is required to repair or replace the component to prevent unscheduled downtime and / or poor yield). This rapidly develops into an unhealthy behavior UHE of the component during the second time period TP2 (i.e., indicating that the component has deteriorated severely to the extent that production is affected and immediate maintenance action is required). The third time period TP3 indicates the healthy state HE of the component. Thus, the transition from the second time period TP2 to the third time period TP3 can indicate that a maintenance action has been performed to replace or repair the component. The last time period TP4 is another deterioration behavior DG period because the component starts to deteriorate again.

[0048] Currently, two alternative methods are commonly used to perform such monitoring and prediction of the health state. First, the estimation of the health state can be automated via supervised machine learning (ML) methods, where a classification ML model receives high-dimensional signals as input from the sensor measurement results and maps these signals to the health state (labels). The classifier is typically trained to give predictions for each data point in the time series without considering the temporal dependence of the data. The labels of the training set can come from domain expertise, performance measurements, or other sources.

[0049] Figure 4 is a flowchart illustrating such a prior art method. The measured sensor data (including unlabeled data 400) undergoes a labeling step 410 to label a (limited) subset of the measured sensor data, thereby obtaining labeled training data 420. The ML classifier 430 then classifies the remaining unlabeled data 400 based on the labeled training data 420 in order to determine the health state 440 of the unlabeled data 400 (and thus determine one or more components to which such data pertains).

[0050] Most supervised learning methods require large amounts of labeled data that are often not available. Since domain experts usually label sensor measurements, obtaining more labels is time-consuming. Additionally, many signals are ambiguous, so it is not possible for either algorithms or domain experts to label these signals with confidence. Additionally, in the process of generating the training set for the ML model: sensor measurements may have a certain number of outliers and discontinuities and be noisy. Therefore, methods that focus on labeling data points are sensitive to this noise and propagate the noise to the prediction.

[0051] A second known method may include applying thresholds (e.g., representing one or more specifications) to each data point via heuristics. However, these thresholds can be inaccurate and cannot handle high-dimensional sensor measurements or ambiguous patterns. Additionally, like the classifier methods described, setting the thresholds is vulnerable to noise in the sensor measurements.

[0052] Accordingly, in either of the two prior art methods described, the prediction is prone to inconsistencies and errors.

[0053] Therefore, it is proposed to process time series data according to the patterns in the time series data and define a graph structure over the patterns. The graph structure can be used to classify the processed time series data using only a limited number of labels and a large number of unlabeled data points. The graph structure can encode the physical properties of sensor degradation via the similarity or distance function used for its construction. Thus, the graph structure can describe the physical properties of a degraded hardware component. If this is unknown, any similarity function can be used.

[0054] Defining a graph structure over the patterns may include or describe, for example, modeling the pairwise relationships between the patterns according to a similarity metric.

[0055] Accordingly, a method for labeling time series data associated with one or more machines is disclosed, the method comprising: obtaining the time series data; segmenting the time series data to obtain a plurality of patterns grouped according to pattern similarity; labeling a subset of the plurality of patterns to obtain a labeled pattern subset, the remaining patterns of the plurality of patterns including unlabeled patterns; defining a graph structure over the patterns, the graph structure describing the similarity between the patterns; and using the graph structure and the labeled pattern subset to classify and / or label the unlabeled patterns to obtain labeled patterns.

[0056] While current methods ignore the temporal dependencies within time series data and only provide predictions for each data point, the proposed method exploits the patterns that occur in the temporal neighborhood of the data points (e.g., to address noise) and provides predictions within that context. Additionally, this property can be imposed by encoding in a similarity graph of the patterns the physical property that similarly shaped patterns should result in similar machine health states. Such a graph can be used to impose smoothness on the label estimation over the graph structure and such a graph can operate on very few labeled instances. Additionally, prior art methods are unable to model the physical manifestations of the degradation of sensor measurements, such as changes in the drift rate in a signal or changes in the variation in a signal, for example.

[0057] Generally, the measured sensor data, which includes unlabeled input time series data from each machine, is segmented into clusters or time series patterns of similar evolution behavior. The similarity between the patterns can then be encoded in a graph. Domain expertise or any other knowledge source (e.g., performance measurements) can be used to apply labels to a small subset of the patterns. These labels can be propagated to the full dataset using semi-supervised algorithms that consider the graph. Human experts (or any other knowledge source) can support the semi-supervised model to improve accuracy. In this way, good accuracy can be obtained with only a limited number of labels in an active learning loop.

[0058] Figure 5 is a flow chart that more detailedly illustrates the proposed method. At step 505, the unlabeled input time series data 500 (e.g., from one or more machines) is segmented or partitioned into a plurality of patterns 510 of similar behavior. Generally, degradation manifests in sensor measurements with corresponding different drift patterns. Drift can be increasing behavior, linear behavior, recurring behavior, sudden change (jump) behavior, or gradual drift behavior. The number of data points between these patterns may differ from each other or vary relative to each other.

[0059] The segmentation step 500 can use any suitable time series segmentation algorithm. The segmentation can be performed in any suitable domain; for example, it can be performed in the time domain, frequency domain, or spatial domain. Examples of suitable algorithms particularly include Gaussian segmentation, hidden Markov models, neural networks for time series segmentation, t-distributed stochastic neighbor embedding (t-SNE), or principal component analysis (PCA) using clustering.

[0060] Specific examples of segmentation algorithms can include performing spatial segmentation with agglomerative clustering and dimensionality reduction defined by Uniform Manifold Approximation and Projection (UMAP). UMAP is a graph-based dimensionality reduction algorithm that uses applied Riemannian geometry to estimate low-dimensional embeddings. The advantage of such an implementation is that it handles limited amounts of data very well and addresses the curse of dimensionality (sensor measurements can have more than 100 dimensions). For example, UMAP is described in "Parametric UMAP Embeddings for Representation and Semisupervised Learning", Sainburg, Tim and McInnes, Leland and Gentner, Timothy Q, Neural Computation, Volume 33, pages 2881-2907, 2021, which is incorporated herein by reference.

[0061] UMAP estimates the nearest neighbor similarity around a data point by defining a region or circle around each data point. The circle for each point includes the points that are its nearest neighbors. For example, the similarity (e.g., the value of a similarity metric) of a data point A can be quantified by defining a circle centered on a specific data point (e.g., data point A) that includes the nearest neighbor data points of data point A. The size of each circle can be defined by the proximity of the neighboring data points of the data point, e.g., such that each circle for each corresponding data point includes a set (same) number of neighboring data points. Those skilled in the art will appreciate that other methods for defining the circle size are possible. The similarity metric value or similarity score for each of the adjacent data points within the circle can be estimated based on the distance from the center (i.e., from data point A in the specific example). In an embodiment, it can be determined that such a similarity score decreases exponentially from the center of the circle to the perimeter of the circle.

[0062] This method can include time series segmentation of applying UMAP to each machine. Applying UMAP to each machine in this way provides a low-dimensional representation that preserves the similarity of data points in the high-dimensional space. Unexpectedly, in the absence of any time information, the resulting representation also adheres to the temporal adjacency of two data points. This can be explained by the radical exponential decay of the similarity of the nearest data points in the UMAP described above. In the hardware degradation signal, data points with temporal adjacency usually have more similar measurements than data points that are temporally far apart. The similarity expands with the exponential decay of the similarity. This means that temporal adjacency is equivalent to spatial adjacency because the signal evolves smoothly. Agglomerative clustering can then be applied to separate the data into time series patterns. In an embodiment, due to the elongated shape of the derived clusters, agglomerative clustering with single linkage can be used. To determine the appropriate number of clusters, for example, the silhouette score or silhouette coefficient can be used.

[0063] The labeling step 515 can include applying rules and / or annotations 517 to a (e.g., small) subset of the pattern 510. For example, depending on the maturity of the domain expertise of a particular hardware component, a domain expert can provide labels as annotations on the time series data or as rules. To provide a specific illustrative example of a rule: it can be defined that when the signal drift rate is higher than a threshold rate, the health state of a particular component is poor and the particular component should be replaced. Other rules can indicate the nature of the aging effect; for example, it can be mandated that the sequence of states must follow three or more consecutive categories, such as: "green" (good) to "orange" (deteriorated) to "red" (poor). In other contexts, performance measurements can be used to indicate labels. For example, machine matching overlap measurements can be used to indicate the health of alignment sensors. The output of this step is the labeled pattern subset 520.

[0064] The labeling step 515 can include, for example, applying the same rule or label to all points of the pattern (cluster). There are several ways to do this. One method includes considering the corresponding representative object or point from each pattern and aggregating the labels of the representative objects or points to estimate one label for the complete pattern. Such aggregation can include, for example, a majority vote within the overall scheme. In a specific example, the center point of each cluster can be defined as the representative object or point.

[0065] At step 525, graph-based semi-supervised learning (SSL) can be performed on data pattern 510 using the partially labeled data 520. Semi-supervised learning is a family of algorithms that exploit a small amount of labeled data and a large amount of unlabeled data to jointly learn the structure of a data set and optimize a supervised objective, such as classifying time series patterns. These algorithms produce more accurate predictions when sufficient unlabeled data is available because they utilize the structure of the unlabeled data when estimating class labels. Graphs provide additional domain information for machine learning algorithms. The purpose of graph-based SSL methods is to impose graph constraints on a loss function and thus ensure or enforce smoothness across the graph.

[0066] The SSL step 525 can include sub-steps illustrated by Figure 6 (a). At step 600, a similarity graph (i.e., a graph indicating pattern similarity according to a similarity metric) is constructed across pattern 510, for example, to describe the relationships between the identified patterns 510 in terms of their similarity. A simplified example of the graph is illustrated in Figure 6 (b), where nodes indicate patterns (the corresponding exemplary patterns are shown next to each node) and the edges between two nodes indicate similarity. The thickness of the edges represents the magnitude of the similarity. Although each of the nodes is associated with a different identified pattern, only a small or relatively small subset of these patterns is initially labeled (e.g., at step 515). At step 610, these initial labels are propagated to all patterns according to the graph.

[0067] For such SSL step 525, there are a number of alternative methods that can be used. The optimal method for a given scenario can depend on the characteristics and / or size of the data. Some exemplary possible methods will be described in more detail later in this specification.

[0068] Returning to Figure 5 , at step 530, the corresponding health state 535 of each pattern is predicted based on the labeled data 527 obtained from the SSL step 525. The method can end at this point, or optionally continue with the following steps to improve the learning.

[0069] At the active learning step 540, a utility score 545 for each pattern can be estimated. Utility is a function that assigns a utility score indicating its information content to each pattern; for example, the labeling of such a pattern affects the estimation of the classification performance. The utility score for each pattern can include a combination of different measures of model uncertainty and pattern diversity. Uncertainty can be margin-based (e.g., the difference between the probabilities of the two most likely classes), entropy-based, or based on the probability of the most likely class. The diversity or representativeness of a pattern can be based on any definition of distance or similarity among the patterns. Graph-theoretic centrality measures such as degree, betweenness, and eigenvector centrality can also be used to indicate diversity and representativeness. An appropriate combination of these quantities may yield hyperparameters, which are learned by hyperparameter tuning techniques such as cross-validation or reinforcement learning.

[0070] At step 550, machines with the most informative patterns can be selected for labeling (step 555) based on the utility score 545. The number of selected machines can be defined based on a threshold or an expert time constraint or based on the difference in utility scores (e.g., in an elbow-like manner, where the elbow method is a heuristic used in clustering to determine the number of clusters in a dataset. These elbow methods are well-known and will not be described further).

[0071] At step 555, a domain expert can label the selected machines. This inserts new domain knowledge, as the domain expert can use additional information (such as interactions with users, overlapping data, or production data) for their labeling. Due to confidentiality issues, such information is usually not available in the application of the proposed method. However, this method is able to use such information in a systematic way. This additional labeling can be added to the partially labeled data 520 used in subsequent iterations of the method.

[0072] More exemplary details of the SSL step 525 will now be described. To construct the similarity graph, first, the distance or similarity between each pair of patterns should be determined. This can be achieved according to any suitable similarity or distance measure.

[0073] Such a similarity measure can be defined based on knowledge of the physics of degradation of the component under monitoring. For example, if for a first sensor, the drift rate of the measured signal defines aging degradation, a distance measure that captures the drift rate (e.g., correlation / covariance / cosine, etc.) can be used in the construction of the graph. Some specific similarity measures and algorithms that can be appropriately used will now be described.

[0074] For example, algorithms that can handle time series of different lengths can be used. These algorithms can include, for example, dynamic time warping (DTW), which is a shape matching algorithm that finds the optimal mapping between two time series by minimizing the cumulative alignment distance. Using DTW, time series of different lengths are naturally handled. As an alternative, the similarity measure can be the computationally accurate but slow time warping edit distance (TWED).

[0075] Traditional similarity functions or measures for time series of the same length can also be used, such as correlation, cross - correlation, Euclidean distance, cosine, edit distance (Levenshtein). These can be used by calculating any of these measures at selected representative points, such as the centroid, center point, or percentile. This solution is fast but less accurate.

[0076] Other similarity measures can include frequency - based similarity measures; for example, similarity functions that capture the dynamic characteristics of time series patterns. For example, similarities (or distances) based on Fourier and wavelet decompositions, spectral density, etc. can be used. Another example can include compression - based similarity measures (e.g., based on information theory).

[0077] Similarity measures can include domain expectations, on which patterns are considered patterns of similar behavior for a particular state, such as a faulty sensor. For example, if the drift rate is crucial for detecting a faulty sensor, angular distances such as correlation or cosine similarity can be used. If the shape of two patterns is crucial, dynamic time warping can be used.

[0078] Once the similarity between patterns is determined, a similarity graph can be constructed, where the graph encodes the structure of the entire set of time series patterns. The graph can be represented by an adjacency matrix W. Each node i represents a time series segment (pattern) and each entry W i,j indicates the weight of the edge connecting node i to another node j. As long as the weight W i,j decays to zero as the distance between the two nodes increases, the weight W i,j can be any function of the similarity / distance between patterns i and j: for example, an exponential, Gaussian, or quadratic kernel.

[0079] A method that can be selected depending on the amount of available data can be used to perform graph-based semi-supervised classification (label propagation step 610). First, such a method can include label propagation via a graph with potential additional sparsity constraints (e.g., sparse dictionary learning or low-rank models). Label propagation spreads the label information of a few available labeled samples to unlabeled samples to estimate their labels using a similarity graph. These methods assume that closer patterns have similar labels. Larger edge weights allow labels to be more easily propagated. Graph neural networks or any other suitable methods can also be used.

[0080] More specifically, based on the similarity graph, the label propagation method can construct an affinity matrix W and its corresponding Laplacian operator S as S = D -1 / 2 - W D -1 / 2 , where D is the diagonal matrix of W. The loss function of label propagation can be local and global consistency. This loss function can include two objectives: 1) a smoothness constraint that enforces consistency on the labels of adjacent data points; and 2) a fitting constraint that enforces that any change from the initial label assignment should be minimized and / or kept small in the final classification.

[0081] Additional sparsity constraints, such as sparse dictionaries and low-rank methods, can be imposed on the graph construction. Label propagation is probabilistic, and thus, for each pattern, all different labels can be considered as a distribution over the labels. Label propagation is a transductive process, meaning that label propagation cannot solve out-of-sample instances.

[0082] Another label propagation method can include generating graph-based pseudo-labels for a neural network. The generation of such pseudo-labels can be similar to clustering. This is also a transductive setting. The graph structure can be used as a clustering method to obtain pseudo-labels for unlabeled data points, and these pseudo-labels for unlabeled data points, together with the labeled samples, can be used to pre-train the neural network. Then the neural network can be fine-tuned using only the available labeled data points.

[0083] As an alternative, for label propagation, if sufficient data is available, a neural network can be used for semi-supervised learning. Since sensor measurements are usually high-dimensional signals with more than 100 dimensions, the graph structure can be used as a regularization to improve generalization to new data. Graph embedding can be written as a loss function such that graph embedding can be regarded as a hidden layer of the neural network. In such a case, the neural network can be regularized on top of the classification loss function (cross-entropy), and an additional loss predicts the graph context:

[0084]

[0085] where λ is a hyperparameter and can be any meaningful transformation of the Laplacian S of the similarity graph (e.g., as described above) when constructing the similarity graph, such as the L2 norm, or a loss function for graph embedding, i.e., the minimization between the distributions of distances in the high-dimensional space and the low-dimensional space. For example, UMAP can be used for the calculation of graph embedding. In such a context, the UMAP similarity estimate is interpreted as a probability, where p 𝑖𝑗 indicates the probability that two nodes i and j are connected in the high-dimensional space and q 𝑖𝑗 indicates the probability that two nodes i and j are connected in the low-dimensional space. The calculation result of UMAP is the loss function , and the loss function can be optimized by gradient descent:

[0086]

[0087] Compared with the above label propagation method, this method is inductive, meaning that this method can be generalized beyond the sample data points. The neural network can be an autoencoder, a CNN, or a simple feedforward network. This method is also closely related to the multi-task autoencoder, where the autoencoder is trained to optimize both the reconstruction error and the similarity of the data points in the original space.

[0088] As an extension of the main concepts disclosed above, these concepts can be used as tools for learning and / or updating domain knowledge in the form of rules. Unlabeled patterns are patterns that domain experts do not know how to relate to the specific states of the hardware components. In other words, the rules of the domain experts cannot cover the complete pattern database and some patterns remain unlabeled. The method disclosed in the present invention can be used to generate new rules.

[0089] To estimate the labels of unlabeled patterns, the classifier described above uses the labels of similar patterns defined by their corresponding graphs. Experiments using different definitions of similarity / distance can help domain experts define which aspects of the signals are crucial rules, such as the drift rate, shape, change. For example, if using the angular distance provides the optimal classification accuracy, the rules should be defined according to the drift rate. The estimated decision boundaries for classification can be used to estimate new rules for the thresholds of updating the rules. Engineers and users can use this knowledge when maintaining and calibrating the machine. Since the generation process of the rules can be described via graphs, the rules are interpretable.

[0090] Figure 7is a flowchart of an active learning method as illustrated. Input time series data 700 and domain knowledge / rules 705 are fed into a rule-based model 710 which includes a clustering / segmentation module 715 and a rule classifier 720. This generates the graph already described, and a label propagation step 725 propagates labels from the labeled data to the unlabeled data based on the graph. Aspects 700 to 725 of this method can be implemented as already described. An implementation learning loop includes the label propagation step 725, an active learning step 730 (e.g., the active learning steps 540, 550 described above), and an optional labeling step 735. In such a labeling step 735, domain experts can insert their knowledge into the graph by labeling patterns. Via a utility function (see Figure 5 : steps 540, 545), the proposed method can receive targeted input regarding rules. The added domain knowledge is stored and utilized in a systematic manner via the graph.

[0091] The output of the labeling step 735 is a new rule 740 which can be used to update the input rules for either this method or any of the other methods disclosed herein.

[0092] The new rule originates from decision boundaries determined during the classification process, which are derived from a graph encoding the physical properties of the degradation process. Domain knowledge is used in the construction of the similarity graph and the labels propagated across the graph. Thus, the classification results can be used to update the domain knowledge (e.g., rules, thresholds, etc.). Thus, decision boundaries for each class are obtained from classification on the graph; where the decision boundaries describe, for example, which patterns are at the edge of each class and / or closest to another class. Based on these patterns, a drift rate (or other measurement) separating the classes can be calculated, and the drift rate (or other measurement) is then used to define the new rule.

[0093] Figure 8 is a flowchart describing the application of the concepts disclosed herein to generate a training set for a machine learning model; e.g., to predict each point. The generated data set can be used to train other machine learning models to provide online predictions daily. While classifying patterns (data point clusters) rather than data points provides more stable results, the classification adds restrictions in the prediction because it does not allow the classification of individual data points. To overcome this, this embodiment is proposed to produce a "gold (exemplary)" or labeled reference data set 800 and use a more conventional machine learning method 805 to classify each instance.

[0094] This method is shown for Figure 7as a supplement to the process and will not repeat the description of elements 700 to 740 for that reason. The ML model 805 receives the pre-labeled data 800 output from the label propagation model 725. The output of the ML model 805 can be used by the active learning step 810, which also uses the production data 815 in addition. The remainder of the process is as described with respect to Figure 7 as described.

[0095] The concepts disclosed herein result in improved models due to more consistent labeling. Domain expert labeling can be lengthy and error-prone. Therefore, the method infers most time series labels and only requests input from domain experts when necessary to define decision boundaries. In this way, the labeling is more consistent. Additionally, less labeled data is needed to achieve the highest model performance.

[0096] Although specific embodiments of the invention have been described above, it should be understood that the invention can be practiced in other ways different from the described ways.

[0097] Although specifically referred to above in the context of the use of embodiments of the invention in optical lithography, it will be appreciated that the invention can be used in other applications (e.g., imprint lithography) and is not limited to optical lithography where the context permits. In imprint lithography, the topography in the patterning device defines the pattern created on the substrate. The topography of the patterning device can be pressed into a resist layer supplied to the substrate, where the resist is cured by applying electromagnetic radiation, heat, pressure, or a combination thereof. After the resist is cured, the patterning device is removed from the resist, leaving a pattern therein.

[0098] As used herein, the terms “radiation” and “beam” encompass all types of electromagnetic radiation, including ultraviolet (UV) radiation (e.g., having a wavelength of or about 365 nm, 355 nm, 248 nm, 193 nm, 157 nm, or 126 nm) and extreme ultraviolet (EUV) radiation (e.g., having a wavelength in the range of 1 nm to 100 nm), as well as particle beams, such as ion beams or electron beams.

[0099] The term “lens” can refer to any one or combination of various types of optical components (including refractive, reflective, magnetic, electromagnetic, and electrostatic types of optical components) where the context permits. Reflective components may be used in devices operating in the UV and / or EUV ranges.

[0100] The breadth and scope of the present invention should not be limited by any of the above exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

[0101] The following aspects are provided:

[0102] 1. A method for labeling time series data associated with one or more machines, the method comprising:

[0103] Obtaining the time series data;

[0104] Segmenting the time series data to obtain a plurality of patterns grouped according to pattern similarity;

[0105] Labeling a subset of the plurality of patterns to obtain a labeled pattern subset, the remaining patterns of the plurality of patterns including unlabeled patterns;

[0106] Defining a graph structure over the patterns, the graph structure describing the similarity between patterns; and

[0107] Using the graph structure and the labeled pattern subset to classify and / or label the unlabeled patterns to obtain labeled patterns.

[0108] 2. The method according to aspect 1, wherein the time series data includes sensor signal data from a plurality of sensors of the one or more machines.

[0109] 3. The method according to aspect 1 or 2, wherein the segmentation step uses a time series segmentation algorithm operable in the time domain, frequency domain, or spatial domain.

[0110] 4. The method according to aspect 3, wherein the time series segmentation algorithm includes at least one of the following: Gaussian segmentation, hidden Markov model, neural network for time series segmentation, t-distributed stochastic neighbor embedding, principal component analysis using clustering, or agglomerative clustering and dimensionality reduction algorithms defined by uniform manifold approximation.

[0111] 5. The method according to any of the preceding aspects, wherein the graph structure encodes physical properties described by the time series data.

[0112] 6. The method according to aspect 5, wherein the physical properties encoded by the graph structure can relate to the degradation of components involved in the time series data.

[0113] 7. The method according to any of the preceding aspects, wherein the graph structure is represented by an adjacency matrix, in which each node represents a pattern and each entry indicates the weight of the edge connecting the nodes, the weight being a function of the similarity between the patterns connected by the corresponding edge.

[0114] 8. The method according to aspect 7, including selecting a similarity metric for quantifying the similarity based on the physics knowledge of the components involved in the time series data.

[0115] 9. The method according to any one of the foregoing aspects, wherein the labeling step is based on domain knowledge and / or rules.

[0116] 10. The method according to any one of the foregoing aspects, wherein the labeling step includes applying the same label to all points of the corresponding pattern.

[0117] 11. The method according to any one of the foregoing aspects, wherein the step of classifying and / or labeling the unlabeled pattern includes applying a semi-supervised learning algorithm to the unlabeled pattern using a subset of labeled patterns.

[0118] 12. The method according to aspect 11, wherein the semi-supervised learning algorithm includes a label propagation algorithm that is operable to propagate the labels of the subset of labeled patterns to the unlabeled pattern according to the graph structure.

[0119] 13. The method according to aspect 12, wherein the label propagation algorithm uses a loss function that enforces consistency on the labels of adjacent patterns and / or mandates that any change in the labels from the subset of labeled patterns should be minimized and / or kept small.

[0120] 14. The method according to aspect 13, wherein the loss function is based on local and global consistency.

[0121] 15. The method according to any one of aspects 12 to 14, wherein the label propagation algorithm includes at least one sparsity constraint.

[0122] 16. The method according to aspect 15, wherein the sparsity constraint is a sparse dictionary learning constraint or a low-rank model constraint.

[0123] 17. The method according to any one of aspects 11 to 16, wherein the semi-supervised learning algorithm generates graph-based pseudo-labels based on the graph structure for training a neural network.

[0124] 18. The method according to any one of aspects 1 to 11, wherein the step of classifying and / or labeling the unlabeled pattern includes applying a neural network to classify the unlabeled pattern based on a subset of labeled patterns, wherein the graph structure is used as regularization.

[0125] 19. The method according to aspect 18, wherein the graph embedding is written as a loss function that is regarded as a hidden layer of the neural network.

[0126] 20. The method according to any one of the foregoing aspects, wherein defining the graph structure includes determining the degree of pattern similarity between each pair of patterns among the plurality of patterns according to a similarity metric.

[0127] 21. The method according to aspect 20, wherein determining the degree of pattern similarity includes using one or more of the following: dynamic time warping algorithm, time warping edit distance algorithm, correlation algorithm, cross - correlation algorithm, Euclidean distance algorithm, cosine algorithm, edit distance algorithm, or frequency - based similarity metric algorithm.

[0128] 22. The method according to any of the preceding aspects, including:

[0129] Determining a utility score indicative of the information content of each pattern;

[0130] Selecting one or more machines in the machine having the corresponding utility score of the pattern indicative of the most abundant information content; labeling the selected machines; and

[0131] Using the labeling in determining the subset of labeled patterns.

[0132] 23. The method according to any of the preceding aspects, including determining new rules for labeling or describing one or more of the patterns based on the determination of the decision boundary obtained in the classification step.

[0133] 24. The method according to any of the preceding aspects, including: generating a labeled reference data set; and using a machine learning model to classify individual data points of the time - series data based on the labeled reference data set.

[0134] 25. The method according to any of the preceding aspects, including using the labeled patterns to determine the device state of the one or more machines.

[0135] 26. The method according to aspect 25, wherein the device state describes the health state of at least one component of the one or more machines.

[0136] 27. The method according to aspect 25 or 26, including scheduling and / or performing maintenance actions on the one or more machines based on the device state.

[0137] 28. The method according to any of the preceding aspects, including using the labeled patterns to determine new rules and / or rule thresholds for labeling time - series data in the labeling step.

[0138] 29. The method according to any of the preceding aspects, including using the labeled patterns to generate labeled training data for a machine learning model.

[0139] 30. The method according to any one of the preceding aspects, wherein the one or more machines include one or more machines used in the manufacture of integrated circuits.

[0140] 31. The method according to any one of the preceding aspects, wherein the one or more machines include one or more lithographic exposure apparatuses.

[0141] 32. A computer program comprising program instructions operable to perform the method according to any one of the preceding aspects when run on a suitable device.

[0142] 33. A non - transitory computer program carrier comprising the computer program according to aspect 32.

[0143] 34. A processing device comprising:

[0144] the non - transitory computer program carrier according to aspect 33; and

[0145] a processor operable to run the computer program comprised on the non - transitory computer program carrier.

[0146] 35. A lithographic system comprising the processing device according to aspect 34.

Claims

1. A method for labeling time series data associated with one or more machines, the method comprising: Obtaining the time series data; Segmenting the time series data to obtain a plurality of patterns grouped according to pattern similarity; Labeling a subset of the plurality of patterns to obtain a labeled pattern subset, with the remaining patterns among the plurality of patterns including unlabeled patterns; Defining a graph structure over the patterns, the graph structure describing the similarity between the patterns; And Using the graph structure and the labeled pattern subset to classify and / or label the unlabeled patterns to obtain labeled patterns.

2. The method according to claim 1, wherein, The time series data includes sensor signal data from a plurality of sensors of the one or more machines.

3. The method according to claim 1 or 2, wherein The segmentation step uses a time series segmentation algorithm operable in the time domain, frequency domain, or spatial domain.

4. The method according to claim 3, wherein, The time series segmentation algorithm includes at least one of the following: Gaussian segmentation, hidden Markov model, neural network for time series segmentation, t-distributed stochastic neighbor embedding, principal component analysis using clustering, or agglomerative clustering and dimensionality reduction algorithms defined by uniform manifold approximation.

5. According to the method described in any one of the preceding claims, wherein, The graph structure encodes physical properties described by the time series data.

6. The method according to claim 5, wherein, The physical properties encoded by the graph structure can relate to the deterioration of components involved in the time series data.

7. The method according to any one of the preceding claims, wherein The graph structure is represented by an adjacency matrix, in which each node represents a pattern and each entry indicates the weight of an edge connecting the nodes, the weight being a function of the similarity between the patterns connected by the corresponding edge of the pattern.

8. The method according to claim 7, comprising: A similarity metric for quantifying the similarity is selected based on the physics knowledge of the components involved in the time series data.

9. The method according to any one of the preceding claims, wherein, The labeling step is based on domain knowledge and / or rules.

10. The method according to any one of the preceding claims, wherein, The step of classifying and / or labeling the unlabeled patterns includes applying a semi-supervised learning algorithm to the unlabeled patterns using the labeled pattern subset.

11. The method according to claim 10, wherein, The semi-supervised learning algorithm includes a label propagation algorithm, which is operable to propagate the labels of the labeled pattern subset to the unlabeled patterns according to the graph structure.

12. The method according to claim 10 or 11, wherein The semi-supervised learning algorithm generates graph-based pseudo-labels based on the graph structure for training a neural network.

13. The method according to any one of claims 1 to 10, wherein, The step of classifying and / or labeling the unlabeled patterns includes applying a neural network to classify the unlabeled patterns based on the labeled pattern subset, where the graph structure is used as regularization.

14. The method according to any one of the preceding claims, wherein, Defining the graph structure includes determining the degree of pattern similarity between each pair of patterns among the plurality of patterns according to a similarity metric.

15. The method according to any one of the preceding claims, comprising: Determining a utility score for each pattern indicating the amount of information of the pattern; Selecting one or more machines in the machine having the corresponding utility score of the pattern indicating the most abundant information; Labeling the selected machines; And Using the labeling when determining the labeled pattern subset.

16. The method according to any one of the preceding claims, comprising: Determining new rules for labeling or describing one or more of the patterns based on the determination of the decision boundary obtained in the classification step.

17. The method according to any one of the preceding claims, comprising using the marked pattern to determine the equipment status of the one or more machines.

18. The method according to claim 17, wherein, The equipment status describes the health status of at least one component of the one or more machines.

19. The method according to claim 17 or 18, comprising scheduling and / or performing maintenance actions on the one or more machines according to the equipment status.

20. The method according to any one of the preceding claims, wherein, The one or more machines include one or more machines used in the manufacture of integrated circuits.

21. A computer program, comprising program instructions that are operable to perform the method according to any one of the preceding claims when run on a suitable device.

22. A non-transitory computer program carrier, comprising the computer program according to claim 21.

23. A processing device, comprising: The non-transitory computer program carrier according to claim 22; and a processor that is operable to run the computer program included on the non-transitory computer program carrier.

24. A lithography system, comprising the processing device according to claim 23.

Citation Information

Patent Citations

  • Method and apparatus for angular-resolved spectroscopic lithography characterization

    US20060033921A1

  • Method and apparatus for angular-resolved spectroscopic lithography characterization

    US20060066855A1

  • Inspection Apparatus, Lithographic Apparatus, Lithographic Processing Cell and Inspection Method

    US20100201963A1

  • Methods and Scatterometers, Lithographic Systems, and Lithographic Processing Cells

    US20110027704A1

  • Metrology Method and Apparatus, Lithographic Apparatus, Device Manufacturing Method and Substrate

    US20110043791A1