Method and apparatus for identifying variables from a plurality of variables that have a dependency on a predetermined variable from the plurality of variables
Patent Information
- Application Number
- DE102024200425
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-17
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for identifying variables from a plurality of variables which have a dependency on a predetermined variable from the plurality of variables, preferably using a random forest, as well as to a device, a computer program and a machine-readable storage medium. State of the art
[0002] Determining the causes of a failure in a system, especially in a production system, is known as 'root cause analysis'.
[0003] The publication by Solé, M., Muntés-Mulero, V., Rana, AI, & Estrada, G. (2017). Survey on models and techniques for root-cause analysis. arXiv preprint arXiv:1701.08546, available online: https: / / arxiv.org / pdf / 1701.08546.pdf, provides an overview of techniques for data-based modeling of system behavior and root cause analysis based on modeling.
[0004] Although the use of machine learning (ML)-based models, such as a random forest, allows for the modeling of complex dependencies, they present another challenge: In contrast to classic, statistical (parameterized) models, whose parameters often allow direct conclusions about the "importance" and "effect" of individual variables for the model or the model's predicted values, it is generally not possible to directly identify these "importance" and "effect" metrics, especially with complex ML-based models. As long as the purpose of a model is only to predict a dependent variable as accurately as possible, this is not a problem, but in the case of root cause analysis, additional metrics for "importance" and "effect" are needed to interpret the model results.These interpretation methods are often referred to as "Interpretable Machine Learning" or "IML" (Molnar et al.: "General Pitfalls of Model-Agnostic Interpretation Methods for Machine Learning Models" | SpringerLink, https: / / link.springer.com / chapter / 10.1007 / 978-3-031-04083-2 4). The additional computational effort associated with these methods can be considerable, which is why their practical application for interactive root cause analyses, where the expertise of process experts can also be incorporated, presents a challenge.
[0005] From the unpublished DE 10 2022 208 394, a method for identifying at least one variable with high importance from a plurality of variables, which has a correlation with a predetermined variable from the plurality of variables, is known. The method begins by providing a data set comprising data points for a plurality of variables for a plurality of products, and selecting the predetermined variable from the plurality of variables. This is followed by preprocessing the data set. This is followed by training a machine learning system on the preprocessed data set, and determining the dependencies of the variables on the predetermined variable based on the trained machine learning system.
[0006] The importance of individual variables in a model can be extracted using various well-known methods (e.g., permutation importance or impurity importance). These methods yield different results from the same model and have different advantages and disadvantages.
[0007] For example, impurity importance is an assessment that is almost “free” when using decision tree models, but which tends to systematically underestimate the importance of categorical variables with a small number of categories (see, for example, Strobl et al: https: / link.springer.com / article / 10.1186 / 1471-2105-8-25).
[0008] An alternative that is not limited to decision trees is the so-called permutation importance, which in certain cases, however, provides unreliable results because, among other things, it determines model errors in extrapolated value ranges where these are naturally high (Hooker et. al: https: / / arxiv.org / abs / 1905.03151).
[0009] However, since both of the above-mentioned importance metrics can deliver very different values in individual cases, it has proven impractical to return both values separately to the user.
[0010] A variety of methods have been developed in recent years for the extraction of “effects”, some of them even model-agnostic, i.e. they can be applied to any type of machine learning model (comparable to permutation importance) and are not limited to specific models (like impurity importance).
[0011] A distinction is made between methods that extract “effects” for individual predictions (“local effects”) and those that do so for an entire data set (“global effects”) (e.g. Molnar: “Interpretable Machine Learning” | https: / / christophm.github.io / interpretable-ml-book / index.html).
[0012] However, all of the above-mentioned methods have in common that it is not obvious to the user below which importance or effect value a variable no longer makes a meaningful contribution to the model or its predictive value. Variables that do not make an important contribution to the model can be removed from the model in order to obtain a less complex and therefore more robust model. Variables that do not have a sufficiently strong lever (“effect”) to influence the value of the dependent variable in the desired way and strength when changed accordingly may not be useful for problem solving. It is therefore important to provide a threshold value for “importance” or “effects” metrics below which variables only contribute to the model as noise or do not offer sufficient leverage. Even if, for example,While there are feature selection algorithms for variable selection that automate the reduction of the variables used based on importance metrics, such as Boruta (Kursa et. al: https: / / www.researchgate.net / publication / 220443685 Boruta - A System for Feature Selection), such methods are sometimes very computationally intensive, often require a large number of observations to function effectively and do not allow the subject matter experts to intervene based on their expertise.
[0013] The same applies to the removal of correlated variables: here, too, it is not obvious to the user below which threshold a correlation measure > 0 may be due exclusively to noise.
[0014] From the unpublished DE 10 2023 202 838.7, a method for determining thresholds for "importance" and the correlation measure is known. This method is based on enriching the data set with a categorical and a numerical random variable. This method can also be used to determine thresholds for "effects" or allows the computationally intensive determination of "effect" metrics to be limited to only important variables.
[0015] It is an object of this invention to improve and simplify the root cause analysis so that a preferably interactive real-time use is possible. Advantages of the invention
[0016] One advantage of the present invention is to provide an importance metric that generates a single, reliable, and easily sortable key figure that allows variables to be sorted by importance. This also makes it possible to further automate a root cause analysis and, for example, to iteratively train additional models based only on the top N variables or to perform additional, computationally intensive calculations to extract the "effects" only for the top N variables. This means that the importance metric makes it possible to sort the variables in the model by importance based on a key figure and, if necessary, to automate further steps (such as removing unimportant variables from the model or calculating additional metrics only for important variables).
[0017] A further advantage of the invention is the introduction of three additional "effect" metrics: An effect size metric, which allows the assessment of the leverage of individual variables. An effect metric, which allows the visualization of the shape and course of the effect of an individual variable on the dependent variable. A third metric, which allows the quantification or visualization of interactions between two variables. Disclosure of the invention
[0018] In a first aspect, the invention relates to a computer-implemented method for identifying at least one variable from a plurality of variables that exhibits a dependency on a predetermined variable from the plurality of variables. The variables each characterize measurements after production steps of a product or the production steps or machine settings of machines that perform one of the production steps. The dependency can be understood as an abstract (causal) relationship through which the identified variable influences the predetermined variable; i.e., the dependency can be understood as a correlation.
[0019] The method begins with providing a data set comprising data points for a plurality of variables for a plurality of products. The data set is preferably a matrix or table. The columns of the matrix are each assigned to one of the variables, while the rows each contain a data point that was recorded for the respective product for the respective variable. A row can also be understood as a series of measurements. The data set can be sparse. It is therefore conceivable that the data set, in particular the matrix, has empty entries along the dimension for the variables as well as along the dimension for the products. This means that there are variables whose measurements or similar have not been measured for the respective product, but other measurements may have already been carried out for this product.
[0020] The products may be semiconductor products such as wafers or frames (a frame can be understood as an exposure field on a wafer, i.e., a repeating arrangement of individual chips on the wafer) or chips or other semiconductor components, which were manufactured, in particular, using the same production machines or in the same factory. The majority of the products may be identical products or differ from one another with regard to certain configurations. It is also conceivable that the majority of the products are different products, which were manufactured, in particular, using the same production machines or in the same factory.
[0021] This is followed by selecting or defining the specified variable from the plurality of variables. The specified variable should be the variable for which, for example, an intolerable deviation from its value range or values outside a defined value range have been observed. The specified variable can be selected due to its abnormal behavior for one product or for a majority of products. A root cause analysis should preferably be performed for the specified variable. It is conceivable that the specified variable is provided by a user.
[0022] With regard to preprocessing of the data set, it should be noted that, particularly in the case where the data set contains a very large number of variables, a manual preselection of those variables can be made which are suspected of having an influence on the given variable.
[0023] This can be followed by preprocessing the data set. The preprocessing step comprises at least extending the provided data set by a numerical and / or categorical random variable characterizing a probability distribution, wherein data points for the random variable are randomly drawn according to the probability distribution and added to the provided data set.
[0024] The step of preprocessing the data set can additionally include the following steps: During preprocessing, those variables and / or products which have a sparsity of data points greater than a specified threshold are deleted, particularly row-by-row and / or column-by-column. Sparseness can be understood as not having data points for every product for a particular variable or not having data points for every variable for a particular product. Sparseness can be specified as a percentage, e.g., how many data points should ideally be present and how many of these are missing. The threshold is, for example, a percentage. For example, the threshold for the data points of the products is ≤ 10% and for the variables < 60%. The advantage of the first step is that it effectively achieves a reasonable sparsity of the data set.It is conceivable that the first preprocessing step differentiates between missing data points that were not recorded but could have been recorded, and missing data points because the product was not manufactured for them. Missing data points that were not recorded but could have been recorded remain empty and are then removed during sparseness filtering if necessary. For missing data points that were not recorded and for which the product was not manufactured for them, "not processed" can be stored as a placeholder data point during the first step. This provides a simple way to account for the significance of missing data and advantageously reduces the sparsity of the data set.
[0025] The step of preprocessing the data set can also include the following additional steps: During preprocessing, additional missing data points of the variables are imputed. During imputation, missing data points are replaced in a first step with a data point of the respective variable that occurs most frequently (categorical variables) or occurs on average (e.g. median) across the majority of products (numerical variables). Based on this imputed data set, a first machine learning system can be trained so that it predicts the originally missing data points based on the given data points. This step can be repeated several times until a termination criterion is reached in order to iteratively improve the quality of the imputed data. An example of such a "multiple imputation" algorithm is MICE (doi:10.18637 / jss.v045.i03>>).The model form used for this first learning system is, in principle, adaptable, but in the context of root cause analysis, it is advisable to choose a form that can handle nonlinearities and correlations between the variables. Alternatively, in the next step, you can also use a machine learning system that can handle incomplete data (Samuele Mazzanti, "Your Dataset Has Missing Values? Do Nothing!" | https: / / towardsdatascience.com / your-datasethas-missing-values-do-nothing-10d1633b3727).
[0026] This is followed by training a second machine learning system on the (preprocessed) dataset. Training can be done by minimizing an MAE or RMSE for regression, or by maximizing an accuracy or kappa for classification. After training is complete, the dependencies of the variables on the given variable are determined based on the second trained machine learning system.
[0027] Importance is preferably determined by aggregating the permutation importance and impurity importance. The importance metric is particularly preferably determined as a lower estimate of the permutation importance (PI) and impurity importance (II) plus a correction factor. The permutation importance (PI) and impurity importance (II) values can be aggregated into an importance metric according to the following formula to enable easy sorting and selection: Importance metric=Min(PI,II)+0.75*Range(PI,II).
[0028] This proposed importance metric has the advantage that even variables that only receive high importance according to one metric appear prominently and quite high in the ranking for the user. This can, for example, prevent categorical variables with few categories that tend to have a lower II value from slipping too far in the importance metric. At the same time, one does not have to rely solely on the less stable PI, which is prone to greater fluctuation due to the fact that it is partially determined by extrapolation. This provides a robust and variable-type-independent metric that allows for better conditional analysis of the system.
[0029] Preferably, the effect size metric is determined using a range operator applied to the accumulated local effects or ALE value to enable easy sorting and selection: Effect Size Metric=Range(ALE)
[0030] This proposed effect size metric has the advantage that no assumptions need to be made regarding the nature of the effect (as would be the case, for example, with the slope of a linear regression). At the same time, when calculating the influence of changed variable values, only small changes are made to these values, thus avoiding extrapolation and suppressing contributions from other variables. The ALE values now provide an estimate of the change in the dependent variable as a function of each independent variable, provided there are no strongly correlated variables.Calculating the rank of these ALE values provides an estimate of the maximum change in the dependent variable across the entire range of the independent variable. The range (ALE) values of all independent variables are comparable because they are in units of the dependent variable alone, and the unit of the independent variable is irrelevant (the effect size metric). For a calculation of ALE values, see Apley, DW & Zhu, J. 2020, "Visualizing the effects of predictor variables in black box supervised learning models," Journal of the Royal Statistical Society. Series B, Statistical methodology, vol. 82, no. 4, pp. 1059-1086.
[0031] Advantageously, variables whose importance and / or effect size are less than or equal to those of the random variable are discarded. The remaining variables can be displayed in descending order according to their importance and / or effect metrics, or they can be presented in a 2D scatter plot with these two metrics. A user can identify one or more variables from the order whose importance in the model or whose effect on the target variable appears to be the strongest.
[0032] A root cause analysis can then be performed for these variables. This allows both the values of the dependent variable and the progression of the ALE values across the value range of a variable to be displayed and analyzed. Furthermore, 2nd-order ALE values for the interaction of the selected variable with other variables can be evaluated (for performance reasons, 2nd-order ALE values or analogous effect size values are not calculated for all possible interactions by default; this can be done for individual or all variables as needed).
[0033] It is proposed that the data points of the random variable be filtered or deleted from the dataset according to a predefined sparseness factor for each product. The advantage of this is that the sparseness of the original dataset can be maintained. Thus, the random variable hardly changes any characteristics of the dataset.
[0034] It should be noted that training is preferably performed through the following steps: dividing the dataset into a training, test, and validation dataset. It is also conceivable that the dataset is further divided into a prediction dataset, which includes non-existent data points for the given variable. This is followed by the creation of the machine learning system, whereby the machine learning system is trained with selected hyperparameters on the training data, the hyperparameters of the machine learning system are evaluated on the validation dataset, and the trained machine learning system is evaluated on the test dataset.
[0035] The second machine learning system is proposed to be a random forest. A random forest is an ensemble of decision trees constructed in slightly different ways, with each tree being provided with a different subset of variables during training. This means that each tree receives a different subset of the training data and leaves behind a subset that is not used ("out-of-bag samples"). The random forest preferably contains 1500-3000 trees. Hyperparameter tuning of the trees is particularly preferred based on the "out-of-bag samples," which requires less validation data.
[0036] A majority vote of all trees forms the final model, and the predictor variables that contribute most to purity gain across all trees are ranked with the highest importance (impurity importance). In addition to impurity importance, permutation importance can be used.
[0037] The advantage of the Random Forest as a data-based modeling technique is that complex, non-linear correlations, especially between categorical variables, can be discovered without prior expert assumptions. Another advantage is that experiments have shown that the Random Forest exhibits low overfitting and thus reliable and robust behavior. A particular advantage is that the Random Forest can handle non-normalized data, missing data, as well as continuous and categorical data. This makes the invention particularly useful for semiconductor production, where measurement data is typically sparse. In combination with the data set preprocessing procedures described above, a dependency analysis is provided that can be reliably used for a wide range of different applications.
[0038] Furthermore, it is proposed that during preprocessing of the data set after the first step, a correlation is additionally determined pairwise between a large number of variables. For pairs exhibiting a correlation that is, for example, greater than a predetermined threshold, only one variable of the respective pair is selected and the second variable is removed from the data set. It has been shown that the robustness of the method can be significantly increased by deleting strongly correlated pairs. Advantageously, variables whose correlation is less than or equal to a correlation between the random variable and other variables are also removed from the data set. This allows the dependencies to be determined even more reliably.
[0039] It is further proposed that the correlation be determined based on normalized mutual information. The normalized mutual information is given by: 2 ∗ I(X;Y) / (H(X) + H(Y)), where I is the Shannon mutual information and H is the entropy.
[0040] Furthermore, it is proposed that, in the case of classification, class balancing, such as upsampling an underrepresented class or category, be performed during training. Class balancing has the advantage that the variables are more evenly distributed in the training data, resulting in a balanced training dataset, which can be used to generate meaningful models for classifications with unbalanced classes (e.g., 95% good parts, 5% bad parts).
[0041] Furthermore, it is proposed that variables that characterize different components of a common product be aggregated. For example, multiple chips can be aggregated into a frame and / or multiple frames into a wafer. An aggregation method such as mean, median, P10, or P90, etc., can be applied to the data points of the aggregated products. This approach has the advantage that data sets that are too large for RAM memory can be compressed accordingly. It is also conceivable that the aggregation is performed across multiple variables across multiple products.
[0042] Furthermore, it is proposed that the data points were recorded in a semiconductor factory. In particular, the data points of the variables are inline measurements and / or PCM measurements and / or wafer-level tests and / or characterize a wafer processing history. The wafer processing history describes, for example, which tool was used to process the wafer and / or which recipe was used. In particular, the variables of the wafer history can be categorical variables, e.g., chamber A or B of the tool, or similar tools for the individual production steps.
[0043] In a further aspect of the invention, the trained second machine learning system according to the first aspect of the invention can be used to predict the variables for future production steps, in particular to predict measurements that may be obtained and, if necessary, to then decide whether the product should be further processed.
[0044] In further aspects, the invention relates to a device and a computer program, each of which is configured to carry out the above methods, and to a machine-readable storage medium on which this computer program is stored.
[0045] Embodiments of the invention are explained in more detail below with reference to the accompanying drawings. In the drawings: Fig. 1 schematically shows a flow diagram of an embodiment of the invention; Fig. 2 schematically shows an embodiment for controlling a manufacturing system; and Fig. 3 schematically shows a training device. Description of the embodiments
[0046] Fig. Figure 1 schematically shows a method for identifying or finding at least one variable from a plurality of variables that exhibits a dependency on a predetermined variable from the plurality of variables. For example, deflections in a first measurement V T observed that cannot be easily explained by the usual suspect measurements, such as THK GD Variations. Using the method according to Fig. 1 variables, in particular their associated measurements, which have an influence on the first measurement can be found in order to find a cause for the anomaly of the first measurement.
[0047] The method begins with providing (S21) a data set comprising data points for a plurality of variables for a plurality of products, such as semiconductor components. In this exemplary embodiment, the data set is available as a matrix or table. Subsequently, the variable is selected from the plurality of variables that exhibits atypical behavior or a deviation, as exemplified above: the measurement V T .
[0048] In general, the following variables are conceivable for semiconductor production: inline tests (layer thicknesses, depths / widths of structures, etc.), PCM test data (captured test measurements from individual component tests), wafer level test (EWS), wafer histories (e.g., which production machine processed which wafer), and / or categorical variables. Categorical variables can be, for example, chamber A or B of a tool, or similar tool variables such as a recipe.
[0049] Preferably, the data set is available as a table or is formatted as a table, with the columns representing the measurements / variables and the rows representing the wafers, frames, or chips (both the aggregation level (wafer, frame, or chip) can be configured, as can the aggregation method (mean, median, P10, etc.). This means that each column is assigned a variable, and each row is assigned a wafer, frame (=litho shot), or chip.
[0050] This is followed by optional preprocessing (S22) of the dataset. The preprocessing step (S22) first involves artificially augmenting the dataset with random variables and then preprocessing or imputing the artificially augmented dataset.
[0051] The step of artificially expanding the data set involves several intermediate steps. In the first intermediate step of artificially expanding, two random variables are preferably added to the data set: a numerical random variable and a categorical random variable, preferably with only two categories. The numerical random variable can describe a normal distribution with a mean of 0 and a standard deviation of 1. The categorical random variable can describe a uniform distribution of two categories. Other distributions are conceivable.
[0052] In the second intermediate step of step (S22), a filter grid search is performed, which applies a larger number of possible filter settings (different thresholds for the permitted sparseness in each row and column, as well as different order in which the filters are applied to rows or columns) to the data set. The filter settings are each configured to create the effect of filtering the typically high-dimensional data sets with high sparseness. Depending on the (specified) sparseness to be achieved and / or depending on a number of usable samples and variables and / or depending on the (specified) sample / variable ratio to be achieved, the user can select suitable filter settings from the various resulting data sets.
[0053] This means that after the filter grid search step, a multitude of smaller datasets are available. These datasets were created from the original (typically incomplete) dataset by applying different thresholds to remove variables (columns) or observations (rows). These smaller datasets contain fewer columns and / or rows and fewer gaps. Preferably, the user will now select one of these datasets based on the information about the remaining rows, columns, and resulting gaps.
[0054] After the artificial expansion of the data set is completed, a first step of preprocessing the data set follows, in which a row- and / or column-wise deletion of those variables and / or products occurs which have a sparsity of the data points greater than a predefined threshold.
[0055] Additionally or alternatively, a step of removing correlated variables can be performed. To do this, a correlation value can be calculated between all remaining columns, each pairwise complete, using all pairwise complete rows. Using a correlation matrix (which also contains the correlation values of the two random variables with all others as a reference), the user can set a threshold for the correlation value, above which one of every two correlated variables is removed (the one with the higher average correlation with all other variables is removed, as it carries less additional information). Alternatively, the user can manually remove one of two correlated variables based on their expertise.
[0056] In the second step of preprocessing, missing data points of the variables are imputed. For performance reasons, the imputation is carried out in several stages, i.e. the quality of the imputed values is improved iteratively. In the first stage of imputation, missing data points are replaced with a data point of the respective variable that occurs most frequently (categorical variables) or occurs on average across the majority of products (numerical variables). In the second stage of imputation, a first machine learning system is trained so that it can predict the imputed values from the first stage with greater accuracy. The replaced data points are then replaced with new data points from the first machine learning system. This process is repeated until a termination criterion is reached (e.g. number of iterations or relative change from the impute value n to the impute value n+1).Preferably, a model is trained only on the training part of the dataset, which can handle nonlinearities and correlations between variables (e.g., a random forest), and then imputed values in both the training and test datasets. Alternatively, multiple datasets can be imputed independently, and then multiple models can be trained on them in the next step.
[0057] After the dataset has been processed according to the preprocessing step (S22), a second machine learning system is trained (S23) on the preprocessed dataset. For training, the dataset can be split, as is known, into train / test data, i.e., a train dataset on which the second machine learning system learns to predict the target variable, and a test dataset to test the trained second machine learning system. It should be noted that the dataset split can alternatively be performed before the imputation step and only subsequently afterward.
[0058] In a preferred embodiment of training S23, hyperparameter tuning and model training are performed. Preferably, the second machine learning system is a random forest. The random forest is an ensemble of a large number of decision trees, each of which can be trained with a different subset of variables and samples (in the case of classification, class balancing occurs after sample selection), and a majority vote across all trees forms the final model.
[0059] Two hyperparameters can be optimized for the random forest: the number of variables used in a decision tree and the minimum number of observations in a final node. The search space for the former is automatically adjusted based on the total number of variables, and tuning is performed automatically based on the out-of-bag error (kappa for classification and RMSE for regression). The latter can be manually adjusted by a user based on the results of an initial model, although automated hyperparameter tuning would also be conceivable here. Typical model performance metrics for training and testing are provided to the user.
[0060] After the training step (S23), a verification (S24) of the trained model on the test dataset can optionally be carried out.
[0061] In the event that the target variable is not available in the extended dataset, but all other variables or a large number of the other variables are available, the target variable can be predicted and added to the extended dataset.
[0062] Subsequently, a test (S25) of the trained second machine learning system can be performed, particularly by the user. The test may have the goal of starting another run of the training S23 with adjusted settings (e.g., a reduced selection of variables or different thresholds, etc.).
[0063] The assessment can be based on one or more of the criteria listed below: 1. The test can be performed with regard to model quality. This can include an assessment of overfitting and, if necessary, a corresponding adjustment of the hyperparameter "minimum number of samples for leaf nodes of the decision trees." 2. The test can be performed for a relative prediction error with respect to the target variable. Here, the prediction error is plotted against the target variable to assess whether the error changes systematically with the value of the target variable. If this is the case, this may indicate that important variables were not part of the model, and this can be taken into account when adjusting the variable selection. 3. The review can be performed based on a variable importance or effect size ranking. First, the variables that were excluded due to sparseness or correlation can be reviewed by a domain expert. Next, the variables that are part of the model are examined. The importance or effect size metrics of the two random variables serve as a reference, and those variables with lower metric values than the random variables can be removed from the dataset after expert review. 4. To examine variables excluded due to correlation, a mutual information-based correlation matrix can also be used. This helps to understand which variables are highly correlated and, in addition to verifying general correlation patterns, allows to understand which second variable caused a first variable to be removed due to excessive correlation. If, from the domain expert's perspective, the "wrong" variable was removed from two correlated variables, the user can correct the model's decision. Furthermore, the correlation values of the random variables with "true" variables can be used to adjust the threshold for removing correlated variables for a subsequent run, if necessary.
[0064] After step S25 is completed, step S26 follows. This step identifies the top N variables. The proposed importance metric, which is an aggregation of permutation importance and impurity importance, is used to identify the top N variables. The user can freely choose N, or all N variables above the importance of the random variable can be automatically determined based on the random variables. Furthermore, the user can adjust model parameters to control over- / underfitting.
[0065] After step S27 has been completed, step S27 follows. This determines the importance and effects metrics for the top N variables determined from step S26. For the top N variables and the two random variables, an importance and an effects metric are calculated on the test dataset. Since this is only done for a subset of variables, more "expensive" methods can be used here than for determining the top N, and an additional error estimate can be calculated. The user can view the metrics, the raw data, and the effect in the model only for the top N variables and, using expert knowledge, verify their plausibility or derive suitable measures or experiments.
[0066] In a preferred embodiment of the method according to Fig. 1, the categorization quality of the second machine learning system is evaluated after step S23 for fine-tuning. For example, variables with a low determined importance or effect value can be removed from the training dataset according to step S27, and the sequence of steps S23 to S27 can be repeated. Another possible way to evaluate the second machine learning system is to consider prediction errors of the machine learning system and, depending on the error severity, retrain the second machine learning system.
[0067] Variables with a low importance or effect value can be removed based on a predefined threshold. The threshold can be predefined.
[0068] Particularly preferred for evaluation is a result of the procedure according to Fig. 1 interactively visualized, allowing an expert to conduct a targeted evaluation via the visualization. This also has the advantage that expert knowledge is implicitly incorporated through the interactive visualization.
[0069] The dependencies resulting from step S25 can then be used to optimize production steps and enable faster and more informed design of experiments (DoE) when coordinating machines during the start-up phase of new production lines.
[0070] For example, if a measurement V T outside a specified or tolerable range, the Variable Importance Ranking and / or the Effect Size Ranking from step S25 can be used to determine which variables from the data set correspond to the corresponding variable of the measurement V Tcorrelate most strongly or can influence them most strongly. From these correlating variables, it can then be deduced to what extent a production process that has an influence on the correlating variables needs to be adjusted. This makes it possible to trace which production step led to the erroneous measurements. That is, the procedure according to Fig. 1 can be used for a root cause analysis. The adjustment can be achieved, for example, by adjusting process parameters, preferably by adjusting these process parameters accordingly in a control system.
[0071] It is also conceivable that the adjustment of the process parameters depends on an absolute deviation of the given variable from the specified or tolerable value range and can be carried out depending on an importance value of the variable importance ranking from step S25 of the correlating variables and optionally based on a physical domain model that characterizes dependencies between the production steps and correlating variables. It should be noted that, depending on the above-mentioned procedure, variables can be identified that are redundant (because they are highly correlated), and thus their associated tests can be removed. This leads to a reduction in the list of required tests and thus to a beneficial reduction in measurement time.
[0072] Fig. 2 shows an embodiment in which the control system 40 with the adapted process parameters is used to control a manufacturing machine 11 of a manufacturing system 200, in that this manufacturing machine 11 controls the actuator 10. The manufacturing machine 11 can be, for example, a machine for punching, sawing, drilling, milling, and / or cutting, or a machine that performs a step of semiconductor manufacturing, such as chemical coating processes, etching and cleaning processes, physical coating and cleaning processes, ion implantation, crystallization, or temperature processes (diffusion, annealing, reflow, etc.), photolithography, or chemical-mechanical planarization.
[0073] Sensor 30 may, for example, be a measuring sensor or detector which, for example, detects properties of manufactured products 12a, 12b, wherein the data points detected thereby preferably correspond to the method according to Fig. 1 are provided.
[0074] Fig. 3 schematically shows a training device 500 comprising a provider 51, which provides instances from a training data set. The instances are fed to the machine learning system 52 to be trained, which determines therefrom, for example, a classification or regression as output variables. Output variables and instances are fed to an evaluator 53, which determines therefrom updated hyper- / parameters, which are transmitted to the parameter memory P and replace the current parameters there. The evaluator 53 is in particular configured to carry out steps S23 of the method according to Fig. 1 to execute.
[0075] The methods executed by the training device 500 can be implemented as a computer program stored on a machine-readable storage medium 54 and executed by a processor 55.
[0076] The term "computer" encompasses any device capable of executing specified computational instructions. These computational instructions can be in the form of software, hardware, or a combination of software and hardware. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] DE 10 2022 208 394
[0005] DE 10 2023 202 838.7
[0014] Zitierte Nicht-Patentliteratur
[0000] Solé, M., Muntés-Mulero, V., Rana, A. I., & Estrada, G. (2017). Survey on models and techniques for root-cause analysis. arXiv preprint arXiv:1701.08546
[0003] https: / / arxiv.org / pdf / 1701.08546.pdf
[0003] Molnar et. al.: „General Pitfalls of Model-Agnostic Interpretation Methods for Machine Learning Models“ | SpringerLink, https: / / link.springer.com / chapter / 10.1007 / 978-3-031-04083-2 4
[0004] Strobl et. al: https: / link.springer.com / article / 10.1186 / 1471-2105-8-25
[0007] Hooker et. al: https: / / arxiv.org / abs / 1905.03151
[0008] Molnar: „Interpretable Machine Learning“ | https: / / christophm.github.io / interpretable-ml-book / index.html
[0011] Kursa et. al: https: / / www.researchgate.net / publication / 220443685 Boruta - A System for Feature Selection
[0012] Algorithmen ist MICE (doi:10.18637 / jss.v045.i03>
[0025] Samuele Mazzanti, „Your Dataset Has Missing Values? Do Nothing!“ | https: / / towardsdatascience.com / your-datasethas-missing-values-do-nothing-10d1633b3727
[0025] Apley, D.W. & Zhu, J. 2020, „Visualizing the effects of predictor variables in black box supervised learning models“, Journal of the Royal Statistical Society. Series B, Statistical methodology, vol. 82, no. 4, pp. 1059-1086
[0030]
Claims
[1] Computer-implemented method for identifying at least one variable from a plurality of variables which has a dependency on a predetermined variable from the plurality of variables, wherein the variables each characterise measurements after production steps of a product (12a, 12b) or the production steps or machine settings of machines (11) which carry out one of the production steps, comprising the following steps: Providing (S21) a data set comprising data points for a plurality of variables for a plurality of products (12a, 12b) in each case, and providing the predetermined variable from the plurality of variables; Preprocessing (S22) of the data set comprising extending the provided data set by at least one random variable characterizing a probability distribution, wherein data points for the random variable are randomly drawn according to the probability distribution and added to the provided data set; Training (S23) a second machine learning system, in particular a random forest, on the preprocessed data set; and determining (S26) the dependencies of the variables on the predetermined variable based on the second trained machine learning system, wherein the variables that have a dependency less than or equal to a dependency of the random variable are discarded, characterized by , that the dependencies are determined using an importance metric and / or an effect metric, and that the importance metric is determined using an aggregation of a permutation importance and an impurity importance, and that the effect metric is determined using an accumulated local effects. [2] Method according to claim 1, characterized by that the importance metric is determined using the formula of a summation of: Min(PI, II) + 0.75Range(PI, II), where the permutation importance is given by PI and the impurity importance by II. [3] Method according to one of the preceding claims, characterized by that the effect metric is determined using a range operator applied to the accumulated local effects value. [4] Method according to one of the preceding claims, wherein the predetermined variable exhibits an abnormal behavior, in particular values outside a specified value range. [5] Method according to one of the preceding claims, wherein the data points were acquired in a semiconductor factory, in particular the data points of the variables characterize inline measurements and / or PCM measurements and / or wafer-level tests and / or a wafer processing history. [6] Method according to one of the preceding claims, wherein a production process for manufacturing the products (12a, 12b) is adapted depending on the determined dependencies of the variables and in particular depending on a value range to be achieved for the predeterminable variable. [7] Device which is arranged to carry out the method according to one of the preceding claims. [8] A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 6. [9] A machine-readable storage medium on which the computer program according to claim 8 is stored.
Citation Information
Patent Citations
Method and apparatus for identifying variables from a plurality of variables that exhibit a dependency on a given variable from the plurality of variables.
DE102022208394A1
Method and apparatus for identifying variables from a plurality of variables that exhibit a dependency on a given variable from the plurality of variables.
DE102023202838A1