Method and apparatus for monitoring an anomaly score in semiconductor manufacturing
The WAUC method in semiconductor manufacturing identifies defective chip ensembles by detecting data distribution drifts, reducing false positives and enhancing user-friendliness in quality control, addressing the inefficiencies of univariate WLT in early defect detection.
Patent Information
- Application Number
- DE102024200321
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-01-15
- Publication Date
- 2025-06-26
AI Technical Summary
Existing wafer-level tests (WLT) in semiconductor chip production often fail to identify all defective or abnormal chips, leading to higher disposal costs in later manufacturing stages due to univariate measurements, and there is a need for early identification of defects to reduce these costs.
A method using a weighted area under the curve (WAUC) is applied to detect data distribution drifts by comparing a data distribution of anomaly values with a reference distribution, marking ensembles of components for further testing if a drift threshold is exceeded, and optionally issuing warnings or stopping the test device.
This approach significantly reduces the number of markings and enhances user-friendliness by identifying ensembles with potential defects, thereby reducing false positives and improving the efficiency of quality control in semiconductor manufacturing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for manufacturing quality testing in component production, a computer program and a machine-readable storage medium, as well as a device for data processing for manufacturing quality testing in component production. State of the art
[0002] In multi-stage manufacturing processes, it can be important, particularly for cost reasons, to be able to identify and sort out defective components that were produced in one process step and are to be used in later stages of the manufacturing process at an early stage. This is particularly true for the production of semiconductor chips. In the latter case, wafer-level tests (WLT) are carried out after a wafer has been produced, during which all chips on the wafer are individually subjected to several tests that indicate the final performance of the chip. Chips that have passed the WLT test can then be assembled into a package together with chips of the same or a different type as the end result of the at least two-stage manufacturing process. The costs of disposing of chips in later process stages are higher than in earlier stages. Therefore, it is important to identify defective or abnormal chips early in WLT.Univariate measurements in WLT often fail to identify all defective or abnormal chips in the WLT. DE102023200852.1 proposes a method for determining an anomaly value for the chips within the WLT. The described anomaly value can be determined from multivariate WLT sensor measurements and can provide a measure of whether an individual chip that has passed previous WLT still has an anomaly.
[0003] The area under the receiver operating characteristic curve (AUC) is an aggregate measure of the performance of a binary classification model across all possible classification thresholds. In arXiv:2107.02990, a weighted AUC is applied to a model predicting deviations between two distributions. Disclosure of the invention
[0004] According to one aspect of the present invention, a computer-implemented method for manufacturing quality control in component manufacturing, in particular semiconductor manufacturing, is provided based on a detection of a drift of data points in a data distribution f beob disclosed. Production quality testing can, in particular, involve testing at the wafer level, i.e., quality testing within the scope of wafer-level tests (WLF). A drift can be a change or deviation of a data distribution compared to a reference distribution. The reference distribution can, for example, have been determined at a specific point in time or can refer to measurements at a specific point in time. The data distribution can have been determined by associated measurements at a later point in time than the reference data distribution or can refer to measurements at this later point in time. The data distribution f beoba frequency distribution of anomaly values. An anomaly value can be a value that indicates whether one or more sensor measurements on a component or a value derived from the sensor measurement(s) on the component lies outside predefined / specified limits. An anomaly value of a component can be a value that indicates a measure or probability that the component is abnormal, i.e., may have a defect. An anomaly value can also indicate whether or that a measured value lies outside a predefined interval for this measurement. For example, the component may not have shown any abnormalities in previously performed, e.g., univariate, test measurements on the component. For example, an anomaly value can be determined by a machine learning system, whereby the machine learning system can receive multivariate sensor measured values from test measurements of a component as input and output an anomaly value.It is also possible for anomaly values to be determined from one or more sensor measurements on a component using further analytical or numerical methods. For example, anomaly values can be determined via the deviations from calculated or simulated values. This can be done, for example, based on previously measured base values that may have been determined from sensor measurements on a component. An anomaly value of a component can, for example, provide a measure of the extent to which, for example, derived and / or combined physical or chemical properties of the component deviate from a reference component. A reference component can be a component that meets the properties specified in a specification. The specification can, for example, be a document in which, for example, the electrical, chemical, mechanical, etc. properties of a non-defective component are precisely defined. The data distribution f. beobcomprises the frequency distribution of anomaly values in measurements on an ensemble of components. Furthermore, a reference data distribution f ref a frequency distribution of anomaly values from an ensemble of reference components. Preferably, a (reference) component can be a chip on a wafer, wherein the ensemble of (reference) components in this case can preferably comprise all chips located on a wafer. Alternatively, the ensemble of (reference) components can comprise the chips in a LOT. In one method step, the cumulative distribution function F ref the reference data distribution f ref The cumulative distribution function F ref can be determined from the reference data distribution by calculating the integral Fref(x)=∫−∞xfref(x')dx'. In a further step, a drift detection value is determined as the weighted area under the curve (AUC), where the weighted area under the curve is determined by the product of at least the cumulative distribution function F ref the reference data distribution and the data distribution f beobintegrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value. In a subsequent step, the determined drift detection value is then compared with a specified drift threshold. The ensemble of components is then marked for further testing if the drift detection value exceeds the drift threshold. Marking can be done, for example, by providing the ensemble of components with a tag, by marking an associated ensemble number in a table, or by a robot / user controlling and sorting out the ensemble. Additionally or alternatively, the test measuring device used to perform the sensor measurements on the ensemble of components from which the anomaly values were determined can be marked for further testing.By marking an ensemble of components as conspicuous, the number of markings and thus also the user-friendliness of an application for domain experts can be improved compared to marking individual components based on an individual anomaly value relating to a single component. This can make the method more user-friendly than a method that would mark the component in question based on individual anomaly values, e.g. individual, component-specific anomaly values, when a predetermined threshold value is exceeded or not reached and could optionally issue a (threshold) alarm in each case. Compared to the latter, exemplary method, the number of markings carried out in the method proposed here can be significantly reduced and thus, among other things, user-friendliness can be significantly increased.
[0005] Optionally, in a further step, a warning message can be issued to a system controlling a production testing system and / or to a user of the production testing system. The warning message can include a visual (lamp, display on a control element / computer screen) and / or acoustic alarm. Furthermore, in response to receiving a warning message, a production testing machine, e.g., a test measuring device, can optionally be stopped for further testing of this machine or the test measuring device.
[0006] Preferably, the individual components from the ensembles of components considered here have been classified as non-defective by previous measurements on the individual components as part of the quality checks of the individual components.
[0007] The weighted AUC determined in the process steps described here can be used, for example, as a measure of the overlap of the data distribution f beoband the reference data distribution f ref The AUC can be determined by calculating the integral ∫−∞∞Fref(x)fobs(x)dx. To determine the weighted AUC, a threshold-dependent weight function can be added to the product in the integral as an additional factor. The weighted AUC can then be determined by ∫−∞∞Fref(x)fobs(x)w(x)dx, a threshold-dependent weight function.
[0008] The AUC as well as the WAUC offer the advantage that drifts, e.g. shifts, of the data distribution f beob or specific data points within the data distribution compared to the data distribution f refTowards lower anomaly values, the construction of AUC or WAUC leads to lower drift detection values. By construction, the drift detection values then remain below the specified drift threshold. Thus, with the method described here, only those drifts - drifts in the data distribution f beob towards higher anomaly values - which may be associated with an anomaly in the ensemble of components or which may indicate an anomaly in the ensemble.
[0009] According to a preferred embodiment, the predetermined drift threshold value can be determined by means of the steps described below. In one step, a plurality of N calibration distributions can be received, wherein each calibration distribution can be a frequency distribution of anomaly values of a respective ensemble of reference components. Different calibration distributions can refer to different ensembles of reference components. For example, an ensemble of reference components can be given by the chips of a wafer or the chips of a LOT. In a further step, a cumulative distribution function F ref,kalfrom the N calibration distributions. Furthermore, in a subsequent step, a drift detection value can be determined for each of the N calibration distributions by calculating the weighted area under the curve with one of the N calibration distributions and the cumulative distribution function F ref,kal The determined N drift detection values may differ from each other and reflect the variability of the anomaly values between different wafers. In a subsequent step, the drift threshold may have been determined from the distribution of the determined N drift detection values as a quantile of this distribution.
[0010] Advantageously, the calibration described above utilizes the natural structure of the component ensembles, which is determined by the chips belonging to a wafer. By determining the drift threshold according to the steps described above, an accepted false positive rate for the detection of anomalous data drift is specified. The steps described above advantageously allow the drift threshold to be set in a transparent manner adapted to the semiconductor domain. In particular, this method allows for the easy inclusion of domain expert preferences and knowledge when determining the quantile.
[0011] Preferably, the quantile of the distribution of the determined N drift detection values, which indicates the drift threshold, can be given by one of the percentiles between the 90th percentile and the 100th percentile. For example, the drift threshold can be given by one of the percentiles between the 94th percentile and the 100th percentile, e.g., by the 95th percentile or the 99th percentile.
[0012] The choice of the percentile can advantageously be made by a domain expert based on his domain knowledge or his preferences.
[0013] According to a preferred embodiment, when determining a drift detection value as a weighted area under the curve, the product of at least the cumulative distribution function F ref the reference data distribution and the data distribution f beob A threshold-dependent weight function can be added as a further factor.
[0014] Adding a threshold-dependent weight function when calculating the weighted area under the curve makes it possible to suppress the contribution of threshold effects. For example, threshold effects could occur when artifacts in the data distribution f beob be overweighted when calculating the area under the curve, for example if these artifacts are in the area of the increase of the cumulative distribution function F ref Such threshold effects can advantageously be suppressed by introducing a threshold-dependent weight function. A threshold-dependent weight function can also be used to compensate for moderate shifts in the data distribution f beob have no significant impact on the WAUC statistics.
[0015] According to a preferred embodiment, the threshold-dependent weight function can be obtained by fitting a distribution to the cumulative distribution function, and using this fitted distribution as the basis for determining the threshold-dependent weight function. This distribution can preferably be given by the distribution function of a log-normal distribution. Alternatively, it is possible to fit other distributions instead of the distribution function of a log-normal distribution.
[0016] Often, the anomaly values can be approximated by a log-normal distribution. However, the reference distribution of anomaly values may contain some artifacts, such as additional "bumps" or artifacts at the edges, such as longer tails. To make the WAUC more robust with respect to such artifacts, e.g., to avoid reducing the detection sensitivity for shifts toward larger anomaly values due to artifacts in the reference distribution, a log-normal distribution can be fitted to the cumulative distribution function of the reference distribution, and this fitted cumulative distribution function can be used in the weighting function w when calculating the WAUC. Advantageously, the sensitivity of drift detection can be further increased by reducing the exponent in the fitted log-normal distribution used in the weighting function.Conversely, the sensitivity of drift detection can be reduced by increasing the exponent in the fitted log-normal distribution.
[0017] Preferably, the threshold-dependent weight function can be given by a power of the fitted distribution.
[0018] For example, the threshold-dependent weight function w(x) can be given by F¯refn(x) be given, where F ref (x) denotes the fitted distribution and n is a natural or positive real number to whose power the fitted distribution F ref (x). This ensures that only drifts of the data distribution f beob contribute to the drift detection value towards higher anomaly values.
[0019] According to a preferred embodiment, depending on the drift detection, i.e. if the drift detection value exceeds the drift threshold, a flag can be set which characterizes that the anomaly detection model, in particular a machine learning system for determining an anomaly value of a component from sensor measurements on this component or another numerical or analytical method for determining an anomaly value of a component from sensor measurements on this component, is outdated, in particular should be readjusted depending on the measurements, or that the ensemble of components is abnormal, or that the test machine is defective. According to a preferred embodiment, an ensemble of components is given by all chips on a wafer or by all chips in a LOT. In this case, a set flag can indicate, among the other possibilities mentioned above, that the wafer comprising the ensemble of the measured components orwhich is the substrate for the ensemble of measured chips, may be abnormal.
[0020] Optionally, in a further process step, an alarm can be triggered if the flag is set. This alarm can be a warning message displayed on a screen for a user or operator in production quality control, a visual and / or audible alarm signal, and / or an alarm signal sent to a control unit of a test device in quality control. In the latter case, the sent alarm signal can result in the operation of the test device that performed the quality control measurements on the wafer from which the corresponding anomaly values were determined being stopped.
[0021] According to a preferred embodiment, an anomaly value can be determined from multivariate sensor measurements, wherein the multivariate measurements are wafer-level test (WLT) measurements on semiconductor components, in particular on chips on a wafer.
[0022] WLT tests often only consider univariate deviations from a test specification. This can be improved, for example, by a machine learning system that considers the multivariate relationship between the values measured in WLT and outputs an anomaly value for each chip in WLT. A corresponding machine learning system can, for example, be trained on WLT data that is known to have produced a large proportion of good chips and can be used to predict an anomaly value for newly produced chips based on WLT measurements.
[0023] Furthermore, the invention relates to a computer program with machine-readable instructions which, when executed on one or more computers, cause the computer(s) to perform one of the methods described above and below. The invention also encompasses a machine-readable data carrier on which the above computer program is stored, as well as a computer or data processing device equipped with the aforementioned computer program and / or the aforementioned machine-readable data carrier.
[0024] Embodiments of the invention are explained in more detail below with reference to the accompanying drawings. In the drawings: Fig. 1 schematically shows an information flow overview of a method described here; Fig. 2 schematically shows a further information flow overview of a method described here; Fig. 3 a data processing device comprising means for carrying out a method described here. Description of the embodiments
[0025] Fig. 1 schematically shows an information flow overview of a computer-implemented method 100 for manufacturing quality control in component manufacturing based on a detection of a drift of data points in a data distribution f beob . This can particularly involve manufacturing quality control in semiconductor manufacturing. The data distribution f beob can include a frequency distribution of anomaly values for an ensemble of components. A component can be, for example, a chip, and an ensemble of components can include all chips located on a wafer or in a LOT. A reference data distribution f refcan be given by a frequency distribution of anomaly values of an ensemble of reference components, e.g., a reference wafer or a reference LOT. The method 100 can comprise the steps described below. In step 101, the reference data distribution f ref the corresponding cumulative distribution function F ref of the reference data distribution. In step 102, a drift detection value is determined as a weighted area under the curve (WAUC). The weighted area under the curve can be determined by the product of at least the cumulative distribution function F ref the reference data distribution and the data distribution f beobintegrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value. In step 103, the drift detection value is compared with a predetermined drift threshold. The ensemble of components can then be marked for further testing in step 104 if the drift detection value exceeds the drift threshold. Additionally or alternatively, the test measuring device used to perform the measurements on the ensemble of components from which the anomaly values were determined can be marked for further testing.
[0026] Fig. Figure 2 schematically shows a further information flow overview of a computer-implemented method 100 for manufacturing quality control in component production. The method steps 101, 102, 103 and 104 have already been described in connection with the Fig. 1 has been described. Fig. 2 shows further method steps 201, 202, 203, and 204, which may relate to determining the predetermined drift threshold value, which is compared in step 103 with the determined drift detection value. Method steps 201, 202, 203, and 204 may have been performed before performing method steps 101 and 102 or may run parallel to the execution of steps 101 and 102. The drift threshold value determined after performing steps 201, 202, 203, and 204 should in any case be available in step 103 for comparison with the determined drift detection value. In step 201, a plurality of N calibration distributions are received. Each calibration distribution is a frequency distribution of anomaly values of a respective ensemble of reference components. In step 202, a cumulative distribution function F is determined. ref,kalfrom the N calibration distributions. Step 203 involves determining a drift detection value for each of the N calibration distributions by calculating the weighted area under the curve with each of the N calibration distributions and the cumulative distribution function F ref,kal Then, in step 204, the drift threshold is determined from the distribution of the determined N drift detection values as a quantile of this distribution.
[0027] Fig. 3 shows an embodiment of a data processing device 10 comprising at least one processor 30 and at least one machine-readable storage medium 20, wherein the machine-readable storage medium 20 contains instructions which, when executed by the processor 30, cause the data processing device 10 to carry out a method according to one of the aspects of the invention.
[0028] The term "computer" encompasses any device capable of executing specified computational instructions. These computational instructions can be in the form of software, hardware, or a combination of software and hardware.
[0029] In general, a plurality can be understood as indexed, meaning that each element of the plurality is assigned a unique index, preferably by assigning consecutive integers to the elements included in the plurality. Preferably, when a plurality comprises N elements, where N is the number of elements in the plurality, the elements are assigned integers from 1 to N. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] DE 102023200852.1
[0002]
Claims
[1] Computer-implemented method (100) for manufacturing quality control in component manufacturing, in particular semiconductor manufacturing, based on detection of a drift of data points in a data distribution f beob , where the data distribution f beob is a frequency distribution of anomaly values, where the data distribution f beob the frequency distribution of anomaly values in measurements on an ensemble of components, where a reference data distribution f ref is a frequency distribution of anomaly values from an ensemble of reference components, the method comprising the steps of: - Determine the cumulative distribution function F ref the reference data distribution (101), - Determining a drift detection value as a weighted area under the curve, where the weighted area under the curve is determined by the product of at least the cumulative distribution function Fref the reference data distribution and the data distribution f beob integrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value (102), - comparing the drift detection value with a predetermined drift threshold value (103), - Marking the ensemble of components for further verification if the drift detection value exceeds the drift threshold (104). [2] Method (100) according to claim 1, wherein the predetermined drift threshold value was determined by the following steps: - receiving a plurality of N calibration distributions, each calibration distribution being a frequency distribution of anomaly values of an ensemble of reference components (201), - Determine a cumulative distribution function F ref,kal from the N calibration distributions (202), - Determination of one drift detection value for each of the N calibration distributions by calculating the weighted area under the curve with one of the N calibration distributions and the cumulative distribution function F ref,kal (203), - Determining the drift threshold from the distribution of the determined N drift detection values as a quantile of this distribution (204). [3] The method (100) of claim 2, wherein the quantile is given by one of the percentiles between the 90th percentile and the 100th percentile. [4] Method (100) according to one of the preceding claims, wherein when determining a drift detection value as a weighted area under the curve, the product of at least the cumulative distribution function F ref the reference data distribution and the data distribution f beob a threshold-dependent weight function is added as a further factor. [5] Method (100) according to claim 4, wherein the threshold-dependent weight function is obtained by fitting a distribution to the cumulative distribution function and this fitted distribution serves as a basis for determining the threshold-dependent weight function. [6] The method (100) of claim 5, wherein the threshold-dependent weight function is given by a power of the fitted distribution. [7] Method (100) according to one of the preceding claims, wherein depending on the drift detection a flag is set which characterizes that the anomaly detection model is outdated, in particular is readjusted depending on the measurements, or the wafer is abnormal, or the test machine is defective. [8] Data processing device (10) comprising means for carrying out the method (100) according to one of claims 1 to 7. [9] A computer program comprising instructions which, when executed by a computer (10), cause the computer (10) to carry out the method (100) according to any one of claims 1 to 7. [10] A computer-readable storage medium (20) comprising instructions which, when executed by a computer (10), cause the computer (10) to carry out the method (100) according to any one of claims 1 to 7.
Citation Information
Patent Citations
US000011544634B2
CN000116991137A