Method and apparatus for calculating and monitoring an anomaly score in semiconductor manufacturing
The method addresses the challenge of early identification of defective semiconductor chips by calculating a weighted AUC using FFT to detect data point drift in anomaly value distributions, effectively reducing disposal costs and optimizing resource allocation.
Patent Information
- Application Number
- DE102023213320
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-26
AI Technical Summary
In multistage production processes, particularly in semiconductor manufacturing, it is challenging to identify defective or abnormal semiconductor chips early on, as univariate measurements during wafer level tests often fail to detect all defective chips, leading to higher costs for disposing of defective chips at later production stages.
A computer-implemented method for manufacturing quality testing that detects data point drift in a data distribution of anomaly values by calculating a weighted area under the curve (AUC) using a cumulative distribution function derived from reference and current data distributions, with the aid of Fast Fourier Transform (FFT) for efficient calculation.
This method enables early identification of defective chips by effectively measuring the overlap between reference and current data distributions, reducing false positives, and allowing only significant drifts towards higher anomaly values to be flagged, thus optimizing resource allocation and reducing disposal costs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for manufacturing quality testing in component production, a computer program and a machine-readable storage medium, as well as a device for data processing for manufacturing quality testing in component production. State of the art
[0002] In multi-stage manufacturing processes, it can be important, particularly for cost reasons, to be able to identify and sort out defective components that were produced in one process step and are to be used in later stages of the manufacturing process at an early stage. This is particularly true for the production of semiconductor chips. In the latter case, wafer-level tests (WLT) are carried out after a wafer has been produced, during which all chips on the wafer are individually subjected to several tests that indicate the final performance of the chip. Chips that have passed the WLT test can then be assembled into a package together with chips of the same or a different type as the end result of the at least two-stage manufacturing process. The costs of disposing of chips in later process stages are higher than in earlier stages. Therefore, it is important to identify defective or abnormal chips early in WLT.Univariate measurements in WLT often fail to identify all defective or abnormal chips in the WLT. DE102023200852.1 proposes a method for determining an anomaly value for the chips within the WLT. The described anomaly value can be determined from multivariate WLT sensor measurements and can provide a measure of whether an individual chip that has passed previous WLT still has an anomaly.
[0003] The area under the receiver operating characteristic curve (AUC) is an aggregate measure of the performance of a binary classification model across all possible classification thresholds. In arXiv:2107.02990, a weighted AUC is applied to a model predicting deviations between two distributions.
[0004] From https: / / doi.org / 10.2307 / 2347084 a method for determining a kernel density estimator using the Fast Fourier Transform (FFT) is known. Disclosure of the invention
[0005] According to one aspect of the present invention, a computer-implemented method for manufacturing quality control in component manufacturing, in particular semiconductor manufacturing, is provided based on a detection of a drift of data points in a data distribution f beobdisclosed. Production quality testing can, in particular, involve testing at the wafer level, i.e., quality testing within the scope of wafer-level tests (WLF). A drift can be a change or deviation of a data distribution compared to a reference distribution. The reference distribution can, for example, have been determined at a specific point in time or can refer to measurements at a specific point in time. The data distribution can have been determined by associated measurements at a later point in time than the reference data distribution or can refer to measurements at this later point in time. The data distribution f beoba frequency distribution of anomaly values. An anomaly value can be a value that indicates whether one or more sensor measurements on a component or a value derived from the sensor measurement(s) on the component lies outside predefined / specified limits. An anomaly value of a component can be a value that indicates a measure or probability that the component is abnormal, i.e., may have a defect. An anomaly value can also indicate whether or that a measured value lies outside a predefined interval for this measurement. For example, the component may not have shown any abnormalities in previously performed, e.g., univariate, test measurements on the component. For example, an anomaly value can be determined by a machine learning system, whereby the machine learning system can receive multivariate sensor measured values from test measurements of a component as input and output an anomaly value.It is also possible for anomaly values to be determined from one or more sensor measurements on a component using further analytical or numerical methods. An anomaly value for a component can, for example, provide a measure of the extent to which derived and / or combined physical or chemical properties of the component deviate from a reference component. A reference component can be a component that meets the properties specified in a specification. The specification can, for example, be a document in which the electrical, chemical, mechanical, etc. properties of a non-defective component are precisely defined. The data distribution f. beob comprises the frequency distribution of anomaly values in measurements on an ensemble of components. Furthermore, a reference data distribution f refa frequency distribution of anomaly values from an ensemble of reference components. Preferably, a (reference) component can be a chip on a wafer, wherein the ensemble of (reference) components in this case can preferably comprise all chips located on a wafer. Alternatively, the ensemble of (reference) components can comprise the chips in a LOT. In one method step, the cumulative distribution function F ref the reference data distribution f ref The cumulative distribution function F ref can be determined from the reference data distribution by calculating the integral Fref(x)=∫−∞xfref(x')dx'. In a further step, a drift detection value is determined as the weighted area under the curve (AUC), where the weighted area under the curve is determined by the product of at least the cumulative distribution function F refthe reference data distribution and the data distribution f beobintegrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value. In a subsequent step, the determined drift detection value is then compared with a specified drift threshold. The ensemble of components is then marked for further testing if the drift detection value exceeds the drift threshold. Marking can be done, for example, by providing the ensemble of components with a tag, by marking an associated ensemble number in a table, or by a robot / user controlling and sorting out the ensemble. Additionally or alternatively, the test measuring device used to perform the sensor measurements on the ensemble of components from which the anomaly values were determined can be marked for further testing.A method described above and below is characterized in that the data distribution f. beob and / or the reference data distribution f ref is / are approximately determined as a kernel density estimator, whereby the kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution is determined on the basis of the frequency distribution of the respective anomaly values using an FFT (Fast Fourier Transform) based method.
[0006] Optionally, in a further step, a warning message can be issued to a system controlling a production testing system and / or to a user of the production testing system. The warning message can include a visual (lamp, display on a control element / computer screen) and / or acoustic alarm. Furthermore, in response to receiving a warning message, a production testing machine, e.g., a test measuring device, can optionally be stopped for further testing of this machine or the test measuring device.
[0007] Preferably, the individual components from the ensembles of components considered here have been classified as non-defective by previous measurements on the individual components as part of the quality checks of the individual components.
[0008] The weighted AUC determined in the process steps described here can be used, for example, as a measure of the overlap of the data distribution f beoband the reference data distribution f ref The AUC can be determined by calculating the integral ∫−∞∞Fref(x)fobs(x)dx. To determine the weighted AUC, a threshold-dependent weight function can be added to the product in the integral as an additional factor. The weighted AUC can then be determined by ∫−∞∞Fref(x)fobs(x)dx, where w(x) denotes a threshold-dependent weight function.
[0009] The AUC as well as the WAUC offer the advantage that drifts, e.g. shifts, of the data distribution f beob or specific data points within the data distribution compared to the data distribution f refTowards lower anomaly values, the construction of AUC or WAUC leads to lower drift detection values. By construction, the drift detection values then remain below the specified drift threshold. Thus, with the method described here, only those drifts - drifts in the data distribution f beob towards higher anomaly values - which may be associated with an anomaly in the ensemble of components or which may indicate an anomaly in the ensemble.
[0010] The use of the FFT-based method for the approximate determination of the data distribution and the reference data distribution ensures, in particular, a fast, efficient calculation of the weighted AUC. This can be particularly relevant in production quality control in semiconductor manufacturing, where a data distribution f beobcan be given by an ensemble of HL components, e.g., by all chips on a wafer. A wafer can contain thousands of chips and thus the data distribution can contain thousands of anomaly values. Alternatively, a data distribution f beob The anomaly values of an ensemble of components, which can be defined by all chips in a LOT, can comprise hundreds of thousands of chips. A LOT can comprise, for example, 24 wafers. A method described above and below enables a fast, resource-efficient, and scalable calculation of the drift in the data distribution of anomaly values, particularly in the case of (reference) data distributions that comprise several thousand or several hundred thousand data points.
[0011] According to a preferred embodiment, the kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution can be determined using the FFT-based method by carrying out the following steps. In a first step, a discrete approximation to the data distribution or the reference data distribution can be determined using a histogram from the anomaly values of the data distribution or the reference data distribution on a grid of predetermined grid points. The associated FFT can be determined for the discrete approximation of the data distribution or reference data distribution, each in the form of a histogram. The predetermined grid points, by which in particular the width of the bins of the respective histogram can be determined, can each be equally spaced grid points.In a further step, the Fourier transform of the respective kernel density estimator can be determined by the product of the FFT of the respective discrete approximation to the data distribution or the reference data distribution and the Fourier transform of the kernel. The kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution can then be determined by the inverse FFT of the determined Fourier transform.
[0012] The above approach scales particularly well and allows for time- and resource-efficient testing of components in an ensemble in semiconductor manufacturing. For example, data distributions of, for example, anomaly values of an ensemble of components in the case of high-precision manufacturing can easily consist of hundreds of thousands to millions of data points.
[0013] Preferably, the kernel density estimator may have a Gaussian kernel. Alternatively, the kernel density estimator may have another possible kernel.
[0014] According to a preferred embodiment, the cumulative distribution function F ref be determined by using the kernel density estimator that represents the reference data distribution f ref approximated, cumulatively integrated.
[0015] By cumulative integration of the kernel density estimate f ref , a quick numerical estimate of the cumulative distribution function F ref be obtained.
[0016] Preferably, a one-dimensional interpolating spline with constant extrapolation to the cumulative distribution function F ref be adjusted.
[0017] For efficient and accurate calculation of the weighted AUC, the cumulative distribution function F refbe known or applied for any positive values, since the anomaly values in an observed data distribution f beob can be much larger than the anomaly values in the reference data distribution, which are used to calculate the cumulative distribution function F ref To calculate the weighted AUC, F ref on the support, e.g. the grid points, the data distribution f beob The determination of a one-dimensional interpolating spline with constant extrapolation to the cumulative distribution function F ref ensures that F ref on the support of data distribution f beob is known and thus the weighted AUC can be determined efficiently and accurately. To determine the reference data distribution f refIf a Gaussian kernel is used, the cumulative distribution function to be determined is particularly smooth and a good approximation can be found by fitting a spline.
[0018] According to a preferred embodiment, the kernel density estimator of the data distribution f beob at a number of specified support points. The one-dimensional interpolating spline, which is fitted to the cumulative distribution function F ref may have been adjusted, can still be evaluated at the specified support points of the data distribution. The weighted area under the curve can then be determined by multiplying the product of the interpolating spline evaluated at the specified support points with the kernel density estimator of the data distribution f determined at the specified support points. beoband formed with a weighting function, and this product is numerically integrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value at the specified support points of the data distribution. Preferably, the weighting function has been determined, or alternatively provided, in such a way that it can also be easily evaluated at the specified support points of the data distribution.
[0019] This allows the weighted area under the curve to be calculated efficiently and precisely for data distributions f beob and f ref with a large number of data points and a scalable solution for monitoring drift in anomaly values is ensured.
[0020] According to a preferred embodiment, depending on the drift detection, i.e. if the drift detection value exceeds the drift threshold, a flag can be set which characterizes that the anomaly detection model, in particular a machine learning system for determining an anomaly value of a component from sensor measurements on this component or another numerical or analytical method for determining an anomaly value of a component from sensor measurements on this component, is outdated, in particular should be readjusted depending on the measurements, or that the ensemble of components is abnormal, or that the test machine is defective. According to a preferred embodiment, an ensemble of components is given by all chips on a wafer or by all chips in a LOT. In this case, a set flag can indicate, among the other possibilities mentioned above, that the wafer comprising the ensemble of the measured components orwhich is the substrate for the ensemble of measured chips, may be abnormal.
[0021] Optionally, in a further process step, an alarm can be triggered if the flag is set. This alarm can be a warning message displayed on a screen for a user or operator in production quality control, a visual and / or audible alarm signal, and / or an alarm signal sent to a control unit of a test device in quality control. In the latter case, the sent alarm signal can result in the operation of the test device that performed the quality control measurements on the wafer from which the corresponding anomaly values were determined being stopped.
[0022] According to a preferred embodiment, an anomaly value can be determined from multivariate sensor measurements, wherein the multivariate measurements are wafer-level test (WLT) measurements on semiconductor components, in particular on chips on a wafer.
[0023] WLT tests often only consider univariate deviations from a test specification. This can be improved, for example, by a machine learning system that considers the multivariate relationship between the values measured in WLT and outputs an anomaly value for each chip in WLT. A corresponding machine learning system can, for example, be trained on WLT data that is known to have produced a large proportion of good chips and can be used to predict an anomaly value for newly produced chips based on WLT measurements.
[0024] According to a preferred embodiment, the predetermined drift threshold value can be determined by means of the steps described below. In one step, a plurality of N calibration distributions can be received, wherein each calibration distribution can be a frequency distribution of anomaly values of a respective ensemble of reference components. Different calibration distributions can refer to different ensembles of reference components. For example, an ensemble of reference components can be given by the chips of a wafer or the chips of a LOT. In a further step, a cumulative distribution function F ref,kalfrom the N calibration distributions. Furthermore, in a subsequent step, a drift detection value can be determined for each of the N calibration distributions by calculating the weighted area under the curve with one of the N calibration distributions and the cumulative distribution function F ref,kal The determined N drift detection values may differ from each other and reflect the variability of the anomaly values between different wafers. In a subsequent step, the drift threshold may have been determined from the distribution of the determined N drift detection values as a quantile of this distribution.
[0025] Advantageously, the calibration described above utilizes the natural structure of the component ensembles, which is determined by the chips belonging to a wafer. By determining the drift threshold according to the steps described above, an accepted false positive rate for the detection of anomalous data drift is specified. The steps described above advantageously allow the drift threshold to be set in a transparent manner adapted to the semiconductor domain. In particular, this method allows for the easy inclusion of domain expert preferences and knowledge when determining the quantile.
[0026] Preferably, the quantile of the distribution of the determined N drift detection values, which indicates the drift threshold, can be given by one of the percentiles between the 90th percentile and the 100th percentile. For example, the drift threshold can be given by one of the percentiles between the 94th percentile and the 100th percentile, e.g., by the 95th percentile or the 99th percentile.
[0027] The choice of the percentile can advantageously be made by a domain expert based on his domain knowledge or his preferences.
[0028] According to a preferred embodiment, when determining a drift detection value as a weighted area under the curve, the product of at least the cumulative distribution function F ref the reference data distribution and the data distribution f beob A threshold-dependent weight function can be added as a further factor.
[0029] Adding a threshold-dependent weight function when calculating the weighted area under the curve makes it possible to suppress the contribution of threshold effects. For example, threshold effects could occur when artifacts in the data distribution f beob be overweighted when calculating the area under the curve, for example if these artifacts are in the area of the increase of the cumulative distribution function F ref Such threshold effects can advantageously be suppressed by introducing a threshold-dependent weight function. A threshold-dependent weight function can also be used to compensate for moderate shifts in the data distribution f beob have no significant impact on the WAUC statistics.
[0030] According to a preferred embodiment, the threshold-dependent weight function can be obtained by fitting a distribution to the cumulative distribution function, and using this fitted distribution as the basis for determining the threshold-dependent weight function. This distribution can preferably be given by the distribution function of a log-normal distribution. Alternatively, it is possible to fit other distributions instead of the distribution function of a log-normal distribution.
[0031] Often, the anomaly values can be approximated by a log-normal distribution. However, the reference distribution of anomaly values may contain some artifacts, such as additional "bumps" or artifacts at the edges, such as longer tails. To make the WAUC more robust with respect to such artifacts, e.g., to avoid reducing the detection sensitivity for shifts toward larger anomaly values due to artifacts in the reference distribution, a log-normal distribution can be fitted to the cumulative distribution function of the reference distribution, and this fitted cumulative distribution function can be used in the weighting function w when calculating the WAUC. Advantageously, the sensitivity of drift detection can be further increased by reducing the exponent in the fitted log-normal distribution used in the weighting function.Conversely, the sensitivity of drift detection can be reduced by increasing the exponent in the fitted log-normal distribution.
[0032] Preferably, the threshold-dependent weight function can be given by a power of the fitted distribution.
[0033] For example, the threshold-dependent weight function w(x) can be given by F¯refn(x) be given, where F̅̅ ref (x) denotes the fitted distribution and n is a natural number to whose power the fitted distribution F̅̅̅̅ ref (x). This ensures that only drifts of the data distribution f beob contribute to the drift detection value towards higher anomaly values.
[0034] Furthermore, the invention relates to a computer program with machine-readable instructions which, when executed on one or more computers, cause the computer(s) to perform one of the methods described above and below. The invention also encompasses a machine-readable data carrier on which the above computer program is stored, as well as a computer or data processing device equipped with the aforementioned computer program and / or the aforementioned machine-readable data carrier.
[0035] Embodiments of the invention are explained in more detail below with reference to the accompanying drawings. In the drawings: Fig. 1 schematically shows an information flow overview of a method described here; Fig. 2 schematically shows a further information flow overview of a method described here; Fig. 3 a data processing device comprising means for carrying out a method described here. Description of the embodiments
[0036] Fig. 1 schematically shows an information flow overview of a computer-implemented method 100 for manufacturing quality control in component manufacturing based on a detection of a drift of data points in a data distribution f beob . This can particularly involve manufacturing quality control in semiconductor manufacturing. The data distribution f beob can include a frequency distribution of anomaly values for an ensemble of components. A component can be, for example, a chip, and an ensemble of components can include all chips located on a wafer or in a LOT. A reference data distribution f refcan be given by a frequency distribution of anomaly values of an ensemble of reference components, e.g., a reference wafer or a reference LOT. The method 100 can comprise the steps described below. In step 101, the reference data distribution f ref the corresponding cumulative distribution function F ref of the reference data distribution. In step 102, a drift detection value is determined as a weighted area under the curve (WAUC). The weighted area under the curve can be determined by the product of at least the cumulative distribution function F ref the reference data distribution and the data distribution f beobintegrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value. In step 103, the drift detection value is compared with a predetermined drift threshold. The ensemble of components can then be marked for further testing in step 104 if the drift detection value exceeds the drift threshold. Additionally or alternatively, the test measuring device by which the measurements were carried out on the ensemble of components from which the anomaly values were determined can be marked for further testing. The method 100 is characterized in particular in that the data distribution f beob and / or the reference data distribution f refis / are approximately determined as a kernel density estimator. The kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution is determined based on the frequency distribution of the respective anomaly values using an FFT-based method.
[0037] Fig. Figure 2 schematically shows a further information flow overview of a computer-implemented method 100 for manufacturing quality control in component production. The method steps 101, 102, 103 and 104 have already been described in connection with the Fig. 1 has been described. Fig. Figure 2 shows further method steps 201, 202, and 203, which concern the determination of the kernel density estimators using an FFT-based method. These steps can preferably be carried out before the cumulative distribution function is determined in method step 101, since the reference distribution f refshould be known. In step 201, a discrete approximation to the data distribution or the reference data distribution is determined using a histogram on a grid of specified grid points. For each discrete approximation, the FFT of this discrete approximation is determined. In step 202, the Fourier transform of the respective kernel density estimator is determined by the product of the FFT of the respective discrete approximation and the Fourier transform of the kernel. The kernel can be a Gaussian kernel, for example. In step 203, the kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution is determined by the inverse FFT of the determined Fourier transform.
[0038] Fig.3 shows an embodiment of a data processing device 10 comprising at least one processor 30 and at least one machine-readable storage medium 20, wherein the machine-readable storage medium 20 contains instructions which, when executed by the processor 30, cause the data processing device 10 to carry out a method according to one of the aspects of the invention.
[0039] The term "computer" encompasses any device capable of executing specified computational instructions. These computational instructions can be in the form of software, hardware, or a combination of software and hardware.
[0040] In general, a plurality can be understood as indexed, meaning that each element of the plurality is assigned a unique index, preferably by assigning consecutive integers to the elements included in the plurality. Preferably, when a plurality comprises N elements, where N is the number of elements in the plurality, the elements are assigned integers from 1 to N. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] DE 102023200852.1
[0002] Cited non-patent literature
[0000] https: / / doi.org / 10.2307 / 2347084
[0004]
Claims
[1] Computer-implemented method (100) for manufacturing quality control in component manufacturing, in particular semiconductor manufacturing, based on detection of a drift of data points in a data distribution f beob , where the data distribution f beob is a frequency distribution of anomaly values, where the data distribution f beob the frequency distribution of anomaly values in measurements on an ensemble of components, where a reference data distribution f ref is a frequency distribution of anomaly values from an ensemble of reference components, the method comprising the steps of: - Determine the cumulative distribution function F ref the reference data distribution (101), - Determining a drift detection value as a weighted area under the curve, where the weighted area under the curve is determined by the product of at least the cumulative distribution function Fref the reference data distribution and the data distribution f beob integrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value (102), - comparing the drift detection value with a predetermined drift threshold value (103), - Marking the ensemble of components for further verification if the drift detection value exceeds the drift threshold (104), characterized by that the data distribution f beob and / or the reference data distribution f ref is / are approximately determined as a kernel density estimator, whereby the kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution is determined on the basis of the frequency distribution of the respective anomaly values using an FFT-based method. [2] Method according to claim 1, wherein the kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution is determined by means of the FFT-based method in which - a discrete approximation to the data distribution or the reference data distribution is determined by a histogram on a grid of given grid points and the FFT of this discrete approximation is determined (201), - the Fourier transform of the respective kernel density estimator is determined by the product of the FFT of the respective discrete approximation and the Fourier transform of the kernel (202), - the kernel density estimator of the data distribution and / or the kernel density estimator of the reference data distribution is determined by inverse FFT of the determined Fourier transform (203). [3] Method (100) according to one of the preceding claims, wherein the kernel density estimator has a Gaussian kernel. [4] Method (100) according to one of the preceding claims, wherein the cumulative distribution function F ref is determined by using the kernel density estimator, which represents the reference data distribution f ref approximated, cumulatively integrated. [5] Method (100) according to one of the preceding steps, wherein a one-dimensional interpolating spline with constant extrapolation is applied to the cumulative distribution function F ref is adjusted. [6] Method (100) according to claim 5, where the kernel density estimator of the data distribution f beob is determined at a number of specified support points, where the one-dimensional interpolating spline fitted to the cumulative distribution function F ref is adapted, is evaluated at the given support points of the data distribution, where the weighted area under the curve is determined by multiplying the product of the interpolating spline evaluated at the specified support points with the kernel density estimator of the data distribution f determined at the specified support points beob and is formed with a weight function and is numerically integrated over the range from the smallest occurring anomaly value to the largest occurring anomaly value in the given support points of the data distribution. [7] Method (100) according to one of the preceding claims, wherein depending on the drift detection a flag is set which characterizes that the anomaly detection model is outdated, in particular is readjusted depending on the measurements, or the wafer is abnormal, or the test machine is defective. [8] A data processing device comprising means for carrying out the method according to any one of claims 1 to 7. [9] A computer program comprising instructions which, when executed by a computer (10), cause the computer (10) to carry out the method (100) according to any one of claims 1 to 7. [10] A computer-readable storage medium (20) comprising instructions which, when executed by a computer (10), cause the computer (10) to carry out the method (100) according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and apparatus for evaluating a statistically distributed measured value when examining an element of a photolithography process
DE102018211099A1