Method and apparatus for restoring position of semiconductor component on wafer

JP2023036038A5Pending Publication Date: 2025-06-26ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022137602
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-09-01
Filing Date
2022-08-31
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The loss of traceability of semiconductor components on wafers after cutting and packaging leads to challenges in mapping semiconductor parts to their original positions, complicating the matching between wafer-level and final tests, and resulting in yield losses and complex process control.

Method used

A method involving mapping rules and machine learning algorithms, such as linear regression and the Hungarian algorithm, to establish a one-to-one correspondence between wafer-level and final test results, using a cost matrix to optimize the mapping process and restore the position of semiconductor components on the wafer.

Benefits of technology

Enables improved process control and yield estimation by restoring the traceability of semiconductor components, allowing for better prediction and analysis of defects, and enhancing the mapping between wafer-level and final tests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for obtaining a mapping rule to link test results from different tests of the same semiconductor element, in which the yield loss during semiconductor element manufacturing is taken into account.SOLUTION: A method includes the steps of adjusting a model, e.g., a linear regression model, and using this model to predict the test data (S23), calculating a cost matrix based on the prediction (S24), applying a Hungarian algorithm to this cost matrix to obtain new mapping rules (S25), and repeating these steps many times.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for restoring the positions of semiconductor components on a wafer on which those semiconductor components were placed after cutting the semiconductor components from the wafer, and an apparatus configured to implement this method.

Background Art

[0002] Background Art In the case of the packaging process of semiconductor components (especially power MOS), the traceability of the semiconductor components with respect to the original wafer of the semiconductor components and the original positions of the semiconductor components on the wafer are lost. Specifically, this means that once the wafer is cut or diced (in English, "diced" = the process of separating semiconductor components from the wafer) and packaged, the positions of each semiconductor component on the wafer can no longer be obtained. The vendor of the packaging process can provide at least a rough matching between the individual semiconductor components in the final test (in English, "Final Test" = the test process of semiconductor components after packaging) and the semiconductor components on the wafer in the wafer-level test (the test process before packaging). However, this still results in thousands of semiconductor components that cannot be mapped for multiple wafers. Since this is essentially a combinatorial problem, the complexity of the solution to this problem becomes factorial because there are a factorial number of different possibilities of arranging the semiconductor components so that they match the correct order, where n is the number of semiconductor components.

[0003] For ASIC semiconductor components, a solution to this combination problem exists. During wafer-level testing, a unique identifier is stored in the memory of the ASIC semiconductor component, allowing the final test to be mapped to the wafer-level test after packaging. However, this is impossible for semiconductor components like power MOS due to the lack of memory.

[0004] A further challenge in matching final tests with wafer-level tests is that individual semiconductor components are sorted after wafer-level testing, and in the process, information about which specific semiconductor components were sorted is lost. This complicates the matching problem because, at this point, it is necessary to determine that there are many more wafer-level tests that may belong to the final test result, and which wafer-level tests do not have a corresponding final test result. [Overview of the project] [Problems that the invention aims to solve]

[0005] Advantages of the invention The advantage of the present invention, as described in independent claim 1, is that, according to the present invention, possible mappings can be determined between semiconductor components related to wafer-level test results and packaged semiconductor components related to final test results, and in this case, metadata that is added retrospectively, such as unique identifiers, is not required.

[0006] An advantage of this invention is that it further takes into account potential yield losses that may occur during manufacturing. This means that not all semiconductor components are mapped to the test. Based on the mapping of yield losses and results, additional traceability information can be obtained that enables estimation of where on the wafer the yield loss occurred. This, in turn, enables improved process control (e.g., root cause analysis of defective areas).

[0007] Other independent claims describe yet other aspects of the invention. Dependent claims describe advantageous developments. [Means for solving the problem]

[0008] Disclosure of the invention According to a first aspect, the present invention relates to a method, particularly a computer implementation, for determining mapping rules that map variables belonging to a second set of variables to a plurality of variables belonging to a first set of first variables. The first set contains more variables than the second set. The mapping rule can map a plurality of second variables to a plurality of first variables in a unique manner; that is, at most one second variable is mapped to a first variable by the mapping rule, and preferably the other way around. In other words, the mapping rule is configured to map second variables to a subset of first variables. Preferably, the number of variables in the subset of first variables is the same as the number of variables in the second set. The mapping rule then maps each second variable to a first variable, in which case some first variables are not mapped to second variables because the first set has more variables than the second set. However, it is also possible that the subset is smaller than the second set.

[0009] Here, a set can be understood as a form in which individual variables are grouped together. Preferably, the first set and the second set are distinct sets that do not have any common variables. Preferably, one subscript is mapped to each variable in the first set and the second set. All the subscripts in the first set and the second set can be considered as an index set. That is, it can be considered as a set containing elements to which variables in the first set or the second set are subscripted with sequential numbers. In this case, the mapping rule maps each subscript belonging to the second index set to the first index set. Therefore, the mapping rule describes which first variable belongs to which second variable, and preferably the reverse as well. The mapping rule can be provided as a list or a table, etc.

[0010] This method begins with the steps of initializing a mapping rule and preparing a first set and a second set, where the first set has more variables than the second set. The initial mapping rule can be chosen randomly or as an identity mapping. Preferably, the mapping rule is initialized as follows: that is, the mapping rule considers all second variables in its mapping, i.e., the second variable is mapped to only an amount of the first variable equivalent to the amount of the second variable that exists, or the first variable is mapped to all second variables. Alternatively, other initial mapping rules can be considered, for example, pre-configured, already partially correct mappings.

[0011] Next, a step is taken to randomly select several first variables, in which case the number of first variables selected is equal to at least the following number, i.e., the first set has this many more variables than the second set. When this step is performed by a computer, random selection can be done using a pseudorandom number generator. That is, generally, random selection is performed without any preference for any of the first variables. However, it is conceivable that certain variables may be selected with a higher probability, for example, because they are unusual with respect to the remaining variables in the first set.

[0012] Following this, steps a) to e) described below are repeated. This iteration can be performed for a predetermined maximum number of iterations, or an interruption criterion can be defined, in which case the iteration is interrupted if the interruption criterion is met. The interruption criterion is, for example, the minimum change in the mapping rule.

[0013] a) A step to create a dataset, the dataset having a first variable that does not include the currently selected first variable, and a second variable that is mapped to these first variables according to mapping rules. The currently selected variable is the first variable that was selected when steps a) to e) were performed last time, or, when these steps are performed for the first time, the currently selected variable is a randomly selected first variable.

[0014] This dataset can also be called the training dataset, in which case the mapped second variable is the so-called "label" of the first variable. It should be noted that this step is optional, because subsequent steps using this dataset only require information about the current mapping rule between the first and second variables, which can be provided by either this dataset or the current mapping rule. The current mapping rule is the mapping rule that exists in the current iteration of steps a) to e), i.e., the mapping rule used when the dataset was last created.

[0015] b) This step involves training the machine learning system so that it depends on a first variable to determine the corresponding second variable in the dataset. Here, training can be understood as adjusting the parameters of the machine learning system so that the predictions made by the machine learning system come as close as possible to the second variable ("label") in the dataset. Optimization can be performed with respect to the cost function, which preferably represents the mathematical difference between the output of the machine learning system and the labels. Optimization is preferably performed by gradient descent. The machine learning system can be one or more decision trees, neural networks, support vector machines, etc. Training can be continued until further training would only slightly improve the machine learning system, i.e., until the second stop criterion is met.

[0016] c) A step of obtaining a cost matrix, where the entries in the cost matrix represent the distance between the predictions of the machine learning system and a second variable according to a mapping rule, and in particular, the distance between the predictions of the machine learning system and all the variables in the second set. This distance can be obtained using the L2 norm. Other distance measures are also possible. The structure of the cost matrix can be as follows: that is, the rows and columns can be mapped to the first variable or the predictions of the machine learning system that depend on the first variable and to the second variable, where the entries represent the distance between the individual corresponding variables in the rows and columns.

[0017] d) This step optimizes the mapping rule in relation to the cost matrix so that the mapping rule yields the minimum total cost based on the entries in the cost matrix. The total cost corresponds to the sum of the entries in the cost matrix required when performing the mapping of variables from the first set to the second set according to the current mapping rule from the cost matrix. In other words, the sum is optimized, and in particular minimized, for the entries selected from the cost matrix in relation to the mapping rule. Note that the entries are selected in relation to the mapping rule as follows: that is, the entries in the individual columns and rows of the cost matrix that are mapped to the first and second variables which are mapped to each other according to the mapping rule are selected according to the mapping rule.

[0018] e) This step involves selecting a first variable that is not mapped to one of the second variables even by the optimized mapping rules. Preferably, the selected first variable is stored in a list, which is updated each time steps a) to e) are performed again.

[0019] The mapping rule obtained in the last iteration of step d) is the final mapping rule output in the optional step. Similarly, for example, the first variable selected in the last iteration of step e) can be output for the purpose of determining which first variables do not have a corresponding second variable.

[0020] These variables can be scalars or vectors, such as time series, and can be detected by sensors or indirectly obtained sensor data. Preferably, the first and second variables are one or more measurement results from one or more different measurements performed on one of a group of objects. That is, each variable is mapped to one of the group of objects. In the step of creating the dataset, it is also possible to use only a predetermined number of measurement results from the group of measurement results for the second variable. The mapping rule can indicate which first and second variables are measurement results from the same object. Particularly preferably, at a first time point, at least one measurement of the object has been performed for the first variable, and at a second time point, a measurement has been performed for the second variable, where the second time point is after the first time point. The second time point can be determined after modifications or changes have been made to the object. Differences in the number of various measurement results can be losses due to sorting or other factors in the measurement process of the object.

[0021] The proposed approach involves optimizing the mapping rules using a cost-minimizing algorithm under a given cost matrix. This could involve, for example, using the Hungarian algorithm applied to the cost matrix. The Hungarian algorithm (also known as the Kuhn-Munkres algorithm) is a method for solving weighted assignment problems. Alternatively, a greedy implementation of the algorithm could be used for cost minimization.

[0022] Furthermore, it is proposed that the machine learning system is a regression model, and this regression model determines a second variable depending on a first variable and the parameters of the regression model, in which case the parameters of the regression model are adjusted during training.

[0023] Regression is used to model the relationship between a dependent variable (often also called the explained variable) and one or more independent variables (often also called explanatory variables). Regression can parameterize relatively complex functions, so that the data are best reproduced according to a determined mathematical criterion. For example, the general least squares method calculates a unique straight line (or hyperplane), which minimizes the sum of squared deviations between the actual data and this line (or hyperplane), i.e., the sum of squared residuals.

[0024] It is further proposed that the first variable and the second variable each represent a product when the product is manufactured according to different manufacturing process steps. For example, in this case, when a certain manufacturing process step is completed, a second time point can be set. This product can be any product manufactured in a manufacturing factory. In particular, when a product is manufactured, the traceability to the preceding process steps of the product is lost (so-called "loose load"), for example, when it is no longer possible to directly map the product from loose loads such as screws to a certain manufacturing batch. What is assumed here is that the first variable represents a component, especially a part, the second variable represents the final product, and the mapping rule describes which components have been processed to form which products, or which parts have been incorporated into which products. For example, this is the case when it is no longer possible to take out the parts inside the product without breaking them in order to read the serial number. In such a case, according to the present invention, the manufacturing batch of the parts can be mapped based on the measurement of the product. The difference in the number of various variables in both sets can be regarded as the yield loss during manufacturing.

[0025] The first and second variables can be measurement results / test results, or other characteristics of products, components, etc. The first variable and the second variable particularly describe the same measurement / characteristic of products, components, etc., although they are slightly different from each other, for example, due to manufacturing tolerances.

[0026] Further proposed is that the first variable is the first test result or measurement result of semiconductor elements on a wafer, and the second variable is the second test result or measurement result of this semiconductor element after the semiconductor element is cut out from the wafer. The semiconductor element can be part of an electrical component grown on the wafer, for example, a transistor group of an integrated circuit. The test result can also be related to the entire semiconductor component. In this case, it has been found that linear regression is particularly effective for the machine learning system to find the best mapping rule. This is because linear regression here assumes a linear relationship that makes reasonable assumptions about the mapping of test results. Linear regression is a special case of regression. In linear regression at that time, a linear function is assumed. That is, only the relationship that the dependent variable is a linear combination of regression coefficients (but not necessarily independent variables) is considered.

[0027] Further proposed is that the first test result is a wafer-level test result, and the second test result is a final test result. Preferably, the final test result is less than the wafer-level test result. These tests are, for example, voltage tests and / or contact connection tests.

[0028] Further proposed is that the semiconductor element is manufactured on a plurality of different wafers. This is because it has been found that this method can reach multiple wafers within an appropriate calculation time and can even find the correct mapping rule.

[0029] A further proposal is to determine, based on mapping rules, which second test result belongs to which first test result, and then, depending on the first test result to which it belongs, determine the position of the semiconductor component within the wafer. This enables position recovery, for the first time, which uniquely traces semiconductor components from the last manufacturing process step to the preceding process step in semiconductor manufacturing. A similar procedure can be performed for selected first variables that are not mapped to one of the second variables by the optimized mapping rules, in order to track which semiconductor components have been sorted or removed. Accordingly, the manufacturing process steps can be modified so that it is no longer necessary to sort semiconductor components produced at the corresponding location on the wafer.

[0030] In further embodiments, the present invention relates to apparatuses and computer programs configured to carry out (for carrying out) the methods described herein, as well as machine-readable storage media in which such computer programs are stored.

[0031] Next, embodiments of the present invention will be described in detail with reference to the accompanying drawings. [Brief explanation of the drawing]

[0032] [Figure 1] This diagram provides a schematic overview of the packaging process. [Figure 2] This figure schematically illustrates one embodiment of the flowchart of the present invention. [Figure 3] This is a schematic diagram of the training equipment. [Modes for carrying out the invention]

[0033] Description of the Examples In the packaging process of semiconductor components or semiconductor elements, traceability of the elements to their original wafers and their original positions on the wafers is generally lost. This is because, after the semiconductor elements are cut out, mixing of individual semiconductor elements can occur, resulting in the loss of their positions on the wafer without unique marking of the components. This is schematically illustrated in Figure 1. Wafer 10 contains multiple semiconductor components or semiconductor elements 11. At this stage, each semiconductor element 11 has a known position on wafer 10. Generally, several tests, also called wafer-level tests, are performed on the semiconductor elements 11 at this stage. Following this, wafer 10 is cut, thereby separating the semiconductor elements 11 from each other. Cutting can be done with a saw 12 or a laser. Finally, the cut semiconductor elements are packaged and incorporated, for example, into a microcontroller 13. In this case, by this stage at the latest, information about which wafer 10 and at what position within that wafer 10 the semiconductor element was initially positioned is lost. Generally, multiple tests, also called final tests, are performed again on the microcontroller 13 equipped with the semiconductor element 11. However, because mixing occurs due to the cutting of the wafer 10, it is not easy to uniquely trace back which wafer 10 each semiconductor element 11 of the microcontroller 13 was placed on, and which wafer-level test corresponds to which final test, that is, to confirm that the test results are for the same semiconductor element. The semiconductor element can be, for example, an integrated circuit (hereinafter also referred to as a chip), a sensor, or other microelectronic module.

[0034] The objective of this invention is to restore traceability after the packaging process in semiconductor manufacturing processes. Such mapping enables further contributions such as better process control or earlier prediction of final chip characteristics. In addition, the analysis of the causes of deviations measured during final testing on the chip plane can be extended to the wafer production process. Furthermore, this enables a deeper understanding of the process, leading to better process control and ultimately improved quality.

[0035] A mapping algorithm is proposed, consisting of an alternating sequence of optimizing regression parameters (in the regression from wafer-level tests to final test data) and subsequently optimizing the mapping of test partners. The current mapping of the final test chip is used as the "regression label" in each iteration.

[0036] The present invention further utilizes a cost-minimizing algorithm that can find the optimal one-to-one correspondence under a pre-defined cost matrix. Regression errors are used to construct a suitable cost matrix, which is done by calculating a suitable distance size (e.g., L2 norm) between the final test predictions of the trained regressor and the regression labels. Based on this cost matrix, the algorithm rearranges the chips in the final test to minimize the regression loss. Depending on the characteristics of the data, the regressor or regression model can be arbitrarily selected (e.g., linear regression for linear dependencies).

[0037] Figure 2 schematically shows a flowchart 20 of the method for determining the mapping rule, which maps the test results of the final test to the corresponding test results of the wafer-level test. After this method is completed, a mapping rule is obtained that maps the relevant test results of the wafer-level test to the final test. In other words, this describes the relevant test results originating from the same semiconductor component.

[0038] This method begins with step S21a. In this step, the mapping rules are initialized. Furthermore, in this step, the test results for wafer-level testing (WLT) and final testing (FT) are prepared. Due to yield loss, there may be only a small fraction of the FT test results compared to the WLT test results.

[0039] Step S21a is followed by step S21b. In this step, the yield loss is determined, which is obtained, for example, from the ratio of the WLT test result to the FT test result. Depending on the yield loss, a subset of the WLT test results is randomly selected. This subset corresponds, for example, to the set of losses of semiconductor components.

[0040] Next, step S22 follows. In this step, a training dataset is created, which has WLT test results and FT test results mapped to these test results according to mapping rules, and a subset selected in step S21b is removed from the WLT test results.

[0041] It should be noted that in this embodiment, the practice of removing test results was employed so that the training dataset had an equal number of WLT test results and FT test results. However, it is equally possible to add FT test results instead of removing them. This addition can be done, for example, based on heuristics.

[0042] Step S23 follows the completion of step S22. In this step, the regressor f is trained so that it can determine the final test, which is mapped according to the training dataset and depends on the wafer-level test (WLT): f(WLT) = FT. The regressor f can be a linear regression model. The regressor is trained by tuning the parameters of the regressor f by known methods, for example, by minimizing the regression error in the training dataset.

[0043] Step S24 follows after the regressor has been trained. Here, a cost matrix is ​​created. The rows and columns are mapped to wafer-level tests and final tests, respectively. Entries to the cost matrix are obtained from the training data by using the L2 norm between regressor predictions, depending, for example, on the corresponding WFT test results for each row and the corresponding FT test results for each column, and stored in the cost matrix. Since the number of test results differs, the cost matrix has a rectangular shape.

[0044] Once step S24 is complete, step S25 continues with the optimization of the mapping rules. This optimization is performed by applying the Hungarian algorithm to the cost matrix to obtain improved mapping rules based on the cost matrix.

[0045] Next, step S26 is performed, in which a first variable is selected that is not mapped to one of the second variables even by the optimized mapping rule. In this case, these selected first variables form a subset of the first variables that are not mapped to the corresponding FT test results.

[0046] If the interruption criteria are not met, steps S22 to S25 are repeated. It should be noted that repeating step S22 means that the subset selected in the previous iteration in step S26 is removed from the training dataset. The interruption criteria can be a predetermined maximum number of iterations.

[0047] If the interruption criteria are met, this method can terminate and the mapping rules can be output.

[0048] In an optional step following step S25, the position of the semiconductor component 11 on the wafer 10 is restored by using this mapping rule. Here, based on the mapping rule, the WLT test results can be identified by working backward from the FT test results. Generally, in addition to the WLT test results, the location on the wafer where each test was performed is stored, thereby restoring the exact location on the wafer where the corresponding semiconductor element was manufactured. Consequently, the positions of semiconductor components whose WFT test results are not mapped to FT test results by the mapping rule can also be restored. That is, the semiconductor components sorted after the WLT test and their individual positions are identified. What is assumed here is that, depending on the position restoration in step S25 and / or step S26, control signals are triggered to control physical systems such as manufacturing machines, especially wafer processing machines, for example, computer-controlled machines. For example, if the FT test results are not optimal, the control signals can be used to adjust the preceding manufacturing steps accordingly, thereby obtaining better FT test results thereafter.

[0049] Figure 3 schematically shows the apparatus 30 for carrying out the method shown in Figure 2.

[0050] The device includes a preparation unit 51 that prepares a training dataset according to step S22. The training data is then supplied to a regressor 52, which determines the output variables. The output variables and training data are supplied to a determination unit 53, which determines updated parameters for the regressor 52 from them, and these parameters are transmitted to a parameter memory P, where they replace the current parameters. The determination unit 53 is configured to perform step S23.

[0051] The steps performed by the device 30 can be implemented as a computer program, stored in the machine-readable storage medium 54, and executed by the processor 55.

[0052] The term "computer" includes any device for processing pre-configurable computational rules. These computational rules may be provided in the form of software, hardware, or a combination of software and hardware.

Claims

1. A method for obtaining a mapping rule for mapping a plurality of first variables belonging to a first set consisting of first variables to a second variable belonging to a second set consisting of second variables, a step of initializing the mapping rule (S21a); a step of preparing the first set and the second set (S21a), wherein the first set has more variables than the second set (S21a); a step of randomly selecting a plurality of first variables (S21b), wherein the number of the selected first variables corresponds to at least the following number, that is, the first set has more variables than the second set by that number (S21b); a step of repeatedly performing steps a) to d), that is, a) training the machine learning system so that the machine learning system obtains the second variables respectively mapped according to the mapping rule depending on the first variables of the first set that do not include the selected first variables (S23); b) a step of obtaining a cost matrix (S24), wherein the entries of the cost matrix represent the distance between the predictions of the machine learning system depending on the first variables of the first set and the second variables of the second set (S24); c) optimizing the mapping rule depending on the cost matrix so that the mapping of the first variables to the second variables according to the mapping rule results in the minimum total cost based on the entries of the cost matrix (S25); d) a step of selecting a first variable that is not mapped to any one of the second variables even by the optimized mapping rule (S26); a step of repeatedly performing; A method comprising.

2. The step of optimizing the mapping rule (S25) is performed using the Hungarian algorithm or the greedy implementation method, The method according to claim 1.

3. The number of columns and rows of the cost matrix is equal to the number of the first variables and the second variables, and in the step of obtaining the cost matrix (S24), the distance between the predictions of the machine learning system for all the first variables with respect to all the second variables is obtained respectively, The method according to claim 2.

4. The machine learning system is a regression model (53), and the regression model (53) obtains the second variable depending on the first variable and the parameters of the regression model (53). The method according to claim 1.

5. The first variable and the second variable each represent a product when the product is manufactured according to different manufacturing process steps, and the mapping rule indicates which of the variables in the first set and the second set represent the same product. The number of the first variables randomly selected in the step of randomly selecting (S21b), or the number of the first variables selected in the step of selecting (S26) corresponds to the yield loss during manufacturing. The method according to claim 1.

6. The first variable is a first test result of a semiconductor element on a wafer, the second variable is a second test result of the semiconductor element after the semiconductor element is cut out from the wafer, and the mapping rule indicates which first and second test results are derived from the same semiconductor element. The method according to claim 1.

7. The first test result is a wafer-level test result, and the second test result is a final test result. The method according to claim 6.

8. The semiconductor element is manufactured on a plurality of different wafers. The method according to claim 1.

9. Depending on the mapping rule, determine which second test results belong to which first test results, then, depending on the first test results to which they belong, determine the positions of the semiconductor elements within the wafer, and determine the positions of the semiconductor elements for which the second test results do not exist. The method according to claim 6.

10. The semiconductor element is a power MOSFET (in English, "power MOSFET"). The method according to claim 6.

11. An apparatus (30) configured to implement the method according to any one of claims 1 to 10.

12. A computer program comprising instructions for causing a computer to perform the method according to any one of claims 1 to 10 when the computer program is executed by the computer.

13. A machine-readable storage medium storing the computer program according to claim 12.