Method and system for modeling semiconductor process
By grouping semiconductor process steps and training machine learning models, the problems of high cost and low efficiency of semiconductor process modeling in the prior art are solved, and more accurate and efficient process modeling and analysis are achieved.
Patent Information
- Application Number
- CN202411473732.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-23
- Filing Date
- 2024-10-22
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to accurately and efficiently model semiconductor processes, resulting in high costs of pre-estimating and analyzing process results when considering a variety of process conditions.
By defining the input data of the subprocess steps and the measurement steps, grouping the subprocess steps correspond to multiple modules and training a machine learning model, including the first submodel and the second submodel, predicting the characteristics of the semiconductor device.
Achieve more accurate and efficient modeling of semiconductor processes, providing improved explanatory properties and higher quality analysis, reducing costs.
Smart Images

Figure CN120030971A_ABST
Abstract
Description
[0001] This application is based on and claims priority from Korean Patent Application No. 10-2023-0164858 filed on November 23, 2023, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein in its entirety by reference. Technical Field
[0002] Various example embodiments relate to methods and / or systems for modeling semiconductor processes, and more particularly, to methods and / or systems for modeling semiconductor processes by using learning, such as machine learning. Background Art
[0003] Predicting the results by analyzing the semiconductor process in advance can improve the reliability of the characteristics of the semiconductor device and / or can shorten the period of developing and studying the semiconductor device. Such predictions may generally be referred to as technology based on computer-aided design (TCAD).
[0004] However, with the recent progress in semiconductor technology and the increase in integration, it may require high costs (such as time and / or computing resources) to pre-estimate and analyze the process and its results under various conditions / factors of the process. Therefore, there is a growing demand or desire for a technology that accurately models semiconductor processes while providing improved analytical properties. Summary of the invention
[0005] Various example embodiments provide methods and / or systems for modeling semiconductor processes that may perform modeling of semiconductor processes more accurately and / or efficiently and / or may provide improved interpretative properties and / or higher quality analysis.
[0006] According to some example embodiments, a method for modeling a semiconductor process is provided, the method comprising: obtaining a measurement value based on input data defining a sub-process step and a measurement step; grouping the sub-process steps into corresponding to a plurality of modules, respectively, based on the measurement step; and training a machine learning model based on the grouped sub-process steps to predict characteristics of a semiconductor device. The machine learning model comprises a first sub-model and a second sub-model, the first sub-model being configured to output at least one feature value based on the plurality of modules, and the second sub-model being configured to output an output value representing an estimated value corresponding to each of the plurality of modules and a characteristic of the semiconductor device based on the feature value.
[0007] Optionally or additionally, a method for modeling a semiconductor process is provided, the method comprising: receiving input data defining sub-process steps and measurement steps; based on the measurement steps, grouping the sub-process steps into groups corresponding to a plurality of modules; calculating at least one eigenvalue corresponding to each sub-process step; outputting an estimated value corresponding to each of the plurality of modules based on the eigenvalues by using a sub-model, outputting an output value representing at least one characteristic of the semiconductor device based on the eigenvalues by using the sub-model; and training the sub-model based on at least one of the estimated value and the output value.
[0008] Optionally or additionally, according to various example embodiments, a system for modeling a semiconductor process is provided, the system comprising: at least one processor configured to execute machine-readable instructions, the machine-readable instructions, when executed by the at least one processor, causing the system to: receive input data defining sub-process steps and measurement steps, and provide a machine learning model, the machine learning model being configured to predict characteristics of a semiconductor device based on the input data, wherein the machine learning model calculates a feature value corresponding to each sub-process step, and outputs a first output value predicting at least one characteristic of the semiconductor device based on the feature value and a second output value predicting a change in the at least one characteristic of the semiconductor device based on a change in a physical characteristic. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Embodiments will be more clearly understood from the following detailed description in conjunction with the accompanying drawings.
[0010] Figure 1 is an illustration of a modeling process according to some example embodiments.
[0011] Figure 2 is a block diagram of the structure of a machine learning model according to some example embodiments.
[0012] Figure 3 is an illustration of a modeling process according to some example embodiments.
[0013] Figure 4 is a diagram of grouping according to some example embodiments.
[0014] Figure 5 is a diagram of a filtering process according to some example embodiments.
[0015] Figure 6 is a diagram of eigenvalue calculation according to some example embodiments.
[0016] Figure 7 is a diagram illustrating feature map generation according to some example embodiments.
[0017] Figure 8is an illustration of feature map analysis according to some example embodiments.
[0018] Fig. 9 is a diagram of loss terms based on estimated values according to some example embodiments.
[0019] Fig.10 is a diagram of a method of predicting characteristics of a semiconductor device according to some example embodiments.
[0020] Fig.11 is a diagram of a loss term based on an output value according to some example embodiments.
[0021] Fig.12 is a diagram illustrating an example of utilizing a loss term according to some example embodiments.
[0022] Fig.13 is a flow chart of a modeling method according to some example embodiments.
[0023] Fig.14 is a flowchart of a method for calculating a feature value according to some example embodiments.
[0024] Fig.15 is a flow chart of a training method according to some example embodiments.
[0025] Fig.16 is a flow chart of a training method according to some example embodiments.
[0026] Fig.17 is a block diagram of a processor that provides a machine learning model according to some example embodiments.
[0027] Fig.18 is a block diagram of a modeling system according to some example embodiments. DETAILED DESCRIPTION
[0028] Hereinafter, various exemplary embodiments will be described in detail with reference to the accompanying drawings.
[0029] Figure 1 is an illustration of a modeling process according to some example embodiments.
[0030] Reference Figure 1 The machine learning model 200 for modeling a semiconductor process may receive input data and may output output data indicating an estimated value and / or an output value of a semiconductor device, and may receive at least one actual measurement value obtained by at least one process equipment 100 and / or a measurement value obtained from a database.
[0031] In some example embodiments, the machine learning model 200 may include an artificial neural network (ANN) for modeling a semiconductor process. ANN may be referred to as and / or may include various computing systems conceived from biological neural networks constituting animal brains. As an example, the machine learning model 200 may include a convolutional neural network (CNN). However, the ANN included in the machine learning model 200 is not limited thereto and may be implemented in various ways. For example, the machine learning model 200 may optionally or additionally include a recurrent neural network (RNN). For example, the machine learning model 200 may be implemented based on one or more of a long short-term memory (LSTM) technology, a gated recurrent unit (GRU) technology, an attention technology, and the like. Unlike a typical algorithm that performs a task according to predefined conditions (such as rule-based programming), an ANN may learn to perform a task by considering multiple samples (or examples) (e.g., multiple pieces of input data). The ANN may have a structure in which artificial neurons (or neurons) are connected to each other, and the connection between neurons may be referred to as a synapse. Neurons may process received signals and send processed signals to other neurons via synapses. The output of a neuron may be referred to as an activation. Neurons and / or synapses may have variable weights, and the influence of a signal processed by a neuron may be increased or decreased according to the weights. In particular, the weight associated with each neuron may be referred to as a bias.
[0032] However, regarding the machine learning method for modeling semiconductor processes, the machine learning model 200 can be implemented based on various learning methods and / or algorithms, not limited to ANN. For example, the machine learning model 200 can also or alternatively be implemented based on random forest technology.
[0033] The input data may include data defining sub-process steps (hereinafter referred to as process steps) and measurement steps. The sub-process steps may be referred to as process steps (or work steps) for manufacturing semiconductor devices. For example, the sub-process steps may indicate a photolithography process, an etching process, an ion implantation process (such as beamline implantation and / or plasma doping implantation), a planarization process (such as a chemical mechanical polishing (CMP) process), a wet process step, a deposition process (such as a chemical vapor deposition process), an annealing process (such as a laser annealing process), a baking process, etc. The measurement step may represent an operation for verifying whether one or more sub-process steps have been correctly performed. For example, the measurement step may be an operation for testing the structural characteristics and / or electrical characteristics of the semiconductor device after a series of sub-process steps have been performed, and in some cases may include one or more of critical dimension (CD) measurement, scribing test measurement, ellipsometry measurement, film thickness measurement, etc.
[0034] The process equipment 100 may receive input data and output a measurement value Mv as a result of actually performing the process. For example, the measurement value Mv may be or may include or indicate a value of a threshold voltage and / or a magnitude of a current (signed or unsigned) of a manufactured semiconductor device, and / or a resistance (such as a sheet resistance of a layer of a semiconductor device and / or a contact resistance of a semiconductor device and / or a via resistance). However, the measurement value Mv is not limited thereto and may include various values indicating characteristics of a semiconductor device. On the other hand, the measurement value Mv may also indicate data in which many actual results of previous processes are stored. For example, the measurement value Mv may indicate historical data obtained from a database (e.g., a domain knowledge database) in which actual process results are recorded.
[0035] As described below, the machine learning model 200 may receive input data and output an estimated value. The estimated value may indicate a value that predicts the result of a measurement step. For example, the estimated value may be or may include a value corresponding to the result of a series of sub-process steps. The machine learning model 200 may output an output value for predicting the characteristics of a semiconductor device. The output value may include a value that predicts the characteristics (e.g., electrical characteristics and / or physical characteristics (such as threshold voltage)) of a semiconductor device to be manufactured by a sub-process step of the input data. On the other hand, the output value may include a value for predicting a change in the characteristics of a semiconductor device according to a change in the physical characteristics of a semiconductor process. For example, when some conditions (e.g., doping concentration) of a sub-process step are changed, the output value may be a value that predicts the characteristics of a semiconductor device (e.g., a change in threshold voltage) accordingly. The modeling method according to some example embodiments may model a semiconductor process through the training of a machine learning model 200 based on the above-mentioned estimated value, output value, and measurement value Mv.
[0036] Figure 2 is a block diagram of the structure of a machine learning model 200 according to some example embodiments.
[0037] Reference Figure 2 , the machine learning model 200 may include a first sub-model 210 and a second sub-model 220. The input data may be grouped (or divided, or partitioned), and the machine learning model 200 may output estimated values and output values as described above based on the grouped input data.
[0038] The first sub-model 210 may include a plurality of modules. Each of the plurality of modules may be referred to as a unit and / or a block that performs at least one function, and may correspond to, for example, a unit in which a series of computational processes occur. Input data may be grouped to correspond to the plurality of modules of the first sub-model 210. The input data may include data defining sub-process steps and measurement steps, and the sub-process steps may be grouped (or divided, or partitioned) based on the measurement steps. For example, each module may include one or more sub-process steps.
[0039] The first sub-model 210 may receive data about sub-process steps grouped into sub-process steps corresponding to a plurality of modules. The first sub-model 210 may represent the data of the sub-process steps as specific values so that the sub-process steps corresponding to each of the plurality of modules may be processed by the machine learning model 200. In some example embodiments, the first sub-model 210 may filter the data of the process steps and represent the filtered data as feature values fv. For example, each feature value fv may be or may be based on a value corresponding to each sub-process step.
[0040] The second sub-model 220 may perform various calculations based on the eigenvalue fv, thereby outputting an estimated value and / or an output value. In some example embodiments, the second sub-model 220 may perform a calculation operation on the eigenvalue fv based on the ANN. The second sub-model 220 may output an estimated value for predicting a result of a sub-process step corresponding to each module. For example, each estimated value may be a value corresponding to each module.
[0041] However, the machine learning model 200 is not limited to the above, and may also include other sub-models for performing multiple functions. The first sub-model 210 and / or the second sub-model 220 may also include lower sub-models therein.
[0042] Figure 3 is an illustration of a modeling process according to some example embodiments.
[0043] Reference Figure 3 The input layer 10 may receive input data Input. The input data Input may be or may include data defining sub-process steps and measurement steps, and when the input data Input is received via the input layer 10, the sub-process steps included in the input data Input may be grouped based on the measurement steps.
[0044] In some example embodiments, the first sub-model 210 may include a convolution layer 20. The first sub-model 210 may input the data of the grouped sub-process steps into the convolution layer 20 to perform a specific filtering process. As will be described below, the first sub-model 210 may represent each sub-process step as a specific value by passing the data of each sub-process step through the convolution layer 20 (e.g., a filter). The data of each sub-process step may be output as one or more feature values fv by passing through the convolution layer 20, and thus, the order between the sub-process steps may also be reflected or maintained.
[0045] In some example embodiments, the second sub-model 220 may include the first dense layer 30 and / or the second dense layer 40. The second sub-model 220 may receive the eigenvalue fv, and may input the received eigenvalue fv to the first dense layer 30 and / or the second dense layer 40. The second sub-model 220 may output an estimated value corresponding to the grouped sub-process steps by performing a specific calculation operation based on the eigenvalue fv by using the first dense layer 30. In addition, the second sub-model 220 may output an output value that predicts one or more characteristics of the semiconductor device and / or an output value that predicts a change in one or more characteristics of the semiconductor device according to a change in a physical characteristic of the semiconductor device by performing a specific calculation operation based on the eigenvalue fv by using the second dense layer 40.
[0046] However, the first sub-model 210 and the second sub-model 220 are not limited to the above, and may alternatively or additionally include other layers for performing various operations in addition to the illustrated layers.
[0047] Figure 4 is a diagram of grouping according to some example embodiments.
[0048] Reference Figure 4 , the input data (N_fab+k_met) may be grouped to correspond to a plurality of modules of the first sub-model 210. The input data may include data defining N sub-process steps and k measurement steps (N and k are natural numbers of 2 or greater). For example, the first measurement step may be performed after the first process step, and the second measurement step may be performed after the second process step and the third process step. In a similar manner, the kth measurement step may be performed after the (N-1)th process step and the Nth process step. In this case, the input data may be grouped based on the measurement step. In some example embodiments, more than one measurement step may be performed consecutively after one or more process steps; example embodiments are not limited thereto. In some example embodiments, process steps performed between measurement steps (e.g., two adjacent measurement steps) may be grouped into one group, and data of the group may be set as input to a corresponding module. For example, each of a plurality of modules may receive data of “one or more process steps grouped based on the above criteria”. For example, data of the first process step may be provided to the first module, data on the second process step and the third process step may be provided to the second module, and data on the (N-1)th process step and the Nth process step may be provided to the kth module. Since the process steps are divided based on the measurement steps and the data of the corresponding process steps are provided to the modules, a plurality of modules may correspond to the measurement steps, respectively.
[0049] In the modeling method according to various example embodiments, by distinguishing the measurement step from the process step through the use of modularization, the input to the machine learning model may include only the process step (i.e., the data of the process step). For example, since the performance result of the measurement step is only the result of some process steps, by separating the performance results from the process steps into different levels, modeling that conforms to the real world (e.g., modeling with better consistency) may be feasible, and / or unnecessary calculation amount may be reduced.
[0050] Alternatively or additionally, by separating the measurement steps, the modularization of the modeling method according to various example embodiments can prevent or reduce the possibility of a reduction in the number of data sets due to missing measurements during the measurement steps in the actual process. For example, when the data of the measurement step is also input into the machine learning model, since there is no data for the part where the missing measurement occurs, learning may become infeasible and / or inaccurate, and thus one or more of poor data utilization, poor consistency, or difficulty in obtaining accurate features and results may occur. On the other hand, in the modeling method according to various example embodiments, even if missing measurements occur, since learning is performed for each module, the tolerance for missing measurements can be increased, and by making full use of the input data set, it may be feasible to increase consistency and / or more accurately identify the desired features.
[0051] Alternatively or additionally, since the modeling method according to various example embodiments can calculate estimated values for each module separated in this way, learning close to the actual measured values for each module may be feasible, the accuracy of modeling can be further improved, and / or more improved process conditions or more optimized process conditions can be suggested.
[0052] Figure 5 is a diagram of a filtering process according to some example embodiments. Figure 6 is a diagram of eigenvalue fv calculation according to some example embodiments.
[0053] Reference Figure 5 , the first sub-model 210a may include a convolutional layer 20. The first sub-model 210a may be Figure 2 The first sub-model 210a may correspond to the first sub-model 210 in the embodiment of the present invention and may be or may include or be included in the first sub-model 210. The first sub-model 210a may receive the reference Figure 4The data of the grouping process described. For example, the first sub-model 210a may receive data grouped into data corresponding to each module in the k modules, and each module may perform a convolution multiplication calculation by using the convolution layer 20. For example, the first sub-model 210a may perform a filtering operation. By using the filtering process, the data of each process step included in the module may be represented as various convolution multiplication result values step conv. For example, the convolution multiplication calculation may be or may include a one-dimensional (1D) convolution.
[0054] Reference Figure 6 , the first sub-model 210b can output the feature value fv based on the convolution multiplication result value step conv. The first sub-model 210b can be Figure 2 , and may be an implementation. As a result of performing a convolution multiplication calculation on any one of the process steps, a plurality of convolution multiplication result values step conv may be generated, and the first submodel 210b may output a feature value fv corresponding to any one of the process steps based on a calculation using the convolution multiplication result value step conv. In some example embodiments, the first submodel 210b may include an average pooling layer that performs average value calculation. For example, the average value may be or be based on one or more measures of central tendency (such as, but not limited to, one or more of the mean, median, or mode). Each module may perform an average calculation by using an average pooling layer. Therefore, the first submodel 210b may output the average value of the result value step conv as the feature value fv by performing an average pooling operation. For example, the first module may output the first eigenvalue fv1 as a value corresponding to (or indicating) the first process step; the second module may output the second eigenvalue fv2 as a value corresponding to the second process step, and output the third eigenvalue fv3 as a value corresponding to the third process step; in the same manner, the kth module may output the (N-1)th eigenvalue fvN-1 as a value corresponding to the (N-1)th process step, and output the Nth eigenvalue fvN as a value corresponding to the Nth process step. For example, the eigenvalue fv corresponding to each process step may have a value between about -1 and about 1. However, such operations and / or calculations are not limited to the above, and other calculation operations may be additionally or optionally performed to output the eigenvalue fv. For example, a maximum pooling operation may also be performed on the result value step conv.
[0055] The modeling method according to various example embodiments can reflect the order of process steps included in the module by using a specific filtering operation, and can be expressed as a feature value fv reflecting the characteristics / characteristics of each process step. Optionally or additionally, since how the model parses the input (process step) can be understood by using the feature value fv, the modeling method according to the embodiment can further improve the analytic properties of the model.
[0056] Figure 7 is a diagram illustrating feature map generation according to some example embodiments.
[0057] Reference Figure 7 , the first sub-model 210c may generate a first feature map based on a plurality of feature values fv1 to fvN. The first sub-model 210c may be Figure 2 The first sub-model 210 corresponds to and may be Figure 2 Implementation of the first sub-model 210 in FIG. Figure 5 and Figure 6 As described, by using a filtering process, the data of the grouped process steps may be represented as a feature value fv. For example, the process steps defined in the input data may include a first process step, a second process step, a third process step, ..., an (N-1)th process step, and an Nth process step (N is a natural number of 2 or greater). After a specific filtering process is completed, the first process step may be represented as a first feature value fv1, the second process step may be represented as a second feature value fv2, the third process step may be represented as a third feature value fv3, the (N-1)th process step may be represented as an (N-1)th feature value fvN-1, and the Nth process step may be represented as an Nth feature value fvN. The first sub-model 210c may generate a first feature map by using a plurality of feature values fv1 to fvN. The first feature map may be represented, for example, in the form of a graph in which each process step and the feature value corresponding thereto are matched.
[0058] Since the modeling method according to various example embodiments can identify how characteristics of a single process step are reflected as a characteristic value through a characteristic graph, the relationship between which process steps affect which characteristics of a semiconductor device can be identified and the analytical properties of the model can be further improved.
[0059] Figure 8 is a diagram of an example of an analysis feature graph according to some example embodiments.
[0060] Reference Figure 7 and Figure 8, the first sub-model 210c may generate a second feature map in which two or more pieces of input data are modeled and compared with each other. For example, the first sub-model 210c may output feature values fva to fvc indicating “data of each process step (step a to step c) corresponding to the manufacture of a substrate (such as wafer A)”, and the feature values fva to fvc may be displayed as a graph on the second feature map. In the same manner, the first sub-model 210c may output feature values fva' to fvc' indicating “data of each process step (step a to step c) corresponding to the manufacture of a substrate (such as wafer B)”, and the feature values fva' to fvc' may be displayed as a graph on the second feature map. The global process (e.g., the process of step a to step c) for manufacturing wafers A and B may be the same, but the detailed process conditions of each step may be different. For example, there may be an in-silico experimental design between the processes of wafer A and wafer B. For example, when the process conditions of process step a for manufacturing wafer A are different from those of process step a for manufacturing wafer B (for example, when the doping concentration is different during one or more doping processes), the characteristic value fva may be different from the characteristic value fva'. Likewise, the characteristic value corresponding to each process step of each wafer may change differently.
[0061] The first sub-model 210c may represent the characteristic values fva to fvc of the process steps of wafer A and the characteristic values fva' to fvc' of the process steps of wafer B in the second characteristic graph, and may represent a graph of various electrical characteristics ET comparing the electrical characteristics of wafer A with the electrical characteristics of wafer B in the second characteristic graph. For example, the electrical characteristic ET may indicate information about one or more threshold voltages, such as the threshold voltages of various transistors, such as thin gate transistors and / or thick gate transistors, and the difference between the threshold voltages of specific transistors of wafer A and wafer B may be shown as a line graph (ET graph). The correlation between the difference in the characteristic values of each process step of the two wafers and the graph of the change in the electrical characteristic ET may be identified. For example, as shown in the figure, it can be understood that in a step where the difference in the characteristic values between the process steps is relatively large, the change in the electrical characteristic ET is relatively large.
[0062] In some example embodiments, by analyzing correlations through the use of feature maps, for example, by identifying portions where changes in the electrical characteristic ET are significant, the modeling method may also identify differences in the effects of process steps (e.g., process steps related to photomask operations and / or annealing processes) on electrical performance between an N-channel field effect transistor (FET) (NFET) process and a P-channel FET (PFET) process and / or between a thin gate process and a thick gate process.
[0063] As described above, the modeling method according to some example embodiments may generate a feature map based on the feature value of each of the different pieces of input data, and by showing the change in electrical characteristics in the feature map, information for identifying correlation by comparing the differences between process steps and the change in electrical characteristics may be provided.
[0064] Fig. 9 is a diagram of loss terms based on estimated values according to some example embodiments.
[0065] Reference Fig. 9 The second sub-model 220a may receive the feature values fv1 to fvN corresponding to the process steps, respectively, and output the estimated values est1 to estk. Figure 2 . As described above, the input data defining the N process steps and the k measurement steps may be grouped to correspond to the k modules based on the measurement steps. As described above, the data of the grouped process steps provided to the plurality of modules (e.g., the first process step provided to the first module, the second to fourth process steps provided to the second module, and the (N-1)th process step and the Nth process step provided to the kth module) may be represented as feature values fv1 to fvN after the filtering operation.
[0066] In some example embodiments, the second submodel 220a may perform a specific operation in units of modules based on the received feature values fv1 to fvN. Therefore, the second submodel 220a may output estimated values est1 to estk corresponding to each module, respectively, as calculation results. In other words, the second submodel 220a may output estimated values est1 to estk corresponding to the measurement steps (the first measurement step to the kth measurement step), respectively. The estimated value may represent a value predicting the result value of each measurement step as described above, that is, a value predicting the result of the sub-process step corresponding to each module. For example, the second submodel 220a may output a first estimated value est1 corresponding to the first module (that is, predicting the result of the process step included in the first module), a second estimated value est2 corresponding to the second module (that is, predicting the results of the second process step to the fourth process step), and a kth estimated value estk corresponding to the kth module. On the other hand, for example, the second submodel 220a may include a first fully connected layer (for example, Figure 3 The first dense layer 30 in FIG. 1 ). The first fully connected layer may receive feature values fv1 to fvN via N nodes, and output estimated values est1 to estk via k nodes.
[0067] In some example embodiments, the machine learning model 200a may compare the estimated values est1 to estk with the first measurement value Mv1. Figure 1 and Figure 2 The machine learning model 200 corresponds to, and can be Figure 1 and Figure 2 Implementation of the machine learning model 200 in. The first measurement value Mv1 may include a first actual measurement value mes1 to a kth actual measurement value mesk as the measurement result of each measurement step. For example, since the process steps are divided into k modules based on the measurement steps, the second sub-model 220a may output k estimated values (a first estimated value est1 to a kth estimated value estk), and the machine learning model 200a may obtain (or receive) the first measurement value Mv1 including k measurement values (a first actual measurement value mes1 to a kth actual measurement value mesk) from the outside, and the k measurement values are the actual measurement results of each of the k measurement steps.
[0068] The machine learning model 200a may compare the estimated values est1 to estk as the predicted values of the model with the first measured value Mv1 as the actual measured value. For example, the machine learning model 200a may calculate the difference between the first estimated value est1 and the first actual measured value mes1, and in the same way, each difference (such as the difference between the second estimated value est2 and the second actual measured value mes2, the difference between the kth estimated value estk and the kth actual measured value mesk) may be calculated. The machine learning model 200a may calculate a first loss L in which all differences are added together. Met . This can be done by using the first loss L Met The machine learning model 200a is trained as input to various loss functions so that the estimated values est1 to estk are close to the first measured value Mv1. For example, the machine learning model 200a may adjust the weights and / or biases between the nodes of the first fully connected layer in the second sub-model 220a.
[0069] In some example embodiments, some of the k measurement steps may not be measured. In other words, due to the missing measurements, some of the first measurement value Mv1 and some of the first actual measurement value mesk to the kth actual measurement value mesk may not have values. Therefore, the machine learning model 200a may not be trained with respect to the portion with missing measurements. However, since the machine learning model 200a is divided into modules and trained, learning can be performed on the portion where the actual measurement values exist, so even if there is a portion where the measurement is missing, data can be accumulated and accurate predictions can be performed as the training proceeds. In other words, even if there are missing measurements, the machine learning model 200a can be trained without data parameter loss, so that the consistency and accuracy of the machine learning model 200a can be improved. In addition, since the machine learning model 200a is divided into modules and trained for each module, prediction and verification for each module may be feasible, complex analysis may be feasible, and the analytical properties of the machine learning model 200a may be further improved.
[0070] Fig.10 is a diagram of a method of predicting characteristics of a semiconductor device 200 b according to some example embodiments.
[0071] Reference Fig.10 , the second sub-model 220b may receive the feature values fv1 to fvN corresponding to each process step, and output a first output value output1 and a second output value output2. Figure 2 The second sub-model 220 in corresponds to and may be Figure 2Implementation of the second sub-model 220 in . The first output value output1 may be a value for predicting the characteristics of a semiconductor device, and may be, for example, a value for predicting the characteristics (e.g., threshold voltage) of a semiconductor device to be manufactured by the process steps (the first process step to the Nth process step) of the input data. The second output value output2 may be a value for predicting the change in the characteristics of the semiconductor device according to the change in physical characteristics. For example, when the conditions (e.g., doping concentration) of some process steps (e.g., doping process) among the process steps (the first process step to the Nth process step) are changed, the second output value output2 may be a value for predicting the characteristics of the semiconductor device (e.g., change in threshold voltage). However, the change in the characteristics of the semiconductor device according to the change in physical characteristics may not be limited to the change in threshold voltage according to doping concentration, but may also include various physical characteristics. For example, the physical characteristics of the semiconductor process may include characteristics / conditions related to the structure, and / or characteristics / conditions related to pressure, etc. For example, the machine learning model 200b may output the second output value output2 to identify the change in threshold voltage according to the change in gate length of the transistor, and / or identify the change in warpage according to the change in pressure. The machine learning model 200b may be used with Figure 1 and Figure 2 The machine learning model 200 corresponds to, and can be Figure 1 and Figure 2 An implementation of the machine learning model 200 in .
[0072] On the other hand, for example, the second sub-model 220b may include a second fully connected layer (eg, Figure 3 The second dense layer 40 in FIG. 4 ). The second fully connected layer may receive feature values fv1 to fvN via N nodes, and output a first output value output1 and a second output value output2.
[0073] In some example embodiments, the first output value output1 and / or the second output value output2 may include a plurality of predicted values. For example, the first output value output1 may include a value predicting various characteristics of the semiconductor device, and the second output value output2 may include a value predicting changes in various characteristics of the semiconductor device according to changes in physical characteristics. In other words, the output of the second fully connected layer may not be limited to two nodes, but may be variously adjusted according to the target characteristic or the number of changes in the characteristic.
[0074] The modeling method according to various exemplary embodiments can predict the change of the characteristics of the semiconductor device according to the change of various physical characteristics, and predict the change of the characteristics of the semiconductor device according to the process steps. In addition, the modeling method according to various exemplary embodiments can predict and analyze the actual results of the manufacturing process by reflecting the change of various physical characteristics in real time, thereby improving the accuracy of modeling and providing a powerful analysis tool.
[0075] Fig.11 is a diagram of a loss term based on an output value according to some example embodiments.
[0076] Reference Fig.10 and Fig.11 The machine learning model 200c can compare the first output value output1 and the second output value output2 as its predicted values with the second measurement value Mv2 and the third measurement value Mv3 as the actual measurement values. Figure 1 and Figure 2 The machine learning model 200 corresponds to, and can be Figure 1 and Figure 2 . The second measurement value Mv2 may be a value indicating a characteristic of a semiconductor device, which has actually been processed according to the process steps (first process step to Nth process step) defined in the input data. For example, the second measurement value Mv2 may be an actual measurement value of a characteristic (e.g., a threshold voltage) of a semiconductor device, which is manufactured using the process steps (first process step to Nth process step) of the input data. The third measurement value Mv3 may be a value obtained by actually measuring a change in a characteristic of a semiconductor device according to a change in a physical characteristic (e.g., doping concentration, structure, pressure, etc.) of a semiconductor process. For example, the third measurement value Mv3 may be a database of semiconductor device processes, that is, data empirically obtained from many accumulated processes. In one example, the machine learning model 200c may obtain (or receive) a change value of a characteristic of a semiconductor device corresponding to a change in a physical characteristic of a target (or arbitrarily set) from a database that integrates many semiconductor process results (i.e., from domain knowledge data) as the third measurement value Mv3.
[0077] In some example embodiments, the machine learning model 200c may compare the first output value output1 as its predicted value with the second measurement value Mv2 corresponding to the first output value output1. For example, the machine learning model 200c may calculate the difference between the first output value output1 and the second measurement value Mv2. When the first output value output1 includes multiple predicted values, the machine learning model 200c may calculate the second loss L obtained by "adding the difference between each predicted value and the corresponding actual measurement value". ETIn the same manner, the machine learning model 200c may compare the second output value output2 as its predicted value with the third measurement value Mv3 “corresponding to the second output value output2”. When the second output value output2 includes multiple predicted values, the machine learning model 200c may calculate the third loss L obtained by “adding the difference between each predicted value and the corresponding actual measurement value (i.e., obtained from the database)”. Phy . This can be done by using the second loss L ET and / or third loss L Phy The machine learning model 200c is trained as input to various loss functions so that the first output value output1 and the second output value output2 are close to the second measurement value Mv2 and the third measurement value Mv3, respectively. For example, the machine learning model 200c can adjust the weights and / or biases between the nodes of the second fully connected layer in the second sub-model 220b.
[0078] In this way, the modeling method according to various example embodiments can not only learn and predict the actual results of the manufacturing process, but also learn and predict the changes in the characteristics of the semiconductor device according to the changes in various physical characteristics / conditions of the semiconductor process. For example, in the modeling method according to various example embodiments, since the model not only learns the data mathematically, but also reflects various physical characteristics in the modeling in real time, the accuracy of the model can be improved, and the analytical ability of the model in terms of prediction can also be enhanced.
[0079] Fig.12 is a diagram illustrating an example of utilizing a loss term according to some example embodiments.
[0080] Reference Fig.11 and Fig.12, the machine learning model 200c may reflect the physical characteristics and / or electrical characteristics of the semiconductor device in the model based on the domain knowledge data. For example, the result of measuring the characteristics of the semiconductor device manufactured based on arbitrary input data may be represented as a first targeting graph targeting graph 1. The results of multiple wafers included in a batch may be represented on the targeting graph as shown. The x-axis of the targeting graph may represent a value "indicating at least one physical characteristic or electrical characteristic in NFET", and the y-axis may represent a value "indicating at least one physical characteristic or electrical characteristic in PFET". For example, the x-axis of the targeting graph may indicate a threshold voltage value of a specific NFET included in the semiconductor device, and the y-axis of the targeting graph may indicate a threshold voltage value of a specific PFET included in the semiconductor device. In order to achieve the target characteristics (product characteristics or target specifications), the characteristics of the semiconductor device manufactured by using the process steps may need to meet the specifications corresponding to the center (thick dashed line center) of the targeting frame TB. As shown in the first targeting graph, the specifications of the semiconductor device manufactured based on the input data may be represented by the first point p1. For example, the actual specifications may not meet the target specifications.
[0081] In some example embodiments, the machine learning model 200c may utilize a database related to the above process (e.g., domain knowledge data) to reflect changes in physical characteristics of the semiconductor process. For example, dose or dopant sensitivity data (e.g., domain knowledge data of changes in threshold voltage according to changes in doping concentration) may be applied to the machine learning model 200c. Fig.11 As described above, based on the dose sensitivity data empirically obtained from a large amount of accumulated process data, when the doping concentration and / or dopant distribution changes by a specific amount, the machine learning model 200c can correspondingly obtain information about the degree of change in the threshold voltage value. Therefore, the machine learning model 200c can calculate the third loss L obtained by "comparing the output value of the machine learning model 200c with the actual measurement value (dose sensitivity data)". Phy , and can be obtained by using the third loss L Phy As the loss term of the loss function, the machine learning model 200c can be trained to converge to the domain knowledge data.
[0082] In some example embodiments, the machine learning model 200c trained based on the above method can be used in a targeting process. For example, in order to meet the target specification, some of the conditions of the sub-process steps defined as input data can be adjusted. For example, in order to meet the specifications corresponding to the center of the targeting frame TB as described above, the doping concentration of the NFET and / or PFET can be changed. The machine learning model 200c can predict the changed characteristics of the semiconductor device in response to the change in doping concentration (e.g., the change in threshold voltage), and therefore, the result can be shown as a second targeting graph targeting graph 2. As shown in the second targeting graph targeting graph 2, the specification of the semiconductor device can be implemented as a second point p2, and is located at the center of the targeting frame TB, so the center targeting can be successful.
[0083] For example, as mentioned above Figures 10 to 12 As described, the machine learning model 200c can predict changes in characteristics of a semiconductor device by reflecting changes in physical characteristics of a semiconductor process in real time, and based on the prediction, can provide information for explanation / analysis.
[0084] Fig.13 is a flow chart of a modeling method according to some example embodiments.
[0085] Reference Fig.13 , the modeling method according to various example embodiments may include a plurality of steps or operations S100, S110, S120, S130, S140, S150 and S160 as shown herein. In the following, the modeling method is described with reference to the previous figures. Fig.13 , and the repeated description given with reference to the previous figures is omitted.
[0086] In operation S100, the machine learning model 200 may receive input data. The input data may be data defining a process step and a measurement step. The process step may include a sub-process step for manufacturing a semiconductor device, and the measurement step may include a step for identifying whether the sub-process step has been correctly performed.
[0087] In operation S110, input data may be classified into corresponding modules by grouping. Sub-process steps may be classified into corresponding modules based on measurement steps. For example, input data may be divided based on measurement steps, and sub-process steps performed between measurement steps (e.g., two adjacent measurement steps) may be grouped into corresponding modules. Thus, multiple modules may correspond to measurement steps, respectively.
[0088] In operation S120, the machine learning model 200 may represent the data of the sub-process steps as feature values (eg, by performing the above-mentioned filtering process). For example, the data of each sub-process step may be output as a feature value fv via the convolution layer 20, and thus the order between the sub-process steps may also be reflected.
[0089] In operation S130, the machine learning model 200 may output an estimated value for predicting a result value of the measurement step. The machine learning model 200 may output an estimated value corresponding to each module of the plurality of modules (i.e., for each module) by performing various calculations based on the feature value. In other words, the estimated value may be a value for predicting a result of a sub-process step included in each module. In some example embodiments, the machine learning model 200 may include a sub-model for calculating the estimated value based on the feature value, and the sub-model may include a fully connected layer.
[0090] In operation S140, the machine learning model 200 may output a value (eg, an output value) related to the prediction of the characteristics of the semiconductor device based on the feature value. In some example embodiments, the machine learning model 200 may include a sub-model that calculates the output value based on the feature value, and the sub-model may include one or more fully connected layers.
[0091] In operation S150, the machine learning model 200 may be trained based on the estimated value and the output value. For example, the machine learning model 200 may train a sub-model by comparing the output value and the estimated value "as the estimated value of the machine learning model 200" with the actual measured value (or accumulated measured value).
[0092] Some example embodiments may manufacture a device based on the output of the machine learning model 200; however, example embodiments are not limited thereto. For example, in operation S160, a semiconductor device may be manufactured based on the sub-model. For example, impurities may be doped or injected into the semiconductor device based on the sub-model.
[0093] Fig.14 is a flowchart of a method for calculating an eigenvalue fv according to some example embodiments.
[0094] Reference Fig.13 and Fig.14 The modeling method according to various exemplary embodiments may calculate the characteristic value based on various methods, and may generate a characteristic graph based on the characteristic value. The operation S120 of calculating the characteristic value may include a plurality of operations S121, S122, and S124. Figures 5 to 7 describe Fig.14 , and the repeated description given with reference to the previous figures is omitted.
[0095] In operation S121, the machine learning model 200 may perform a convolution calculation for performing a specific filtering operation. The machine learning model 200 may represent "sub-process steps grouped into corresponding to a plurality of modules" as a specific value (e.g., a feature value) based on the convolution calculation. The order between the sub-process steps may be reflected in the data through filtering processing. For example, the machine learning model 200 may include a convolution layer. The machine learning model 200 may input the data of the sub-process steps into the convolution layer based on a plurality of modules, and may output a calculation result value.
[0096] In some example embodiments, in operation S122, the machine learning model 200 may output a feature value based on the result value of the convolution calculation. The machine learning model 200 may calculate the average value of the result value step conv by performing an average pooling operation. For example, the machine learning model 200 may include an average pooling layer. The machine learning model 200 may input the result value of the convolution calculation to the average pooling layer and output the feature value.
[0097] In some example embodiments, in operation S124, the machine learning model 200 may generate a feature graph based on the feature value. The machine learning model 200 may match each sub-process step with the feature value corresponding thereto and represent the result in the form of a graph. Based on the graph, a correlation between the sub-process step and the characteristics of the semiconductor device may be identified.
[0098] Fig.15 is a flow chart of a training method according to some example embodiments.
[0099] Reference Fig. 9 and Fig.15 The modeling method according to various example embodiments may perform training based on the estimated values. The machine learning model 200 may be trained by comparing the estimated values with the actual measured values.
[0100] In operation S151, the machine learning model 200 may compare an estimated value corresponding to each of the multiple modules with an actual measurement value. The actual measurement value may be an actual measurement result of each measurement step, and thus the actual measurement value may correspond to each of the multiple modules. The machine learning model 200 may obtain (or receive) an estimated value and compare the obtained estimated value with the measurement value corresponding thereto. In other words, the machine learning model 200 may calculate an estimated value for each module and compare the estimated values with each other.
[0101] In operation S152, the machine learning model 200 may calculate a first loss L obtained by comparing the estimated value with the measured value (eg, the actual measured value). Met For example, the machine learning model 200 may output a first loss L obtained by adding the difference between each of the multiple estimated values and each measured value.Met .
[0102] In operation S153, the machine learning model 200 may use the first loss L Met As input to various loss functions, the machine learning model 200 can be trained so that the estimated value is close to the measured value (for example, the machine learning model 200 can be trained so that the first loss L Met is reduced). In some example embodiments, a sub-model of the machine learning model 200 that calculates the estimated value based on the feature value may be trained. For example, the machine learning model 200 may adjust the weights and / or biases between the nodes of the first fully connected layer used to calculate the estimated value. Therefore, the machine learning model 200 may be trained for each module.
[0103] Fig.16 is a flow chart of a training method according to some example embodiments.
[0104] Reference Fig.10 , Fig.11 and Fig.16 , the modeling method according to various example embodiments may perform training based on the output value of the model.
[0105] In operation S141, the machine learning model 200 may generate a first output value for predicting a characteristic of a semiconductor device to be manufactured by using a sub-process step included in the input data. In some example embodiments, the machine learning model 200 may include a sub-model that calculates the first output value based on the characteristic value. For example, the machine learning model 200 may predict a threshold voltage value of a specific transistor included in the semiconductor device to be manufactured.
[0106] In operation S154, the machine learning model 200 may compare the first output value with the actual measured value. The machine learning model 200 may compare its predicted value (first output value) with a value (measured value) indicating a characteristic of a semiconductor device actually manufactured according to the process steps defined in the input data.
[0107] In operation S155, the machine learning model 200 may calculate a second loss L as a result of comparing the first output value with the actual measurement value. ET For example, the machine learning model 200 may output the difference between the predicted threshold voltage value and the threshold voltage value of the actually manufactured semiconductor device as the second loss L ET .
[0108] In operation S142, the machine learning model 200 may generate a second output value that predicts a change in a characteristic of the semiconductor device according to a change in the physical characteristic. In some example embodiments, the machine learning model 200 may include a sub-model that calculates the second output value based on the characteristic value. For example, the machine learning model 200 may calculate a change in a threshold voltage and / or some other electrical characteristic (such as an on-current of a specific transistor) according to a change in doping concentration and / or a change in dopant distribution.
[0109] In operation S156, the machine learning model 200 may compare the second output value with the actual measured value. The machine learning model 200 may compare its predicted value (second output value) with a value (measured value) indicating the amount of change in the characteristics of the semiconductor device that actually occurred, based on the change in the physical characteristics of the semiconductor process. For example, the measured value may be data empirically obtained from many accumulated processes. In various example embodiments, the machine learning model 200 may obtain (or receive) "the amount of change in the characteristics of the semiconductor device corresponding to the change in the physical characteristics of the semiconductor process to be identified (e.g., doping concentration, structure, pressure, etc.)" as a measured value from a database (e.g., domain knowledge data) that integrates many accumulated semiconductor process results.
[0110] In operation S157, the machine learning model 200 may calculate a third loss L as a result of comparing the second output value with an actual measurement value (eg, domain knowledge data). Phy For example, the machine learning model 200 may output the difference between the predicted threshold voltage change and the actual measured threshold voltage value change of the semiconductor device (eg, obtained from domain knowledge data) as the third loss L Phy .
[0111] In operation S158, the second loss L Met and / or third loss L Phy The machine learning model 200 is trained as an input to various loss functions so that the predicted values of the machine learning model 200 are close to the measured values (for example, the machine learning model 200 can be trained so that the second loss L Met and / or third loss L Phy In some example embodiments, a sub-model of the machine learning model 200 that calculates the first output value and the second output value based on the feature value may be trained. For example, the machine learning model 200a may adjust the weights and / or biases between the nodes of the second fully connected layer used to calculate the second output value and / or the third output value.
[0112] Fig.17 is a block diagram of a processor that provides a machine learning model 311 according to some example embodiments.
[0113] Reference Figure 1 and Fig.17 , at least one processor 310 may provide a machine learning (ML) model 311 for modeling a semiconductor process. The machine learning model 311 may be used with Figure 1 The machine learning model 200 in FIG. hereinafter is described with reference to the previous diagram. Fig.17 , and the repeated description given with reference to the previous figures is omitted.
[0114] At least one processor 310 may execute a program module including system executable commands. The program module may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. For example, at least one processor 310 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a neural processing unit (NPU). However, this is an example, and at least one processor 310 is not limited thereto. At least one processor 310 may receive input data defining sub-process steps and measurement steps. At least one processor 310 may divide the process steps included in the input data into multiple modules based on the above modularization. At least one processor 310 may execute a machine learning model 311 to output an estimated value and / or an output value representing the characteristics of a semiconductor device. In addition, at least one processor 310 may receive a measurement value obtained from a database, and may train the machine learning model 311 by performing processing.
[0115] In some example embodiments, the machine learning model 311 may include an ANN and / or an algorithm for performing a modeling operation on a semiconductor process. The machine learning model 311 may output an estimated value based on input data. The estimated value may be or may correspond to a value for predicting a result of a measurement step, and the machine learning model 311 may predict "a result of executing a sub-process step included in one module" and may output an estimated value. The machine learning model 311 may output an output value for predicting a characteristic of a semiconductor device based on the input data. The output value may include a value for predicting a change in a characteristic of a semiconductor device according to a change in a physical characteristic.
[0116] Fig.18 is a block diagram of a modeling system 400 , according to some example embodiments.
[0117] Reference Fig.17 and Fig.18 The semiconductor process modeling system 400 may include at least one processor 410, an artificial intelligence (AI) accelerator 420, a memory 430, and a hardware (HW) accelerator 440, and the at least one processor 410, the AI accelerator 420, the memory 430, and the hardware accelerator 440 may communicate with each other via a bus 450. The at least one processor 410 may communicate with Fig.17In some example embodiments, at least one processor 410, the AI accelerator 420, the memory 430, and the hardware accelerator 440 may also be included in one semiconductor chip. In addition, in some example embodiments, at least two of the at least one processor 410, the AI accelerator 420, the memory 430, and the hardware accelerator 440 may also be included in each of the two or more semiconductor chips mounted on the substrate.
[0118] At least one processor 410 may execute commands. For example, at least one processor 410 may also execute an operating system by executing commands stored in memory 430, and / or may execute an application running on the operating system. In some example embodiments, at least one processor 410 may instruct an AI accelerator 420 and / or a hardware accelerator 440 to perform a task by executing a command, and may also obtain a result of performing the task from the AI accelerator 420 and / or the hardware accelerator 440. In some example embodiments, at least one processor 410 may include an application-specific instruction set processor (ASIP) customized for a specific purpose, and may also support a dedicated instruction set.
[0119] The memory 430 may have any structure for storing data. For example, the memory 430 may also be or alternatively include a volatile memory device (such as one or more of a dynamic random access memory (RAM) (DRAM), a static RAM (SRAM)), or may include a non-volatile memory device (such as a flash memory and a resistive RAM (RRAM)). At least one processor 410, AI accelerator 420, and hardware accelerator 440 may transfer data (e.g., Figure 1 The input data, measured values, estimated values, and output values in the memory 430 are stored in the memory 430, or data can be read from the memory 430 (for example, Figure 1 input data, measurements, estimates, and output values in ).
[0120] The AI accelerator 420 may be referred to as hardware designed for AI applications. In some example embodiments, the AI accelerator 420 may include an NPU for implementing a neuromorphic structure, and the AI accelerator 420 may generate output data by processing input data provided by at least one processor 410 and / or hardware accelerator 440, and may provide the output data to at least one processor 410 and / or hardware accelerator 440. In some example embodiments, the AI accelerator 420 may be programmable and may be programmed by at least one processor 410 and / or hardware accelerator 440.
[0121] The hardware accelerator 440 may be referred to as hardware designed to perform a specific task at high speed. For example, the hardware accelerator 440 may be designed to perform data transformations (such as demodulation, modulation, encoding, and decoding) at high speed. The hardware accelerator 440 may be programmable and may be programmed by at least one processor 410 and / or the hardware accelerator 440.
[0122] In some example embodiments, the AI accelerator 420 may execute the machine learning model described above with reference to the previous diagrams (e.g., Figure 2 Machine learning model 200). Alternatively or additionally, for example, the AI accelerator 420 may execute “a model that outputs a feature value based on the above input data” and “a model that predicts an estimated value of each of a plurality of modules or predicts characteristics of a semiconductor device”. The AI accelerator 420 may generate an output including useful information by processing input parameters, feature maps, etc. In addition, in some example embodiments, at least some of the models executed by the AI accelerator 420 may also be executed by at least one processor 410 and / or a hardware accelerator 440.
[0123] Any element and / or functional block disclosed above may include a processing circuit (such as hardware including a logic circuit, a hardware / software combination (such as a processor that executes software), or a combination thereof), or may be implemented in a processing circuit (such as hardware including a logic circuit, a hardware / software combination (such as a processor that executes software), or a combination thereof). For example, the processing circuit may more specifically include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application specific integrated circuit (ASIC), etc. The processing circuit may include an electrical component such as at least one of a transistor, a resistor, a capacitor, etc. The processing circuit may include an electrical component such as a logic gate (including at least one of an AND gate, an OR gate, a NAND gate, a NOT gate, etc.).
[0124] Although various example embodiments have been specifically shown and described with reference to the various example embodiments, it will be understood that various changes in form and details may be made therein without departing from the spirit and scope of the appended claims. The example embodiments are not necessarily mutually exclusive. For example, some example embodiments may include one or more features described with reference to one or more of the accompanying drawings, and may also include one or more other features described with reference to one or more of the other accompanying drawings.
Claims
1. A method for modeling a semiconductor process, the method comprising: obtaining measurement values based on input data defining sub-process steps and measurement steps; Based on the measurement steps, the sub-process steps are grouped into corresponding to a plurality of modules respectively; as well as training a machine learning model to predict at least one characteristic of the semiconductor device based on the grouped sub-process steps and the measured values, Among them, machine learning models include: a first sub-model configured to output at least one feature value based on the plurality of modules, and The second sub-model is configured to output an estimated value corresponding to each of the plurality of modules and an output value representing the at least one characteristic of the semiconductor device based on the at least one feature value.
2. The method according to claim 1, wherein: The first sub-model includes a convolutional layer configured to receive input data corresponding to the plurality of modules.
3. The method according to claim 2, wherein: The first sub-model outputs a value based on an average of output values of the convolutional layer as a feature value corresponding to each sub-process step.
4. The method according to claim 1, wherein: The first sub-model generates a feature map corresponding to each sub-process step based on the at least one feature value.
5. The method according to claim 1, in, The second sub-model includes a first fully connected layer configured to output the estimated value, The measurements include a first measurement value, the first measurement value being a result of performing each measurement step, and The steps to train a machine learning model include: Training is performed to reduce a first loss obtained by comparing the estimated value and the first measured value.
6. The method according to claim 1, wherein: The output values include: a first output value for predicting the at least one characteristic of the semiconductor device; and The second output value is used to predict the change amount of at least one characteristic of the semiconductor device according to the change of at least one physical characteristic of the semiconductor process, wherein The second sub-model includes a second fully connected layer configured to output the output value.
7. The method according to claim 6, wherein: The measured values include a second measured value obtained by actually measuring the at least one characteristic of the semiconductor device, and The step of training the machine learning model includes: training to reduce a second loss obtained by comparing the first output value and the second measurement value.
8. The method according to claim 6, wherein: The measured values include a third measured value obtained by actually measuring a change in the at least one characteristic of the semiconductor device, and The step of training the machine learning model includes: training to reduce a third loss obtained by comparing the second output value and the third measurement value.
9. The method according to claim 6, wherein: The at least one physical property represents a doping concentration, and The variation amount of the characteristic indicates the variation amount of at least one electrical characteristic of the semiconductor device.
10. A method for modeling a semiconductor process, the method comprising: receiving input data defining sub-process steps and measurement steps; Based on the measurement steps, the sub-process steps are grouped to correspond to a plurality of modules; calculating at least one characteristic value corresponding to each sub-process step; outputting an estimation value corresponding to each module of the plurality of modules based on the at least one feature value by using the sub-model; outputting an output value representing at least one characteristic of the semiconductor device based on the at least one feature value by using the sub-model; as well as A sub-model is trained based on at least one of the estimated value and the output value.
11. The method according to claim 10, wherein: The step of calculating the at least one eigenvalue includes performing a convolution calculation on input data corresponding to each of the plurality of modules.
12. The method according to claim 10, wherein: The step of training the sub-model includes calculating a first loss obtained by comparing the estimated value with an actual measurement value of each measurement step, and training the sub-model to reduce the first loss.
13. The method according to claim 10, wherein: The output values include: a first output value for predicting the at least one characteristic of the semiconductor device; and The second output value is used to predict the amount of change in the characteristics of the semiconductor device according to the change in at least one physical characteristic.
14. The method according to claim 13, wherein: The step of training the sub-model includes comparing the second output value with an actually measured variation of the at least one characteristic of the semiconductor device to calculate a third loss, and training the sub-model to reduce the third loss.
15. A system for modeling a semiconductor process, the system comprising: at least one processor configured to execute machine-readable instructions that, when executed by the at least one processor, cause the system to: receive input data defining sub-process steps and measurement steps, and provide a machine learning model configured to predict at least one characteristic of a semiconductor device based on the input data, Among them, the machine learning model is configured as: Calculate the characteristic values corresponding to each sub-process step, and Based on the feature value, a first output value predicting at least one characteristic of the semiconductor device and a second output value predicting a change amount of the at least one characteristic of the semiconductor device according to a change in at least one physical characteristic are output.
16. The system of claim 15, wherein: The at least one processor is configured to: Based on the measurement steps, the sub-process steps are grouped into corresponding modules, and An estimate corresponding to each module of the plurality of modules is calculated.
17. The system of claim 15, wherein: The machine learning model is configured to perform at least one convolution calculation on the data of each sub-process step.
18. The system of claim 16, wherein: The at least one processor is configured to: receiving a first measurement value based on a result of executing the measuring step, and The machine learning model is trained so that a first loss obtained by comparing the estimated value with the first measured value is reduced.
19. The system of claim 15, wherein: The at least one processor is configured to: receiving a third measurement value based on an amount of change in a characteristic of the semiconductor device actually measured, and The machine learning model is trained so that a third loss obtained by comparing the second output value with the third measurement value is reduced.
20. The system of claim 15, wherein: The at least one physical property is indicative of a doping concentration, and The variation indicates a variation of at least one electrical characteristic of the semiconductor device.
Citation Information
Patent Citations
Composition for Encapsulating Semiconductor Device and Semiconductor Device Encapsulated Using the Same
KR1020230164858A