Simulation Method, Electronic Device, and Storage Medium for Semiconductor Devices
Through the neural network model training method, training samples are generated and parallel prediction is performed, which solves the problem of slow calculation speed of CMP model, realizes faster modeling and simulation, and improves accuracy.
Patent Information
- Application Number
- CN202510267351.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Existing chemical mechanical polishing (CMP) models are slow to calculate when predicting metal layer thickness and cannot effectively use graphics processors or central processors for parallel acceleration, resulting in the long-term modeling and simulation processes.
The neural network model training method is adopted to generate training samples and update parameters using the stochastic gradient descent algorithm, and parallel prediction is performed in combination with the simulation results of the CMP physical model, replacing the traditional time-step physical model.
It significantly accelerates the establishment and simulation process of CMP physical model, reduces time consumption, and improves simulation accuracy.
Smart Images

Figure CN119782826B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure mainly relate to integrated circuits, and more particularly, to a simulation method, an electronic device, and a storage medium for semiconductor devices. Background Art
[0002] Currently, the prediction of the metal thickness after the Chemical Mechanical Polishing (CMP) back-end process usually highly depends on the CMP model.
[0003] The underlying traditional CMP model is mainly a time-step physical model based on the physical process of CMP itself. The characteristics of this model are that it fully simulates the real physical process of CMP, has a relatively high modeling accuracy, and can output all relevant variables in the intermediate process, helping users better understand and adjust the CMP process recipe and predict the results after CMP grinding. However, the disadvantage is that the calculation speed is relatively slow. Summary of the Invention
[0004] According to an exemplary embodiment of the present disclosure, a simulation solution for semiconductor devices and a training method for a neural network model are provided to at least partially overcome the above or other potential defects.
[0005] According to one aspect of the present disclosure, a method for generating training samples is provided. The method includes: generating a first number of feature vectors based on data points representing various graphic features on a wafer, where the feature vectors are combinations of data representing different types of graphic features in the data points; generating a second number of physical vectors based on physical information related to the CMP of the wafer, where the physical vectors are combinations of data representing different types of physical information; and performing a second number of simulations by a CMP physical model based on the feature vectors and the physical vectors to generate simulation values corresponding to the first number of feature vectors of the surface height of the wafer in each simulation as training samples.
[0006] In some embodiments, generating a first number of feature vectors based on data points representing various graphic features on a wafer includes: determining data within each step range based on the value range and corresponding step size of the data points of each graphic feature; and orthogonalizing the data within each step range to generate a feature data table, where the feature data table includes data of the first number of feature vectors.
[0007] In some embodiments, generating a second quantity of physical vectors based on physical information related to CMP of a wafer includes: determining data for respective step ranges based on value ranges and corresponding step sizes of data points of respective physical vectors; and orthogonalizing the data within respective step ranges to generate a physical information data table, where the physical information data table includes data of the second quantity of physical vectors.
[0008] In some embodiments, the CMP physical model performing a second quantity of simulations based on eigenvectors and physical vectors includes: respectively generating a first simulation value of the surface height of a non-grooved structure and a second simulation value of the surface height of a grooved structure; and determining a simulation value of the surface height of the wafer based on the first simulation value and the second simulation value.
[0009] In some embodiments, the graphical features at least include: the density of the patterns within the lattice points on the wafer; the width of the patterns; and the spacing between the patterns.
[0010] In some embodiments, the physical information at least includes the following items: information on the deposited material; maximum dishing defect information and minimum dishing defect information for each step; slurry information; long-tail effect parameter information; duration information; nominal pressure information.
[0011] According to a second aspect of the present disclosure, there is provided a method for training a neural network model. The method includes: training a neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated by a CMP physical model based on graphical features and physical information for determining the surface height of the wafer after polishing; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the respective updated parameters as the trained neural network model, where the trained neural network model outputs a predicted value of the surface height.
[0012] In a third aspect of the present disclosure, there is provided an electronic device. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the device to perform operations, the operations including: training a neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated by a CMP physical model based on graphical features and physical information for determining the surface height of the wafer after polishing; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the respective updated parameters as the trained neural network model, where the trained neural network model outputs a predicted value of the surface height.
[0013] In some embodiments, training a neural network model using training samples to determine the gradient of the loss function of the neural network model includes: calculating the gradient of the loss function using a stochastic gradient descent algorithm.
[0014] In some embodiments, updating the parameters of the neural network model based on the gradient of the loss function includes: in response to a predetermined number of trainings being performed, performing a gradient update on the gradient of the loss function; and updating the parameters of the neural network model based on the gradient update.
[0015] In some embodiments, training a neural network model using training samples to determine the gradient of the loss function of the neural network model includes: training the neural network model using a part of the training samples to perform gradient descent backpropagation to update the parameters; using the trained neural network model to predict the remaining part of the training samples to generate a predicted value of the surface height of the polished wafer; and in response to the comparison result between the predicted value and the simulated value of the CMP physical model satisfying a predetermined condition, determining the corresponding trained neural network model as the final neural network model.
[0016] In some embodiments, the comparison result between the predicted value and the simulated value satisfying a predetermined condition includes: the root mean square error between the predicted value and the simulated value being less than a predetermined threshold.
[0017] In some embodiments, the method further includes: determining the loss function value of the neural network model; and wherein, the training satisfying a predetermined condition includes: the difference between the loss function values of two consecutive iterations being less than a predetermined difference threshold; or the absolute value of the ratio of the loss function values of two consecutive iterations being greater than a predetermined ratio threshold. In some embodiments, the training samples are generated in parallel on multiple servers so that the training of the neural network model is performed in parallel.
[0018] In a fourth aspect of the present disclosure, there is provided a neural network model generated according to the method of the third aspect.
[0019] In a fifth aspect of the present disclosure, there is provided a simulation model, which includes: a CMP physical model configured to determine a simulated value of the surface height of a polished wafer based on input data, where the input data includes graphic feature data and physical information; and a neural network model according to the fifth aspect, the neural network model being trained using the simulated value as a training sample and configured to perform parallel prediction on multiple sets of input data during the simulation process of the CMP physical model to generate multiple predicted values of the surface height of the polished wafer corresponding to the multiple sets of input data, respectively serving as the corresponding simulated values of the surface height.
[0020] In some embodiments, parallel prediction of multiple sets of input data during the simulation of the CMP physical model includes: using multiple graphics processors to perform parallel acceleration on the calculation process of the neural network model.
[0021] In a sixth aspect of the present disclosure, there is provided a simulation method for a semiconductor device, which is executed using the simulation model of the fifth aspect. The method includes: predicting a simulated value of the surface height of the polished wafer by the neural network model in the simulation model based on input data, where the input data includes graphic feature data and physical information.
[0022] In some embodiments, predicting a simulated value of the surface height of the polished wafer by the neural network model in the simulation model based on input data includes: the neural network model performing parallel prediction on multiple sets of input data to generate multiple predicted values of the surface height of the polished wafer corresponding to the multiple sets of input data, respectively serving as the corresponding simulated values of the surface height.
[0023] In some embodiments, predicting a simulated value of the surface height of the polished wafer by the neural network model in the simulation model based on input data includes: the neural network model respectively generating a first predicted value of the surface height of the non-groove structure and a second predicted value of the surface height of the groove structure; and the neural network model determining the simulated value of the wafer surface height based on the first predicted value and the second predicted value.
[0024] In some embodiments, it further includes: determining the condition of the surface of the wafer during the polishing process through the CMP physical model in the simulation model.
[0025] In a seventh aspect of the present disclosure, there is provided a neural network model, including: an input layer configured to receive graphic feature data on the wafer and physical information related to CMP of the wafer; an intermediate layer configured to execute the method of the sixth aspect to determine the surface height value of the polished wafer; and an output layer configured to output the surface height value.
[0026] In some embodiments, the method further includes: an embedding layer configured to map the input physical information to generate a physical vector.
[0027] In an eighth aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above method according to the present disclosure.
[0028] It will be understood from the following description that the technical solution of the present disclosure can significantly accelerate the establishment and simulation process of the CMP physical model, reduce the related time consumption, and improve the simulation accuracy.
[0029] The Summary of the Invention section is provided to introduce, in simplified form, a selection of concepts that will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify key or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;
[0031] Figure 2 A flowchart showing a method for generating training samples according to some embodiments of the present disclosure;
[0032] Figure 3 A flowchart showing a method for training a neural network model according to some embodiments of the present disclosure;
[0033] Figure 4 A schematic diagram showing a neural network architecture according to some embodiments of the present disclosure;
[0034] Figure 5 A block diagram showing a computing device capable of implementing multiple embodiments of the present disclosure.
[0035] In the various drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The principles of the present disclosure will now be described with reference to various exemplary embodiments shown in the drawings. It should be understood that the description of these embodiments is only for enabling those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that, where feasible, like or identical reference numerals may be used in the figures, and like or identical reference numerals may represent like or identical functions. Those skilled in the art will readily recognize that alternative embodiments of the structures and methods described herein may be employed without departing from the principles of the invention described herein.
[0037] As used herein, the term "comprising" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless otherwise specified, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0038] CMP is a method that combines chemical and physical methods to remove materials. It has become an indispensable step in the modern integrated circuit (IC) industry. Specifically, CMP uses chemical corrosion and mechanical force to flatten silicon wafers or other substrate materials during processing. CMP can achieve nanometer-level material removal, making the wafer surface flat.
[0039] The CMP model, as the name implies, is a physical model for simulating the CMP process. The model outputs some important physical indicators of each grid point of the wafer after the CMP process, such as the thickness of the trench structure, the thickness of the non-trench structure, and the thickness of the metal material after CMP. In simple terms, it is a physical model that predicts the surface morphology of the wafer after the CMP process. CMP is one of the eight major semiconductor processes, connecting the previous and the next. Other processes following CMP will be affected by the CMP process. In other words, when it is necessary to evaluate and control the process results after CMP, it is best to know the impact of the CMP process, which is of great significance to the full process control and yield improvement of the entire semiconductor manufacturing. In addition, the CMP process itself is an indispensable and necessary step, because the CMP process can be better adjusted through the CMP model. For example, the results of the CMP model can be used to infer which process parameter settings need to be modified and adjusted, and whether there are some unreasonable aspects in the design of the Graphic Data System (GDS), which may cause hotspots, that is, defects.
[0040] As mentioned above, the traditional CMP model has high modeling accuracy, but the disadvantage is that the calculation speed is slow. Because the time-step model is calculated step by step, the result of each time step depends on the result of the previous time step, and it cannot be accelerated in parallel using a graphics processing unit (GPU) or CPU. In the process of searching for the optimal solution for parameters, the running time of each attempt cannot be reduced. Therefore, building a better CMP model requires repeated adjustments and runs, which consumes a lot of time and energy. In addition, in the full chip simulation stage, if a larger chip is encountered, the amount of data increases significantly, and the simulation of the CMP physical model based on the time step will be very slow, usually taking several hours.
[0041] A known CMP time-step physical model is composed of a mixture of several physical and mathematical formulas. The core feature of the model is that the polishing time for each step needs to be set according to the actual CMP process, assumed to be t1, t2, t3, …, and the unit time dt is set, where dt is usually a relatively small number, and the smaller dt is, the closer the CMP physical model is to the real CMP physical process. Because the core assumption of this CMP physical model is that at an extremely short moment dt, the polishing rate of the platen on the CMP material remains nearly constant. When dt approaches 0 infinitely, the polishing rate approaches a fixed value at this moment (the idea of differentiation). Therefore, if a relatively accurate physical model needs to be established, usually dt needs to be set very small, but too small dt will cause a sharp increase in the amount of calculation. Suppose there are several time steps (step) in total, and the duration of each step is t1, t2, …, then the total time is their sum, assumed to be T. Then, to perform a complete CMP calculation, the results of T / dt time steps need to be calculated. Usually, T will be hundreds of seconds (such is the real physical process). If dt is set to 0.001 s, then hundreds of thousands of time steps need to be calculated. And when calculating each time step, usually not just one point participates in the operation. During the modeling stage, a DOE experiment needs to be carried out. The DOE experiment algorithm generates Q test points for R iterations. The larger the values of Q and R are set, the closer the finally found model parameters are to the global optimal solution. During the simulation stage, it depends on the scale of the entire chip. When it is relatively large, there will be tens of millions of points that need to be simulated. And a well-established model usually needs to be used several times, and there may be many entire chips that need to be simulated using this model. All of the above illustrate the computational complexity of the current CMP physical model. Therefore, it is very necessary to perform acceleration.
[0042] In view of this, the present disclosure provides an improved solution.
[0043] Some embodiments of the present disclosure provide an improved method for training a neural network model. The method includes: training a neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated by a CMP physical model based on graphic features and physical information for determining the surface height of the polished wafer; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the updated parameters as the trained neural network model, where the trained neural network model outputs a predicted value of the surface height.
[0044] Some embodiments of the present disclosure also provide an improved simulation model and a neural network model.
[0045] Embodiments of the present disclosure can significantly accelerate the establishment and simulation process of the CMP physical model, reduce the associated time consumption, and improve the simulation accuracy.
[0046] Embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings.
[0047] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As Figure 1 shown, the example environment 100 includes a computing device 110 and a client 120.
[0048] In some embodiments, the computing device 110 can interact with the client 120. For example, the computing device 110 can receive an input message from the client 120 and output a feedback message to the client 120. In some embodiments, the input message from the client 120 can be design layout data. The computing device 110 can perform corresponding mathematical operations on the design layout data and output the corresponding operation results to the client 120.
[0049] In some embodiments, the computing device 110 can include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant PDA, a media player, etc.), a consumer electronic product, a small computer, a large computer, cloud computing resources, etc.
[0050] It should be understood that describing the structure and function of the example environment 100 only for exemplary purposes is not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. This environment is merely illustrative and not used to limit the application environment of the embodiments of the present disclosure.
[0051] To more clearly explain the principle of the solution of the present disclosure, the following will be described in more detail with reference to Figures 2 - 5 .
[0052] First, the CMP physical model is introduced. The essence of the CMP physical model is composed of several physical formulas mixed together, rather than a black box that cannot be described by mathematical language. As is well known, deep learning models have extremely strong fitting capabilities and can even fit black box functional relationships that cannot be described by explicit mathematical formulas at all. Therefore, based on deep learning, a suitable deep learning model can be established to learn and approximate the CMP physical model.
[0053] Next, specifically elaborate on how to construct this deep learning model. The CMP physical model usually consists of three steps because the actual CMP process can often be divided into three stages, namely: the bulk silicon process (bulk) stage, in which the topmost metal (usually copper) is mainly polished; the touch down (tchd) stage, which is defined as the exposure of the buffer layer between the metal layer and the bottommost oxide; and the barrier remove (br) stage, which is defined as the complete polishing of the buffer layer within this time period, with most points touching the bottommost oxide material. There are different physical parameters in each stage.
[0054] The bulk silicon process generally refers to the period from the start of polishing to the end of the bulk silicon process stage, during which only the topmost metal material is polished, and only metal is polished in this stage. Before polishing, a deposition process is carried out in CMP, depositing different materials. Usually, the topmost layer is metal (generally copper), the middle layer is the buffer layer, and the bottommost layer is the oxide. What is touched here is the buffer layer material, that is, in this stage, it involves the transition from only polishing metal to starting to touch the buffer layer material. Regarding the buffer layer removal, as the name implies, in this layer, mainly the buffer layer is removed. Ideally, the bulk silicon process only polishes metal completely. The tchd stage ends exactly when the buffer layer is touched. The br stage can completely polish the buffer layer material. However, in reality, it is not so ideal. In the bulk silicon process stage, there may also be some areas where the material has changed from metal to the buffer layer. The same is true for tchd. It is impossible to control all points to touch the buffer layer material at the same moment, and this is even more so for the br stage. It is very difficult to ensure that each point perfectly polishes the buffer layer without touching the lower oxide layer.
[0055] In some embodiments, based on the above-mentioned real CMP process information, for the data input into the deep learning neural network, the following information may be included: vectors of Graphic Data System (GDS) features, specifically density vectors, space vectors, width vectors, etc.; physical information vectors. The physical information vectors may specifically include deposition material information, which may specifically include deposition material thickness information, nominal polishing rate information of the deposition material, maximum dishing information of each step, minimum dishing information, slurry information, long range effect parameter information, duration information, nominal pressure information, and so on, parameters used in the CMP physical model. Here, Dishing refers to the difference in height between the trench structure and the non-trench structure, and Erosion refers to the difference in height of the non-trench structure of each grid point relative to the reference value thickness. The long range effect can be understood as that the current density, width, space, etc. conditions in a region not only affect the polishing result of the current region, but also affect the polishing results of its surrounding regions.
[0056] The input of the neural network model can be determined based on the above information. Further details about the Figure 4 architecture form of the neural network model will be introduced later.
[0057] For the neural network model, it needs to be trained with training samples. The method for generating training samples is described below.
[0058] First, refer to Figure 2 , Figure 2 which shows a flowchart of a method 200 for generating training samples according to some embodiments of the present disclosure.
[0059] At block 202, a first number of feature vectors are generated based on the data points representing various graphic features on the wafer. The feature vectors are combinations of data representing different types of graphic features in the data points.
[0060] In some embodiments, generating a first number of feature vectors based on data points representing various graphical features on a wafer may include: determining data within each step range respectively based on the value range of the data points of each graphical feature and the corresponding step size; and orthogonalizing the data within each step range to generate a feature data table, where the feature data table includes data of the first number of feature vectors. In some embodiments, the graphical features at least include: the density of the patterns within the grid points on the wafer; the width of the patterns; and the space between the patterns. For example, the graphical features density, width, and space have a significant impact on the height after CMP polishing. They are significant independent variables constituting the CMP polishing height prediction model, and each of these three variables has its own certain range. In addition, the step size can be divided within the range according to actual needs. This will be described in detail later.
[0061] The data points are used to represent the data of the measurement points (grid points) on the wafer. The data points can be provided by the user. For example, a data point represents features such as density, width, and space, for example, represented as (density, width, space). Here, for example, density is 0.5, width is equal to 30 nm, and space is equal to 40 nm. Among them, density, width, and space each represent a graphical feature, that is, density, width, and space. The combination of the data used to represent these features can form a vector. For example, for density, width, and space, a combination of their respective values can be called a vector. For example, the previously mentioned (0.5, 30 nm, 40 nm). Combinations of different values of these variables can form multiple vectors. They can be combined in various ways.
[0062] One way is the orthogonal design verification method (Design of Experiments, abbreviated as DOE). The DOE design method, that is, DOE data orthogonalization, is an experimental design method used to explore and verify the influence of factors on the results. In DOE, usually the experiment is divided into multiple combinations, and each combination controls one factor and measures its influence on the results. In this way, a more comprehensive understanding of the influence of factors on the results can be obtained, and the optimal factor combination can be determined.
[0063] In some embodiments, the DOE design can be divided into two parts. One part is the GDS feature design, that is, the design of density / width / space mentioned above.
[0064] The ranges and step sizes of density / width / space can be set first. Since these three variables all have physical meanings, there is usually a relatively definite range in most processes. The common ranges of these three variables can be statistically analyzed. For example, it is known that density is a number between 0 and 1, width and space are all positive numbers, usually a number between 0 and 50 (unit: um), and usually the maximum value does not exceed hundreds. Thus, data can be actually generated according to the actual situation (such as training resources, whether there are a large number of GPUs to accelerate training or not).
[0065] Suppose there are k data points, and the densities are 0.1, 0.2, 0.3, ……, d_k respectively. Then the density vector is [0.1, 0.2, 0.3, …, d_k]. Similarly, the other vectors can be obtained.
[0066] The DOE mentioned above is to make the neural network training sufficient so as to more accurately track the CMP physical model. Specifically, orthogonal DOE data generation can be performed on density / width / space. The so-called orthogonality means that, assuming that density generates N possible value results, width generates M possible value results, and space generates P possible value results, then the number of data entries in the final data set is data entries. The characteristics of the data here are: for example, randomly checking the generated data set, there are several data entries with a certain value of width. Specifically, how many there are, of course, is data entries. In fact, it is an exhaustive combination, ensuring that each width has all possible density and space value results. By making the data orthogonal, the comprehensiveness of the data and the reliability of the results can be ensured.
[0067] At block 204, a second quantity of physical vectors is generated based on physical information related to CMP of the wafer. The physical vectors are combinations of data representing different types of physical information.
[0068] In some embodiments, the physical information at least includes the following items: information about the deposited material; maximum dishing defect information and minimum dishing defect information for each step; slurry information; long-tail effect parameter information; duration information; nominal pressure information.
[0069] In some embodiments, generating a second quantity of physical vectors based on physical information related to CMP of the wafer may include: determining data for each step range respectively based on the value range and corresponding step size of the data points of each physical vector; and performing data orthogonality on the data within each step range to generate a physical information data table, where the physical information data table includes data of a second quantity of physical vectors.
[0070] In some embodiments, the physical information can be expanded into vectors with lengths equal to those of density, space, and width. In other words, density, space, and width have equal lengths, and this physical information can be expanded into vectors with lengths respectively equal to those of density, space, and width. These vectors can be combined to form a matrix, which is the input and denoted as X. This X can be used as the input to the neural network model.
[0071] Assume that the length of the generated GDS feature data table is k. Since the physical parameters are also input variables, they are also designed by DOE, and a combined table of DOE is generated according to the actual process conditions. For example, the pressure can be selected from several cases such as low pressure, medium pressure, and high pressure. Assume that there are a total of L combined tables of DOE for the physical parameters, then the CMP physical simulation model needs to be run L times in total.
[0072] At block 206, the CMP physical model performs a second number of simulations based on the feature vectors and physical vectors to generate, in each simulation, a simulated value of the surface height of the wafer corresponding to the first number of feature vectors as training samples.
[0073] In some embodiments, the corresponding number of simulations is performed based on the number of physical vectors.
[0074] In some embodiments, each time the designed GDS feature data and one of the L physical model combined tables are input, a result table with a length of L×k (k is the number of data points) will be generated. Since these multiple tables are independent of each other, they can be generated in parallel, greatly accelerating the generation speed. Specifically, assume that there are several features participating in the orthogonalization, and an orthogonal table is obtained after the orthogonalization. Each row of data in the orthogonal table contains the GDS features (density / space / width) and process parameters, which are input into the physical model for calculation to obtain the calculation results, that is, the corresponding SThickNT and SThickT. Repeat this several times until a corresponding SThickNT and SThickT are generated for each row of data.
[0075] In some embodiments, for the convenience of operation and processing in a neural network, data can be merged (or spliced) through an embedding operation, or a mapping operation. The specific method is as follows: Assume that the process parameter a takes a value of 0.5, then the vector of the process parameter a is [0.5, 0.5, …] (actually replicated). A vector with a length equal to density, and the process vector b is similar, so each process vector can be obtained. Similar to the previous operation of splicing density and width, these process vectors can also be spliced to the back, thus forming a final input rectangle X. Assume that density, space, and width are used as GDS features, and the process parameters a, b, and c are used as process features, and there are k data points in total. Then the final input matrix is a matrix with k rows and 6 columns. Since there are 3 GDS features and 3 process parameters, there are a total of 6 columns. Splicing the data is more convenient in engineering implementation. In addition, since the input matrix will perform matrix multiplication with the parameter matrix of the neural network, it is easy to incorporate the process information into the input, with high efficiency.
[0076] In some embodiments, the second number of simulations performed by the CMP physical model based on the feature vector and the physical vector may include: generating a first simulation value of the surface height of the non-groove structure and a second simulation value of the surface height of the groove structure respectively; and determining a simulation value of the surface height of the wafer based on the first simulation value and the second simulation value.
[0077] It should be understood that the embodiments of the present disclosure are not limited thereto. For example, in some embodiments, the vectors may not be spliced, but may be input into the neural network model as one-dimensional vectors.
[0078] It should be noted that the table after DOE orthogonalization only has inputs (inputs to the neural network), that is, there are no columns related to S. The columns related to S refer to the labeled data columns necessary for training the neural network. Therefore, the DOE orthogonal table needs to be input into the physical model for simulation first to obtain the simulation results of the physical model. Then the simulation results of the physical model become the labeled data columns. The function of the DOE orthogonal table (after generating the labeled data columns through physical model simulation) is to supply the neural network for training and to set aside a part as a validation data set to verify whether the training is in place.
[0079] In some embodiments, the result table may be: SThickNT / SThickT / SDishing / SErosion, where SThickNT and SThickT are output by the neural network, and the difference between the two is dishing; the reference surface height (a certain constant) minus SThickNT is SErosion.
[0080] In addition, parallel acceleration was mentioned earlier. In this case, there can be two scenarios for parallel acceleration.
[0081] One scenario is that there is a trained neural network, which replaces the physical model because the neural network can achieve parallel acceleration. For example, there are k data points, so the vector length is k. Suppose the Genetic Algorithm (GA) needs to iterate 100 times and the population size is 1000. This means that for each iteration, simulation results of k data points need to be generated, and it needs to be generated 1000 times. So, 1000 results of length k need to be generated in each round. When generating these 1000 results of length k in each round, since these 1000 populations are independent of each other, the GPU can be fully used for parallel acceleration. However, due to the computational complexity of the CMP physical model, it is difficult to deploy it on the GPU, and at most, ordinary multi-threading in the computer can be used for implementation. For example, 20 threads are used for acceleration, that is, these 20 threads are used to calculate these 1000 results of length k (equivalent to each thread completing 50, and everyone completes it together). However, the neural network is naturally suitable for the GPU and is very easy to implement. The acceleration ability of the GPU is not comparable to that of ordinary multi-processes. The GPU can implement tens of thousands of threads or even more, while ordinary multi-threading depends on the capabilities and scale of the CPU, but the CPU itself cannot have many threads. In addition, when the neural network generates the results of length k, since the neural network obtains the results in one step, it can independently predict the k data points, but the CMP physical model needs to consider environmental information and cannot operate independently. The neural network model can directly obtain the calculation results in full parallel through parallel deployment (on the GPU).
[0082] Another scenario is when generating samples for training the neural network. At this time, since the samples are independent of each other, many machines can be used to generate the results and then collect them. For example, if 10,000 samples need to be generated for training, one way is to generate these 10,000 samples one by one, and another way is to start 100 servers, each server generates partial results, and then collect them (a total of 10,000) and then conduct training.
[0083] The following is described in combination with Figure 4 for description. Figure 4A schematic diagram of a deep learning model according to some embodiments of the present disclosure is shown, specifically, a schematic diagram of a neural network architecture 400. The architecture includes: an input layer 402; the input layer includes a physical parameter information table and GDS feature vectors (which can be called graphic vectors), including a Width vector, a Space vector, and a Density vector. An embedding layer 404; usually, physical information parameters are all scalars, and through embedding, they are mapped into vectors of the same length as density / width / space, and then these vectors can be all concatenated to form an input matrix X. A first convolutional layer 406, connected to the embedding layer 404; a pooling layer 408, connected to the first convolutional layer 406; a second convolutional layer 410, connected after the pooling layer; a first fully connected layer 412, connected after the second convolutional layer 410; a first regularization layer 414, connected after the first fully connected layer 412; a second fully connected layer 418, connected after the first regularization layer 414; a second regularization layer 420, connected after the second fully connected layer 418, a third fully connected layer 422, connected after the second regularization layer 420; the output layer of the third fully connected layer 422 is connected to two fully connected layers, namely, a fourth fully connected layer 424 and a fifth fully connected layer 426. Each of the last two fully connected layers is connected to an output layer, namely, an SThickNT output layer 428 and an SThickT output layer 430. Among them, the SThickNT output layer 428 outputs a simulated value of the surface height of the non-groove structure, and the SThickT output layer 430 outputs a simulated value of the surface height of the groove structure. The simulated value of the surface height of the wafer can be determined based on the simulated value of the surface height of the non-groove structure and the simulated value of the surface height of the groove structure.
[0084] In some embodiments, the simulated value H1 of the height of the non-groove structure on the wafer can be determined; the simulated value H2 of the height of the groove structure on the wafer can be determined; and the simulated value of the surface height of the circle can be determined based on the simulated value H1, the simulated value H2, and the density of the graphics in the lattice region on the wafer. Specifically, the surface height is equal to the height of the non-groove (Nontrench) structure multiplied by (1 - density) plus the product of the height of the groove (Trench) structure and density. Expressed by the formula as: ; where density represents the density of the graphics in the lattice region on the wafer, and its range is between 0 and 1.
[0085] As Figure 4As shown, the architecture also includes a Relu activation function 416. The Relu activation function, namely the Rectified Linear Unit (ReLU) activation function, is one of the most commonly used activation functions in convolutional neural networks (CNNs) and many deep learning models. Its main role is to introduce non-linearity, enabling the model to learn and represent more complex features.
[0086] As mentioned before, multiple result tables of length k are generated. In addition, after determining the training samples and constructing the neural network model, the neural network model can be trained. The following will Figure 3 specifically describe the method for training the neural network model. Figure 3 FIG. shows a flowchart of a training method 300 for a neural network model according to some embodiments of the present disclosure.
[0087] At block 302, the neural network model is trained using the training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated by the CMP physical model based on the graphical features and physical information for determining the surface height of the polished wafer.
[0088] In some embodiments, the training samples can be pre-generated, for example, stored in a predetermined database and available for the neural network model to train.
[0089] In some embodiments, the training samples can be generated during the training process of the neural network model.
[0090] In some embodiments, training the neural network model using the training samples to determine the gradient of the loss function of the neural network model includes: calculating the gradient of the loss function using the stochastic gradient descent algorithm.
[0091] In some embodiments, updating the parameters of the neural network model based on the gradient of the loss function includes: performing gradient update on the gradient of the loss function when a predetermined number of trainings (e.g., 100 times) are performed; and updating the parameters of the neural network model based on the gradient update.
[0092] In some embodiments, training a neural network model using training samples to determine the gradient of the loss function of the neural network model may include: training the neural network model using a portion of the training samples to perform gradient descent backpropagation to update parameters; using the trained neural network model to predict the remaining portion of the training samples to generate a predicted value of the surface height of the polished wafer; in the case where the comparison result between the predicted value and the simulation value of the CMP physical model meets a predetermined condition, the corresponding trained neural network model may be determined as the final neural network model. That is, in this way, the training degree of the neural network can be determined. For example, assume that DOE orthogonally generates 10 million data, and the simulation results of these 10 million data are simulated using a physical model. Then take 70% (it can also be 80% or 90%) of them as training samples for training. After each of the 7 million training samples is input into the neural network for gradient descent backpropagation to update the parameters, the training is completed. Next, use the trained neural network to predict the remaining 30% of the data, and then calculate the rmse between the neural network predicted value and the physical model simulation value. If this rmse meets a certain standard, it can be considered that the training is in place. If not, DOE can be made more detailed to generate more data and continue training, repeating the above steps.
[0093] For a neural network, in many cases, "input" and "training sample" are essentially the same. A training sample refers to some data that allows the neural network to calculate, and then the neural network updates its parameters based on the calculation error through backpropagation. Compared with the training sample, the dataset to be predicted (input) simply lacks the true corresponding values and thus cannot be used by the neural network to update and adjust its own parameters. But essentially, they are all the matrix X mentioned above. In addition, the concept of "input" is broader than that of "training sample". If an input has feedback, it is a training sample because it can be used to adjust the neural network parameters, which is also "training". If an input has no feedback, then it is a dataset to be predicted because it is impossible to know whether the prediction result is good or bad, or whether the result of the neural network is accurate enough. Thus, it cannot be used for training. Generally speaking, an input describes a matrix or vector that conforms to the regulations of the input layer of the neural network, while a training sample refers to a set of data points, each of which has a standard reference value that can be used to evaluate the quality of the neural network prediction. The non-standard answer part of this dataset can be extracted as an input and input into the neural network. The neural network gives the corresponding output, and then by comparing the standard answer and the output, the parameters are updated through backpropagation. Therefore, an input can also be regarded as a part of the training sample.
[0094] As mentioned above, after obtaining the physical information, a matrix can be generated based on the physical information and the graphical feature information as the input of the neural network model.
[0095] In some embodiments, an orthogonal array is generated based on DOE, and simulations are performed using a CMP physical model to obtain training samples. The training samples include the information that needs to be input into the neural network (features of the GDS and results of process parameters), as well as the simulation results given by the physical model. As mentioned above, the GDS features and process parameter columns can be extracted as the input X matrix and input into the neural network. The neural network will obtain a calculation result, and then the loss function is calculated and backpropagated by comparing the neural network result and the simulation result of the physical model to update the network parameters. Finally, a trained neural network is obtained. For example, there is an existing data point that includes GDS features and process information. This point can be input into the neural network and the physical model, and the output result of the neural network will be very close to the output result of the physical model. Therefore, the training is effective, and a neural network that has learned the physical model is obtained.
[0096] Return to Figure 2 Continuing the description, in some embodiments, training a neural network model using training samples to determine the gradient of the loss function of the neural network model includes: the gradient of the loss function can be calculated using the stochastic gradient descent algorithm.
[0097] In some embodiments, the gradient of the loss function can be calculated using the stochastic gradient descent algorithm.
[0098] In some embodiments, training a neural network model using training samples to determine the gradient of the loss function of the neural network model may include: training the neural network model using a part of the training samples to perform gradient descent backpropagation to update the parameters; using the trained neural network model to predict the remaining part of the training samples to generate a predicted value of the surface height of the polished wafer.
[0099] In fact, the essence of neural network training is to first perform simulation calculations using the neural network, then compare the labeled data, calculate the loss of the objective function, perform gradient descent, and backpropagate to update the parameters.
[0100] At block 304, the parameters of the neural network model can be updated based on the gradient of the loss function.
[0101] In some embodiments, updating the parameters of the neural network model based on the gradient of the loss function may include: performing gradient updates on the gradient of the loss function after a predetermined number of training times, such as 100 times; and updating the parameters of the neural network model based on the gradient update. It should be understood that the number of training times of prediction can vary according to actual needs, and the embodiments of the present disclosure do not limit this.
[0102] In some embodiments, the stochastic gradient descent algorithm is used to calculate the gradient and continuously update the neural network parameters.
[0103] At block 306, in response to the training satisfying a predetermined condition, a neural network model corresponding to the respective updated parameters is determined as a trained neural network model, where the trained neural network model outputs a predicted value of the surface height.
[0104] In some embodiments, the method further includes: determining a loss function value of the neural network model; and where the training satisfying the predetermined condition includes: the difference between the loss function values of two consecutive iterations is less than a predetermined difference threshold; or the absolute value of the ratio of the loss function values of two consecutive iterations is greater than a predetermined ratio threshold.
[0105] In some embodiments, the random gradient descent algorithm is used to calculate the gradients and continuously update the neural network parameters until the loss curve approaches a steady state, at which point the training can be stopped.
[0106] In some embodiments, as mentioned previously, training the neural network model using training samples to determine the gradients of the loss function of the neural network model may include: training the neural network model using a portion of the training samples for gradient descent backpropagation to update the parameters; using the trained neural network model to predict the remaining portion of the training samples to generate a predicted value of the surface height of the polished wafer. In the case where the comparison result between the predicted value and the simulation value of the CMP physical model satisfies a predetermined condition, the corresponding trained neural network model can be determined as the final neural network model. Where the comparison result between the predicted value and the simulation value satisfying the predetermined condition may include: the root mean square error between the predicted value and the simulation value is less than a predetermined threshold.
[0107] In some embodiments, it may further include: determining a loss function value of the neural network model; and where the training satisfying the predetermined condition includes: the difference between the loss function values of two consecutive iterations is less than a predetermined difference threshold; or the absolute value of the ratio of the loss function values of two consecutive iterations is greater than a predetermined ratio threshold.
[0108] In some embodiments, the training samples are generated in parallel on multiple servers so that the training of the neural network model is performed in parallel.
[0109] In some embodiments, a trained neural network model can be generated according to the above neural network training method.
[0110] In some embodiments, a neural network model is further provided, including: an input layer configured to receive graphic feature data on the wafer and physical information related to CMP of the wafer; an intermediate layer configured to perform the above semiconductor device simulation method to determine the surface height value of the polished wafer; and an output layer configured to output the surface height value.
[0111] In some embodiments, the neural network model further includes: an embedding layer configured to map the input physical information to generate a physical vector.
[0112] Through the above embodiments, a trained deep learning neural network model for tracking the CMP physical model is provided, and how to utilize it will be elaborated later. In some embodiments, the neural network model can be used together with the CMP model.
[0113] In some embodiments, a simulation model is further provided, and the model includes: a CMP physical model configured to determine a simulated value of the surface height of the polished wafer based on input data, where the input data includes graphic feature data and physical information; and the neural network model as described above, which is generated by training with the simulated value as a training sample and is configured to perform parallel prediction on multiple sets of input data during the simulation process of the CMP physical model to generate multiple predicted values of the surface height of the polished wafer corresponding to the multiple sets of input data, respectively serving as the corresponding simulated values of the surface height.
[0114] In some embodiments, performing parallel prediction on multiple sets of input data during the simulation process of the CMP physical model includes: using multiple graphics processors to perform parallel acceleration on the calculation process of the neural network model.
[0115] In some embodiments, a simulation method is further provided, and the method is executed using the above simulation model. The method includes: predicting, by the neural network model in the simulation model, a simulated value of the surface height of the polished wafer based on input data, where the input data includes graphic feature data and physical information.
[0116] In some embodiments, predicting, by the neural network model in the simulation model, a simulated value of the surface height of the polished wafer based on input data may include: the neural network model performing parallel prediction on multiple sets of input data to generate multiple predicted values of the surface height of the polished wafer corresponding to the multiple sets of input data, respectively serving as the corresponding simulated values of the surface height.
[0117] In some embodiments, predicting, by the neural network model in the simulation model, a simulated value of the surface height of the polished wafer based on input data includes: the neural network model respectively generating a first predicted value of the surface height of the non-groove structure and a second predicted value of the surface height of the groove structure; and the neural network model determining a simulated value of the surface height of the wafer based on the first predicted value and the second predicted value.
[0118] In some embodiments, it further includes: determining the condition of the surface of the wafer during the grinding process through the CMP physical model in the simulation model. For users, they ultimately also need a physical model because sometimes they need to output the intermediate process to help them confirm some relevant details of the process. For example, if a manufacturer needs to change the process parameter settings of their base station, they are likely to need to know the grinding conditions at some intermediate moments to assist them in adjusting the parameters. The neural network cannot provide the intermediate process, so the physical model can meet this need.
[0119] According to the above description, it can be seen that the embodiments of the present disclosure also provide a simulation method. This method can be implemented through the above simulation model.
[0120] In some embodiments, simulation can be performed through the CMP physical model according to the DOE design principle to obtain the simulation results of SThickNT / SThickT / SDishing / SErosion of the CMP physical model.
[0121] In some embodiments, the neural network model in the simulation model predicts the simulation value of the surface height of the wafer after grinding based on the input data, where the input data includes graphic feature data and physical information.
[0122] As can be seen from the above, the current CMP physical model will have relatively large time and computing power resource consumption during training, mainly because when running a parameter combination of a certain experiment, due to the characteristics of the time step model, a very large number of calculations need to be generated. The embodiments of the present disclosure provide a neural network model that has learned this complex physical relationship. This model receives the same input as the CMP physical model but can give the output result extremely quickly. In fact, the time for the CMP physical model to perform one experiment may be dozens of seconds, while the time for this neural network model to give the output is only a few milliseconds, so the operation speed and efficiency will be greatly improved.
[0123] Similarly, for the CMP physical model, the fitting adopts the GA genetic algorithm. The GA genetic algorithm needs to set two parameters: the population size (pop size) and the number of iterations (max iter) of the experiment. When the single calculation time is relatively long, these two parameters cannot be set too large, otherwise the total time consumed will be very long. The GA algorithm is used to solve the parameters in the CMP physical model. The GA algorithm needs to conduct continuous experiments, and each experiment requires the CMP physical model to give the simulation results of many experimental individuals. The iteration of the GA algorithm only requires an objective function, and here the objective function is set as the squared error between the measured data points (true label data) and the simulation results given by the simulation model. The computational complexity of the GA algorithm is relatively large (if you want to find a suitable solution). If you want to obtain a better solution, the GA algorithm needs to obtain a very large number of experimental results. The experimental result refers to that given a parameter combination, using the CMP physical model with this parameter combination to give the simulation values of all true data points, and then comparing the difference between the simulation values given by the model and the true label values, which is an experimental result.
[0124] The parameter combination refers to the parameter combination generated by the undetermined parameters in the CMP physical model. For example, if there are 10 undetermined parameters set in the CMP physical model, these 10 undetermined parameters are the parameter combination. Since each parameter has many possible values, the number of this combination is very large. For each additional undetermined parameter, the number of all possible cases of the parameter combination increases exponentially. The total number of searches of the GA algorithm is the population size multiplied by the number of iterations. It takes 10s to evaluate the CMP physical model once, while it only takes 0.1s to evaluate the deep learning model once. It can be seen that it is very necessary to use deep learning to accelerate.
[0125] Regarding the acceleration of solving the physical model parameters, as described before, because given a data set, the speed at which the physical model obtains the simulation results of each point in this data set is much slower than that of the neural network. And during the iteration of the GA algorithm, it is necessary to repeatedly simulate the data set to obtain the simulation results. If only one simulation is calculated, the difference between 10s and 0.1s is not significant. However, if 10,000 simulations need to be calculated, the difference will be very large. The same is true for accelerating the simulation. For a given data set, the neural network calculates much faster than the physical model, which is no problem. In addition, as the scale of the data set increases, the advantage of the neural network becomes more and more obvious. For example, for only 100 data points, the physical model takes 3s and the neural network takes less than 0.1s. If it is necessary to simulate 10 million data points, the physical model may take seconds, while the calculation time of the neural network does not increase in this way and may only take 2s.
[0126] During the entire use of the CMP tool, as mentioned above, there are mainly two scenarios. One is parameter solving, and the other is global simulation. When solving parameters, the user provides dozens of real data points collected to solve the parameters of the physical model. When performing global simulation, the user can provide tens of millions of data points to calculate the simulation values of each data point through simulation.
[0127] By minimizing the objective function, a set of parameter solutions can be found, which can make the results given by the simulation model as close as possible to the real values. If deep learning acceleration is not introduced, the objective function here is the squared error between the simulation results and the real results of the physical simulation model. If deep learning is introduced, it becomes the squared error between the prediction results of the deep learning network and the real results.
[0128] The physical model is a large model (very complex, time-consuming in calculation, large in computing resources and time-consuming to solve the optimal parameters, but more accurate, and can output the simulation values of the CMP intermediate process. The neural network model has a very fast operation speed, and because the serial process is removed, the calculation of the neural network can be fully accelerated in parallel. Simply put, even without parallel acceleration (using GPU), the pure neural network prediction calculation speed is much faster than the physical model. Now, precisely because of the calculation characteristics of the two, the neural network can be further accelerated in parallel using GPU, which can greatly improve the speed. Therefore, the finally used approximate replacement model can greatly improve the parameter solving efficiency.
[0129] In the embodiments of the present disclosure, through the approximate replacement of the neural network model, the single operation time is greatly reduced. Therefore, a very large population size and the number of iterations can be set to find the optimal parameter combination, and the total time does not increase compared to before, greatly increasing the probability of finding a better solution.
[0130] For simulation, the above embodiments provide a neural network model that can respond extremely quickly and is a model that is easy to accelerate using hardware such as GPU (compared to the CMP physical model). Therefore, even if the amount of full-chip data to be simulated is huge, even tens of millions of rows, and the original CMP physical model takes several hours, the neural network model in the embodiments of the present disclosure only takes a few minutes.
[0131] As is well known, deep learning neural network models have extremely strong fitting capabilities. However, one characteristic of deep learning neural networks is that they require a very large amount of data. Therefore, the drawback is that a large amount of training data is required, and in many actual situations, there is not so much labeled data available for training. In the CMP process, due to practical limitations (such as cost), manufacturers cannot provide the data volume required to train a neural network. Usually, only 50 - 100 data points can be provided, while for a neural network, hundreds of thousands or even millions of data points are often required, which is obviously impossible to achieve. In addition, as mentioned before, the CMP physical model is a serial computing model, that is, the result at each moment also depends on the result of the previous step, and this computing characteristic makes it impossible to perform parallel computing.
[0132] However, in this application, instead of directly using a neural network model to fit the real CMP process (where the real measurement points are often very few) with real measurement data, a different approach is taken. The neural network is made to track and imitate the CMP physical model. Based on the known CMP physical model, theoretically, an infinite amount of labeled data can be generated, thus solving the problem of neural network model training. Additionally, in this application, the neural network does not directly replace the CMP physical model. It only uses its extremely fast computing speed to replace the requirement for the return result of the objective function in the GA algorithm. Instead of directly using the CMP physical model to return the test results required by the GA algorithm, a trained neural network model is used to give the test results. The ultimate goal remains unchanged, which is still to obtain a parameter combination that makes the CMP physical model fit better. However, in this application, a neural network is cleverly used to accelerate the process.
[0133] Some embodiments of the present disclosure provide a method for generating training samples, a method for training a neural network, and a method for semiconductor device simulation. It should be noted that the examples given in the above embodiments are only for illustrating the solutions of the embodiments of the present disclosure and are not used to limit the solutions of the present disclosure.
[0134] In some embodiments of the present disclosure, by creatively using the idea of approximate substitution, a neural network is cleverly introduced to accelerate the parameter solution and simulation speed of the physical model. The embodiments of the present disclosure can utilize the advantages of the neural network and are completely feasible both theoretically and practically.
[0135] The present disclosure creatively introduces a deep learning / machine learning model to greatly accelerate the establishment and simulation process of the CMP physical model, significantly reducing the associated time and effort consumption, and having extremely high practicality in actual use.
[0136] The deep learning / machine learning models introduced in this disclosure do not change the original accuracy and results of the CMP physical model, but only greatly assist in its establishment and simulation through appropriate utilization, and even further improve its accuracy. This disclosure uses deep learning models to greatly reduce the simulation and establishment of the CMP physical model, reducing the time cost. The ingenious use of deep learning models in this disclosure can help improve the accuracy of the CMP physical model and help find more appropriate combinations of model parameters.
[0137] It should be understood that the embodiments shown in the drawings are only for schematically showing the solutions of some embodiments of this disclosure and are not intended to limit this disclosure. The embodiments of this disclosure can also have various other forms.
[0138] An electronic device is also disclosed in the embodiments of this disclosure. The electronic device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, the instructions when executed by the processor causing the device to perform operations, the operations including: training a neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated by the CMP physical model based on graphical features and physical information for determining the surface height of the polished wafer; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the respective updated parameters as the trained neural network model, where the trained neural network model outputs an enhanced formula for determining the surface height.
[0139] A computer-readable storage medium is also disclosed in the embodiments of this disclosure, on which a computer program is stored, and the program when executed by the processor implements the method of the embodiments of this disclosure described above.
[0140] Figure 5 A schematic block diagram of an electronic device according to some exemplary embodiments of this disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of this disclosure described and / or claimed herein.
[0141] As Figure 5As shown, device 500 includes a CPU 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0142] Multiple components in device 500 are connected to the I / O interface 505. The multiple components include: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0143] The various processes and treatments described above, such as methods 200, 300, can be executed by the CPU 501. For example, in some embodiments, methods 200, 300 can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more steps of the methods 200, 300 described above can be executed.
[0144] The solution according to the embodiments of the present disclosure can be a method, a device, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present disclosure are carried. The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable program instructions can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network.
[0145] The embodiments of the present disclosure have been described above. The above description is exemplary and only represents alternative embodiments of the present disclosure. It is not exhaustive and does not limit the present disclosure. Although the claims in this application are formulated for specific combinations of features, it should be understood that the scope of the present disclosure also includes any novel feature or any novel combination of features that are explicitly or implicitly disclosed herein or any generalization thereof, regardless of whether it relates to the same solution in any of the currently claimed claims. It should be understood that new claims may be formulated for these features and / or combinations of these features during the examination of this application or in any further application derived therefrom.
[0146] The choice of terms used herein is intended to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other ordinary technicians in the technical field to understand the embodiments disclosed herein. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for generating training samples, comprising: Generating a first number of feature vectors based on data points for representing various pattern features on a wafer, where the feature vectors are combinations of data representing different types of pattern features among the data points; Generating a second number of physical vectors based on physical information related to chemical mechanical polishing of the wafer, where the physical vectors are combinations of data representing different types of physical information; And Performing, by a chemical mechanical polishing physical model, the second number of simulations based on the feature vectors and the physical vectors to generate, in each simulation, simulation values corresponding to the first number of feature vectors of the surface height of the wafer as training samples.
2. The method according to claim 1, wherein generating a first number of feature vectors based on data points for representing various pattern features on a wafer comprises: Respectively determining data within each step range based on the value range and corresponding step size of the data points of each of the pattern features; And Orthogonalizing the data within each step range to generate a feature data table, where the feature data table includes data of the first number of feature vectors.
3. The method according to claim 1, wherein generating a second number of physical vectors based on physical information related to chemical mechanical polishing of the wafer comprises: Respectively determining data within each step range based on the value range and corresponding step size of the data points of each of the physical vectors; And Orthogonalizing the data within each step range to generate a physical information data table, where the physical information data table includes data of the second number of physical vectors.
4. The method according to claim 1, wherein performing, by a chemical mechanical polishing physical model, the second number of simulations based on the feature vectors and the physical vectors comprises: Respectively generating a first simulation value of the surface height of a non-groove structure and a second simulation value of the surface height of a groove structure; And Determining the simulation value of the surface height of the wafer based on the first simulation value and the second simulation value.
5. The method according to claim 1, wherein the pattern features at least include: The density of patterns within lattice points on the wafer; The width of the patterns; and The spacing between patterns.
6. The method according to claim 1, wherein the physical information at least includes the following items: Information on deposition materials; Maximum dishing defect information and minimum dishing defect information for each step; Abrasive slurry information; Long tail effect parameter information; Duration information; Nominal pressure information.
7. A method for training a neural network model, comprising: Training a neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated by the method according to any one of claims 1 to 6; Updating the parameters of the neural network model based on the gradient of the loss function; and In response to the training satisfying a predetermined condition, determining the neural network model corresponding to the updated parameters as the trained neural network model, where the trained neural network model outputs a predicted value of the surface height.
8. The method according to claim 7, wherein training the neural network model using training samples to determine the gradient of the loss function of the neural network model includes: Calculating the gradient of the loss function using a stochastic gradient descent algorithm.
9. The method according to claim 8, wherein updating the parameters of the neural network model based on the gradient of the loss function includes: Performing gradient update on the gradient of the loss function in response to a predetermined number of trainings; And Updating the parameters of the neural network model based on the gradient update.
10. The method according to claim 7, wherein training the neural network model using training samples to determine the gradient of the loss function of the neural network model includes: Training the neural network model using a part of the training samples to perform gradient descent backpropagation to update the parameters; Using the trained neural network model to predict the remaining part of the training samples to generate predicted values of the surface height of the polished wafer; And In response to the comparison result between the predicted value and the simulated value of the chemical mechanical polishing physical model satisfying a predetermined condition, determining the corresponding trained neural network model as the final neural network model.
11. The method according to claim 10, wherein the comparison result between the predicted value and the simulated value satisfying a predetermined condition includes: The root mean square error between the predicted value and the simulated value is less than a predetermined threshold.
12. The method according to claim 7, further comprising: Determining the loss function value of the neural network model; And wherein, the training satisfying a predetermined condition includes: The difference between the loss function values of two consecutive iterations is less than a predetermined difference threshold; or The absolute value of the ratio of the loss function values of two consecutive iterations is greater than a predetermined ratio threshold.
13. The method according to any one of claims 7 to 12, wherein the training samples are generated in parallel on multiple servers so that the training of the neural network model is performed in parallel.
14. A neural network model generated by the method according to any one of claims 7 to 13.
15. A simulation model, comprising: A chemical mechanical polishing physical model configured to determine a simulated value of the surface height of the polished wafer based on input data, wherein the input data includes graphic feature data and physical information; And The neural network model according to claim 14, which is generated by training using the simulated value as a training sample, and is configured to perform parallel prediction on multiple sets of input data during the simulation process of the chemical mechanical polishing physical model to generate multiple predicted values of the surface height of the polished wafer corresponding to the multiple sets of input data, respectively as the corresponding simulated values of the surface height.
16. The simulation model according to claim 15, wherein performing parallel prediction on multiple sets of input data during the simulation process of the chemical mechanical polishing physical model includes: Using multiple graphics processors to perform parallel acceleration on the calculation process of the neural network model.
17. A simulation method for a semiconductor device, which is performed using the simulation model described in claim 15, the method comprising: Predicting a simulated value of the surface height of the polished wafer by a neural network model in the simulation model based on input data, where the input data includes pattern feature data and physical information.
18. The method according to claim 17, wherein predicting a simulated value of the surface height of the polished wafer by a neural network model in the simulation model based on input data comprises: The neural network model performing parallel prediction on multiple sets of input data to generate multiple predicted values of the surface height of the polished wafer corresponding to the multiple sets of input data, respectively as the corresponding simulated values of the surface height.
19. The method according to claim 17, wherein predicting a simulated value of the surface height of the polished wafer by a neural network model in the simulation model based on input data comprises: The neural network model respectively generating a first predicted value of the surface height of the non-groove structure and a second predicted value of the surface height of the groove structure; And The neural network model determining a simulated value of the surface height of the wafer based on the first predicted value and the second predicted value.
20. The method according to claim 17, further comprising: Determining the condition of the surface of the wafer during the polishing process through a chemical mechanical polishing physical model in the simulation model.
21. A neural network model, comprising: An input layer configured to receive pattern feature data on a wafer and physical information related to chemical mechanical polishing of the wafer; An intermediate layer configured to perform the method according to any one of claims 17 to 20 to determine a surface height value of the polished wafer; And An output layer configured to output the surface height value.
22. The method according to claim 21, further comprising: An embedding layer configured to map the input physical information to generate a physical vector.
23. An electronic device, comprising: A processor; And A memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the device to perform actions, the actions including: Training a neural network model using training samples to determine a gradient of a loss function of the neural network model, where the training samples are generated by the method according to any one of claims 1 to 6; Updating parameters of the neural network model based on the gradient of the loss function; and In response to the training satisfying a predetermined condition, determining the neural network model corresponding to the updated parameters as a trained neural network model, where the trained neural network model outputs an enhanced formula for determining the surface height.
24. A computer-readable storage medium having machine-executable instructions stored thereon, which when executed by a processor cause the processor to implement the method according to any one of claims 1 to 13 and claims 17 to 20.
Citation Information
Patent Citations
Method, device and medium for simulating wafer processing process
CN118862368A