Simulation Method, Electronic Device, and Storage Medium for Semiconductor Devices

By using deep learning neural network to adjust the initial formula in the simulation method of semiconductor devices, the problem of insufficient prediction accuracy of metal layer thickness in the CMP back-stage process in the prior art is solved, and higher simulation accuracy and prediction fitting capabilities are achieved.

CN119761276BActive Publication Date: 2025-07-01QUANXIN INTELLIGENT MFG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510247200.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-01
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

When the prior art predicts the thickness of the metal layer after the chemical mechanical polishing (CMP) process, the modeling accuracy is insufficient and it is impossible to accurately predict the situation after grinding.

Method used

A simulation method for semiconductor devices is provided, by determining the simulation value of the surface height of the wafer after grinding based on the time step and the grinding rate, and determining the correction terms related to the grinding rate using a deep learning neural network, adjusting the initial formula to improve the simulation accuracy.

Benefits of technology

It significantly improves the simulation accuracy and prediction fitting ability of the simulation model, and can more accurately predict the thickness of the metal layer after CMP grinding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761276B_ABST
    Figure CN119761276B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a simulation method, an electronic device, and a storage medium for semiconductor devices. The method includes: determining a simulated value of the surface height of a polished wafer based on a time step and a corresponding polishing rate; determining a correction term related to the polishing rate based on a comparison between the simulated value and a target value, so that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the surface of the polished wafer. The technical solution of the present disclosure can significantly improve the simulation accuracy and prediction fitting ability of the simulation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure mainly relate to integrated circuits, and more particularly, to a simulation method, an electronic device, and a storage medium for semiconductor devices. Background Art

[0002] Currently, the prediction of the metal layer thickness after the post-chemical mechanical polishing (CMP) process typically relies heavily on the CMP model.

[0003] In traditional solutions, a pure physical model is derived through intuitive physical mechanisms. However, in the face of certain complex situations and process conditions, the modeling accuracy is somewhat insufficient, making it impossible to accurately predict the situation after CMP grinding. Summary of the Invention

[0004] According to an exemplary embodiment of the present disclosure, a simulation solution for semiconductor devices is provided to at least partially overcome the above or other potential defects.

[0005] According to one aspect of the present disclosure, a simulation method for a semiconductor device is provided. The method includes: determining a simulated value of the surface height of a wafer after grinding based on a time step and a corresponding grinding rate; determining a correction term related to the grinding rate based on a comparison between the simulated value and a target value, such that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the surface of the wafer after grinding.

[0006] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the device to perform operations, the operations including: determining a simulated value of the surface height of a wafer after grinding based on a time step and a corresponding grinding rate; determining a correction term related to the grinding rate based on a comparison between the simulated value and a target value, such that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the surface of the wafer after grinding.

[0007] In some embodiments, determining a correction term related to the grinding rate based on a comparison between the simulated value and the target value includes: determining the correction term using a deep learning neural network based on graphic feature information, process information, and information on an initial formula for calculating the surface height to correct the initial formula.

[0008] In some embodiments, the correction terms at least include polynomials related to at least one of the following: the width of the patterns within the grid points on the wafer; the spacing between the patterns; and the sum of the perimeters of the respective patterns within the grid points.

[0009] In some embodiments, determining correction terms using a deep learning neural network to correct an initial formula includes: generating a random number at the start of each iteration of the neural network; selecting a source of the correction terms for correcting the initial formula based on a comparison of the random number with a random probability threshold; and determining updated correction terms using the neural network based on the selection to correct the initial formula.

[0010] In some embodiments, determining updated correction terms using the neural network based on the selection to correct the initial formula includes: in response to a first comparison result of the random number with the random probability threshold, selecting a symbol from a symbol library as a correction term to correct the initial formula using the neural network, where the symbol library is generated by taking the Cartesian product combination of a mathematical symbol library and a variable library, and the variable library contains graphic feature information and process information; and in response to a second comparison result of the random number with the random probability threshold, using the neural network to make a prediction to generate a new correction term and adding the new correction term to the symbol library.

[0011] In some embodiments, determining correction terms using a deep learning neural network to correct an initial formula includes: generating a reward term based on the simulated value and the measured value of the surface height of each grid point; and iterating an array composed of the initial formula, the correction terms, the updated formula, and the reward term as an element of a training sample in the neural network to generate an element of a new training sample.

[0012] In some embodiments, generating a reward term based on the simulated value and the measured value of the surface height of each grid point includes: determining a first difference between the standard deviation of the simulated value of the surface height of each grid point and the standard deviation of the measured value; determining a second difference between the RMSE of the simulated value of the surface height of each grid point and the RMSE of the measured value; determining the range between the simulated value and the measured value of the surface height of each grid point; and generating a reward term based on the first difference, the second difference, and the range.

[0013] In some embodiments, the iteration includes a prediction phase and a training phase, and wherein: in response to an element generated in each round of iteration in the prediction phase satisfying a first predetermined condition, terminating the round of iteration and taking all elements in the round of iteration as a trajectory; and in response to the number of total trajectories obtained during each round of iteration reaching a predetermined trajectory threshold, terminating the prediction phase.

[0014] In some embodiments, the first predetermined condition includes: the number of elements generated in each round of iteration reaching a first predetermined threshold; or the value of RMSE in the current iteration being less than a second predetermined threshold.

[0015] In some embodiments, determining a correction term using a deep learning neural network to correct an initial formula includes: in a training phase, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network; and updating the parameters of the neural network based on the gradient of the loss function.

[0016] In some embodiments, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network includes: using a policy gradient algorithm in the field of reinforcement learning to perform backpropagation on the training samples to determine the gradient of the loss function.

[0017] In some embodiments, determining a correction term using a deep learning neural network to correct an initial formula further includes: determining the value of the loss function; in response to the value of the loss function being lower than a predetermined threshold, determining the correction term output by the neural network corresponding to the value of the loss function as a correction term related to the grinding rate; and determining an updated formula based on the determined correction term to determine the height value of the surface.

[0018] In some embodiments, the iteration in the training phase stops when the following conditions are met: the difference between the values of the loss function in two consecutive iterations is less than a predetermined difference threshold; the absolute value of the ratio of the values of the loss function in two consecutive iterations is greater than a predetermined ratio threshold; or the number of iterations or the iteration time reaches a maximum iteration threshold.

[0019] In some embodiments, the random probability threshold changes from an initial first probability threshold to a second probability threshold lower than the first probability threshold during the iteration.

[0020] In a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to the first aspect of the present disclosure.

[0021] In a fourth aspect of the present disclosure, a simulation model is provided, configured to execute the method according to the first aspect of the present disclosure.

[0022] In a fifth aspect of the present disclosure, a neural network is provided, including: a plurality of input layers, configured to receive feature information of a pattern on a wafer, process information, and an initial formula for calculating the surface height of the wafer after grinding; an intermediate layer, configured to execute the method according to the first aspect of the present disclosure to determine an updated formula for calculating the simulation value of the surface height; and an output layer, configured to output the updated formula.

[0023] In a fifth aspect of the present disclosure, there is provided a method for training a neural network model, including: predicting by the neural network model based on input data to generate training samples, where the input data includes graphic feature information, process information, and an initial formula for determining the surface height of a polished wafer; training the neural network model using the training samples to determine the gradient of the loss function of the neural network model; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training meeting a predetermined condition, determining the neural network model corresponding to the updated parameters of the neural network model as the trained neural network model, where the trained neural network model outputs an enhanced formula for determining the surface height.

[0024] In some embodiments, predicting by the neural network model based on input data to generate training samples includes: predicting by the neural network model based on the initial formula and an initial correction term to generate an updated formula and a reward term, where the reward term represents a feedback value generated based on the difference between the simulation value and the target value of simulating the surface height of the polished wafer by the neural network model; forming an array composed of the initial formula, the initial correction term, the updated formula, and the reward term as the first training sample; and iterating in the neural network model based on the first training sample to generate multiple training samples.

[0025] In some embodiments, the reward term is generated by the following method: determining a first difference between the standard deviation of the simulation values of the surface height of each grid point on the wafer and the standard deviation of the target value; determining a second difference between the root mean square error of the simulation values of the surface height of each grid point and the root mean square error of the target value; determining the range between the simulation value and the target value of the surface height of each grid point; and generating the reward term based on the first difference, the second difference, and the range.

[0026] In some embodiments, in response to the samples generated in each iteration meeting a first predetermined condition, terminating the iteration, and taking all the samples in this iteration as a trajectory; and in response to the number of total trajectories obtained during each iteration process reaching a predetermined trajectory threshold, terminating the iteration.

[0027] In some embodiments, training the neural network model using the training samples to determine the gradient of the loss function of the neural network model includes: using the policy gradient algorithm in the field of reinforcement learning to perform backpropagation on the training samples to determine the gradient of the loss function.

[0028] In some embodiments, it further includes: determining the loss function value of the neural network model; and wherein, the training meeting a predetermined condition includes: the difference between the loss function values of two consecutive iterations is less than a predetermined difference threshold; or the absolute value of the ratio of the loss function values of two consecutive iterations is greater than a predetermined ratio threshold.

[0029] In some embodiments, the neural network model makes predictions based on input data to generate training samples, including: generating a random number at the beginning of each iteration of the neural network; selecting the source of the initial correction term for correcting the initial formula based on the comparison between the random number and the random probability threshold; and determining an updated correction term based on the selection to correct the initial formula using the neural network to generate an updated formula.

[0030] In some embodiments, determining an updated correction term based on the selection to correct the initial formula using the neural network to generate an updated formula includes: in response to a first comparison result between the random number and the random probability threshold, selecting a symbol from the symbol library as the initial correction term to correct the initial formula using the neural network model, where the symbol library is generated by taking the Cartesian product of mathematical symbols and a variable library, and the variable library contains graphical feature information and process information; and in response to a second comparison result between the random number and the random probability threshold, making a prediction using the neural network model to generate an initial correction term to correct the initial formula using the neural network model to generate a new correction term, and adding the new correction term to the symbol library.

[0031] In the sixth aspect of the present disclosure, a neural network model generated according to the method of the fifth aspect is provided.

[0032] It will be understood from the following description that the technical solution of the present disclosure can significantly improve the simulation accuracy and prediction fitting ability of the simulation model.

[0033] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or main features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0035] Figure 2 A flowchart showing a simulation method for semiconductor devices according to some embodiments of the present disclosure;

[0036] Figure 3 A schematic diagram showing a cross-section of a wafer with trenches according to some embodiments of the present disclosure;

[0037] Figure 4 A schematic diagram showing a neural network architecture according to some embodiments of the present disclosure;

[0038] Figure 5 A block diagram showing a computing device capable of implementing multiple embodiments of the present disclosure.

[0039] In the various figures, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners

[0040] The principles of the present disclosure will be described below with reference to various exemplary embodiments shown in the accompanying drawings. It should be understood that the description of these embodiments is only for enabling those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that, where feasible, similar or identical reference numerals may be used in the figures, and similar or identical reference numerals may represent similar or identical functions. Those skilled in the art will readily recognize that alternative embodiments of the structures and methods described herein may be employed without departing from the principles of the invention described herein.

[0041] As used herein, the term "comprising" and its variations mean open-ended inclusion, that is, "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an exemplary embodiment" and "an embodiment" mean "at least one exemplary embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.

[0042] CMP is a method of removing materials that combines chemistry and physics, and has become an indispensable step in the modern integrated circuit (IC) industry. Specifically, CMP uses chemical etching and mechanical force to planarize the silicon wafer or other substrate materials during the processing. CMP can achieve material removal at the nanoscale level, making the wafer surface planar.

[0043] The CMP model, as the name implies, is a physical model for simulating the CMP process. The model outputs some important physical indicators of each grid point of the wafer after the CMP process, such as the thickness of the trench structure, the thickness of the non-trench structure, and the thickness of the metal material after CMP. In simple terms, it is a physical model that predicts the surface morphology of the wafer after the CMP process. CMP is one of the eight major semiconductor processes, connecting the previous and the next. Other processes following CMP will be affected by the CMP process. In other words, when it is necessary to evaluate and control the process results after CMP, it is best to know the impact of the CMP process, which is of great significance to the full process control and yield improvement of the entire semiconductor manufacturing. In addition, the CMP process itself is an indispensable and necessary step, because the CMP process can be better adjusted through the CMP model. For example, the results of the CMP model can be used to infer which process parameter settings need to be modified and adjusted, and whether there are some unreasonable aspects in the design of the Graphic Data System (GDS), which may cause hotspots, that is, defects.

[0044] The metal layer is deposited on the bottom of the groove. As the name implies, the bottom of the groove is the bottom of the groove. In fact, the process flow is to first go through the etching step, that is, to etch the wafer to produce grooves, or to produce the bottom of the groove, and then to carry out the deposition process, which will deposit various materials on the wafer. Because of the etching, and the different areas have different groove depths due to different GDS designs, naturally the stacking height of the materials in each area after deposition will also be different. For example, the most common metal deposition material for CMP is copper. Usually in the last step of deposition, that is, before entering CMP, the material on the top of the wafer is copper, that is, the first step of grinding is copper. Because of the existence of etching, different areas have grooves of different depths and different deposition thicknesses. Therefore, for the surface of the wafer after the CMP process, it is necessary to evaluate the surface morphology, not only to understand how much change has occurred in the CMP step, but also to understand the changes and fluctuations that already existed before CMP.

[0045] In fact, CMP essentially refers to polishing the surface of a wafer immersed in a polishing liquid with a polishing head on a chuck. Simply put, GDS is a data format of a real entity, the wafer, in a computer program. When the data of a wafer is input into a program for observation, description, and analysis, it needs to be abstracted into a type of computer data, that is, a GDS file. The GDS file clearly depicts information such as the structure and design of this wafer. The wafer in the real world is the GDS in a virtual computer. From a more professional perspective, the GDS layout is a layout file given by circuit designers, which contains various graphics, that is, patterns, and semiconductor manufacturing is to engrave a wafer into the appearance of the layout design through various processes.

[0046] Usually, the value at the bottom of the trench is directly regarded as a fixed value, or simply an etching table is made using some empirical values or mathematical formulas, and the value is obtained by looking up some limited values in the etching table. There are great limitations in both accuracy and adjustability. Sometimes, in order to fit the measured metal layer thickness value, the CMP model is changed, resulting in the CMP model sacrificing the fitting accuracy of the dishing and erosion indicators. Here, dishing refers to the difference in height between the trench structure and the non-trench structure, and erosion refers to the difference in height of the non-trench structure at each grid point relative to the reference thickness value.

[0047] In view of this, the present disclosure provides an improved solution.

[0048] Some embodiments of the present disclosure provide an improved simulation method for semiconductor devices. The method includes: determining a simulated value of the height of the polished wafer surface based on a time step and a corresponding polishing rate; determining a correction term related to the polishing rate based on the comparison between the simulated value and a target value, so that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the polished wafer surface.

[0049] Some embodiments of the present disclosure also provide a training method for a neural network model. The method includes: predicting by the neural network model based on input data to generate training samples, where the input data includes graphic feature information, process information, and an initial formula for determining the surface height of the polished wafer; training the neural network model using the training samples to determine the gradient of the loss function of the neural network model; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the parameters of the updated neural network model as the trained neural network model, where the trained neural network model outputs an enhanced formula for determining the surface height.

[0050] The technical solution of the present disclosure can significantly improve the simulation accuracy and prediction fitting ability of the simulation model.

[0051] Embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings.

[0052] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. As Figure 1 shown, the example environment 100 includes a computing device 110 and a client 120.

[0053] In some embodiments, the computing device 110 can interact with the client 120. For example, the computing device 110 can receive an input message from the client 120 and output a feedback message to the client 120. In some embodiments, the input message from the client 120 can be, for example, layout data. The computing device 110 can perform corresponding mathematical operations on the layout data and output the corresponding operation results to the client 120.

[0054] In some embodiments, the computing device 110 can include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant PDA, a media player, etc.), a consumer electronic product, a minicomputer, a mainframe computer, cloud computing resources, etc.

[0055] It should be understood that describing the structure and function of the example environment 100 only for exemplary purposes is not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. This environment is merely illustrative and is not used to limit the application environment of the embodiments of the present disclosure.

[0056] To more clearly explain the principle of the solution of the present disclosure, the following will be described in more detail with reference to Figure 2 to describe in more detail.

[0057] Figure 2 A flowchart of a simulation method for a semiconductor device according to some embodiments of the present disclosure is shown.

[0058] At block 202, a simulated value of the surface height of the polished wafer is determined based on the time step and the corresponding grinding rate.

[0059] Several key elements in some embodiments of the present disclosure are introduced first:

[0060] Element 1: The features of GDS (also known as graphic features), mainly referring to features such as density, width, space, and perimeter. At each position on an entire GDS, there are different GDS features (different density and width values), and different GDS features will greatly affect the topography height of each position after the etching process. In some embodiments, the density of the pattern, the width of the pattern, the space between patterns, and the perimeter of the lattice points can be referred to as feature size information. The average space between patterns within a lattice point is called space, the average width of the pattern itself is called width, and the same applies to the perimeter. The sum of the side lengths of all patterns within a lattice point is the perimeter of this lattice point (perimeter).

[0061] In GDS, there are many patterns (or called graphics), which can be microscopic structures such as circuits. In a region of a fixed size, the ratio of the total area of the existing patterns to the area of this region is the density. So density is a number between 0 and 1. Width is a descriptive statistical indicator for measuring the width of the pattern itself within this region. The patterns are usually rectangles or combinations of rectangles. Usually, all the patterns within the scanned region are obtained, and their width values are finally averaged or the median is taken as the width value of this region). Space describes the space between patterns within this region because the patterns are arranged at a certain interval.

[0062] Element 2: Prediction target / simulation object. What is simulated is the value of the surface height of each point on the GDS after the CMP process. In addition, the trench height after the etching process, that is, the value at the bottom of the trench, is denoted as TB.

[0063] Element 3: Process information table (recipe), which records process information. Different manufacturers will have different process settings, specifically including information such as the duration of the process, start time, end time, unit time step, etc.; and various materials, thicknesses obtained from the previous deposition process, and the target thickness after etching and polishing, etc.

[0064] The CMP physical model refers to using the features extracted from the GDS, usually density, linewidth (abbreviated as width), linespace (abbreviated as space), perimeter combined with process parameters (usually pressure, removal rate of different materials remove_rate, process time time, etc.) to form a physical formula to simulate the surface height (Surface Height) after the CMP process.

[0065] In some embodiments, the polishing rate is determined based on the feature information of the patterns on the wafer and the parameters related to the polishing rate.

[0066] In some embodiments, a simulated value H1 of the height of the non-trench structure on the wafer can be determined; a simulated value H2 of the height of the trench structure on the wafer can be determined; and the simulated value of the surface height of the wafer can be determined based on the simulated values H1, H2 and the density of the patterns within the grid area on the wafer. Specifically, the surface height is equal to the height of the non-trench structure multiplied by (1 - density) plus the product of the height of the trench structure and density. Expressed by the formula: H1 * density + H2 * (1 - density); where density represents the density of the patterns within the grid area on the wafer, and its range is between 0 and 1.

[0067] The following will be described with reference to Figure 3 this. Figure 3 FIG. shows a schematic cross-section of a wafer with trenches according to some embodiments of the present disclosure.

[0068] As mentioned previously, Dishing is the difference between the height of the trench structure and the height of the non-trench structure, and the specific formula is SThickNT – SthickT. As Figure 3 shown, where D represents Dishing, SThickNT represents the height of the non-trench structure, and SthickT represents the height of the trench structure. In this figure, Dishing is actually the depth of the trench. As Figure 3 shown, the area shown in this wafer includes multiple trenches, that is, an array of trenches. The left side is the shown control area without a pattern structure, which is a blank test control area specifically used for comparison with the patterned area. And all the grid areas in the model in the embodiments of the present disclosure represent areas with patterns, as Figure 3 shown in the right part of

[0069] Erosion refers to the difference between the height of the non-trench structure of each grid point and the reference thickness value, and the specific formula is field reference thickness (the reference thickness in the control area, Figure 3 shown as the reference height in

[0070] Returning to Figure 2Continue the description. In some embodiments of the present disclosure, the CMP process is based on a basic physical model and theory. In some embodiments of the present disclosure, the model formula is extended based on the above theory. The theory is that within a very short time t (usually less than or equal to 0.1 s), denoted as dt, the polishing rate is constant, and the trench height at the next moment can be expressed by the following formula:

[0071] SThickNT(t=t1) = SThickNT(t = t0) – dt*R_NT(t = t0);

[0072] The non-trench height can be expressed by the following formula:

[0073] SThickT(t=t1) = SThickT(t = t0) – dt*R_T(t = t0);

[0074] In the above formula, R_NT(t = t0) represents the polishing rate of the non-trench structure at t0, and the meaning of R_T can be obtained in the same way.

[0075] It can be seen that the calculation of R_NT and R_T is the core of the entire CMP physical model. According to the basic model, the polishing rate Rate is related to the actual pressure, and the actual pressure is related to the density.

[0076] For the non-trench structure, the most primitive physical mechanism formula for polishing is: R_NT = P / (1 - density). For the trench structure, R_T = P / density. However, in fact, it needs to be corrected, that is, it needs to be corrected with a correction term.

[0077] In some embodiments, the correction term includes at least a polynomial related to one of the following: the width of the pattern within the grid points on the wafer; the space between patterns; and the sum of the perimeters of each pattern within the grid points.

[0078] In fact, for the non-trench structure, R_NT at a certain moment is as follows:

[0079] R_NT = P / (1 - density) * Width correction term * Space correction term * Perimeter correction term.

[0080] For the trench structure, R_T at a certain moment is as follows:

[0081] R_T = P / density * Width correction term * Space correction term * Perimeter correction term.

[0082] There are no very clear gold standards and undisputed formulas for these correction items before. Different model developers and different manufacturers have their own different presets. Specifically, the Width correction item refers to a polynomial with Width as the core of change (or a polynomial related to Width), and the same applies to the Space correction item and the Perimeter correction item. This polynomial can be a polynomial in the form of linear summation, or it can be a complex polynomial doped with forms such as exponential and logarithmic forms.

[0083] In some embodiments of the present disclosure, the most advanced deep reinforcement learning technology is used to complete the derivation of these correction items based on actual data, so as to complete the expansion of the CMP basic physical formula. In some embodiments, the actual data may include graphic feature information.

[0084] In some embodiments, the graphic feature information may include the following information: the density of the graphics within the grid points on the wafer; the width of the graphics; the spacing between the graphics; and the initial height of the bottom of the trench, the total perimeter of the graphics within the grid points.

[0085] In the usage scenario of the entire model, the so-called point refers to a grid point. By extracting the information inside a grid point, a data point can be obtained. The so-called measurement point is also this grid point. Strictly speaking, a grid point can be said to be within a rectangular area delimited based on a specific size (the size is usually between 5um*5um and 20um*20um).

[0086] In some embodiments, the simulated value of the surface height can be determined based on the initial data and initial parameters in the CMP model. In some embodiments, the initial data may include the measured values of several measurement points (recording the position of each point in GDS, its corresponding GDS feature, and the height of the trench after etching).

[0087] From the above description, it can be known that the simulated value of the surface height of the polished wafer can be determined based on the time step and the corresponding polishing rate. For example, it can be calculated based on the initial physical formula. In fact, it can be determined by any method of determining the initial simulated value through a simulation model.

[0088] In some embodiments, a deep learning neural network is constructed to determine the simulated value of the surface height through randomly generated initial parameters and initial data. A simulated value closer to the measured value can be obtained through the iteration of the CMP model.

[0089] The determined simulated value of the height may have a large gap from the actual measured value. A simulated value closer to the measured value can be obtained through an improved CMP model.

[0090] At block 204, a correction term related to the polishing rate is determined based on a comparison between the simulated value and the target value, such that the difference between the simulated value corresponding to the correction term and the target value becomes smaller.

[0091] In some embodiments, an improved CMP model is obtained through a deep learning neural network to determine a correction term related to the polishing rate. In other words, an improved physical formula can be obtained to calculate the simulated value of the surface height.

[0092] In some embodiments, based on the graphic feature information, process information, and information of the initial formula for calculating the surface height, a deep learning neural network can be used to determine a correction term to correct the initial formula.

[0093] In some embodiments, other neural networks can also be employed to determine a correction term related to the polishing rate.

[0094] Deep Learning specifically refers to machine learning based on deep neural network models and methods. It has evolved on the basis of algorithmic models such as statistical machine learning and artificial neural networks, in combination with the development of contemporary big data and high computing power. The most important technical feature of deep learning is its ability to automatically extract features, and the extracted features are also called deep features or deep feature representations. Compared with manually designed features, deep features have stronger and more robust representation capabilities. A deep neural network is the model basis for deep learning to automatically extract features, and a deep neural network is essentially a nested series of non-linear transformations.

[0095] Regarding neural networks, some basic definitions are introduced first: State, defined as the current physical formula, denoted as S0; Action, the behavioral action performed based on the current state S0, specifically defined as which mathematical calculation symbols and variables to select; S1, the state at the next moment, which changes after the action a is performed on the basis of S0, that is, the new physical formula. R is defined as the Reward, which means that when the action a is taken based on the state S0 and transferred to the new state S1, the obtained reward is R. The specific content of R reflects whether the new formula performs better in terms of data compared to the old formula, and whether the Root Mean Square Error (RMSE) is lower. If it is lower, it is a positive reward, and if it deteriorates, it is a negative reward (actually a penalty). R is a term in reinforcement learning. Reinforcement learning is essentially a series of concept models related to mathematics and does not have practical physical significance. Reinforcement learning is a specialized field and a branch of deep learning.

[0096] In some embodiments, a deep learning neural network can be built based on a Recurrent Neural Network (RNN) structure as an agent for deriving formulas. It should be understood that the neural network in the embodiments of the present disclosure is not limited to RNN, and other neural networks can also be used. The content received at the input end of the RNN neural network can be the features of GDS, process parameter information, and initial formula information. The network includes several fully connected layers, pooling layers, convolutional layers, etc. The characteristic of RNN is that each fully connected layer contains a gate structure. Compared with an ordinary neural network that receives the output of the previous layer as input to calculate the value of the activation function to determine whether to activate, the gate structure of RNN also considers information from even earlier layers as auxiliary inputs to participate in the calculation of the activation function. The final output of this neural network can be, for example, symbols (such as including +-* / , square root, square, logarithm and exponent with base e, and related GDS feature variables, etc.), or a complete formula.

[0097] In some embodiments, using a deep learning neural network to determine a correction term to correct an initial formula may include: generating a reward term based on the simulation values and measured values of the surface heights of each grid point; and iterating an array composed of the initial formula, the correction term, the update formula, and the reward term as elements of a training sample in the neural network to generate elements of a new training sample.

[0098] In some embodiments, training samples can be generated in the following manner: the neural network model makes predictions based on the initial formula and the initial correction term to generate an update formula and a reward term, where the reward term represents a feedback value generated based on the difference between the simulation value and the target value of the surface height of the polished wafer simulated by the neural network model; an array composed of the initial formula, the initial correction term, the update formula, and the reward term is used as a first training sample; and multiple training samples are generated by iterating based on the first training sample in the neural network model.

[0099] In some embodiments, the reward term can be generated in the following manner: determining a first difference between the standard deviation of the simulation values of the surface heights of each grid point on the wafer and the standard deviation of the target values; determining a second difference between the root mean square error of the simulation values of the surface heights of each grid point and the root mean square error of the target values; determining the range between the simulation values and the target values of the surface heights of each grid point; and generating a reward term based on the first difference, the second difference, and the range.

[0100] In some embodiments, when the samples generated in each iteration meet a first predetermined condition, the iteration of this round is terminated, and all the samples in this round of iteration are used as a trajectory; when the total number of trajectories obtained in each round of iteration reaches a predetermined trajectory threshold, the iteration is terminated.

[0101] In some embodiments, a neural network model is trained using training samples to determine the gradient of the loss function of the neural network model. Specifically, a policy gradient algorithm in the field of reinforcement learning can be used to perform backpropagation on the training samples to determine the gradient of the loss function.

[0102] In some embodiments, the iteration using a deep learning neural network may include a prediction phase and a training phase, and: when the elements generated in each iteration of the prediction phase satisfy a first predetermined condition, the iteration is terminated, and all the elements in the iteration are used as a trajectory; and when the number of total trajectories obtained during each iteration reaches a predetermined trajectory threshold, the prediction phase is terminated. This is further described later.

[0103] In some embodiments, the first predetermined condition may include: the number of elements generated in each iteration reaches a first predetermined threshold; or the value of RMSE in the current iteration is less than a second predetermined threshold.

[0104] In some embodiments, the agent is trained by the process and learning method of policy gradient in the field of reinforcement learning (essentially continuously iterating the RNN so that it can predict a complete formula). The initially available data may include a number of measurement data points under a certain process. The content of each data point is the XY coordinates, density / width / space / perimeter GDS feature information, and the corresponding ThickNT and ThickT after the CMP process, that is, the measured surface height and the initial ThickNT and ThickT values before the CMP process. Suppose there are N measurement points (data points). Each data point has a surface height result. A so-called data point actually refers to a lattice point. As mentioned before, the smallest unit in the whole model is a lattice point, and several lattice points form a complete chip. Similarly, a dataset is composed of several data points, and the full chip dataset refers to the set composed of each data point represented by each lattice point on the full chip.

[0105] During the process of simulating using this neural network, the agent can be initialized first. Initialize the basic formula, that is, initialize the ground height Thick1 to be equal to the initial height Thick0 before grinding minus the set unit time multiplied by the actual grinding rate, where the actual grinding rate is equal to the sum of the nominal grinding rates of each material multiplied by their corresponding correction factors. The specific formulas for the correction factors are slightly different for non-grooved structures and grooved structures. The original physical formula for the NT structure is: P / (1 - density). The correction factor (correction term) is, for example: * width * space * perimeter; the original formula for the T structure is: P / density. The correction factor is, for example: * width * space * perimeter. It should be understood that the correction factors shown here are only illustrative, and in fact, it is very likely not as simple as several multiplications shown, but may have various complex polynomial structures.

[0106] The state represented by the original physical formula is the initial state S0. The RNN neural network can be initialized. This neural network structure includes an input layer, which receives the feature information of GDS, the initial formula, and the set process information table. The process information table may include the set nominal pressure P0, the set slurry ratio Slu0, the total number of layers of materials to be polished in CMP, the deposition thickness T of each layer of material, the initial trench height and non-trench height, and the nominal grinding rate R_material of each material, the hardness pcoef of the pad, the total grinding time t0, and the set unit time dt. CMP uses a grinding head to assist the slurry in grinding the wafer. The slurry has different chemical concentrations, resulting in different effects.

[0107] The parameters of the neural network are random (when initialized). The parameters of the neural network are mainly the weights that make up the network. The neural network can be regarded as a combination of many simple formulas, and finally a complex integrated mathematical formula is obtained. There are many undetermined parameters in this formula. When not trained, the parameter initialization is a random value. As training progresses, the gradient is calculated using backpropagation with the training data, and the values of the parameters are continuously updated. These parameters can also be called weights.

[0108] Therefore, the neural network will randomly give a new formula symbol. A new CMP physical formula is obtained according to the new formula symbol. For example, specifically, the alternative action library can be obtained by taking the Cartesian product of various basic mathematical operation symbols and the variable library (the variable library includes process parameters and GDS features), such as "+density, -width, *space, / perimeter, +ln(Slu0)", etc., which is the action set in this disclosure. In the conventional case, when using reinforcement learning for formula expansion, the action setting is often basic mathematical symbols, such as addition, subtraction, multiplication, division, and some numbers. However, in some embodiments of this disclosure, another definition is provided for formula expansion. Because if there are only mathematical symbols, it is impossible to perform too complex combinations, but the Cartesian product of mathematical symbols, variables, and parameters is a relatively large combination library, with many choices, and better expansion results can be obtained. Take the simplest example. Suppose the initial formula is 1 + density, and the action given by the neural network is *width, so the formula becomes (1 + density) * width. Here, *width is a correction term.

[0109] The physical formula calculates the polished trench height and non-trench height based on N measurement points and calculates the surface height according to the formula described above. This value is the simulation value, that is, the RMSE between the simulated surface height and the true measured surface height can be calculated.

[0110] The neural network can be used to predict the next mathematical symbol. Its output is the next mathematical symbol, and its function is to expand the physical formula. The output of the physical formula is the relevant simulation result. For example, if the current formula is P / (1 - density) * width * space * perimeter, and the next symbol given by the neural network is -width, then the formula becomes P / (1 - density) * width * space * perimeter - width. The new formula is also an expanded physical model, and this physical model can calculate SThickNT (the polished height of the NT structure) and SThickT (the polished height of the T structure). *width * space * perimeter - width can also be called a correction term.

[0111] Deep reinforcement learning itself is just a framework that describes a general form of model construction. In some embodiments of the present disclosure, this model is used to expand the formula. Thus, in combination with the characteristics of the present disclosure, the surface topography of the wafer after the CMP process is predicted using density, width, space, and CMP physical process parameters. Here, within the framework of reinforcement learning, the present disclosure defines that an action is a combination of relevant features and some process parameters with common mathematical symbols.

[0112] In some embodiments, a reward term can be generated based on the simulation values and measurement values of the surface height of each grid point, including: determining a first difference between the standard deviation of the simulation values and the standard deviation of the measurement values of the surface height of each grid point; determining a second difference between the RMSE of the simulation values and the RMSE of the measurement values of the surface height of each grid point; determining the range between the simulation values and the measurement values of the surface height of each grid point; and generating a reward term based on the first difference, the second difference, and the range. It should be noted that only the first difference and the second difference can be determined according to actual needs, without determining the extreme values. Thus, the reward term is generated only based on the first difference and the second difference. For example, the standard deviation and range of the surface height before and after CMP in the simulation data can be calculated simultaneously (here, two ranges and two standard deviations are calculated, respectively before and after simulation. The calculation before simulation is based on measurement data, and the calculation after simulation is based on the simulation result. The difference between the two ranges (specifically, the one before simulation minus the one after simulation) forms a variable, denoted as G0. Similarly, the difference between the standard deviations before and after simulation can also form a variable, denoted as G1). The range is the difference between the Max (maximum value) and the Min (minimum value). The range of the surface height before CMP in the simulation data refers to the difference between the maximum value and the minimum value among the data of each measurement point.

[0113] In one prediction, the reinforcement learning neural network gives an expansion term to supplement the original formula. So there are two formulas (the original formula and the supplemented formula), and each formula can perform simulation calculations on these several data points, that is, two simulation calculations. Each simulation calculation also generates an RMSE. After performing the simulation calculations of the original formula and the new formula, there are two RMSEs, denoted as rmse0 and rmse1, and the difference between them can be denoted as G2. The two ranges are denoted as max_diff0 and max_diff1 respectively, and the two standard deviations are std_error0 and std_error1 respectively. One reward (penalty) can be calculated from these 6 indicators.

[0114] Then G0, G1, and G2 constitute the reward for this prediction (this prediction is an action). The specific effect is whether the RMSE decreases (compared to the initial basic formula, because the formula derived from the neural network shows better fitting performance on the data), and whether G0 and G1 are relatively large positive numbers, because if G0 and G1 are relatively large positive numbers, it conforms to the actual situation of the CMP process. Simply put, the surface should be more "flat" after CMP polishing (the flat performance means that the standard deviation and range of the heights of different position points become smaller), and the larger the better.

[0115] It is hoped that the RMSE is as small as possible. Therefore, if (rmse0 - rmse1) is greater than 0, it proves that the extended formula has better accuracy and gives a reward (positive feedback), otherwise it is a punishment (negative feedback). This is the main reward and punishment mechanism (that is to say, there can be three goals, one main goal and two secondary goals. The main goal is to reduce the RMSE, and the secondary goals are to reduce the data difference before and after simulation. The secondary goals can also be regarded as "regularization terms". The significance of the regularization term is to avoid the phenomenon that the neural network completely ignores the objective reality in order to achieve the main goal. Usually, users do hope to obtain a formula with more accurate fitting, but it should be able to fit the objective characteristics and facts of the CMP process to a certain extent. This is also an innovation point of this disclosure, not only considering the fitting accuracy. Combining the actual process of CMP, a more reasonable and ideal result of the CMP process is to reduce the range and standard deviation of the data. That is to say, if (max_diff0 – max_diff1)>0, some additional rewards (positive feedback) are given, otherwise some punishments (negative feedback) are given, and the same applies to the standard deviation. Ultimately, it is hoped to find a new formula that can reduce the RMSE, range, and standard deviation. In the above embodiments, the RMSE, range, and standard deviation are used to form the reward. The embodiments of this disclosure are not limited to this and can be variously changed as needed. For example, the reward can also be formed only based on these two of the RMSE and the standard deviation.

[0116] In some embodiments, the specific reward formula is:

[0117] R = a * (rmse0–rmse1) + b * (max_diff0–max_diff1) + c * (std_error0–std_error1);

[0118] Let rmse0–rmse1 = G2, then R = a * G2 + b * G0 + c * G1;

[0119] Among them, a, b, and c are positive numbers between 0 and 1. The specific numbers can be adjusted, depending on which goal is expected. Generally, a is a relatively large value because the most core goal is to fit more accurately. Through the above process, all the elements required to generate a deep reinforcement learning neural network are the initial state S0, the action taken a0, the transferred state S1, and the corresponding reward R. Continuously let the agent (RNN neural network) make predictions, that is, many different <S0, a0, S1, R> tuples will be continuously generated. This tuple is the sample for training the agent (optimizing the parameters through backpropagation of the neural network). The tuple is obtained after the neural network makes a prediction once. After making several such predictions, several such tuples can be obtained.

[0120] Based on a obtained neural network, it can be trained based on a number of training samples. First, it is necessary to clarify how to train. Here, the definition of the loss function is proposed. Define the loss function as J(θ), where θ refers to the set of all hyperparameters in the entire RNN network, that is, a specific θ corresponds to a neural network with determined parameters. Here, the definition of the trajectory is introduced. The symbol of the trajectory is recorded as , and what is called a trajectory in this disclosure refers to a set composed of several <S0, a0, S1, R>. In this disclosure, it is necessary to determine what counts as a complete trajectory, or how many <S0, a0, S1, R> tuples should be included in a complete strategy. To avoid overfitting of the formula, the following two methods are adopted here to determine whether to stop making predictions (essentially also training) to determine the strategy, which are respectively:

[0121] Condition 1: Whether the number of predictions (the number of <S0, a0, S1, R>) reaches 10 times. If it reaches, stop. It should be understood that this value is exemplary and can be changed according to needs.

[0122] Condition 2: Whether the RMSE value of the current formula is less than 1. If it reaches, stop. Generally speaking, the unit for measuring the surface height is nm. A CMP formula is considered a relatively accurate model if its prediction accuracy of the surface height reaches an error within 2 - 5 nm.

[0123] If any of the above two conditions is met, it is considered that the termination can be carried out. That is, during initialization, make a prediction once, record and update <S0, a0, S1, R> in the database, continue the second prediction, continue to update and record the new <S0, a0, S1, R>, and so on. Continue like this until one of the above Conditions 1 or 2 is met and stop (whichever condition is met first is taken as the criterion). Thus, a complete strategy is obtained. That is, a complete strategy is obtained through prediction.

[0124] Generally speaking, the meaning of a policy is its Chinese meaning as a noun, which can be interpreted as "receiving an input A and generating a deterministic output B. Under a policy, no matter how many times the input A is executed, the result obtained must always be B and will not suddenly become C at a certain time". In some embodiments of the present disclosure, a neural network agent is created, and the framework and methods of reinforcement learning are used to continuously adjust and improve the parameters of the neural network. When the parameters of a neural network are all determined, receiving an input will definitely obtain a definite output, and the prediction results will not change after several such operations. This is the so-called policy. Therefore, it can be said that a policy refers to a neural network under certain parameters. The essence of reinforcement learning is to continuously adjust the parameters of the neural network. (Neural networks with different parameters may produce different outputs when facing the same input, so different parameters mean different policies). The essence is to continuously adjust the output obtained when facing the same input, that is, to adjust the policy. The R in reinforcement learning is used to evaluate the quality of the policy, whether a good decision has been made to bring benefits, or a bad decision has been made to bring punishment. A neural network usually contains several parameters, which can be represented by theta (a mathematical symbol, θ), so θ is usually used to represent the policy in the reinforcement learning framework).

[0125] In the actual calculation process, multiple trajectories are continuously sampled, and the length of each trajectory is T (that is, it is necessary to take T steps, from 0 to T, and there is a corresponding reward or punishment for each step. Therefore, to evaluate the comprehensive rewards and punishments of a trajectory, the sum of the rewards for each step of this trajectory needs to be considered). Neural network training and parameter update are performed every once in a while. This every once in a while is actually continuous sampling, that is, using the current neural network to continuously make predictions and generate several trajectories with a length of T steps.

[0126] It should be noted that in some embodiments of the present disclosure, there are actually two purposes. One is to expand the current formula so that the current formula can achieve the expected fitting accuracy (rmse meets the standard) for the existing real data, and the other is to train a stable neural network. It is not only used for this time (referring to being able to expand a formula that fits the current data), but also this stable neural network can give an extended result that meets the standard for other new data.

[0127] In some embodiments, determining a correction term to correct an initial formula using a deep learning neural network may include: in a training phase, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network; and updating the parameters of the neural network based on the gradient of the loss function. In some embodiments, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network may include: using a policy gradient algorithm in the field of reinforcement learning to perform backpropagation on the training samples to determine the gradient of the loss function.

[0128] According to the definition of the policy gradient algorithm, the training process is to calculate the policy gradient of the loss (objective) function to update the neural network parameters. The calculation of the policy gradient metric is as follows:

[0129] The derivative of the loss function J with respect to the parameter θ is:

[0130]

[0131] where E is the symbol for mathematical expectation, T refers to the length of the trajectory, t refers to a certain step in the trajectory, t = 0 is the initial of the trajectory, t = 1 is the first newly generated step of the trajectory, and R refers to the reward.

[0132] P in the above formula is not an independent letter, is a complete term, and the meaning of P is probability (Probability). The meaning of this term is the probability of the output trajectory (trace) under the current neural network parameters. The so-called trajectory is a term in the Markov decision process (MDP). Specifically, it is a continuous process that includes the initial state S0, the action a0 taken, the next state S1, the action a1 taken, another state S2, and the action a2. And so on, until a certain time T. As for how long the time T is, that is, how long the trajectory is, it is uncertain. In fact, the length of the formula is not infinite and there will be a specific maximum length limit. Of course, it varies for different manufacturers. Assuming the maximum length limit for formula expansion is set to 10, it means that the neural network can be expanded at most 10 times on the basis of the basic formula. If the set maximum expansion times are reached, it will stop. That is to say, the time T is 10, and a complete trajectory contains 10 Ss and corresponding actions.

[0133] The Log technique refers to the mathematical properties of logarithms. If y = x1 * x2, that is, there is a function where the independent variables x1, x2 and the dependent variable y have a multiplicative relationship. However, for the convenience of subsequent calculations and mathematical derivations, it is desired to transform it into an additive form. That is, the Log technique is used. Take the logarithm of both sides of the equation, log(y) = logx1 + logx2. And the mathematical properties of logarithms ensure that it does not change the monotonicity of the function before and after taking the logarithm. That is, if there are x1 and x2 that can make y reach the maximum value, then these x1 and x2 can also definitely make log(y) reach the maximum value.

[0134] Based on J(θ), the theoretical neural network parameter update formula is as follows: Assume that currently in the k-th training stage, α is the learning rate, which is a common parameter in neural networks. Then for the (k + 1)-th training, it can be expressed as follows:

[0135]

[0136] The above formula for calculating the gradient is a theoretical formula. In fact, it can be approximately replaced by the following formula:

[0137]

[0138] D is the set of collected trajectories:

[0139] The update of the parameter is from ,

[0140] to become ;

[0141] represents the gradient, and represents the estimated value. That is, represents the approximate calculation formula for the estimation of the gradient.

[0142] Reinforcement learning is a sub-branch field of deep learning. The policy gradient algorithm is a common classic algorithm in reinforcement learning. This formula and the transformation process, that is, the derivation process of the objective function of the policy gradient algorithm, are well-known and commonly used in the industry.

[0143] In some embodiments, a random probability threshold can be set. During iteration, the random probability threshold changes from an initial first probability threshold to a second probability threshold lower than the first probability threshold. That is, the random probability threshold takes the maximum value at the beginning of the iteration, and gradually becomes smaller as the iteration progresses. This will be further described later.

[0144] In some embodiments, throughout the entire process of neural network simulation, the number of operations is denoted as M, and the initial value of M is 0. That is, the random selection probability is e, and the initial value of e is 1.0. During this process, training samples for the reinforcement learning neural network are continuously generated, that is, several data combinations <S0, a0, S1, R> with different results are continuously produced. It is hoped to produce results in various different situations as much as possible, so that the trained reinforcement learning neural network will be more powerful. However, for a neural network, whether it has been trained or not, a given input will definitely obtain a definite output. When the network parameters remain unchanged, the same input leads to the same output. It is hoped that the samples cover all possible inputs that the neural network continuously encounters, thereby obtaining all possible outputs. In addition, if a random selection is made each time, it is possible that many of the choices made are not good choices. In fact, many possible combinations are not good combinations. If it is completely random, a large amount of time will be wasted in generating these bad combinations. The desired final effect is that the neural network continuously receives inputs for prediction to obtain an output. It should have the experience learned before, filter out many completely unnecessary attempts, and also maintain a certain degree of randomness, so as not to be completely rigid (unable to break out of limited paths and unable to generate new samples).

[0145] That is to say, there are two ways to select an action (action). One is to randomly select one from the symbol library, and the other is to predict one using the current neural network. The former has no experience and is purely random, while the latter will tend to give actions that make R (reward) larger as the neural network is gradually trained. Generally speaking, the first random selection is "irrational", and the second using neural network prediction is "rational" (especially as the neural network is gradually trained and the neural network becomes more and more "intelligent"). Since the training of the neural network is essentially a mathematical optimization problem, a common possible situation in mathematical optimization problems is encountering local optima, and retaining randomness can break this local optimum dilemma.

[0146] In some of the above embodiments, e will continuously decrease. Because at the beginning, the neural network also randomly generates parameters, so the results given by the neural network can also be said to be "random selections" to some extent. There is no essential difference between the two. However, as the training progresses, the neural network gradually becomes smarter, e gradually decreases but does not completely become 0, which exactly conforms to the previous explanation.

[0147] The meaning of M is the number of operations, that is, the cumulative number of times of performing input, output, generating data combinations, and saving these operations. Its starting value e is 1, which means that at the beginning, random selection is always carried out, so as to ensure the generation of a large number of new combinations and a large number of different samples. However, as the operations proceed, the neural network will take samples from the sample library for training every once in a while. The neural network will gradually become somewhat tendentious and intelligent, and will select the formula that can obtain better prediction results as the output. On this basis, it is necessary to derive formulas. It is not necessary to continue wasting time trying to derive formulas that do not conform well at the beginning. Therefore, e will continue to decrease. For example, at the beginning, it is not known which combinations are good, so all combinations are selected for trial. After several rounds of trials, a general concept can be obtained. Some combinations are relatively good and have "potential", while some combinations are very poor and there is no need to continue trying. Therefore, in order to improve efficiency, as the trials proceed, combinations with no potential can be directly excluded without wasting time on derivation and evaluation on this basis. At the same time, for those good combinations with potential, it is possible to break away from the original idea and try a new derivation at a certain step to obtain a result better than before. Therefore, a certain degree of randomness will still be maintained to avoid falling into the dilemma of local optimum.

[0148] The above described how to actually update the network parameters. Assume the total number of times is M. Assume that now gradient descent is performed every 200 times to update the network parameters, and assume that the step size of a trajectory is specified as 20, that is, each time the neural network samples (200 / 10) = 10 trajectories, and every time 10 trajectories are sampled, gradient descent is used to update the network parameters using the results just sampled. M is the total number of times (since it is not certain when the training is in place, the value of M cannot be determined). This e is related to M. At the beginning, e is set to 1. As M continues to increase, e gradually decreases. Assume that after M times, the neural network is trained to a certain extent. If the standard has been reached, then the training is completed. In addition, a maxM can be set, assumed to be 10000. What if M reaches 10000 but has not reached the described stable state? The state at this time must be that the loss function curve oscillates and cannot be stabilized, which indicates that new samples with random selection need to be generated to break the "deadlock" (oscillation usually indicates that the sampled trajectories are not comprehensive enough and the information contained is not enough). So the above process can be repeated again, reset M, reset e, and then continue to sample to generate trajectories (this is more likely to generate random new samples) and conduct training.

[0149] In some embodiments, determining a correction term to correct an initial formula using a deep learning neural network may include: generating a random number at the beginning of each iteration of the neural network, where the range of the random number may be from 0 to 1, and the present disclosure is not limited thereto and may vary according to needs; selecting a source of the correction term for correcting the initial formula based on a comparison between the random number and a random probability threshold, that is, selecting whether to generate a correction term through the prediction of the neural network or selecting a symbol from a symbol library for correction; and determining an updated correction term using the neural network based on the selection result to correct the initial formula.

[0150] In some embodiments, determining an updated correction term using the neural network based on the selection to correct the initial formula may include: when the comparison between the random number and the random probability threshold is a first comparison result (for example, the random number p is less than the random threshold probability e), selecting a symbol from the symbol library as the correction term to correct the initial formula using the neural network, where the symbol library is generated by performing a Cartesian product combination on a mathematical symbol and a variable library, and the variable library contains graphic feature information and process information; when the comparison between the random number and the random probability threshold is a second comparison result (for example, the random number p is greater than the random threshold probability e), predicting using the neural network to generate an initial correction term to correct the initial formula using the neural network model to generate a new correction term, and adding the new correction term to the symbol library.

[0151] In some embodiments, the neural network and related variables may be initialized, and the neural network may be used to perform predictions. A random number p may be generated before the prediction. As mentioned above, if p > e, the neural network is used for prediction; otherwise, random sampling is performed from the symbol library. And <S0,a0,S1,R> is obtained, and the process is repeated to obtain a trajectory, and the trajectory is recorded in the database. The above process is repeated using the neural network to obtain other different trajectories. In this process, for example, every 10 times, 70% of the trajectories are randomly selected from the data for backpropagation calculation (according to the gradient calculation formula described above), and the parameters of the neural network are updated. At the same time, the loss value of the loss function may be recorded. Additionally, when M gradually increases, e is also gradually decreased (for example, the minimum is 0.1, and it may vary according to needs). The specific change formula is:

[0152] e = e0 – tanh(M) + 0.1;

[0153] Among them, tanh is a common function. For example, tanh = f(x), where x is a variable. When x gradually increases starting from 0, f(x) gradually increases starting from 0 and finally approaches 1 infinitely. Specifically, tanh is a common activation function used in the connection between layers of a neural network. Each layer of the neural network has many neurons. The essence of a neuron is a combination of several parameters. Each layer calculates based on the input of this layer plus the parameters contained in all the neurons in this layer. The result of the calculation is input into the activation function of this layer, and the result output by the activation function is output to the next layer as the input of the next layer. The formula of the tanh activation function is:

[0154]

[0155] Where x represents a variable, and in some embodiments of the present disclosure, it may represent the number of iterations.

[0156] In some embodiments, as mentioned above, the value of the loss function is also determined to judge whether continuous training can be ended. When the value of the loss function is lower than a predetermined threshold, it indicates that a reliable neural network has been obtained and training can be ended. At this time, the correction term output by the neural network corresponding to the value of the loss function can be determined as the correction term related to the grinding rate; and an update formula can be determined based on the determined correction term to determine the height value of the surface.

[0157] In some embodiments, the iteration in the training phase can stop when the following conditions are met: the difference between the values of the loss function in two consecutive iterations is less than a predetermined difference threshold; the absolute value of the ratio of the values of the loss function in two consecutive iterations is greater than a predetermined ratio threshold; or the number of iterations or the iteration time reaches the maximum iteration threshold.

[0158] Generally speaking, when the loss curve gradually approaches a steady state or reaches the set maximum number of iterations, training is stopped to obtain a trained neural network.

[0159] The loss curve is a curve plotted by calculating the value of the loss function during each training. Assuming 100 times of training (referring to the operation of extracting several samples to calculate the gradient and update the network parameters), a value of the loss function will be obtained each time. For example, after 100 times, 100 values of the loss function can be obtained. Then a curve can be plotted. The horizontal axis is the number of training times, and the vertical axis is the loss value. As training progresses, the function value of the loss will continuously decrease and finally reach a steady state. The theoretical basis for the loss curve to reach a steady state is that the labeled data is a finite number and will not keep increasing, so the amount of information contained in this data must be fixed. The training of the neural network is essentially to continuously use this information to adjust its own network parameters. That is to say, there will always be a time when learning is completed, that is, there will always be a time when no new knowledge can be learned.

[0160] The neural network of the embodiments of the present disclosure can make several steps of predictions based on the CMP basic physical formula to obtain a final enhanced CMP physical formula. It should be understood that the trajectories at intervals of 10 times and 70% mentioned here are only exemplary, and the present disclosure is not limited thereto, and various changes can be made according to needs. For example, the common range can be from 10 to 100 times, and it can be between 50% and 80%.

[0161] Backpropagation calculation is a well-known calculation for neural network training in the industry. Generally speaking, as long as it is the training of a neural network, backpropagation calculation is used for training. Specifically, backpropagation refers to calculating the gradient using samples and updating the parameters of the neural network using the gradient. The purpose of this method is to change and adjust the parameters of the neural network. The difference between a randomized neural network and a trained neural network lies in the different parameters. For the same network structure, the results predicted by different parameters are completely different.

[0162] In some embodiments of the present disclosure, the neural network can accept a formula as input and obtain a corresponding output. The output result is a mathematical symbol (but not limited to addition, subtraction, multiplication, and division. In fact, it is a combination of symbols and other terms (such as variables)). The predicted mathematical symbol (new action) is written behind the current formula to obtain a new formula. Repeat the above steps to train the neural network, so the neural network continuously expands the initial formula and finally obtains a new formula.

[0163] Regarding the specific architecture of the RNN neural network in the present disclosure, as Figure 4 shown. Figure 4 A schematic diagram of a neural network architecture according to some embodiments of the present disclosure is shown.

[0164] As Figure 4 shown: This neural network is of the NV1 structure, that is, multiple inputs and one output. And there are also 2 fully connected layers, one convolutional layer, and one pooling layer between the neurons h1, h2, and h3. The number of neurons in the fully connected layer is 128. It should be understood that the RNN network structure described here is only schematic and different structures can be adopted according to actual needs.

[0165] Each neuron corresponds to a different input, and its expression is:

[0166]

[0167] where x1 - x3 refer to the inputs. For example, x1 is GDS feature information, x2 is process information, and x1 and x2 are constant values; x3 represents the formula of the previous step.

[0168] U, W, b, and V are all neural network parameters, and their meanings are well-known. These are the ones that will be corrected and changed through continuous training, and all these network parameters are random values at the beginning. t represents the current moment. In some embodiments, U is the weight matrix from the input layer to the hidden layer; V is the weight matrix from the hidden layer to the output layer; W is the weight of the previous value of the hidden layer as the input for this time. N represents the Nth layer. The structure of the neural network is composed of many layers to form a neural network, and each layer has many neurons.

[0169] Figure 4 In it, the formula for h1 is expressed as: the output result of the first layer of the hidden layer is equal to the function U(x1) plus the subsequent terms. U is a composite function because there are several neurons in one layer of the neural network, and each neuron has parameters. It is impossible to write out the formula for each neuron when listing the formula. The final input of this layer is the result obtained by comprehensively calculating all neurons together. Therefore, a function symbol U is used to represent this comprehensive function.

[0170] The training of the neural network requires many tuples like <s0, a, s1, r>, where s0 represents the current formula, which is the input of the neural network, that is, Figure 4 h0 in it. Actually, when recording, in addition to recording a_t (the action a taken at time t), the action a_t - 1 taken at the previous moment will also be recorded. Assume that the action of a_t - 1 is to multiply by space. Then, conversely, the formula at the previous moment of S0 is equal to the formula at S0 divided by space, and it can be deduced by inversion.

[0171] The formula here actually reflects the core characteristic of the neural network with the RNN structure. Its main meaning is that the output y at a certain moment also depends on the inputs at the previous t - 1 moments. N represents how many hidden layers are constructed to remember the previous moments. For example, when N is equal to 5, it means that 5 hidden layers need to be constructed to remember. Obviously, the larger N is, the more is remembered, but at the same time, the network structure becomes more complex, the training calculation amount becomes larger, and the training difficulty becomes greater. Therefore, the value of N is set according to the specific situation. T can actually be understood as equal to N. Writing T is just to accurately convey the core characteristic of the neural network with the RNN structure, that is, to consider the previous output. The most ideal situation is to consider the outputs of all previous moments. However, since N is a definite number, actually what is calculated here is to consider the outputs of these N time steps from T - N to T. Here, the time step refers to making a prediction using the neural network (this time step is different from the time step involved in calculating the surface height previously). Making a prediction is one time step. Assume N = 5, which means that when the neural network makes a prediction each time, it will find the results of the previous five predictions and input them into the corresponding hidden layer for calculation to obtain the final output of the current step.

[0172] As can be seen from the above description, the process of making the agent (RNN neural network) make predictions in some embodiments can be generally described as follows:

[0173] First, there is a dataset, which is usually provided by the users who use the model (i.e., those semiconductor manufacturing factories in reality). The user will give the GDS features (density, width, etc.), process parameter settings of each data point, and the actual ground results of each collected point. This is what is usually referred to as the "label" in deep learning (machine learning). After having the labels, a deep reinforcement learning neural network can be constructed using the label data. The network parameters (weights) of this neural network are initially random. The neural network is set to receive an initial formula as input and output the next mathematical symbol. Thus, the training begins. Each time the neural network generates an output, this output plus the input constitutes a new formula, which can calculate the ground surface height. Then, the calculated ground height is compared and calculated with the label data to obtain an error value. Similarly, the old formula can also calculate the ground height to obtain the error value with the label (which is used to measure the entire dataset). Since the goal is to make the formula obtain a good fitting accuracy at each point as much as possible, then each point should be considered. That is, when calculating the error of the formula, not only a certain point is calculated, but all points in the label data are calculated. Therefore, the above-mentioned metrics, such as the range and standard deviation, are statistical metrics for a group of samples rather than for a certain data point. It is necessary to determine whether the error value has decreased or increased due to the addition of a new mathematical symbol (that is, whether the new formula given by the neural network is good or bad), so as to obtain a reward or punishment (positive feedback or negative feedback), and then there is a data combination. This combination includes the following contents: the current formula (as the input of the neural network, labeled as the current state S0), the behavior or action taken. The behavior taken is the output of the neural network, which is a mathematical symbol, labeled as a), the new formula (that is, the new formula obtained by adding the output to the old formula, labeled as S1), and the feedback R obtained. If the new formula fits the data better, it is positive feedback, otherwise it is negative feedback. Then such a group of data combinations are the training samples of the deep reinforcement learning neural network. Many attempts can be made, and thus many such data combinations are obtained. Such data combinations are the samples used to train the deep reinforcement learning neural network. These samples can be provided to the deep reinforcement learning neural network to train the deep reinforcement learning neural network and continuously adjust the network parameters (weights). When the loss curve is stable, the training is completed.

[0174] In the above embodiments, a deep learning neural network is mainly used as an example for illustration. It should be understood that the methods of the embodiments of the present disclosure are not limited thereto, and various changes can be made according to actual needs to determine the correction terms, and then an enhanced CMP physical formula can be obtained to determine the surface height value.

[0175] At block 206, in response to the difference satisfying a predetermined condition, the simulation value corresponding to the correction term is determined as the height value of the surface of the polished wafer.

[0176] As mentioned before, the neural network can be continuously trained to obtain a trained neural network. The trained neural network can perform several steps of prediction based on the CMP basic physical formula to obtain a final enhanced CMP physical formula. Thus, the formula can be used to accurately simulate the height of the wafer surface, and the simulation value can be used as the height value of the wafer surface.

[0177] In some embodiments, the predetermined condition may refer to satisfying at least one of the following: being less than a predetermined threshold; the number of iterations or the iteration time reaching a predetermined threshold; or the ratio between the values of two consecutive iterations being greater than a predetermined threshold, etc. The embodiments of the present disclosure are not limited thereto, and other predetermined conditions can be adopted according to needs.

[0178] The present disclosure also provides a simulation model configured to execute the method according to the above embodiments.

[0179] The present disclosure also provides a neural network, including: a plurality of input layers configured to receive the feature information of the patterns on the wafer, process information, and an initial formula for calculating the surface height of the polished wafer; an intermediate layer configured to execute the method according to the embodiments of the present disclosure to determine an updated formula for calculating the simulation value of the surface height; and an output layer configured to output the updated formula.

[0180] Embodiments of the present disclosure also disclose an electronic device. The electronic device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, and the instructions, when executed by the processor, cause the device to perform actions, the actions including: determining a simulation value of the surface height of the polished wafer based on a time step and a corresponding polishing rate; determining a correction term related to the polishing rate based on a comparison between the simulation value and a target value so that the difference between the simulation value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulation value corresponding to the correction term as the height value of the surface of the polished wafer.

[0181] Embodiments of the present disclosure also disclose a computer-readable storage medium having a computer program stored thereon, and the program, when executed by a processor, implements the method for semiconductor device simulation according to the embodiments of the present disclosure.

[0182] Embodiments of the present disclosure also disclose a neural network model. This neural network model can be obtained by the training method of the neural network model in the foregoing embodiments.

[0183] Some embodiments of the present disclosure provide a method for semiconductor device simulation. It should be noted that the examples given in the above embodiments are only for illustrating the solutions of the embodiments of the present disclosure and do not limit the solutions of the present disclosure.

[0184] Some embodiments of the present disclosure introduce a method of symbolic regression to expand formulas and models. Some embodiments of the present disclosure introduce the current state-of-the-art deep reinforcement learning model for solving and optimizing symbolic regression. Some embodiments of the present disclosure are based on expanding the basic CMP physical formula rather than starting from scratch, and consider the complexity of the formula and the fitting degree to the process, which is feasible and ensures the rationality of the theory. Some disclosed embodiments derive and expand the physical formula based on the current deep reinforcement learning and symbolic regression techniques, which can better improve the accuracy and prediction fitting ability of the model. Symbolic regression technology refers to constructing a model (here a reinforcement learning neural network) to predict mathematical symbols.

[0185] The model of the embodiments of the present disclosure can actually be regarded as a flexible and variable "engineer", which can innovatively and flexibly respond to different processes in actual situations from the most basic and error-free CMP process perspective, and create models suitable for their respective situations.

[0186] Some embodiments of the present disclosure extend and expand the original physical model, so the new model still has physical meaning. Some embodiments of the present disclosure can intelligently learn when facing different processes and different measurement data. This learning is based on the most general CMP basic physical formula, and considers the model complexity and the process fitting degree, and can give the most suitable CMP physical model for various processes.

[0187] It should be understood that the embodiments shown in the drawings are only for schematically showing the solutions of some embodiments of the present disclosure and do not limit the present disclosure. The embodiments of the present disclosure can also have various other forms.

[0188] Figure 5A schematic block diagram of an electronic device according to some exemplary embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0189] As Figure 5 shown, the device 500 includes a CPU 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0190] A plurality of components in the device 500 are connected to the I / O interface 505. The plurality of components include: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0191] Each of the processes and processes described above, such as the method 200, can be executed by the CPU 501. For example, in some embodiments, the method 200 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more steps of the method 200 described above can be executed.

[0192] Solutions according to embodiments of the present disclosure may be methods, apparatuses, systems, and / or computer program products. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure. The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable program instructions may be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network.

[0193] The embodiments of the present disclosure have been described above. The above description is exemplary and only represents alternative embodiments of the present disclosure, which is not exhaustive and does not limit the present disclosure. Although the claims in this application have been formulated for specific combinations of features, it should be understood that the scope of the present disclosure also includes any novel feature or any novel combination of features that are explicitly or implicitly disclosed herein or any generalization thereof, regardless of whether it relates to the same solution in any of the currently claimed rights. The applicant hereby notifies that new claims may be formulated for these features and / or combinations of these features during the examination of this application or in any further application derived therefrom.

[0194] The choice of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary technicians in the technical field to understand the embodiments disclosed herein. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A simulation method for a semiconductor device, comprising: Determine a simulated value of a surface height of the wafer after grinding based on the time step and the corresponding grinding rate; Determining a correction term related to the grinding rate based on a comparison between the simulation value and the target value so that the difference between the simulation value corresponding to the correction term and the target value becomes smaller, specifically comprising: determining the correction term based on graphic feature information, process information, and information of an initial formula for calculating the surface height using a deep learning neural network to correct the initial formula; as well as In response to the difference satisfying a predetermined condition, a simulation value corresponding to the correction term is determined as a height value of the surface of the ground wafer.

2. The method according to claim 1, wherein the correction term comprises at least a polynomial related to one of the following: The width of the pattern within the grid points on the wafer; The spacing between the graphics; and The sum of the perimeters of the figures within the grid.

3. The method according to claim 1, wherein determining the correction term using a deep learning neural network to correct the initial formula comprises: generating a random number at the beginning of each iteration of the neural network; Selecting a source of a correction term for correcting the initial formula based on a comparison of the random number with a random probability threshold; as well as Based on the selection, an updated correction term is determined using the neural network to correct the initial formula.

4. The method of claim 3, wherein determining an updated correction term based on the selection using the neural network to correct the initial formula comprises: In response to a first comparison result between the random number and the random probability threshold, selecting a symbol from a symbol library as a correction item to correct the initial formula using the neural network, wherein the symbol library is generated by combining mathematical symbols and a variable library by performing a Cartesian product, and the variable library includes the graphic feature information and the process information; and In response to a second comparison result between the random number and the random probability threshold, the neural network is used to make a prediction to generate an initial correction term, the initial formula is corrected using the neural network to generate a new correction term, and the new correction term is added to the symbol library.

5. The method according to claim 1, wherein determining the correction term using a deep learning neural network to correct the initial formula comprises: Generate reward items based on simulated and measured values ​​of surface height at each grid point on the wafer; as well as The array consisting of the initial formula, the correction term, the update formula, and the reward term is used as an element of the training sample and iterated in the neural network to generate new elements of the training sample.

6. The method according to claim 5, wherein generating a reward item based on the simulated value and the measured value of the surface height of each grid point on the wafer comprises: Determine a first difference between a standard deviation of a simulated value and a standard deviation of a target value of a surface height at each grid point; Determine a second difference between a root mean square error of a simulated value of the surface height at each grid point and a root mean square error of a target value; Determine the range between the simulated value and the target value of the surface height at each grid point; as well as The reward item is generated based on the first difference, the second difference, and the range.

7. The method of claim 5, wherein the iterations include a prediction phase and a training phase, and wherein: In response to an element generated in each round of iteration of the prediction phase satisfying a first predetermined condition, terminating the round of iteration and treating all elements in the round of iteration as a trajectory; as well as In response to the number of total trajectories obtained during each round of iterations reaching a predetermined trajectory threshold, the prediction phase is terminated.

8. The method according to claim 7, wherein the first predetermined condition comprises: The number of elements generated in each round of iteration reaches a first predetermined threshold; or The root mean square error in the current iteration is less than a second predetermined threshold.

9. The method according to claim 7, wherein determining the correction term using a deep learning neural network to correct the initial formula comprises: In the training phase, the neural network is trained using the training samples in the trajectory to determine the gradient of the loss function of the neural network; as well as The parameters of the neural network are updated based on the gradient of the loss function.

10. The method of claim 9, wherein training the neural network using the training samples in the trajectory to determine the gradient of the loss function of the neural network comprises: Using a policy gradient algorithm in the field of reinforcement learning, backpropagation is performed on the training samples to determine the gradient of the loss function.

11. The method according to claim 9, wherein determining the correction term by using a deep learning neural network to correct the initial formula further comprises: Determining a value of the loss function; In response to the value of the loss function being lower than a predetermined threshold, determining a correction term output by the neural network corresponding to the value of the loss function as a correction term associated with the grinding rate; as well as An updated formula is determined based on the determined correction term to determine a height value of the surface.

12. The method according to claim 9, wherein the iteration of the training phase stops when the following conditions are met: The difference between the values ​​of the loss function of the two previous and subsequent iterations is less than a predetermined difference threshold; The absolute value of the ratio of the loss function values ​​of the two previous and subsequent iterations is greater than a predetermined ratio threshold; or The number of iterations or iteration time reaches the maximum iteration threshold.

13. The method according to claim 3 or 4, wherein the random probability threshold changes from an initial first probability threshold to a second probability threshold lower than the first probability threshold during the iteration.

14. A simulation model configured to perform the method according to any one of claims 1 to 13.

15. A semiconductor device simulation model based on a neural network, comprising: A plurality of input layers configured to receive feature information of patterns on a wafer, process information, and an initial formula for calculating a surface height of the wafer after grinding; an intermediate layer configured to execute the method of any one of claims 1 to 13 to determine an update formula for calculating a simulation value of the surface height; as well as The output layer is configured to output the update formula.

16. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, the instructions causing the device to perform actions when executed by the processor, the actions comprising: Determine a simulated value of a surface height of the wafer after grinding based on the time step and the corresponding grinding rate; Determining a correction term related to the grinding rate based on a comparison between the simulation value and the target value so that the difference between the simulation value corresponding to the correction term and the target value becomes smaller, specifically comprising: determining the correction term based on graphic feature information, process information, and information of an initial formula for calculating the surface height using a deep learning neural network to correct the initial formula; and In response to the difference satisfying a predetermined condition, a simulation value corresponding to the correction term is determined as a height value of the surface of the ground wafer.

17. The electronic device according to claim 16, wherein the correction term comprises at least a polynomial related to one of the following: The width of the pattern within the grid points on the wafer; The spacing between the graphics; and The sum of the perimeters of the figures within the grid.

18. The electronic device according to claim 16, wherein determining the correction term by using a deep learning neural network to correct the initial formula comprises: generating a random number at the beginning of each iteration of the neural network; Selecting a source of a correction term for correcting the initial formula based on a comparison of the random number with a random probability threshold; as well as Based on the selection, an updated correction term is determined using the neural network to correct the initial formula.

19. The electronic device of claim 18, wherein determining an updated correction term based on the selection using the neural network to correct the initial formula comprises: In response to a first comparison result between the random number and the random probability threshold, selecting a symbol from a symbol library as a correction item to correct the initial formula using the neural network, wherein the symbol library is generated by combining mathematical symbols and a variable library by performing a Cartesian product, and the variable library includes the graphic feature information and the process information; and In response to a second comparison result between the random number and the random probability threshold, the neural network is used to make a prediction to generate an initial correction term, the initial formula is corrected using the neural network to generate a new correction term, and the new correction term is added to the symbol library.

20. A computer-readable storage medium having machine-executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Interval type index forecasting method based on robust interval extreme learning machine

    CN104537167A

  • Method for optimizing timing control process parameters in chemical mechanical polishing

    TWI221435B