Neural network model training method, electronic equipment and storage medium
By using deep learning neural network to adjust the initial formula of the CMP model in the simulation method of semiconductor devices, the problem of insufficient prediction accuracy of metal layer thickness in the CMP back-stage process in the prior art is solved, and higher simulation accuracy and prediction fitting capabilities are achieved.
Patent Information
- Application Number
- CN202510248126.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
When the prior art predicts the thickness of the metal layer after the chemical mechanical polishing (CMP) process, the modeling accuracy is insufficient and it is impossible to accurately predict the situation after grinding.
A simulation method of semiconductor devices is adopted to determine the simulation value of wafer surface height based on time steps and grinding rate, and to determine correction terms related to grinding rate using a deep learning neural network, and adjust the initial formula to improve simulation accuracy.
It significantly improves the simulation accuracy and predictive fitting ability of the simulation model, and can more accurately predict the thickness of the metal layer after CMP grinding.
Smart Images

Figure CN120180890A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure mainly relate to integrated circuits, and more particularly, to a simulation method, an electronic device, and a storage medium for semiconductor devices. Background Art
[0002] Currently, the prediction of the metal layer thickness after the back-end process of Chemical Mechanical Polishing (CMP) typically depends heavily on the CMP model.
[0003] In traditional solutions, a pure physical model is derived through intuitive physical mechanisms. However, in the face of certain complex situations and process conditions, the modeling accuracy is somewhat insufficient, making it impossible to accurately predict the situation after CMP grinding. Summary of the Invention
[0004] According to an exemplary embodiment of the present disclosure, a simulation solution for semiconductor devices and a training solution for a neural network model are provided to at least partially overcome the above or other potential defects.
[0005] According to one aspect of the present disclosure, a simulation method for a semiconductor device is provided. The method includes: determining a simulated value of the surface height of a wafer after grinding based on a time step and a corresponding grinding rate; determining a correction term related to the grinding rate based on a comparison between the simulated value and a target value, so that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the surface of the wafer after grinding.
[0006] In a second aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the device to perform operations, the operations including: determining a simulated value of the surface height of a wafer after grinding based on a time step and a corresponding grinding rate; determining a correction term related to the grinding rate based on a comparison between the simulated value and a target value, so that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the surface of the wafer after grinding.
[0007] In some embodiments, determining a correction term related to the grinding rate based on a comparison between the simulated value and the target value includes: determining, using a deep learning neural network, a correction term based on graphic feature information, process information, and information of an initial formula for calculating the surface height to correct the initial formula.
[0008] In some embodiments, the correction term at least includes a polynomial related to at least one of the following: the width of the pattern within the grid points on the wafer; the spacing between patterns; and the sum of the perimeters of the patterns within the grid points.
[0009] In some embodiments, determining the correction term using a deep learning neural network to correct the initial formula includes: generating a random number at the beginning of each iteration of the neural network; selecting the source of the correction term for correcting the initial formula based on the comparison of the random number with a random probability threshold; and determining an updated correction term using the neural network based on the selection to correct the initial formula.
[0010] In some embodiments, determining an updated correction term using the neural network based on the selection to correct the initial formula includes: in response to a first comparison result of the random number with the random probability threshold, selecting a symbol from a symbol library as the correction term to correct the initial formula using the neural network, where the symbol library is generated by taking the Cartesian product combination of mathematical symbols and a variable library, and the variable library contains graphic feature information and process information; and in response to a second comparison result of the random number with the random probability threshold, making a prediction using the neural network to generate a new correction term and adding the new correction term to the symbol library.
[0011] In some embodiments, determining the correction term using a deep learning neural network to correct the initial formula includes: generating a reward term based on the simulated value and the measured value of the surface height of each grid point; and iterating the array composed of the initial formula, the correction term, the updated formula, and the reward term as elements of the training sample in the neural network to generate elements of a new training sample.
[0012] In some embodiments, generating a reward term based on the simulated value and the measured value of the surface height of each grid point includes: determining a first difference between the standard deviation of the simulated value of the surface height of each grid point and the standard deviation of the measured value; determining a second difference between the RMSE of the simulated value of the surface height of each grid point and the RMSE of the measured value; determining the range between the simulated value and the measured value of the surface height of each grid point; and generating a reward term based on the first difference, the second difference, and the range.
[0013] In some embodiments, the iteration includes a prediction phase and a training phase, and wherein: in response to the elements generated in each round of iteration in the prediction phase satisfying a first predetermined condition, terminating the round of iteration and taking all the elements in the round of iteration as a trajectory; and in response to the number of total trajectories obtained during each round of iteration reaching a predetermined trajectory threshold, terminating the prediction phase.
[0014] In some embodiments, the first predetermined condition includes: the number of elements generated in each round of iteration reaching a first predetermined threshold; or the value of RMSE in the current iteration being less than a second predetermined threshold.
[0015] In some embodiments, determining a correction term using a deep learning neural network to correct an initial formula includes: in a training phase, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network; and updating the parameters of the neural network based on the gradient of the loss function.
[0016] In some embodiments, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network includes: using a policy gradient algorithm in the field of reinforcement learning to perform backpropagation on the training samples to determine the gradient of the loss function.
[0017] In some embodiments, determining a correction term using a deep learning neural network to correct an initial formula further includes: determining the value of the loss function; in response to the value of the loss function being lower than a predetermined threshold, determining the correction term output by the neural network corresponding to the value of the loss function as a correction term related to the polishing rate; and determining an updated formula based on the determined correction term to determine the height value of the surface.
[0018] In some embodiments, the iteration in the training phase stops when the following conditions are met: the difference between the values of the loss function in two consecutive iterations is less than a predetermined difference threshold; the absolute value of the ratio of the values of the loss function in two consecutive iterations is greater than a predetermined ratio threshold; or the number of iterations or the iteration time reaches a maximum iteration threshold.
[0019] In some embodiments, the random probability threshold changes from an initial first probability threshold to a second probability threshold lower than the first probability threshold during the iteration.
[0020] In a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to the first aspect of the present disclosure.
[0021] In a fourth aspect of the present disclosure, a simulation model is provided, configured to execute the method according to the first aspect of the present disclosure.
[0022] In a fifth aspect of the present disclosure, a neural network is provided, including: a plurality of input layers, configured to receive feature information of patterns on a wafer, process information, and an initial formula for calculating the surface height of the polished wafer; an intermediate layer, configured to execute the method according to the first aspect of the present disclosure to determine an updated formula for calculating the simulated value of the surface height; and an output layer, configured to output the updated formula.
[0023] In a fifth aspect of the present disclosure, a method for training a neural network model is provided, including: training a neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated based on an initial formula for determining the surface height of a polished wafer; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the respective updated parameters as the trained neural network model, where the trained neural network model outputs an enhanced formula for determining the surface height.
[0024] In some embodiments, each training sample is an array composed of: an initial formula; an initial correction term for correcting the initial formula; an updated formula generated based on the initial formula and the initial correction term; and a reward term representing a feedback value generated based on the difference between the simulated value and the target value of the surface height.
[0025] In some embodiments, the reward term is generated by: determining a first difference between the standard deviation of the simulated values of the surface height at each lattice point on the wafer and the standard deviation of the target value; determining a second difference between the root mean square error of the simulated values of the surface height at each lattice point and the root mean square error of the target value; determining the range between the simulated value and the target value of the surface height at each lattice point; and generating the reward term based on the first difference, the second difference, and the range.
[0026] In some embodiments, the neural network model is a reinforcement learning neural network model, and training the neural network model using training samples to determine the gradient of the loss function of the neural network model includes: performing backpropagation on the training samples using a policy gradient algorithm in the field of reinforcement learning to determine the gradient of the loss function.
[0027] In some embodiments, updating the parameters of the neural network model based on the gradient of the loss function includes: in response to a predetermined number of training times, performing gradient update on the gradient of the loss function; and updating the parameters of the neural network model based on the gradient update.
[0028] In some embodiments, it further includes: determining the loss function value of the neural network model; and where the training satisfying the predetermined condition includes: the difference between the loss function values of two consecutive iterations is less than a predetermined difference threshold; or the absolute value of the ratio of the loss function values of two consecutive iterations is greater than a predetermined ratio threshold.
[0029] In a sixth aspect of the present disclosure, a method for generating training samples is provided, including: predicting based on an initial formula and an initial correction term to generate an updated formula and a reward term, where the initial formula is used to determine the surface height of the polished wafer, the initial correction term is used to correct the initial formula, and the reward term represents a feedback value generated based on the difference between the simulated value and the target value of the surface height; forming an array consisting of the initial formula, the initial correction term, the updated formula, and the reward term as a first training sample; and iterating in a neural network model based on the first training sample to generate a plurality of training samples.
[0030] In some embodiments, in response to the samples generated in each iteration satisfying a predetermined condition, terminating the iteration, and taking all the samples in that iteration as a trajectory; and in response to the number of total trajectories obtained during each iteration process reaching a predetermined trajectory threshold, terminating the iteration; and taking the samples in the total trajectories as samples for training the neural network model.
[0031] In some embodiments, predicting based on the initial formula and the initial correction term to generate an updated formula includes: generating a random number at the beginning of each iteration; selecting a source of the initial correction term for correcting the initial formula based on the comparison between the random number and a random probability threshold; and correcting the initial formula based on the selection to generate an updated formula.
[0032] In some embodiments, it further includes: performing a Cartesian product combination on mathematical symbols and a variable library to generate a symbol library, where the variable library contains graphic feature information and process information.
[0033] In some embodiments, correcting the initial formula based on the selection to generate an updated formula includes: in response to a first comparison result between the random number and the random probability threshold, selecting a symbol from the symbol library as the initial correction term to correct the initial formula using the neural network model; and in response to a second comparison result between the random number and the random probability threshold, using the neural network model to make a prediction to generate a predicted initial correction term to correct the initial formula, and adding the predicted initial correction term to the symbol library.
[0034] In a seventh aspect of the present disclosure, a neural network model generated according to the method of the fifth aspect is provided.
[0035] It will be understood from the following description that the technical solution of the present disclosure can significantly improve the simulation accuracy and prediction fitting ability of the simulation model.
[0036] The Summary of the Invention section is provided to introduce a selection of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify the key features or main features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;
[0038] Figure 2 A flowchart showing a method for training a neural network model according to some embodiments of the present disclosure;
[0039] Figure 3 A schematic diagram showing a cross-section of a wafer with grooves according to some embodiments of the present disclosure;
[0040] Figure 4 A schematic diagram showing a neural network architecture according to some embodiments of the present disclosure;
[0041] Figure 5 A block diagram showing a computing device capable of implementing multiple embodiments of the present disclosure.
[0042] In the various figures, the same or corresponding reference numerals denote the same or corresponding parts. Detailed Description of the Embodiments
[0043] The principles of the present disclosure will be described below with reference to various exemplary embodiments shown in the drawings. It should be understood that the description of these embodiments is only for enabling those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that, where feasible, similar or identical reference numerals may be used in the figures, and similar or identical reference numerals may represent similar or identical functions. Those skilled in the art will readily recognize that alternative embodiments of the structures and methods described herein may be employed without departing from the principles of the present invention described herein.
[0044] As used herein, the term "comprising" and its variants mean open-ended inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0045] CMP is a method of removing materials that combines chemistry and physics, and it has become an indispensable step in the modern integrated circuit (IC) industry. Specifically, CMP uses chemical etching and mechanical force to planarize the silicon wafer or other substrate materials during the processing. CMP can achieve material removal at the nanoscale level, making the wafer surface flat.
[0046] The CMP model, as the name implies, is a physical model for simulating the CMP process. The model outputs some important physical indicators of each grid point of the wafer after the CMP process, such as the thickness of the trench structure, the thickness of the non-trench structure, and the thickness of the metal material after CMP. In simple terms, it is a physical model that predicts the surface morphology of the wafer after the CMP process. CMP is one of the eight major semiconductor processes, connecting the previous and the next. Other processes following CMP will be affected by the CMP process. In other words, when it is necessary to evaluate and control the process results after CMP, it is best to know the impact of the CMP process, which is of great significance to the full process control and yield improvement of the entire semiconductor manufacturing. In addition, the CMP process itself is an indispensable and necessary step, because the CMP process can be better adjusted through the CMP model. For example, the results of the CMP model can be used to infer which process parameter settings need to be modified and adjusted, and whether there are some unreasonable aspects in the design of the Graphic Data System (GDS), which may cause hotspots, that is, defects.
[0047] The metal layer is deposited on the bottom of the groove. As the name implies, the bottom of the groove is the bottom layer of the groove. In fact, the process flow first goes through the etching step, that is, etching the wafer to produce grooves, or the bottom of the groove, and then the deposition process is carried out. Deposition will deposit various materials on the wafer. Because it has been etched, and the depth of the grooves produced in different areas is different due to different GDS designs, naturally the stacking height of the materials in each area after deposition will also be different. For example, the most common metal deposition material in CMP is copper. Usually in the last step of deposition, that is, before entering CMP, the material on the top layer of the wafer is copper, that is, the first step of grinding is copper. Because of the existence of etching, there are grooves of different depths in different areas, and the thickness of the deposition is also different. Therefore, for the surface of the wafer after the CMP process, it is necessary to evaluate the surface morphology. It is necessary to know how much change has occurred in the CMP step, and it is also necessary to know the changes and fluctuations that already existed before CMP.
[0048] In fact, CMP essentially refers to polishing the surface of a wafer immersed in a polishing liquid with a polishing head on a chuck. Simply put, GDS is a data format of a real entity, the wafer, in a computer program. When the data of a wafer is input into a program for observation, description, and analysis, it needs to be abstracted into a type of computer data, that is, a GDS file. The GDS file clearly depicts information such as the structure and design of this wafer. The wafer in the real world is the GDS in the virtual computer. From a more professional perspective, the GDS layout is a layout file given by circuit designers, which contains various graphics, that is, patterns, and semiconductor manufacturing is to engrave a wafer into the appearance of the layout design through various processes.
[0049] Usually, the value at the bottom of the trench is directly regarded as a fixed value, or simply an etching table is made using some empirical values or mathematical formulas, and the value is obtained by looking up a limited number of values in the etching table. There are great limitations in both accuracy and adjustability. Sometimes, in order to fit and obtain a value that conforms to the measured metal layer thickness, the CMP model is changed, resulting in the CMP model sacrificing the fitting accuracy of the dishing and erosion indexes. Here, dishing refers to the difference in height between the trench structure and the non-trench structure, and erosion refers to the difference in height of the non-trench structure at each grid point relative to the reference value thickness.
[0050] In view of this, the present disclosure provides an improved solution.
[0051] Some embodiments of the present disclosure provide an improved simulation method for semiconductor devices. The method includes: determining a simulated value of the height of the polished wafer surface based on a time step and a corresponding polishing rate; determining a correction term related to the polishing rate based on a comparison between the simulated value and a target value, so that the difference between the simulated value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulated value corresponding to the correction term as the height value of the polished wafer surface.
[0052] Some embodiments of the present disclosure also provide a training method for a neural network model. The method includes: training the neural network model using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated based on an initial formula for determining the surface height of the polished wafer; updating the parameters of the neural network model based on the gradient of the loss function; and in response to the training satisfying a predetermined condition, determining the neural network model corresponding to the corresponding updated parameters as the trained neural network model, where the trained neural network model outputs an enhanced formula for determining the surface height.
[0053] Some embodiments of the present disclosure also provide a method for generating training samples. The method includes: making a prediction based on an initial formula and an initial correction term to generate an updated formula and a reward term, where the initial formula is used to determine the surface height of the polished wafer, the initial correction term is used to correct the initial formula, and the reward term represents a feedback value generated based on the difference between the simulated value and the target value of the surface height; forming an array consisting of the initial formula, the initial correction term, the updated formula, and the reward term as the first training sample; and performing iterations in a neural network model based on the first training sample to generate multiple training samples.
[0054] The technical solution of the present disclosure can significantly improve the simulation accuracy and prediction fitting ability of the simulation model.
[0055] Embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings.
[0056] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the exemplary environment 100 includes a computing device 110 and a client 120.
[0057] In some embodiments, the computing device 110 can interact with the client 120. For example, the computing device 110 can receive an input message from the client 120 and output a feedback message to the client 120. In some embodiments, the input message from the client 120 can be, for example, layout data. The computing device 110 can perform corresponding mathematical operations on the layout data and output the corresponding operation results to the client 120.
[0058] In some embodiments, the computing device 110 can include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant PDA, a media player, etc.), a consumer electronic product, a minicomputer, a mainframe computer, cloud computing resources, etc.
[0059] It should be understood that describing the structure and function of the exemplary environment 100 only for exemplary purposes is not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. This environment is merely illustrative and is not used to limit the application environment of the embodiments of the present disclosure.
[0060] To more clearly explain the principle of the solution of the present disclosure, the following will be described in more detail with reference to Figure 2 to describe in more detail.
[0061] Figure 2 FIG. shows a flowchart of a method for training a neural network model according to some embodiments of the present disclosure.
[0062] At block 202, a neural network model is trained using training samples to determine the gradient of the loss function of the neural network model, where the training samples are generated based on an initial formula for determining the surface height of a polished wafer.
[0063] First, several key elements in some embodiments of the present disclosure are introduced:
[0064] Element 1: The features of the GDS (which can also be called graphical features), mainly referring to features such as density, width, space, and perimeter. At each position on an entire GDS, there are different GDS features (different density and width values), and different GDS features will greatly affect the topography height at each position after the etching process. In some embodiments, the density of the pattern, the width of the pattern, the space between the patterns, and the perimeter of the grid points can be called feature size information. The average space between the patterns within a grid point is called space, the average width of the pattern itself is called width, and the same applies to the perimeter. The sum of the side lengths of all the patterns within a grid point is the perimeter of this grid point.
[0065] In the GDS, there are many patterns (or called graphics), which can be microscopic structures such as circuits. In a fixed-size area, the ratio of the total area of the existing patterns to the area of this region is the density. So the density is a number between 0 and 1. The width is a descriptive statistical indicator for measuring the width of the pattern itself within this region. The patterns are usually rectangles or combinations of rectangles. Usually, all the patterns within the scanned area are obtained, and their width values are obtained, and finally the average value or median value is taken as the width value of this region). The space describes the space between the patterns within this region, because the patterns are arranged at a certain interval.
[0066] Element 2: The prediction target / simulation object. The simulation is the surface height value of each point on the GDS after the CMP process. In addition, the trench height after the etching process, that is, the value at the bottom of the trench, is denoted as TB.
[0067] Element 3: The process information table (recipe), which records process information. Different manufacturers will have different process settings, specifically including information such as the duration of the process, the start time, the end time, the unit time step, etc.; and various materials, thicknesses obtained from the previous deposition process, and the target height after etching and polishing, etc.
[0068] The CMP physical model refers to using the features extracted from GDS, usually density, linewidth (abbreviated as width), line space (abbreviated as space), perimeter, combined with process parameters (usually pressure, removal rate of different materials remove_rate, process time time, etc.) to form a physical formula to simulate the surface height after the CMP process.
[0069] In some embodiments, the polishing rate is determined based on the feature information of the patterns on the wafer and the parameters related to the polishing rate.
[0070] In some embodiments, a simulated value H1 of the height of the non-groove structure on the wafer can be determined; a simulated value H2 of the height of the groove structure on the wafer can be determined; and based on the simulated value H1, simulated value H2, and the density of the patterns within the lattice region on the wafer, the simulated value of the surface height of the wafer can be determined. Specifically, the surface height is equal to the height of the non-trench structure multiplied by (1 - density) plus the product of the height of the trench structure and density. Expressed by the formula: H1 * density + H2 * (1 - density); where density represents the density of the patterns within the lattice region on the wafer, and its range is between 0 and 1.
[0071] The following is described with reference to Figure 3 for illustration. Figure 3 FIG. shows a schematic cross-sectional view of a wafer with grooves according to some embodiments of the present disclosure.
[0072] As mentioned previously, Dishing is the difference between the height of the groove structure and the height of the non-groove structure, and the specific formula is SThickNT – SthickT. As Figure 3 shown, where D represents Dishing, SThickNT represents the height of the non-groove structure, and SthickT represents the height of the groove structure. In this figure, Dishing is actually the depth of the groove. As Figure 3 shown, the region shown in this wafer includes multiple grooves, that is, an array of grooves. The left side is the shown control area without a pattern structure, which is a blank test control area specifically used for comparison with the patterned area. And all the lattice regions in the model of the embodiments of the present disclosure represent areas with patterns, as Figure 3 shown in the right part of
[0073] Erosion refers to the difference between the height of the non-groove structure at each grid point and the reference thickness value. The specific formula is field reference thickness (the reference height shown in Figure 3 as the reference height) - SThickNT (the thickness of the non-groove structure), where the reference height value can be given by the user and is related to the user's respective process settings. Both Dishing and Erosion are important metrics for evaluating process results in the CMP model.
[0074] Returning to Figure 2 Continue the description. In some embodiments of the present disclosure, the CMP process is based on a basic physical model and theory. In some embodiments of the present disclosure, the model formula is extended based on the above theory. The theory is that within a very short time t (usually less than or equal to 0.1 s), denoted as dt, the polishing rate is constant, and the groove height at the next moment can be expressed by the following formula:
[0075] SThickNT(t = t1) = SThickNT(t = t0) – dt * R_NT(t = t0);
[0076] The non-groove height can be expressed by the following formula:
[0077] SThickT(t = t1) = SThickT(t = t0) – dt * R_T(t = t0);
[0078] In the above formula, R_NT(t = t0) represents the polishing rate of the non-groove structure at t0. Similarly, the meaning of R_T can be obtained.
[0079] It can be seen that the calculation of R_NT and R_T is the core of the entire CMP physical model. According to the basic model, the polishing rate Rate is related to the actual pressure, and the actual pressure is related to the density.
[0080] For the non-groove structure, the most primitive physical mechanism formula for polishing is: R_NT = P / (1 - density). For the groove structure, R_T = P / density. However, in practice, it needs to be corrected, that is, a correction term is needed to correct it.
[0081] In some embodiments, the correction term at least includes a polynomial related to at least one of the following: the width of the pattern within the grid point on the wafer; the spacing between patterns; and the sum of the perimeters of each pattern within the grid point.
[0082] In fact, for the non-groove structure, R_NT at a certain moment is as follows:
[0083] R_NT = P / (1 - density) * Width correction term * Space correction term * Perimeter correction term.
[0084] For the trench structure, R_T at a certain moment is as follows:
[0085] R_T = P / density * Width correction term * Space correction term * Perimeter correction term.
[0086] There is no very clear golden standard and undisputed formula for these correction terms at present. Different model developers and different manufacturers have their own different presets. Specifically, the Width correction term refers to a polynomial with Width as the core of change (or a polynomial related to Width), and the Space correction term and Perimeter correction term are the same. This polynomial can be a polynomial in the form of linear summation, or can be doped with complex polynomials such as exponential and logarithmic forms.
[0087] In some embodiments of the present disclosure, the most advanced deep reinforcement learning technology is used to complete the derivation of these correction terms based on actual data, so as to complete the expansion of the basic physical formula of CMP. In some embodiments, the actual data may include graphic feature information.
[0088] In some embodiments, the graphic feature information may include the following information: the density of the graphics within the grid points on the wafer; the width of the graphics; the spacing between the graphics; and the initial height of the bottom of the trench, the total perimeter of the graphics within the grid points.
[0089] In the usage scenario of the entire model, the so-called point refers to a grid point. By extracting the information inside a grid point, a data point can be obtained. The so-called measurement point is also this grid point. Strictly speaking, the grid point can be said to be a rectangular area delimited based on a specific size (the size is usually between 5um * 5um and 20um * 20um).
[0090] In some embodiments, the simulated value of the surface height can be determined based on the initial data and initial parameters in the CMP model. In some embodiments, the initial data may include the measured values of several measurement points (recording the position of each point in GDS, its corresponding GDS features, and the height of the trench after etching).
[0091] From the above description, it can be known that the simulated value of the surface height of the polished wafer can be determined based on the time step and the corresponding polishing rate. For example, it can be calculated based on the initial physical formula. In fact, it can be determined by any way of determining the initial simulated value through a simulation model.
[0092] In some embodiments, a deep learning neural network is constructed, and the simulated value of the surface height is determined by randomly generated initial parameters and initial data. A simulated value closer to the measured value can be obtained through iteration of the CMP model.
[0093] The determined simulated value of the height may have a large gap from the actual measured value. A simulated value closer to the measured value can be obtained through an improved CMP model.
[0094] In some embodiments, the training samples can be pre-generated, for example, stored in a predetermined database and available for training the neural network model.
[0095] In some embodiments, the training samples can be generated during the training process of the neural network model.
[0096] In some embodiments, the training samples can be generated in the following manner: predictions are made based on an initial formula and an initial correction term to generate an updated formula and a reward term, where the initial formula can be used to determine the surface height of the polished wafer, the initial correction term can be used to correct the initial formula, and the reward term can represent a feedback value generated based on the difference between the simulated value and the target value of the surface height; an array composed of the initial formula, the initial correction term, the updated formula, and the reward term is used as the first training sample; and multiple training samples can be generated by iterating based on the first training sample in the neural network model.
[0097] In some embodiments, each training sample is an array composed of the following items: an initial formula; an initial correction term, which is used to correct the initial formula; an updated formula, which is generated based on the initial formula and the initial correction term; and a reward term, which represents a feedback value generated based on the difference between the simulated value and the target value of the surface height. It should be understood that the composition of the training samples in the embodiments of the present disclosure is not limited to this, but can vary according to actual needs.
[0098] In some embodiments, the reward term is generated in the following manner: a first difference between the standard deviation of the simulated values of the surface heights of each grid point on the wafer and the standard deviation of the target value is determined; a second difference between the root mean square error of the simulated values of the surface heights of each grid point and the root mean square error of the target value is determined; the range between the simulated value and the target value of the surface height of each grid point is determined; and the reward term is generated based on the first difference, the second difference, and the range.
[0099] In some embodiments, when the samples generated in each iteration meet a predetermined condition, the iteration can be terminated, and all the samples in this iteration are used as a trajectory; and when the number of total trajectories obtained during each iteration reaches a predetermined trajectory threshold, the iteration is terminated; in addition, the samples in the total trajectories can be used as samples for training the neural network model.
[0100] In some embodiments, predictions can be made based on an initial formula and initial correction terms to generate an updated formula. For example, a random number is generated at the beginning of each iteration; a source of the initial correction terms for correcting the initial formula is selected based on a comparison of the random number with a random probability threshold; and the initial formula is corrected based on the selection to generate an updated formula.
[0101] In some embodiments, a symbol library is also generated for use during the iterative process. For example, the symbol library can be generated by taking the Cartesian product combination of mathematical symbols and a variable library, where the variable library contains graphic feature information and process information.
[0102] In some embodiments, the initial formula can be corrected based on the selection to generate an updated formula. For example, in the case where a first comparison result is obtained by comparing the random number with the random probability threshold, a symbol is selected from the symbol library as an initial correction term to correct the initial formula using a neural network model; and in the case where a second comparison result is obtained by comparing the random number with the random probability threshold, a neural network model is used for prediction to generate a predicted initial correction term to correct the initial formula, and the predicted initial correction term is added to the symbol library to enrich the symbol library.
[0103] In some embodiments, an improved CMP model is obtained through a deep learning neural network to determine correction terms related to the polishing rate. In other words, an improved physical formula can be obtained to calculate the simulated value of the surface height.
[0104] In some embodiments, correction terms can be determined based on graphic feature information, process information, and information of the initial formula for calculating the surface height, and the initial formula is corrected using a deep learning neural network.
[0105] In some embodiments, other methods can also be used to determine correction terms related to the polishing rate.
[0106] Deep Learning specifically refers to machine learning based on deep neural network models and methods. It has evolved based on algorithm models such as statistical machine learning and artificial neural networks, combined with the development of contemporary big data and high computing power. The most important technical feature of deep learning is its ability to automatically extract features, and the extracted features are also called deep features or deep feature representations. Compared with manually designed features, deep features have stronger and more robust representation capabilities. The deep neural network is the model basis for deep learning to automatically extract features, and the deep neural network is essentially a nested series of non-linear transformations.
[0107] Regarding neural networks, first introduce some basic definitions: State, defined as the current physical formula, denoted as S0; Action, the behavioral action performed based on the current state S0, specifically defined as which mathematical calculation symbols and variables to select; S1, the state at the next moment, which changes after performing action a on S0, that is, the new physical formula. R is defined as the reward, which means that after taking action a based on state S0 and transitioning to the new state S1, the obtained reward is R. The specific content of R reflects whether the new formula performs better in terms of data compared to the old formula, and whether the Root Mean Square Error (RMSE) is lower. If it is lower, it is a positive reward; if it deteriorates, it is a negative reward (actually a penalty). R is a term in reinforcement learning. Reinforcement learning is essentially a series of mathematical-related conceptual models and does not have practical physical significance. Reinforcement learning is a specialized field and a branch of deep learning.
[0108] In some embodiments, a deep learning neural network can be built based on a Recurrent Neural Network (RNN) structure as an agent for deriving formulas. It should be understood that the neural network in the embodiments of the present disclosure is not limited to RNN and can also be other neural networks. The content received at the input end of the RNN neural network can be the features of GDS, process parameter information, and initial formula information. The network includes several fully connected layers, pooling layers, convolutional layers, etc. in the middle. The characteristic of the RNN is that each fully connected layer contains a gate structure. Compared with an ordinary neural network that receives the output of the previous layer as input to calculate the value of the activation function to determine whether to activate, the gate structure of the RNN also considers information from even earlier layers as auxiliary inputs to participate in the calculation of the activation function. The final output of this neural network can be, for example, symbols (such as including +-* / , square root, square, logarithm and exponent with base e, and related GDS feature variables, etc.), or a complete formula.
[0109] In some embodiments, using a deep learning neural network to determine a correction term to correct an initial formula may include: generating a reward term based on the simulation values and measurement values of the surface heights of each grid point; and using an array composed of the initial formula, the correction term, the updated formula, and the reward term as elements of the training sample to iterate in the neural network to generate elements of a new training sample.
[0110] In some embodiments, the iteration using a deep learning neural network may include a prediction phase and a training phase, and wherein: when the elements generated in each iteration of the prediction phase satisfy a first predetermined condition, the iteration is terminated, and all the elements in the iteration are taken as a trajectory; and when the number of total trajectories obtained during each iteration reaches a predetermined trajectory threshold, the prediction phase is terminated. Further description thereof will be provided hereinafter.
[0111] In some embodiments, the first predetermined condition may include: the number of elements generated in each iteration reaches a first predetermined threshold; or the value of RMSE in the current iteration is less than a second predetermined threshold.
[0112] In some embodiments, the agent is trained by the process and learning method of policy gradient in the field of reinforcement learning (essentially iterating the RNN continuously so that it can predict a complete formula). The initially available data may include a number of measurement data points under a certain process. The content of each data point is the XY coordinates, density / width / space / perimeter GDS feature information, and the corresponding ThickNT and ThickT after the CMP process, that is, the measured surface height and the initial ThickNT and ThickT values before the CMP process. Suppose there are N measurement points (data points). Each data point has a result of surface height. A so-called data point actually refers to a lattice point. As mentioned above, the smallest unit in the whole model is a lattice point, and several lattice points form a complete chip. Similarly, a data set is composed of several data points, and the full chip data set refers to the set composed of each data point represented by each lattice point on the full chip.
[0113] During the process of simulating using the neural network, the agent may be initialized first. The basic formula is initialized, that is, the ground height Thick1 is initialized to be equal to the initial height Thick0 before grinding minus the set unit time multiplied by the actual grinding rate, where the actual grinding rate is equal to the sum of the nominal grinding rates of each material multiplied by their corresponding correction factors. The specific formulas of the correction factors are slightly different for non-groove structures and groove structures. The original physical formula for the NT structure is: P / (1 - density). The correction factor (correction term) is, for example: *width*space*perimeter; the original formula for the T structure is: P / density. The correction factor is, for example: *width*space*perimeter. It should be understood that the correction factors shown here are only illustrative, and in fact, it is very likely that they are not as simple as several terms multiplied as shown, but may have various complex polynomial structures.
[0114] The state represented by the original physical formula is the initial state S0. The RNN neural network can be initialized. The neural network structure includes an input layer, which receives the feature information of GDS, the initial formula, and the set process information table. The process information table may include the set nominal pressure P0, the set slurry ratio Slu0, the total number of layers of materials to be polished in CMP, the deposition thickness T of each layer of material, the initial trench height and non-trench height, and the nominal polishing rate R_material of each material, the hardness pcoef of the pad, the total polishing time t0, and the set unit time dt. CMP uses a polishing head to assist the slurry in polishing the wafer. The slurry has different chemical concentrations, resulting in different effects.
[0115] The parameters of the neural network are random (at initialization). The parameters of the neural network are mainly the weights that make up the network. The neural network can be regarded as a composite of many simple formulas, and finally a complex integrated mathematical formula is obtained. There are many undetermined parameters in this formula. When not trained, the parameter initialization is a random value. As training progresses, the gradient is calculated by backpropagation using the training data, and the values of the parameters are continuously updated. These parameters can also be called weights.
[0116] Therefore, the neural network will randomly give a new formula symbol. A new CMP physical formula is obtained according to the new formula symbol. For example, specifically, the alternative action library can be obtained by taking the Cartesian product of various basic mathematical operation symbols and the variable library (the variable library includes process parameters and GDS features), such as "+density, -width, *space, / perimeter, +ln(Slu0)", etc. This is the action set in this disclosure. In the conventional case, when using reinforcement learning for formula expansion, the action setting is often basic mathematical symbols, such as addition, subtraction, multiplication, division, and some numbers. However, in some embodiments of this disclosure, another definition is provided for formula expansion. Because if there are only mathematical symbols, it is impossible to perform very complex combinations, but the Cartesian product of mathematical symbols, variables, and parameters is a relatively large combination library, with many choices, and better expansion results can be obtained. Take the simplest example. Suppose the initial formula is 1 + density, and the action given by the neural network is *width, so the formula becomes (1 + density) * width. Here, *width is a correction term.
[0117] This physical formula calculates the polished trench height and non-trench height based on N measurement points and calculates the surface height according to the formula described above. This value is the simulation value, and thus the RMSE between the simulated surface height and the true measured surface height can be calculated.
[0118] A neural network can be used to predict the next mathematical symbol. Its output is the next mathematical symbol, and its function is to expand the physical formula, and the output of the physical formula is the relevant simulation result. For example, the current formula is P / (1 - density)*widt*space*perimeter, and the next symbol given by the neural network is - width, so the formula becomes P / (1 - density)*width*space*perimeter - width. The new formula is also an expanded physical model, and this physical model can calculate SThickNT (the polished height of the NT structure) and SThickT (the polished height of the T structure). *width*space*perimeter - width can also be referred to as the correction term.
[0119] Deep reinforcement learning itself is just a framework that describes a general model construction form. In some embodiments of the present disclosure, this model is used to expand the formula. Thus, combined with the characteristics of the present disclosure, the surface topography of the wafer after the CMP process is predicted using density, width, space, and CMP physical process parameters. Here, under the framework of reinforcement learning, the present disclosure defines that the action is a combination of relevant features and some process parameters with common mathematical symbols.
[0120] In some embodiments, a reward term can be generated based on the simulated values and measured values of the surface height of each grid point, including: determining a first difference between the standard deviation of the simulated values of the surface height of each grid point and the standard deviation of the measured values; determining a second difference between the RMSE of the simulated values of the surface height of each grid point and the RMSE of the measured values; determining the range between the simulated values and the measured values of the surface height of each grid point; and generating a reward term based on the first difference, the second difference, and the range. It should be noted that, according to actual needs, only the first difference and the second difference can be determined, without determining the extreme values. Thus, a reward term is generated only based on the first difference and the second difference. For example, the standard deviation and range of the surface height before and after CMP in the simulation data can be calculated simultaneously (here, two ranges and two standard deviations are calculated, respectively before and after simulation. The calculation before simulation is based on the measured data, and the calculation after simulation is based on the simulation results. The difference between the two ranges (specifically, the range before simulation minus the range after simulation) constitutes a variable, denoted as G0. Similarly, the difference between the standard deviations before and after simulation can also constitute a variable, denoted as G1). The range is the difference between the Max (maximum value) and the Min (minimum value). The range of the surface height before CMP in the simulation data refers to the difference between the maximum value and the minimum value among the data of each measurement point.
[0121] In one prediction, the reinforcement learning neural network gives an extended term to supplement the original formula, so there are two formulas (the original formula and the supplemented formula). Each formula can perform simulation calculations on these several data points, that is, two simulation calculations. Each simulation calculation will also generate an RMSE. After performing the simulation calculations of the original formula and the new formula, there are two RMSEs denoted as rmse0 and rmse1, and the difference between them can be denoted as G2. The two ranges are denoted as max_diff0 and max_diff1 respectively, and the two standard deviations are std_error0 and std_error1 respectively. One reward (penalty) can be calculated from these 6 indicators.
[0122] Then G0, G1, and G2 constitute the reward for this prediction (this prediction is an action). The specific effect is whether the RMSE decreases (compared to the initial basic formula, because the formula derived from the neural network shows better fitting performance on the data), and whether G0 and G1 are relatively large positive numbers, because if G0 and G1 are relatively large positive numbers, it conforms to the actual situation of the CMP process. Simply put, the surface should be more "flat" after CMP grinding (the flatness is manifested as a decrease in the standard deviation and range of the heights of different position points), and the larger the better.
[0123] It is desirable that the RMSE be as small as possible. Therefore, if (rmse0 - rmse1) is greater than 0, it proves that the extended formula has better accuracy and a reward (positive feedback) is given; otherwise, a penalty (negative feedback) is given. This is the main reward and punishment mechanism. (That is to say, there can be three goals, one main goal and two secondary goals. The main goal is to reduce the RMSE, and the secondary goals are to reduce the data difference before and after simulation. The secondary goals can also be regarded as "regularization terms". The significance of the regularization terms is to prevent the neural network from completely ignoring the objective reality in order to achieve the main goal. Usually, users do hope to obtain a formula with better fitting, but it should be able to fit the objective characteristics and facts of the CMP process to a certain extent. This is also an innovation point of this disclosure, not only considering the fitting accuracy. Combining the actual process of CMP, a more reasonable and ideal result of the CMP process is to reduce the range and standard deviation of the data. That is to say, if (max_diff0 – max_diff1)>0, some additional rewards (positive feedback) are given; otherwise, some penalties (negative feedback) are given. The same applies to the standard deviation. Ultimately, it is hoped to find a new formula that can reduce the RMSE, range, and standard deviation. In the above embodiments, the RMSE, range, and standard deviation are used to form the rewards. The embodiments of this disclosure are not limited thereto and can be variously changed according to needs. For example, the rewards can also be formed only based on the RMSE and the standard deviation.)
[0124] In some embodiments, the specific reward formula is as follows:
[0125] R = a*(rmse0 – rmse1) + b*(max_diff0 – max_diff1) + c*(std_error0 – std_error1);
[0126] Let rmse0 – rmse1 = G2, then R = a*G2 + b*G0 + c*G1;
[0127] Where a, b, and c are positive numbers between 0 and 1, and the specific numbers can be adjusted depending on which goal is expected. Usually, a is a relatively large value because fitting more accurately is the most core goal. After the above process, all the elements required for a deep reinforcement learning neural network are generated, namely the initial state S0, the action a0 taken, the transferred state S1, and the corresponding reward R. Continuously let the agent (RNN neural network) make predictions, that is, many different <S0, a0, S1, R> tuples will be continuously generated. This tuple is the sample for training the agent (optimizing the parameters of the neural network through backpropagation). The tuple is obtained after the neural network makes a prediction once. After making several such predictions, several such tuples can be obtained.)
[0128] Based on a neural network obtained, it can be trained based on a number of training samples. First, it is necessary to clarify how to perform the training, and the definition of the loss function is proposed here. Define the loss function as J(θ), where θ refers to the set of all hyperparameters in the entire RNN network, that is, a specific θ corresponds to a neural network with determined parameters. Here, the definition of a trajectory is introduced, and the symbol of the trajectory is recorded as τ. What is called a trajectory in this disclosure refers to a set composed of several <S0, a0, S1, R>. In this disclosure, it is necessary to determine what constitutes a complete trajectory, or how many <S0, a0, S1, R> tuples should be included in a complete strategy. To avoid overfitting of the formula, the following two methods are adopted here to determine whether to stop making predictions (essentially also training) to determine the strategy, namely:
[0129] Condition 1: Whether the number of predictions (the number of <S0, a0, S1, R>) reaches 10 times. If it reaches, stop. It should be understood that this value is exemplary and can be changed according to needs.
[0130] Condition 2: Whether the RMSE value of the current formula is less than 1. If it reaches, stop. Generally speaking, the unit for measuring the surface height is nm. A CMP formula is considered a relatively accurate model if its prediction accuracy of the surface height reaches an error within 2 - 5 nm.
[0131] If any of the above two conditions is met, it is considered that the termination can be carried out. That is, during initialization, make a prediction once, record and update <S0, a0, S1, R> in the database, continue the second prediction, continue to update and record the new <S0, a0, S1, R>, and so on. Continue like this until one of the above Conditions 1 or 2 is met and stop (whichever condition is met first is taken as the criterion), and thus a complete strategy is obtained. That is, a complete strategy is obtained through prediction.
[0132] Generally speaking, the meaning of a policy is its Chinese meaning as a noun, which can be interpreted as "receiving an input A and generating a deterministic output B. Under a policy, no matter how many times the input A is executed, the result obtained must always be B and will not suddenly become C at a certain time". In some embodiments of the present disclosure, a neural network agent is created, and the framework and methods of reinforcement learning are used to continuously adjust and improve the parameters of the neural network. When the parameters of a neural network are all determined, receiving an input will definitely obtain a definite output, and the prediction results will not change after several such operations. This is the so-called policy. Therefore, it can be said that a policy refers to a neural network under certain parameters. The essence of reinforcement learning is to continuously adjust the parameters of the neural network. (Neural networks with different parameters may produce different outputs when facing the same input, so different parameters mean different policies). The essence is to continuously adjust the output obtained when facing the same input, that is, to adjust the policy. The R in reinforcement learning is used to evaluate the quality of the policy, whether a good decision has been made to bring benefits, or a bad decision has been made to bring punishment. A neural network usually contains several parameters, which can be represented by theta (a mathematical symbol, θ), so θ is usually used to represent the policy in the reinforcement learning framework).
[0133] In the actual calculation process, multiple trajectories are continuously sampled, and the length of each trajectory is T (that is, it is necessary to take T steps, from 0 to T, and there is a corresponding reward or punishment for each step. Therefore, to evaluate the comprehensive rewards and punishments of a trajectory, the sum of the rewards for each step of this trajectory needs to be considered). Neural network training and parameter update are performed every once in a while. This every once in a while is actually continuous sampling, that is, using the current neural network to continuously make predictions and generate several trajectories with a length of T steps.
[0134] It should be noted that in some embodiments of the present disclosure, there are actually two purposes. One is to expand the current formula so that the current formula can achieve the expected fitting accuracy (rmse meets the standard) for the existing real data, and the other is to train a stable neural network, not just for this time (referring to being able to expand a formula that fits the current piece of data), but also for other new data, this stable neural network can also give an extended result that meets the standard.
[0135] In some embodiments, determining a correction term using a deep learning neural network to correct an initial formula may include: in a training phase, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network; and updating the parameters of the neural network based on the gradient of the loss function. In some embodiments, training the neural network using training samples in a trajectory to determine the gradient of the loss function of the neural network may include: using a policy gradient algorithm in the field of reinforcement learning to perform backpropagation on the training samples to determine the gradient of the loss function.
[0136] According to the definition of the policy gradient algorithm, the training process is to calculate the policy gradient of the loss (objective) function to update the neural network parameters. The calculation of the policy gradient metric is as follows:
[0137] The derivative of the loss function J with respect to the parameter θ is:
[0138]
[0139] where E is the symbol for mathematical expectation, T refers to the length of the trajectory, t refers to a certain step in the trajectory, t = 0 is the initial of the trajectory, t = 1 refers to the first newly generated step of the trajectory, and R refers to the reward.
[0140] P in the above formula is not an independent letter. P(τ|θ) is a complete term. The meaning of P is probability. This term means the probability of outputting the trajectory τ (trace) under the current neural network parameters. The so-called trajectory is a term in the Markov decision process (MDP). Specifically, it is a continuous process that includes the initial state S0, the action a0 taken, the next state S1, the action a1 taken, another state S2, and the action a2. And so on until a certain time T. As for how long the time T is, that is, how long the trajectory is, it is uncertain. In fact, the length of the formula is not infinite and there will be a specific maximum length limit. Of course, it varies for different manufacturers. Assuming the maximum length limit for formula expansion is set to 10, it means that the neural network can be expanded at most 10 times on the basis formula. If the set maximum expansion times are reached, it will stop. That is to say, the time T is 10, and a complete trajectory contains 10 Ss and corresponding actions.
[0141] The Log technique refers to the mathematical properties of logarithms. If y = x1 * x2, that is, there is a function where the independent variables x1, x2 and the dependent variable y have a multiplicative relationship. However, for the convenience of subsequent calculations and mathematical derivations, it is desired to transform it into an additive form. That is, using the Log technique, take the logarithm of both sides of the equation, log(y) = logx1 + logx2. And the mathematical properties of logarithms ensure that it does not change the monotonicity of the function before and after taking the logarithm. That is to say, if there are x1 and x2 that can make y reach the maximum value, then these x1 and x2 can also definitely make log(y) reach the maximum value.
[0142] Based on J(θ), the theoretical neural network parameter update formula is as follows: Assume that currently in the k-th training stage, α is the learning rate, which is a common parameter in neural networks. Then for the (k + 1)-th training, it can be expressed as follows:
[0143]
[0144] The above formula for calculating the gradient is a theoretical formula. In fact, it can be approximately replaced by the following formula:
[0145]
[0146] D is the set of collected trajectories: The update of the parameter is from to
[0147] denotes the gradient, and represents the estimated value. That is to say, represents the estimated (approximate) calculation formula for the gradient.
[0148] Reinforcement learning is a sub-branch field of deep learning. The policy gradient algorithm is a common classic algorithm in reinforcement learning. This formula and the transformation process, that is, the derivation process of the objective function of the policy gradient algorithm, are well-known and commonly used in the industry.
[0149] In some embodiments, a random probability threshold can be set. During iteration, the random probability threshold changes from an initial first probability threshold to a second probability threshold lower than the first probability threshold. That is, the random probability threshold takes the maximum value at the beginning of the iteration, and gradually decreases as the iteration progresses. This will be further described later.
[0150] At block 204, update the parameters of the neural network model based on the gradient of the loss function.
[0151] In some embodiments, updating the parameters of a neural network model based on the gradient of a loss function may include: after a predetermined number of training operations, for example, 100 times (referring to the operation of extracting several samples to calculate the gradient and update the network parameters), performing gradient update on the gradient of the loss function; and updating the parameters of the neural network model based on the gradient update.
[0152] In some embodiments, throughout the entire process of neural network simulation (or training), the number of operations is denoted as M, and the initial value of M is 0. That is, the random selection probability is e, and the initial value of e is 1.0. During this process, training samples for the reinforcement learning neural network are continuously generated, that is, several data combinations <S0, a0, S1, R> with different results are continuously produced. It is desired to generate results in various different situations as much as possible, so that the trained reinforcement learning neural network will be more powerful. However, for a neural network, whether it has been trained or not, a given input will definitely result in a definite output. When the network parameters remain unchanged, the same input leads to the same output. It is desired that the samples cover all possible inputs that the neural network may encounter, thereby obtaining all possible outputs. Additionally, if random selection is made each time, it is possible that many of the choices made are not good choices. In fact, many possible combinations are not good combinations. If completely random, a large amount of time will be wasted in generating these bad combinations. The desired final effect is that the neural network continuously receives inputs for prediction to obtain an output. It should have the experience learned before, filtering out many completely unnecessary attempts, and also maintain a certain degree of randomness, so as not to become completely rigid (unable to break out of limited paths and unable to generate new samples).
[0153] That is to say, there are two ways to select an action. One is to randomly select one from a symbol library, and the other is to use the current neural network to predict one. The former has no experience and is purely random, while the latter will tend to give an action that makes R (reward) larger as the neural network is gradually trained. Generally speaking, the first random selection is "irrational", and the second using the neural network prediction is "rational" (especially as the neural network is gradually trained, the neural network becomes smarter and smarter). Because the training of the neural network is essentially a mathematical optimization problem, and a common possible situation in a mathematical optimization problem is encountering a local optimum, and retaining randomness can break this local optimum dilemma.
[0154] In some of the above embodiments, e will continuously decrease. Because initially the neural network also has randomly generated parameters, the results given by the neural network can also be said to be "randomly selected" to a certain extent, and there is no essential difference between the two. However, as the training progresses, the neural network gradually becomes smarter, e gradually decreases but does not completely become 0, which exactly conforms to the previous explanation.
[0155] The meaning of M is the number of operations, that is, the cumulative number of times of performing input, output, generating data combinations, and saving these operations. Its initial value e is 1, which means that at the beginning, random selection is always carried out, so as to ensure the generation of a large number of new combinations and a large number of different samples. However, as the operations proceed, the neural network will take samples from the sample library for training at regular intervals, and the neural network will gradually become somewhat tendentious and intelligent, and will select the formula that can obtain better prediction results as the output. It is necessary to derive formulas on this basis. It is not necessary to continue wasting time trying to derive many formulas that are not suitable at the beginning. Therefore, e will continue to decrease. For example, at the beginning, it is not known which combinations are good, so all combinations are selected for trial. After several rounds of trials, a general concept can be obtained. Some combinations are relatively good and have "potential", while some combinations are very poor and there is no need to continue trying. Therefore, in order to improve efficiency, as the trials proceed, combinations with no potential can be directly excluded without wasting time deriving and evaluating on this basis. At the same time, for those good combinations with potential, it is possible to break away from the original idea and try a new derivation at a certain step to obtain a result better than before. Therefore, a certain degree of randomness will still be maintained to avoid falling into the dilemma of local optimum.
[0156] The above described how to actually update the network parameters. Assume the total number of times is M. Assume that gradient descent is performed every 200 times to update the network parameters, and assume that the step size of a trajectory is 20, that is, each time the neural network samples (200 / 10) = 10 trajectories, and every time 10 trajectories are sampled, the gradient descent is used to update the network parameters using the results just sampled. M is the total number of times (since it is not certain when the training is in place, the value of M cannot be determined). This e is related to M. At the beginning, e is set to 1. As M continues to increase, e gradually decreases. Assume that after M times, the neural network is trained to a certain extent. If the standard has been reached, then the training is completed. In addition, a maxM can be set, assumed to be 10000. What if M reaches 10000 but has not reached the described stable state? The state at this time must be that the loss function curve oscillates and cannot be stabilized, which means that new randomly selected samples need to be generated to break the "deadlock" (oscillation usually indicates that the sampled trajectories are not comprehensive enough and the information contained is not enough). So the above process can be repeated again, reset M, reset e, and then continue to sample to generate trajectories (this is more likely to generate random new samples) and perform training.
[0157] In some embodiments, determining a correction term using a deep learning neural network to correct an initial formula may include: generating a random number at the beginning of each iteration of the neural network, where the range of the random number may be from 0 to 1, and the present disclosure is not limited thereto and may vary according to needs; selecting a source of the correction term for correcting the initial formula based on a comparison between the random number and a random probability threshold, that is, selecting to generate a correction term through the prediction of the neural network or selecting a symbol from a symbol library for correction; and determining an updated correction term using the neural network based on the selection result to correct the initial formula.
[0158] In some embodiments, determining an updated correction term using the neural network based on the selection to correct the initial formula may include: when the comparison between the random number and the random probability threshold is a first comparison result (for example, the random number p is less than the random threshold probability e), selecting a symbol from the symbol library as the correction term to correct the initial formula using the neural network, where the symbol library is generated by taking the Cartesian product of mathematical symbols and a variable library, and the variable library contains graphical feature information and process information; when the comparison between the random number and the random probability threshold is a second comparison result (for example, the random number p is greater than the random threshold probability e), using the neural network to make a prediction to generate an initial correction term to correct the initial formula using the neural network model to generate a new correction term, and adding the new correction term to the symbol library.
[0159] In some embodiments, the neural network and related variables may be initialized, and the neural network may be used to perform a prediction. A random number p may be generated before the prediction. As mentioned above, if p > e, the neural network is used for prediction, otherwise, random sampling selection is performed from the symbol library. And <S0, a0, S1, R> is obtained, and the process is repeated to obtain a trajectory, and the trajectory is recorded in the database. The above process is repeated using the neural network to obtain other different trajectories. In this process, for example, every 10 times, 70% of the trajectories are randomly selected from the data for backpropagation calculation (according to the gradient calculation formula described above), and the parameters of the neural network are updated. At the same time, the loss value of the loss function may be recorded. Additionally, when M gradually increases, e is also gradually decreased (for example, the minimum is 0.1, and it may vary according to needs). The specific change formula is:
[0160] e = e0 – tanh(M) + 0.1;
[0161] Among them, tanh is a common function. For example, tanh = f(x), where x is a variable. When x gradually increases starting from 0, f(x) gradually increases starting from 0 and finally approaches 1 infinitely. Specifically, tanh is a common activation function used in the connection between layers of a neural network. Each layer of the neural network has many neurons. The essence of a neuron is a combination of several parameters. Each layer calculates based on the input of this layer plus the parameters contained in all the neurons in this layer. The result of the calculation is input into the activation function of this layer, and the result output by the activation function is output to the next layer as the input of the next layer. The formula of the tanh activation function is:
[0162]
[0163] Where x represents a variable, and in some embodiments of the present disclosure, it can represent the number of iterations.
[0164] In some embodiments, as mentioned above, the value of the loss function is also determined, and thus it is judged whether the continuous training can be ended. When the value of the loss function is lower than a predetermined threshold, it indicates that a reliable neural network has been obtained and the training can be ended. At this time, the correction term output by the neural network corresponding to the value of the loss function can be determined as the correction term related to the grinding rate; and an update formula can be determined based on the determined correction term to determine the height value of the surface.
[0165] In some embodiments, the iteration in the training stage can be stopped when the following conditions are met: the difference between the values of the loss function in two consecutive iterations is less than a predetermined difference threshold; the absolute value of the ratio of the values of the loss function in two consecutive iterations is greater than a predetermined ratio threshold; or the number of iterations or the iteration time reaches the maximum iteration threshold.
[0166] Generally speaking, when the loss curve gradually approaches a steady state or reaches the set maximum number of iterations, the training is stopped, and a trained neural network is obtained.
[0167] The loss curve is a curve plotted by calculating the value of the loss function during each training. Assuming 100 training times, a loss function value will be obtained each time. For example, after 100 times, 100 loss function values can be obtained. Then a curve can be plotted. The horizontal axis is the number of training times, and the vertical axis is the loss value. As the training progresses, the function value of the loss will continuously decrease and finally reach a steady state. The theoretical basis for the loss curve to reach a steady state is that the labeled data is a finite number and will not keep increasing, so the amount of information contained in this data must be fixed. The essence of the training of the neural network is to continuously use this information to adjust its own network parameters. That is to say, there will always be a time when learning is completed, that is, there will always be a time when no new knowledge can be learned.
[0168] The neural network of the embodiments of the present disclosure can make several steps of predictions based on the CMP basic physical formula to obtain a final enhanced CMP physical formula. It should be understood that the intervals of 10 times and 70% of the trajectories mentioned here are only exemplary, and the present disclosure is not limited thereto, and various changes can be made according to needs. For example, common ranges can be from 10 to 100 times, and can be between 50% and 80%.
[0169] Backpropagation calculation is a well-known calculation for neural network training in the industry. Generally speaking, as long as it is neural network training, backpropagation calculation is used for training. Specifically, backpropagation refers to using samples to calculate gradients and using gradients to update the parameters of the neural network. The purpose of this method is to change and adjust the parameters of the neural network. The difference between a randomized neural network and a trained neural network lies in the different parameters. For the same network structure, different parameters will predict completely different results.
[0170] In some embodiments of the present disclosure, the neural network can accept a formula as input and obtain a corresponding output. The output result is a mathematical symbol (but not limited to addition, subtraction, multiplication, and division. In fact, it is a combination of symbols and other terms (such as variables)). The predicted mathematical symbol (new action) is written behind the current formula to obtain a new formula. Repeat the above steps to train the neural network, so that the neural network continuously expands the initial formula and finally obtains a new formula.
[0171] Regarding the specific architecture of the RNN neural network in the present disclosure as Figure 4 shown. Figure 4 The schematic diagram of the neural network architecture according to some embodiments of the present disclosure is shown.
[0172] As Figure 4 shown: This neural network is of the NV1 structure, that is, multiple inputs and one output. And there are also 2 fully connected layers, one convolutional layer, and one pooling layer between the neurons h1, h2, and h3. The number of neurons in the fully connected layer is 128. It should be understood that the RNN network structure described here is only schematic, and different structures can be adopted according to actual needs.
[0173] Each neuron corresponds to a different input, and its expression is:
[0174] h1 = f(Ux1 + Wh0 + b)
[0175] hN = f(Ux N + Wh N -1 + b)
[0176] yt = Vht
[0177] Among them, x1 - x3 refer to the inputs. For example, x1 is the GDS feature information, x2 is the process information, and x1 and x2 are constant values; x3 represents the formula of the previous step.
[0178] U, W, b, and V are all neural network parameters, and their meanings are well-known. These are the ones that will be corrected and changed through continuous training, and these network parameters are all random values initially. t represents the current moment. In some embodiments, U is the weight matrix from the input layer to the hidden layer; V is the weight matrix from the hidden layer to the output layer; W is the weight with the value of the previous hidden layer as the input for this time. N represents the Nth layer. The structure of the neural network consists of many layers that form a neural network, and each layer has many neurons.
[0179] Figure 4 In it, the formula of h1 is expressed as: the output result of the first layer of the hidden layer is equal to the function U(x1) plus the subsequent terms. U is a composite function because there are several neurons in one layer of the neural network, and each neuron has parameters. It is impossible to write out the formula of each neuron when listing the formula. The final input of this layer is the result obtained by comprehensively calculating all neurons together. Therefore, a function symbol U is used to represent this comprehensive function.
[0180] The training of the neural network requires many tuples like <s0, a, s1, r>. Among them, s0 represents the current formula, which is the input of the neural network, that is, Figure 4 h0 in. Actually, when recording, in addition to recording a_t (the action a taken at time t), the action a_t - 1 taken at the previous moment will also be recorded. Assuming the action of a_t - 1 is to multiply by space, then conversely, the formula of the previous moment of S0 is equal to the formula at S0 divided by space, and it can be deduced by reverse.
[0181] The formula here actually reflects the core characteristics of the neural network with the RNN structure. Its main meaning is that the output y at a certain moment also depends on the inputs at the previous t - 1 moments. N represents how many hidden layers are constructed to remember the previous moments. For example, when N is equal to 5, it means that 5 hidden layers need to be constructed to remember. Obviously, the larger N is, the more is remembered, but at the same time, the network structure becomes more complex, the training computational amount becomes larger, and the training difficulty also becomes greater. Therefore, the value of N is set according to specific circumstances. T can actually be understood as equal to N. Writing it as T is just to accurately convey the core characteristics of the neural network with the RNN structure, that is, to consider the previous outputs. The most ideal situation is to consider the outputs at all previous moments. However, since N is a definite number, actually what is calculated here is the outputs at the N time steps from T - N to T. Here, the time step refers to making a prediction using the neural network (this time step is different from the time step involved in calculating the surface height previously). Making a prediction is one time step. Assuming N = 5, it means that when the neural network makes a prediction each time, it will find the results of the previous five predictions and input them into the corresponding hidden layer for calculation to obtain the final output of the current step.
[0182] From the above description, it can be seen that in some embodiments, the process of making the agent (RNN neural network) make a prediction can be generally described as follows:
[0183] Preferably, there is a dataset, which is usually provided by users who use the model (i.e., those semiconductor manufacturing plants in reality). The user will give the GDS characteristics (density, width, etc.), process parameter settings of each data point, and the actual ground results collected for each point, which are the so-called "labels" in deep learning (machine learning). After having the labels, a deep reinforcement learning neural network can be constructed using the labeled data. The network parameters (weights) of this neural network are initially random. The neural network is set to receive an initial formula as input and output the next mathematical symbol. From this, the training starts. Each time the neural network generates an output, this output plus the input forms a new formula, which can calculate the ground surface height. Then, the calculated ground height is compared with the labeled data to obtain an error value. Similarly, the old formula can also calculate the ground height to obtain the error with the label (which is used to measure the entire dataset). Since the goal is to make the formula obtain good fitting accuracy at each point as much as possible, each point should be considered. That is, when calculating the error of the formula, not only a certain point is calculated, but all points in the labeled data are calculated. Therefore, the previously mentioned metrics, such as the range and standard deviation, are statistical metrics for a group of samples rather than for a certain data point. It is necessary to determine whether the error value has decreased or increased due to the addition of a new mathematical symbol (that is, whether the new formula given by the neural network is good or bad), so as to obtain a reward or punishment (positive feedback or negative feedback), and then there is a data combination. This combination includes the following content: the current formula (as the input of the neural network, labeled as the current state S0), the action taken or called the action. The action taken is the output of the neural network, which is a mathematical symbol, labeled as a), the new formula (that is, the new formula obtained by adding the output to the old formula, labeled as S1), and the obtained feedback R. If the new formula fits the data better, it is positive feedback, otherwise it is negative feedback. Then such a data combination is a training sample of the deep reinforcement learning neural network. Many attempts can be made, and thus many such data combinations are obtained. Such data combinations are the samples used to train the deep reinforcement learning neural network. These samples can be provided to the deep reinforcement learning neural network to train the deep reinforcement learning neural network and continuously adjust the network parameters (weights). When the loss curve is stable, the training is completed.
[0184] At block 206, in response to the training satisfying a predetermined condition, the neural network model corresponding to the respective updated parameters is determined as the trained neural network model, wherein the trained neural network model outputs an enhanced formula for determining the surface height.
[0185] As mentioned above, the neural network can be continuously trained to obtain a trained neural network. The trained neural network can make several steps of predictions based on the CMP basic physical formula to obtain a final enhanced CMP physical formula. Thus, this formula can be used to accurately simulate the height of the wafer surface, and the simulation value can be used as the height value of the wafer surface.
[0186] In some embodiments, the predetermined condition may refer to satisfying at least one of the following: the difference between the loss function values of two consecutive iterations is less than a predetermined difference threshold; or the absolute value of the ratio of the loss function values of two consecutive iterations is greater than a predetermined ratio threshold, etc. The embodiments of the present disclosure are not limited thereto, and other predetermined conditions can be adopted as needed.
[0187] In the above embodiments, a deep learning neural network is mainly used as an example for illustration. It should be understood that the manner of the embodiments of the present disclosure is not limited thereto, and various changes can be made according to actual needs to determine the correction term, and then an enhanced CMP physical formula can be obtained to determine the surface height value.
[0188] The present disclosure also provides a simulation model configured to execute the method according to the above embodiments.
[0189] The present disclosure also provides a neural network, including: a plurality of input layers configured to receive the feature information of the patterns on the wafer, the process information, and the initial formula for calculating the surface height of the polished wafer; an intermediate layer configured to execute the method according to the present disclosure configured to execute the method according to the above embodiments to determine an updated formula for calculating the simulation value of the surface height; and an output layer configured to output the updated formula.
[0190] In the embodiments of the present disclosure, an electronic device is also disclosed. The electronic device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, and the instructions, when executed by the processor, cause the device to perform actions, the actions including: determining a simulation value of the surface height of the polished wafer based on the time step and the corresponding polishing rate; determining a correction term related to the polishing rate based on the comparison between the simulation value and the target value so that the difference between the simulation value corresponding to the correction term and the target value becomes smaller; and in response to the difference satisfying a predetermined condition, determining the simulation value corresponding to the correction term as the height value of the surface of the polished wafer.
[0191] The embodiments of the present disclosure also disclose a computer-readable storage medium having a computer program stored thereon, and the program, when executed by a processor, implements the method for semiconductor device simulation according to the embodiments of the present disclosure.
[0192] The embodiments of the present disclosure also disclose a neural network model. The neural network model can be obtained by the training method of the neural network model in the foregoing embodiments.
[0193] Some embodiments of the present disclosure provide methods for semiconductor device simulation. It should be noted that the examples given in the above embodiments are only for illustrating the solutions of the embodiments of the present disclosure and do not limit the solutions of the present disclosure.
[0194] Some embodiments of the present disclosure introduce the method of symbolic regression to expand formulas and models. Some embodiments of the present disclosure introduce the current state-of-the-art deep reinforcement learning model for the solution and optimization of symbolic regression. Some embodiments of the present disclosure are based on the expansion of the basic CMP physical formula rather than starting from scratch, and consider the complexity of the formula and the fitting degree to the process, which is feasible and ensures the rationality of the theory. Some disclosed embodiments derive and expand physical formulas based on current deep reinforcement learning and symbolic regression techniques, which can better improve the accuracy and prediction fitting ability of the model. Symbolic regression technology refers to constructing a model (here a reinforcement learning neural network) to predict mathematical symbols.
[0195] The model of the embodiments of the present disclosure can actually be regarded as a flexible and variable "engineer", which can innovatively and flexibly respond to different processes in actual situations from the most basic and error-free CMP process perspective, and create models suitable for their respective situations.
[0196] Some embodiments of the present disclosure extend and expand the original physical model, so the new model still has physical meaning. Some embodiments of the present disclosure can intelligently learn when facing different processes and different measurement data. This learning is based on the most general CMP basic physical formula, and considers the model complexity and process fitting degree, and can give the most suitable CMP physical model for various processes.
[0197] It should be understood that the embodiments shown in the drawings are only for schematically showing the solutions of some embodiments of the present disclosure and do not limit the present disclosure. The embodiments of the present disclosure can also have various other forms.
[0198] Figure 5 A schematic block diagram of an electronic device according to some exemplary embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0199] As Figure 5As shown, device 500 includes a CPU 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of device 500 can also be stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0200] Multiple components in device 500 are connected to the I / O interface 505. The multiple components include: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0201] Each of the processes and processes described above, such as method 200, can be executed by the CPU 501. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly included in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the CPU 501, one or more steps of method 200 described above can be executed.
[0202] The solution according to an embodiment of the present disclosure can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions for performing various aspects of the present disclosure stored thereon. The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable program instructions can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network.
[0203] The embodiments of the present disclosure have been described above. The above description is exemplary and is only an alternative embodiment of the present disclosure, not exhaustive, and is not used to limit the present disclosure. Although the claims in this application have been formulated for specific combinations of features, it should be understood that the scope of the present disclosure also includes any novel feature or any novel combination of features that are explicitly or implicitly disclosed herein or any generalization thereof, regardless of whether it relates to the same solution in any of the currently claimed claims. The applicant hereby informs that new claims may be formulated for these features and / or combinations of these features during the examination of this application or in any further application derived therefrom.
[0204] The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary technicians in the technical field to understand the embodiments disclosed herein. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A training method for a neural network model, comprising: Training the neural network model using training samples to determine a gradient of a loss function of the neural network model, wherein the training samples are generated based on an initial formula for determining a surface height of a wafer after grinding; Updating the parameters of the neural network model based on the gradient of the loss function; as well as In response to the training satisfying a predetermined condition, the neural network model corresponding to the corresponding updated parameters is determined as a trained neural network model, wherein the trained neural network model outputs an enhanced formula for determining the surface height.
2. The method according to claim 1, wherein each training sample is an array consisting of the following items: The initial formula; An initial correction term, wherein the initial correction term is used to correct the initial formula; an update formula, the update formula being generated based on the initial formula and the initial correction term; as well as A reward item representing a feedback value generated based on a difference between a simulation value and a target value of the surface height.
3. The method according to claim 2, wherein the reward item is generated by: Determine a first difference between a standard deviation of a simulated value of a surface height of each grid point on the wafer and a standard deviation of the target value; Determine a second difference between a root mean square error of a simulated value of the surface height of each grid point and a root mean square error of the target value; Determine the range between the simulated value of the surface height of each grid point and the target value; as well as The reward item is generated based on the first difference, the second difference, and the range.
4. The method of claim 1, wherein the neural network model is a reinforcement learning neural network model, and wherein training the neural network model using training samples to determine the gradient of the loss function of the neural network model comprises: The policy gradient algorithm in the field of reinforcement learning is used to back-propagate the training samples to determine the gradient of the loss function.
5. The method according to claim 1, wherein updating the parameters of the neural network model based on the gradient of the loss function comprises: In response to performing a predetermined number of trainings, performing a gradient update on the gradient of the loss function; as well as The parameters of the neural network model are updated based on the gradient update.
6. The method according to claim 1, further comprising: Determining a loss function value of the neural network model; And wherein, the training satisfies the predetermined conditions including: The difference between the loss function values of two previous and subsequent iterations is less than a predetermined difference threshold; or An absolute value of a ratio of the loss function values of two previous and subsequent iterations is greater than a predetermined ratio threshold.
7. A method for generating a training sample, comprising: Predicting based on an initial formula and an initial correction term to generate an updated formula and a reward term, wherein the initial formula is used to determine the surface height of the wafer after grinding, the initial correction term is used to correct the initial formula, and the reward term represents a feedback value generated based on the difference between a simulation value and a target value of the surface height; Using an array consisting of the initial formula, the initial correction term, the update formula, and the reward term as a first training sample; as well as Iteration is performed in the neural network model based on the first training sample to generate multiple training samples.
8. The method according to claim 7, wherein: In response to the training samples generated in each round of iteration satisfying a predetermined condition, the round of iteration is terminated, and all the training samples in the round of iteration are regarded as a trajectory; as well as In response to the sum of the numbers of each trajectory obtained in each round of iteration reaching a predetermined trajectory threshold, terminating the iteration; as well as The training samples in each of the trajectories are used as training samples for training the neural network model.
9. The method according to claim 8, wherein making predictions based on the initial formula and the initial correction term to generate an updated formula comprises: Generate random numbers at the beginning of each iteration; selecting a source of an initial correction term for correcting the initial formula based on a comparison of the random number with a random probability threshold; as well as The initial formula is modified based on the selection to generate an updated formula.
10. The method according to claim 9, further comprising: The mathematical symbols and the variable library are combined by Cartesian product to generate a symbol library, wherein the variable library contains graphic feature information and process information.
11. The method of claim 9, wherein modifying the initial formula based on the selection to generate an updated formula comprises: In response to a first comparison result between the random number and the random probability threshold, selecting a symbol from the symbol library as an initial correction term to correct the initial formula using the neural network model; as well as In response to a second comparison result between the random number and the random probability threshold, prediction is performed using the neural network model to generate a predicted initial correction term to correct the initial formula, and the predicted initial correction term is added to the symbol library.
12. A neural network model generated according to the method according to any one of claims 1 to 6.
13. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, the instructions causing the device to perform actions when executed by the processor, the actions comprising: Training the neural network model using training samples to determine a gradient of a loss function of the neural network model, wherein the training samples are generated based on an initial formula for determining a surface height of a wafer after grinding; Updating the parameters of the neural network model based on the gradient of the loss function; as well as In response to the training satisfying a predetermined condition, the neural network model corresponding to the corresponding updated parameters is determined as a trained neural network model, wherein the trained neural network model outputs an enhanced formula for determining the surface height.
14. A computer-readable storage medium having machine-executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the method according to any one of claims 1 to 11.
Citation Information
Cited By
Simulation method and device for metal surface deposition morphology and readable storage medium
CN122154629A
A simulation method and device for metal surface deposition morphology and a readable storage medium
CN122154629B