Determination of recipe for manufacturing semiconductor

Machine learning is employed to streamline semiconductor manufacturing recipe development by integrating experimental and virtual simulation results, addressing inefficiencies in traditional manual methods and enhancing process optimization.

JP2025106555APending Publication Date: 2025-07-15LAM RES CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025067233
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-23
Filing Date
2025-04-16
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The complexity of semiconductor manufacturing processes with numerous adjustable parameters and interdependent subsystems leads to lengthy and costly recipe development, often requiring extensive manual experimentation and analysis, which is time-consuming and inefficient.

Method used

A method utilizing machine learning (ML) to determine semiconductor manufacturing recipes by combining experimental results with virtual simulations, training an ML algorithm to create a new recipe based on specified processing requirements, thereby reducing the need for physical tests and shortening the development time.

Benefits of technology

This approach significantly reduces experimental costs and time by enabling faster and more accurate recipe development, allowing for efficient process optimization and improved understanding of process parameters through ML model-based process correlation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106555000001_ABST
    Figure 2025106555000001_ABST
Patent Text Reader

Abstract

To provide methods, systems, and programs for determining a recipe for manufacturing a semiconductor with the use of machine learning (ML) to accelerate the definition of recipes.SOLUTION: A method includes each experiment being controlled by a recipe, from a set of recipes, that identifies parameters for manufacturing equipment. The method includes an operation for performing experiments for processing a component and an operation for performing virtual simulations for processing the component, each simulation being controlled by one recipe from the set of recipes. An ML model is obtained by training an ML algorithm using experiment results and virtual results from the virtual simulations. The method further includes: an operation for receiving specifications for a desired processing of the component; and an operation for creating, by the ML model, a new recipe for processing the component based on the specifications.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Priority Claim] This application claims priority based on U.S. Provisional Patent Application No. 62 / 925,157, filed on October 23, 2019, entitled "Determination of Recipe for Manufacturing Semiconductor". This provisional application is hereby incorporated by reference in its entirety.

[0002] [Related Application] This application is related to U.S. Patent Application No. 16 / 260,870, filed on January 29, 2019, entitled "Fill Process Optimization Using Feature Scale Modeling", and is hereby incorporated by reference in its entirety.

[0003] The subject matter disclosed herein generally relates to methods, systems, and machine-readable storage media for manufacturing semiconductors.

Background Art

[0004] The background art described herein is for presenting the content of the present disclosure generally. The work of the inventors, currently named, is not admitted as prior art to the present disclosure, either expressly or implicitly, to the extent described in this background art section and in aspects of the description that do not fall within the scope of the prior art at the time of filing.

[0005] Chemical reaction apparatuses used in the deposition process development of semiconductor chips tend to have many interdependent subsystems (e.g., sensors, actuators, gas supply, power supply, coordination network). These subsystems are independently controlled by process parameters that follow a set of instructions included in the recipe of the control system. The operations of these subsystems collectively determine the output performance in the product wafer.

[0006] The increasing complexity of current process equipment means that the number of system components has increased to address this complexity, thereby increasing the number of process "knobs" (e.g., adjustable process parameters) in the system. Various system states (pressure, temperature, flow setpoints, etc.) are factors that play a role in the desired output to the wafer, such as the step coverage of the film, the non-uniformity of the film, and the etching depth.

[0007] The proper setting of process parameter values is an important issue in the semiconductor device industry, and it often takes weeks or months of process development to obtain a recipe that can set all these process parameters to obtain components that meet the desired condition criteria. SUMMARY OF THE INVENTION

[0008] A method, system, and computer program for determining a recipe for manufacturing a semiconductor using machine learning (ML) are presented to facilitate the definition of the recipe. A general aspect includes a method including steps for conducting experiments for processing components, each experiment being controlled by one recipe from a set of recipes that specify the parameters of the manufacturing equipment. This method further includes steps for conducting virtual simulations for processing components, each simulation being controlled by one recipe from the set of recipes. The ML model is obtained by training an ML algorithm using the experimental results and the virtual results from the virtual simulations. This method further includes steps for receiving the specification of the desired processing of the components and, based on that specification, creating a new recipe for processing the components by the ML model. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The various attached drawings merely represent exemplary embodiments of the present disclosure and are not considered to limit its scope.

[0010]

Figure 1

[0011]

Figure 2

[0012]

Figure 3

[0013]

Figure 4

[0014]

Figure 5A

[0015]

Figure 5B

[0016]

Figure 6

[0017]

Figure 7

[0018]

Figure 8

[0019]

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0020] Exemplary methods, systems, and computer programs are directed to determining a recipe for manufacturing a semiconductor. The examples are merely illustrative of possible variations. Unless specified otherwise, components and functions are optional, may be combined or subdivided, and the steps may differ in sequence or may be combined or subdivided. In the following description, for purposes of explanation, some specific details are set forth in order to provide a thorough understanding of exemplary embodiments. However, it will be apparent to one skilled in the art that the subject matter may be practiced without these specific details.

[0021] FIG. 1 shows a process for finding a semiconductor manufacturing recipe according to some exemplary embodiments. Current approaches to recipe development are based on historical learning from subject matter experts on the process engineering team.

[0022] Product requirements 102 are provided to subject matter expert 104 who plans experiments (step 106) based on past experience. An example of product requirements 102 is the deposition amount required for a feature.

[0023] The experiments include recipes created by subject matter expert 104 based on product requirements 102. However, finding the correct recipe can be a difficult process, especially when the product requirements are on the order of 1 nm accuracy.

[0024] The experiments are conducted in a laboratory 108, and the results are measured and compared 110 to the original product requirements 102. Some exemplary result criteria include uniform film thickness and resistivity of the conductor.

[0025] A comparison 112 is performed to determine whether the result is suitable to match the product requirement 102. If the result is appropriate, a recipe is found 114. If the result is not appropriate, the method returns to step 106, where a skilled person fine-tunes the recipe based on the result to improve the recipe to approach the product requirement 102.

[0026] As the number of processing parameters in the system increases, it gradually becomes difficult to truly extract and learn the influence of individual factors and their influence on the desired results. Manufacturing tools are complex and have many adjustable parameters such as gas flow, plasma characteristics, and thermal characteristics. For example, a typical atomic layer deposition (ALD) process has 100 - 200 adjustable parameters.

[0027] General experimental techniques such as single variable testing (SVT) and design of experiments (DOE) (e.g., full factorial, screening plan, response surface model, mixture model, Taguchi orthogonal) not only determine the influence of experimental factors on the result but also enable learning the correlations between different factors. However, these methods require many experiments to determine the influence of various processing parameters. To obtain a near-perfect recipe, a process engineer or their team usually needs to conduct many tests over several weeks (or months), increasing the cost of wafer development.

[0028] Furthermore, when testing for one parameter, there is no optimization for other parameters, and the experiment may fall into a local minimum. Thus, when multiple parameters are specified simultaneously, a skilled person may not be able to develop the recipe in the correct direction.

[0029] The limitation of this method is that it is extremely manual, requiring an extended tool time to conduct the tests to implement the method, and additional time for the analysis to make changes for the subsequent series of tests. This results in multiple cycles of learning, which can take months in complex cases. Usually, the factors subject to experimentation do not include all possible variables, and the correlation coefficient of the model suggests the degree of predictability of the model. There are blocking plans and screening plans for testing more factors in the experimental setup, but expertise is required for planning the experimental setup. Ultimately, the result is the same, with many tests required, which are costly and time-consuming.

[0030] Figure 2 represents the use of a simulation tool for conducting virtual experiments using a recipe, according to some exemplary embodiments. The embodiments presented herein describe a method for finding a process recipe using machine learning (ML) and a simulation tool 206.

[0031] Process simulation is used to test an exemplary recipe 202 in order to minimize the costs incurred for physical tests and related methods and to shorten the process development time (reduce the experimental and learning cycles). An example of a simulation tool is SEMulator3D by Lam Research Corporation, which provides a voxel model of a semiconductor process based on the predicted behavior simulated by a behavior model 210. However, the same principle may apply to other simulation tools.

[0032] The behavior model 210 explains the output of the process based on an analytical formula. The behavior model 210 qualitatively captures the process but does not provide guidance for process development in terms of the quantities that define the process parameters.

[0033] For example, regarding the deposition process, assume there are two surfaces B, where deposition selectively occurs on one of the two surfaces and no deposition occurs on the other surface. The model measures the deposition thickness across the entire surface. For example, in some regions, the deposition thickness is 0.72, and in other regions, it is 0.71. The behavior model examines the behavior of particles in the chamber, but rather measures the result of the behavior (e.g., deposition thickness).

[0034] As an example, an ALD system may have hundreds of parameters 204 that define recipe 202. A non-exhaustive list of these parameters 204 includes general tool parameters (e.g., number of ALD cycles, presence or absence of soak, tool process mode, etc.), fluid parameters (e.g., flow rates of various gases, fluid concentration, dilution gas, non-reactive gas for station isolation, etc.), chamber conditioning or precoat parameters (e.g., precoat temperature, precoat time, etc.), pressure parameters (e.g., chamber pressure, reservoir pressure, throttle valve angle, precursor ampule pressure, vacuum clamp pressure, etc.), nucleation chemical parameters (e.g., dose time, dose flow, chemical concentration, cycle, etc.), temperature parameters (e.g., pedestal temperature, ampule temperature, chamber temperature, showerhead temperature, etc.), ALD time adjustment parameters (e.g., dose time of precursor A, purge time, dose time of precursor B, preheat time, etc.), and various other parameters. These parameters may be used as features of the ML model as described below with respect to Figure 4.

[0035] In some examples, there is a linear dependence relationship for the parameters regarding the predicted output, but in certain cases, there is a non-linear dependence relationship. Predicting these dependence relationships and constructing an appropriate model is not easy. Furthermore, the methods used today do not constitute placement dependence (i.e., upstream process-intensive steps).

[0036] The simulation tool 206 constructs a three-dimensional model of what occurs on the substrate when the recipe 202 is executed through the process and generates simulation results measured by the measurement method 214. The measurement method 214 provides the simulation results 212 and includes items such as layer thickness and resistivity. Image analysis may be used to examine the simulation results 212, but other types of measurement methods 214 may also be used.

[0037] The physical model 208 is an explanation of the physical behavior on the substrate and is usually based on first principles, but may also be empirically used using ML and statistical methods with bases in physics and chemistry.

[0038] The physical model 208 incorporates chamber parameters such as pressure, temperature, and species flux (the number of particles crossing a unit area per second). The physical model 208 analyzes these parameters to predict the behavior of physical particles (e.g., species flux) that affect the process. For example, the flux value affects the deposition thickness, meaning that the higher the flux value, the thicker the deposition thickness compared to a lower flow rate value. These parameters may be used as features of the ML model as described below with respect to FIG. 4.

[0039] In some exemplary embodiments, a bridge is generated that links the behavior model 210 and the actual process recipe via the physical model 208 by some correlation methods. These correlation methods include multivariate regression methods, neural networks, decision trees, support vector machines (SVMs), etc. In some exemplary embodiments, the physical model 208 may be a combination of multiple models showing different aspects of the physical model. Further, in some exemplary embodiments, the behavior model 210 may be a combination of multiple models where each model includes different behavior aspects.

[0040] The simulation tool 206 generates simulation results 212 using the behavior model 210 and the physical model 208. Since the simulation results 212 are actually virtual results because no actual experiments have been conducted.

[0041] The simulation results 212 are compared with the product requirements 102 to determine whether the recipe 202 meets the product requirements 102. If the simulation results 212 are met, the working recipe is found. Otherwise, new simulations can be performed with different recipes 202 to continue searching for the correct recipe 202.

[0042] A successful simulation means fewer tests, leading to time and cost savings. For unsuccessful simulations, the results are fed back into the model to improve accuracy in future predictions.

[0043] Figure 3 represents the use of machine learning to facilitate the definition of recipes according to some exemplary embodiments. ML is an application that provides a computer system with the ability to perform tasks without being explicitly programmed by inferring based on patterns found in data analysis. ML relies on data that can learn from the data to make inferences.

[0044] In some exemplary embodiments, the data for ML includes experimental results 306 from actual experiments 302 conducted on semiconductor manufacturing tools and virtual results 308 from simulations 304. The experiments 302 and simulations 304 may use the same or different recipes 202. This method can be applied to multiple semiconductor manufacturing processes such as deposition, etching, and cleaning.

[0045] Performing the actual experiment 302 is costly and time-consuming. However, performing the simulation 304 is much faster and less expensive. Therefore, many simulations 304 may be performed (e.g., 10 to 1000 or more) to obtain many virtual results 308 that can be used to train a machine learning algorithm. For example, the simulation 304 may be performed by changing the values of these parameters of interest to enable prediction of how the parameters of interest can be adjusted to create a new recipe 316.

[0046] In step 310, the ML algorithm is trained using data from the experimental results 306 and the virtual results 308. The result of the training in step 310 is an ML model 314 configured to receive a desired component 312 (e.g., product requirements) and create a new recipe 316. In some exemplary embodiments, the number of experiments is 10 to 100, but other numbers are possible. Further, the number of simulations 304 is 100 to 100,000, but other numbers are possible.

[0047] In some exemplary embodiments, there are few experimental results 306, and since the actual experimental results 306 are more accurate than the virtual results 308 obtained from the simulation 304, the experimental results 306 are given a greater weight than the virtual results 308 for training (step 310). Further, other data such as data obtained from an experiment library may be used for training. More details about ML are provided below in connection with FIG. 5A.

[0048] In searching for the best recipe that satisfies the product requirements 102, the ML model 314 may be used repeatedly several times. To improve the accuracy of the ML model 314, the new recipe 316 may be used in the data used for the experiment (real or virtual) and the training of the ML algorithm.

[0049] The embodiments described in this specification associate chamber set points and sensor data for a simulator behavior model calibrated based on real measurement data (such as image data, film property data, etc.). In some exemplary embodiments, it is possible to use prior knowledge to determine a new design space by methods such as Bayesian inference. Further, guidance by physics-based modeling can be used in conjunction with the behavior model to further improve the accuracy in the correlation process of process parameters and model outputs.

[0050] The influence of process parameters on process performance is inferred based on virtual model results. This not only accelerates the understanding of the model-based process, but also makes it easier for process engineers and technicians to understand by associating the process with tool parameters, and can lead to process optimization on the tool. As a result, the presented solutions reduce experimental costs and time. With less repetition, the process efficiency is higher.

[0051] In some exemplary embodiments, different ML models 314 are used for different semiconductor manufacturing processes. For example, one ML model is created for the deposition process (e.g., using deposition experiments and simulations), another model is created for the etching process, and another model is created for substrate cleaning.

[0052] One advantage of the ML model 314 is that it can explore not only process parameters but also physical parameters in the search for a new recipe 316.

[0053] An example of an application is deposition using a suppression profile. A behavior model calibrated by adjusting the behavior of a deposition-suppression-deposition model or a deposition-etching-deposition model is described in U.S. Patent Application No. 16 / 260,870, filed on January 29, 2019, entitled "Fill Process Optimization Using Feature Scale Modeling", which is incorporated herein by reference. The calibrated behavior model is associated with experimental variables such as dose time, purge time, flow rates of various suppression chemicals, system pressure, wafer temperature, and molecular transport for each recessed feature shape. These variables from a test set (a small number of samples on which experiments are conducted) are used to train an ML model that fills the gap between the process results and the simulation behavior based on one or more ML models calibrated using an optimization method such as the gradient descent method. Key parameters that have a significant impact on the process are extracted and verified experimentally to obtain an ideal recipe based on the sample space on which the model is calibrated. This model takes into account the shape of the structure, which is often overlooked by other analysis methods. Details of this process are provided below in relation to Figure 6.

[0054] Another example of deposition is for the filling and roughness control of 3D NAND WL (word lines). There, the reactant species in the ALD system need to travel not only the full length of the high aspect ratio (HAR) structure but also flow laterally inside the WL. In addition to the challenges related to molecular transport, these processes are typically performed very quickly to match the throughput expected by the customer (less than 1 second per cycle for the precursor dose-purge-reductant dose-purge steps).

[0055] Roughness due to particle growth can lead to pinch-off and void formation and can be adjusted by suppressing growth in specific regions. The model predicts the profile behavior based on reaction-diffusion models in the vertical and lateral structures. Experimental data enables calibration of the model and its association with tool parameters such as precursor dose time, chemical purge time, reducing agent dose time, inhibitor molecule dose time, system pressure, wafer temperature, and the shape of the structure modeled in the simulator. Based on the sample set used for data training, the ML model is used to associate the results of the optimal solution (e.g., void-free film, low roughness film, possibility of post-deposition etch-back, etc.) based on process parameters.

[0056] FIG. 4 shows some of the features 402 used in a machine learning program according to some exemplary embodiments. Features are used by the ML algorithm to represent data. One feature is an individual measurable property of an observed event. The concept of a feature is related to the concept of an explanatory variable used in statistical techniques such as linear regression. In effective processes of ML in pattern recognition, classification, and regression, it is important to select informative, discriminative, and independent features. Features can be of different types such as numerical features, strings, and graphs.

[0057] In some exemplary embodiments, the features 402 of the ML algorithm used to find a process recipe include recipe features 404, experimental result features 406, virtual result features 408, and measurement features 410. Other models may use additional features or some of these features.

[0058] Recipe features 402 include parameters associated with the recipe (workflow, gas flow, chamber temperature, chamber pressure, process duration, radio frequency (RF) values (e.g., frequency, voltage), etc.).

[0059] The experimental result feature 406 and the virtual result feature 408 include values measured from the resulting semiconductor (conformality, lateral ratio, isotropic ratio, deposition depth, global adhesion coefficient, surface-dependent adhesion coefficient, retardation thickness, neutral particle-ion ratio, ion angle distribution function, etc.).

[0060] The measurement feature 410 includes criteria used by measurement methods such as imaging methods (e.g., operating electron microscope (SEM), transmission electron microscope (TEM)), thickness measurement (e.g., fluorescent X-ray analysis (XRF), polarization analysis method), sheet resistance, surface resistivity, stress measurement, and other analysis methods used to determine layer thickness. These other analysis methods include one or more of X-ray diffraction method (XRD), X-ray reflectivity method (XRR), pre-electron diffraction method (PED), electron energy loss spectroscopy (EELS), energy dispersive X-ray spectroscopy (EDS), secondary ion mass spectrometry (SIMS), etc.

[0061] In some exemplary embodiments, the measurement method includes time-series data, and the time-series data includes sensor measurement values taken over time for specific parameters such as how the chamber pressure evolves over time during the manufacturing process.

[0062] Figure 5A illustrates the training and use of a machine learning program according to some exemplary embodiments. In some exemplary embodiments, a machine learning program (MLP), also referred to as a machine learning algorithm or machine learning tool, is used to perform steps related to determining a recipe for manufacturing a semiconductor.

[0063] Machine learning is an algorithm, also referred to herein as a tool, that explores the study and construction of algorithms that can learn from existing data and make predictions about new data. Such machine learning algorithms operate by constructing an ML model 314 from exemplary training data 512 in order to make predictions or decisions based on data represented as an output or assessment (e.g., discovery of a new recipe 316). Exemplary embodiments are presented with respect to several machine learning tools, but the principles presented herein may also apply to other machine learning tools.

[0064] ML has two common modes: supervised ML and unsupervised ML. Supervised ML uses prior knowledge (e.g., examples associating inputs with outputs or results) to learn the relationship between inputs and outputs. The goal of supervised ML is to learn the function that best approximates the relationship between training inputs and outputs given some training data so that the ML model can provide the same relationship when given an input to generate the corresponding output. Unsupervised ML is the training of ML algorithms that use information that is neither classified nor labeled, enabling the algorithm to act on that information without guidance. Unsupervised ML is effective in exploratory analysis because it can automatically identify the structure of the data.

[0065] Common tasks for supervised ML are classification problems and regression problems. A classification problem, also called a categorization problem, aims to classify items into several category values (e.g., is this object an apple or an orange?). A regression algorithm aims to quantify several items (e.g., by providing scores for several input values). Some examples of commonly used supervised ML algorithms are logistic regression (LR), naive Bayes, random forest (RF), neural network (NN), deep neural network (DNN), matrix factorization, and support vector machine (SVM).

[0066] Some common tasks of unsupervised ML include clustering, representation learning, and density estimation. Some examples of commonly used unsupervised ML algorithms are k-means, principal component analysis, and autoencoders.

[0067] In some embodiments, an exemplary machine learning algorithm determines a new recipe 316 for manufacturing a semiconductor. The machine learning algorithm uses training data 512 to find correlations between the identified features 402 that affect the result. In one exemplary embodiment, the features may be of different types and may comprise the features 402 described above with respect to FIG. 4.

[0068] During the training process 310, the ML algorithm analyzes the training data 512 based on the identified features 402 and the configuration parameters 511 defined for training. The result of the training process 310 is an ML model 314 that can be input to create an assessment.

[0069] Typically, training of an ML algorithm involves analyzing large amounts of data (e.g., from gigabytes to terabytes or more) to find data correlations. The ML algorithm uses the training data 512 to find correlations between the identified features 402 that affect the result or assessment (e.g., the new recipe 316). In some exemplary embodiments, the training data 512 includes labeled data that is known data for one or more identified features 402 and one or more results (such as measured values).

[0070] The ML algorithm typically explores many possible functions and parameters before finding what it identifies as the best correlation in the data. Therefore, training may require significant computing resources and time.

[0071] Many ML algorithms have configuration parameters 511, and the more complex the ML algorithm, the more parameters 511 are available to the user. Configuration parameters 511 define the variables of the ML algorithm in the search for the best ML model. Configuration parameters include model parameters and hyperparameters. Model parameters are learned from training data, while hyperparameters are not learned from training data but are instead provided to the ML algorithm.

[0072] Some examples of model parameters include the maximum model size, the maximum number of passes through the training data, the type of data shuffling, regression coefficients, decision tree split positions, and the like.

[0073] In some exemplary embodiments, representative model parameters are scalar attributes or context attributes. Examples of scalar attributes are determined physical model parameters or behavior model parameters such as deposition rate or etch depth. Further, a context attribute is an attribute that depends on other attributes (e.g., context) and may have physical, statistical, and machine learning-based relationships. In particular, a physical context may be the deposition rate with respect to the aspect ratio.

[0074] Examples of statistical context include parameter reduction using principal component analysis (PCA) or linear discriminant analysis (LDA). PCA is a dimensionality reduction method used to reduce the dimensionality of a large set of data by transforming the large set of variables into a small set that contains most of the information of the large set. The small set of data is much easier and faster to explore, visualize, and create analysis data for machine learning algorithms without variables irrelevant to the process.

[0075] Linear discriminant analysis (LDA) is a method used to find a linear combination of features that characterize or separate two or more classes of objects or events. The resulting combination may be used as a linear classifier or for dimensionality reduction prior to classification.

[0076] Examples of context by machine learning are autoencoders, neural networks, or trained regressors. These scalar or context attributes are representative model parameters when verified as representatives of experimental data and are used as inputs for the next modeling task. These parameters are called "virtual results" or "simulation results" when they are the results of simulation work.

[0077] Hyperparameters may include the number of hidden layers in a neural network, the number of hidden nodes in each layer, the learning rate (possibly by various adaptation methods for the learning rate), the regularization parameter, the type of non-linear activation function, and so on. Finding the correct (or best) set of hyperparameters can be a very time-consuming task that requires a large amount of computer resources.

[0078] When the ML model 314 is used to perform an assessment, the specification 518 is provided as input to the ML model 314, and the ML model 314 creates a new recipe 316 as output.

[0079] Figure 5B depicts the use of a machine learning program using active process control according to some exemplary embodiments. In some exemplary embodiments, the purpose of the final recipe is active process control. Active process control is a process correlation method for compensating for incoming changes, downstream changes, or environmental changes. For example, the number of deposition cycles can be increased to compensate for work performed on a larger structure in response to the previous output when working on a particular structure.

[0080] The trained ML model 314 can be arranged to determine which process parameters meet the control target. The input 520 includes the recipe and control specifications for the desired active process control. The resulting new recipe with control parameters 522 may include setpoints that depend on local control requirements during the process execution time.

[0081] FIG. 6 shows an example of a deposition-inhibit-deposition (DID) deposition process using inhibit control enhanced (ICE) fill that can be optimized using a behavior model. In some exemplary embodiments, the behavior model uses an abstraction of the process to predict the details of the structure of components (parts, components) manufactured by one or more semiconductor device manufacturing processes. Examples of behavior models are described in U.S. Pat. No. 9,015,016 and U.S. Pat. No. 9,659,126, which patents are incorporated herein by reference.

[0082] The pre-fill stage 606 shows an unfilled component 602. The component 602 may be formed on one or more layers on a semiconductor substrate and may have one or more layers along the sidewalls and / or bottom of the component 602 as needed. The purpose is to prevent voids in the fill of the component 602.

[0083] Stage 608 shows the component 602 after an initial deposition of a fill material to form a layer of the material 604 that fills the component 602. Examples of the material 604 include tungsten, cobalt, molybdenum, and ruthenium, but the techniques described herein may be used to optimize the fill of any suitable material 604, including other conductors and dielectrics such as oxides (e.g., SiO x , Ab03), nitrides (e.g., SiN, TiN) and carbides (e.g., SiC).

[0084] Stage 610 shows the component 602 after an inhibit process. The inhibit process is a process that has the effect of inhibiting subsequent deposition on the processing surface 614. Inhibition may include various mechanisms that depend on various factors including the surface being processed, the inhibit chemical, and whether the inhibition is a thermal process or a plasma process. In one example, tungsten deposition resulting from tungsten nucleation is inhibited by exposure to a nitrogen-containing chemical. This may include, for example, the generation of activated nitrogen-containing species by a remote plasma generator or a direct plasma generator, or by exposure to ammonia vapor in an example of a thermal (non-plasma) process.

[0085] Examples of the inhibition mechanism can include a chemical reaction between the activation species and the component surface to form a thin layer of a composite material such as tungsten nitride (WN) or tungsten carbide (WC). In some embodiments, the inhibition can include surface effects such as adsorption that passivate the surface without forming a layer of the composite material.

[0086] The inhibition can be characterized by an inhibition depth 616 and an inhibition gradient. That is, the inhibition can vary with depth such that it is greater at the opening of component 602 than at its bottom and extends only partway into component 602. In the illustrated example, the inhibition depth 616 is about half of the total depth of component 602. Additionally, the inhibition treatment is stronger at the upper portion of component 602. Since deposition is inhibited near the opening of component 602, during the second deposition stage Dep2 612, the material preferentially deposits at the bottom of component 602 without depositing, or with little deposition, at the opening of component 602. This can prevent the formation of voids and seams inside the filled component 602. Thus, during the second deposition Dep2 stage 612, the material 604 can be filled in a technique characterized by bottom-up filling rather than conformal first deposition Dep1 filling.

[0087] As deposition continues, the inhibition effect may be removed so that deposition on the mildly treated surface is no longer inhibited. This is represented at stage 612 where the treated surface 614 is in a narrower range than before the Dep2 stage. In the illustrated example, as the second deposition Dep2 progresses, the inhibition is ultimately removed over the entire surface and, as shown at stage 614, the component is completely filled with the material 604.

[0088] Although only one inhibition cycle is shown, this process can include several deposition-inhibition cycles. Behavior modeling is used and the recipe is fine-tuned to control the deposition and inhibition parameters so that voids in the fill are eliminated and the fill material meets the requirements. Measurement methods are used to measure different criteria of deposition and inhibition, including the appearance of voids in the fill.

[0089] FIG. 7 is an etching chamber 700 according to one embodiment. Exciting an electric field between two electrodes is one way to obtain a radio frequency (RF) gas discharge within the etching chamber. When an oscillating voltage is applied between the electrodes, the resulting discharge is called a capacitively coupled plasma (CCP) discharge.

[0090] Plasma 702 may be generated using a stable source gas to obtain various chemically reactive by-products resulting from the ionization of various molecules caused by electron-neutral collisions. The chemical aspects of etching include the reaction of neutral gas molecules and their dissociation by-products with the molecules of the surface to be etched, as well as the production of volatile molecules that can be exhausted. When the plasma is generated, positive ions are accelerated from the plasma across a space charge sheath that separates the plasma from the chamber walls and impinge on the wafer surface with sufficient energy to remove material from the wafer surface. This is known as ion bombardment or ion sputtering. However, some industrial plasmas do not generate ions with sufficient energy to efficiently etch the surface by physical means alone.

[0091] Controller 716 manages the processes of chamber 700 by controlling different elements of the chamber, such as RF generator 718, gas source 722, and gas pump 720. In one embodiment, fluorocarbon gases (such as CF4 and C-C4F8) are used in the dielectric etching process due to their anisotropic etching and selective etching capabilities, but the principles described herein are applicable to other plasma generating gases as well. Fluorocarbon gases are easily ionized into chemically reactive by-products that include smaller molecules and atomic radicals. These chemically reactive by-products etch a dielectric material that can be SiO2 or SiOCH for low dielectric constant devices in one embodiment.

[0092] Chamber 700 represents a processing chamber having an upper electrode 704 and a lower electrode 708. The upper electrode 704 may be grounded or may be connected to an RF generator (not shown). The lower electrode 708 is connected to an RF generator 718 via a matching network 714. The RF generator 718 provides RF power at one, two, or three different RF frequencies. Depending on the desired configuration of the chamber 700 for a particular process, at least one of the three RF frequencies may be turned on or off. In the embodiment shown in FIG. 7, the RF generator 718 provides frequencies of 2 MHz, 27 MHz, and 60 MHz, although other frequencies are possible.

[0093] Chamber 700 includes, on the upper electrode 704, a gas showerhead for injecting gas provided from a gas source 722 into the chamber 700, and a perforated confinement ring 712 through which gas can be exhausted from the chamber 700 by a gas pump 720. In some exemplary embodiments, the gas pump 720 is a turbomolecular pump, although other types of pumps may be used.

[0094] When a substrate 706 is inside the chamber 700, a silicon focus ring 710 is installed next to the substrate 706 so that a uniform RF field exists at the bottom surface of the plasma 702 for uniform etching of the surface of the substrate 706. The embodiment of FIG. 7 shows a triode reactor configuration in which the upper electrode 704 is surrounded by laterally symmetric RF ground electrodes 724. The insulator 726 is a dielectric that separates the ground electrode 724 from the upper electrode 704.

[0095] Each frequency may be selected for a specific purpose in the wafer manufacturing process. In the example of FIG. 7, for RF power provided at 2 MHz, 27 MHz, and 60 MHz, the 2 MHz RF power provides ion energy control, and the 27 MHz and 60 MHz powers provide control of plasma density and ionization patterns of chemicals. This configuration where each RF power can be turned on or off enables specific processes that use ultra-low ion energy on a substrate or wafer, and specific processes where the ion energy needs to be low (less than 700 eV or 200 eV), such as the soft etching of low dielectric constant materials.

[0096] In another embodiment, 60 MHz RF power is used at the upper electrode 704 to obtain ultra-low energy and ultra-high density. This configuration enables chamber cleaning by high-density plasma when the substrate 706 is not in the chamber while minimizing sputtering on the electrostatic chuck (ESC) surface. The ESC surface is exposed when the substrate 706 is absent, and any ion energy on the surface must be avoided. This is the reason why the bottom 2 MHz and 27 MHz power supplies are turned off during cleaning.

[0097] FIG. 8 is a flowchart of a method 800 for determining a recipe for semiconductor manufacturing according to some exemplary embodiments. Various steps of this flowchart are presented and described in sequence, but one of ordinary skill in the art can recognize that some or all of the steps may be performed in a different order, or may be combined or omitted, or may be performed simultaneously.

[0098] In step 802, a plurality of experiments are performed to process components. Each experiment is controlled by one of a plurality of recipes that specify parameters for the manufacturing apparatus.

[0099] From step 802, method 800 proceeds to step 804 for performing a plurality of virtual simulations for processing components. Each simulation is controlled by one of a plurality of recipes.

[0100] Method 800 proceeds from step 804 to step 806 where an ML model is obtained by training an ML algorithm using experimental results and virtual results from virtual simulations.

[0101] Method 800 proceeds from step 806 to step 808 for receiving a specification for a desired processing of a component. In step 810, the ML model creates a new recipe for processing the component based on the specification.

[0102] In one example, the ML model is based on a plurality of features including recipe features, experimental result features, virtual result features, and measurement features.

[0103] In one example, the measurement features include one or more of imaging methods, transmission electron microscopes, thickness measurements, sheet resistance, surface resistivity, stress measurements, and analysis methods used to determine layer thickness, composition, particles, or orientation.

[0104] In one example, the recipe features include workflow, gas flow, chamber temperature, chamber pressure, process duration, and radio frequency (RF) values.

[0105] In one example, the virtual simulation is performed by a simulation tool based on behavioral modeling.

[0106] In one example, the experimental results are values measured from the processing of a component and include values including one or more of aspect ratio, isotropy ratio, deposition depth, global adhesion coefficient, surface-dependent adhesion coefficient, retardation thickness, neutral particle-ion ratio, and ion angular distribution function.

[0107] In one example, each experiment is performed in a semiconductor manufacturing apparatus based on an experiment recipe, and one experiment is performed to measure the effect of changing one parameter value from a previous recipe used in a previous experiment.

[0108] In one example, the processing of the component is for a deposition process using a suppression profile.

[0109] In one example, the processing of the component is for deposition in 3D NAND word line (WL) filling.

[0110] Another general aspect relates to a system comprising a memory containing instructions and one or more computer processes. When the instructions are executed by one or more computer processors, the instructions cause the one or more computer processors to perform a process of conducting a plurality of experiments for processing a component, where each experiment is controlled by one of a plurality of recipes that specify parameters of a manufacturing apparatus, a process of conducting a plurality of virtual simulations for processing the component, where each simulation is controlled by one of a plurality of recipes, a process of obtaining an ML model by training a machine learning (ML) algorithm using experimental results and virtual results of the virtual simulations, a process of receiving a specification for a desired processing of the component, and a process of creating a new recipe for processing the component by the ML model based on the specification.

[0111] In yet another general aspect, a machine-readable storage medium (e.g., a non-transitory storage medium) contains instructions that, when executed by a machine, cause the machine to perform a process of conducting a plurality of experiments for processing a component, where each experiment is controlled by one of a plurality of recipes that specify parameters of a manufacturing apparatus, a process of conducting a plurality of virtual simulations for processing the component, where each simulation is controlled by one of a plurality of recipes, a process of obtaining an ML model by training a machine learning (ML) algorithm using experimental results and virtual results of the virtual simulations, a process of receiving a specification for a desired processing of the component, and a process of creating a new recipe for processing the component by the ML model based on the specification.

[0112] FIG. 9 is a block diagram representing an example of a machine 900 in which an embodiment of one or more of the exemplary processes described herein is performed or controlled. In other embodiments, machine 900 may operate as a stand-alone device or may be connected (e.g., network-connected) to other machines. In a network-connected arrangement, machine 900 may operate as a server machine, a client machine, or both in a server-client network environment. In an example, machine 900 may function as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Further, although only a single machine 900 is shown, the term “machine” may also be construed to include a group of machines that individually or jointly execute a set (or multiple sets) of instructions for implementing one or more of the methods described herein (e.g., by cloud computing, software as a service (SaaS), or other computer cluster architectures).

[0113] The examples described in this specification may include, or be operated by, logic, some components, or mechanisms. A circuit network is a group of circuits implemented in a physical object including hardware (e.g., a simplex circuit, a gate, logic). The components of a circuit network may be flexible with respect to time and with respect to changes in the underlying hardware. A circuit network includes elements that can perform specific operations during operation, either alone or in combination. In an example, the hardware of a circuit network may be fixedly designed (embedded in the hardware) to perform a specific operation. In an example, the hardware of a circuit network may include a physically changeable computer-readable medium (e.g., magnetically, electrically, by invariant dense particles) for encoding instructions for a specific operation, and variable-connected physical components (e.g., an execution device, a transistor, a simplex circuit). When connecting physical components, the basic electrical characteristics of the hardware elements are changed (e.g., from an insulator to a conductor, or vice versa). Instructions enable variable connections to form elements of a circuit network of embedded hardware (e.g., an execution device, or a loading mechanism) to perform part of a specific operation during operation. Accordingly, the computer-readable medium is communicatively connected to other components of the circuit network while the device is operating. In an example, any physical element may be used in one or more elements of one or more circuit networks. For example, an execution device may be used in a first circuit of a first circuit network at a certain point in time during operation, and may be used again in a second circuit of the first circuit network, or may be used again in a third circuit of a second circuit network at different times.

[0114] A machine (e.g., a computer system) 900 may include a hardware processor 902 (e.g., a central processing unit (CPU), a hardware processor core, or a combination thereof), a graphics processing unit (GPU) 903, a main memory 904, and a static memory 906, and some or all of these may communicate with each other through an interlink (e.g., a bus) 908. The machine 900 may further include a display device 910, an alphanumeric input device (e.g., a keyboard) 912, and a user interface (UI) navigation device (e.g., a mouse) 914. In an example, the display device 910, the alphanumeric input device 912, and the UI navigation device 914 may be a touch screen. The machine 900 may further include a mass storage device (e.g., a drive unit) 916, a signal generating device (e.g., a speaker) 918, a network interface device 920, and one or more sensors 921 (such as a global positioning system (GPS) sensor, an orientation magnet, an accelerometer, or another sensor). The machine 900 may include an output control device 928 (e.g., a serial wiring connection (such as a universal serial bus (USB)), a parallel wiring connection, or another wiring connection, or a wireless connection (such as infrared (IR), near field communication (NFC))) that communicates with or controls one or more peripheral devices (e.g., a printer, a card reader).

[0115] The mass storage device 916 may include a machine-readable medium 922 storing one or more sets of data structures or instructions 924 (e.g., software) that implement or are used by one or more of the techniques or functions described herein. The instructions 924 may be present, in whole or at least in part, within the main memory 904, within the static memory 906, within the hardware processor 902, or within the GPU 903 during execution thereof by the machine 900. In an example, one or any combination of the hardware processor 902, the GPU 903, the main memory 904, the static memory 906, and the mass storage device 916 may constitute a machine-readable medium.

[0116] The machine-readable medium 922 is shown as a single medium, but the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized management database or a distributed database, and / or associated caches and servers) configured to store one or more instructions 924.

[0117] The term "machine-readable medium" may include any medium that can store, encode, or execute instructions 924 for execution by the machine 900, and any medium that can cause one or more techniques of the present disclosure to be implemented by the machine 900, or any medium that can store, encode, or execute a data structure used by such instructions 924 or a data structure associated with such instructions 924. Non-limiting examples of machine-readable media may include solid-state memory as well as optical and magnetic media. In an example, the group of machine-readable media includes the machine-readable medium 922 comprising a plurality of particles having an invariant (e.g., stationary) mass. Thus, the group of machine-readable media is not a transient propagation signal. Specific examples of the group of machine-readable media may include non-volatile memory (such as semiconductor memory devices (e.g., erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), and flash memory devices), magnetic disks (such as internal hard disks and removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0118] The instructions 924 may further be transmitted or received over a communication network 926 using a transmission medium via the network interface device 920.

[0119] Throughout this specification, multiple examples may include components, processes, or structures described as one example. Although the individual steps of one or more methods are described as separate steps, one or more of the individual steps may be performed simultaneously, and the steps need not be performed in the order of illustration. Structures and functions presented as separate components in an illustrative configuration may be implemented as an integrated structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other changes, modifications, additions, and improvements fall within the scope of the subject matter of this specification.

[0120] The embodiments described in this specification are described in sufficient detail for those skilled in the art to carry out the teachings of the disclosure. Other embodiments may be used or derived therefrom without departing from the scope of the disclosure, such that structural and logical replacements and changes may be made. Therefore, the forms for carrying out the invention should not be construed in a limiting sense, and the scope of various embodiments is defined only by the scope of the appended claims and the full scope equivalent to the scope of such claims.

[0121] As used herein, the term "or" may be interpreted in either an inclusive or exclusive sense. Also, multiple examples may be provided for resources, processes, or structures described as one example in this specification. Additionally, the boundaries between various resources, processes, modules, engines, and data stores are somewhat arbitrary, and a particular process is described in terms of a particular illustrative configuration. Other functional distributions are conceivable and may fall within the scope of various embodiments of the present disclosure. Generally, structures and functions presented as separate resources in an illustrative configuration may be implemented as an integrated structure or resource. Similarly, structures and functions presented as a single resource may be implemented as separate resources. These and other changes, modifications, additions, and improvements fall within the scope of the embodiments of the present disclosure as set forth in the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method comprising: Performing a plurality of experiments for processing components, each experiment being controlled by one recipe from a plurality of recipes that specify parameters for a manufacturing apparatus; Performing a plurality of virtual simulations for processing the components, each simulation being controlled by one recipe from the plurality of recipes; Obtaining an ML model by training a machine learning (ML) algorithm using experimental results and virtual results of the virtual simulations; Receiving a specification for a desired processing of the components; Creating a new recipe for processing the components by the ML model based on the specification; A method.

2. The method according to claim 1, wherein The ML model is based on a plurality of features including recipe features, experimental result features, virtual result features, and measurement features.

3. The method according to claim 2, wherein The measurement features include one or more of imaging methods, transmission electron microscopes, thickness measurements, sheet resistance, surface resistivity, stress measurements, and analytical methods used to determine layer thickness, composition, particles, or orientation.

4. The method according to claim 2, wherein The recipe features include workflow, gas flow, chamber temperature, chamber pressure, process duration, and radio frequency (RF) values.

5. The method according to claim 2, wherein The ML model includes active process control for determining process parameters to meet a control target, and the input to the ML model includes the recipe and the control target for the desired active process control.

6. The method according to claim 1, wherein The virtual simulation is performed by a simulation tool based on behavior modeling.

7. The method according to claim 1, wherein The experimental results include values measured from the processing of the components, the values including one or more of lateral ratio, isotropy ratio, deposition depth, global adhesion coefficient, surface-dependent adhesion coefficient, retardation thickness, neutral particle-ion ratio, and ion angle distribution function.

8. The method according to claim 1, wherein Each experiment is carried out in a semiconductor manufacturing apparatus based on the recipe for the experiment, and one experiment is carried out to measure the effect of changing the value of one parameter from a previous recipe used in a previous experiment. **Claim 9** The method according to claim 1, wherein processing the component is for a deposition process using a suppression profile. **Claim 10** The method according to claim 1, wherein processing the component is for deposition in 3D NAND word line (WL) filling. **Claim 11** A system comprising: a memory storing instructions; one or more computer processors, the instructions, when executed by the one or more computer processors, cause the system to: perform a plurality of experiments for processing a component, each experiment being controlled by one recipe from a plurality of recipes specifying parameters for a manufacturing apparatus; perform a plurality of virtual simulations for processing the component, each simulation being controlled by one recipe from the plurality of recipes; obtain an ML model by training a machine learning (ML) algorithm using experimental results and virtual results of the virtual simulations; receive a specification for a desired processing of the component; create a new recipe for processing the component by the ML model based on the specification. **Claim 12** The system according to claim 11, wherein the ML model is based on a plurality of features including recipe features, experimental result features, virtual result features, and measurement features. **Claim 13** The system according to claim 12, wherein the measurement features include one or more of an imaging method, a transmission electron microscope, thickness measurement, sheet resistance, surface resistivity, stress measurement, and an analysis method used to determine layer thickness, composition, particles, or orientation. **Claim 14** The system according to claim 12, wherein the recipe features include a workflow, a gas flow, a chamber temperature, a chamber pressure, a process duration, and a radio frequency (RF) value. **Claim 15** The system according to claim 11, The system, wherein the experimental results include values measured from the processing of the component, and the values include one or more of a lateral ratio, an isotropy ratio, a deposition depth, a global adhesion coefficient, a surface-dependent adhesion coefficient, a delay thickness, a neutral particle-ion ratio, and an ion angular distribution function.

16. The system according to claim 11, wherein each experiment is performed in a semiconductor manufacturing apparatus based on the recipe for the experiment, and one experiment is performed to measure the effect of changing the value of one parameter from a previous recipe used in a previous experiment.

17. A tangible machine-readable storage medium, when executed by a machine, performing a plurality of experiments for processing a component, each experiment being controlled by one recipe from a plurality of recipes specifying parameters for a manufacturing apparatus, performing a plurality of virtual simulations for processing the component, each simulation being controlled by one recipe from the plurality of recipes, obtaining an ML model by training a machine learning (ML) algorithm using experimental results and virtual results of the virtual simulations, receiving a specification for a desired processing of the component, creating a new recipe for processing the component by the ML model based on the specification, and including instructions for causing the machine to perform the above.

18. The tangible machine-readable storage medium according to claim 17, wherein the ML model is based on a plurality of features including recipe features, experimental result features, virtual result features, and measurement features.

19. The tangible machine-readable storage medium according to claim 18, wherein the measurement features include one or more of an imaging method, a transmission electron microscope, a thickness measurement, a sheet resistance, a surface resistivity, a stress measurement, and an analysis method used to determine the thickness, composition, particles, or orientation of a layer.

20. The tangible machine-readable storage medium according to claim 18, wherein the recipe features include a workflow, a gas flow, a chamber temperature, a chamber pressure, a process duration, and a radio frequency (RF) value.

21. The tangible machine-readable storage medium according to claim 17, The tangible machine-readable storage medium, the experimental results of which include values measured from the processing of the components, the values including one or more of a lateral ratio, an isotropy ratio, a deposition depth, a global adhesion coefficient, a surface-dependent adhesion coefficient, a retardation thickness, a neutral particle-ion ratio, and an ion angular distribution function.