Semiconductor manufacturing equipment and performance prediction of related processes

A hybrid model combining physics-based and data-driven approaches addresses the inefficiencies of current semiconductor manufacturing process prediction methods, offering faster, more accurate, and cost-effective performance estimation.

JP2025084789APending Publication Date: 2025-06-03LAM RES CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025021203
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2025-02-13
Publication Date
2025-06-03

Smart Images

  • Figure 2025084789000001_ABST
    Figure 2025084789000001_ABST
Patent Text Reader

Abstract

To provide methods, systems, and machine-readable storage media for predicting the performance of deposition, etch, and clean processes in a semiconductor manufacturing tool.SOLUTION: A method includes obtaining a plurality of machine-learning (ML) models related to predicting a performance metric for operation of a semiconductor manufacturing tool. Each ML model utilizes features defining inputs for the ML model. The method further includes: receiving a process definition for manufacturing a product with the semiconductor manufacturing tool; utilizing one or more ML models to estimate a performance of the process definition used in the semiconductor manufacturing tool; and presenting, on a display, results showing the estimate of the performance of the manufacturing of the product. Use of hybrid models of physics-based models and data-driven models improves predictive accuracy of a system by augmenting capabilities of the data-driven models with reinforcement provided by the physics-based models.SELECTED DRAWING: Figure 12
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] <Claims of Priority> This application claims the benefit of priority of U.S. Patent Application No. 62 / 966,378, filed on January 27, 2020, the entire disclosure of which is incorporated herein by reference.

[0002] The subject matter disclosed herein generally relates to methods, systems, and machine-readable storage media for predicting the performance of deposition, etching, and cleaning processes in semiconductor manufacturing tools.

Background Art

[0003] The background description provided here is for the purpose of generally presenting the content of the present disclosure. Within the scope described in this background art section, research by the inventors named at the present time, as well as aspects of the description that cannot be separately regarded as prior art at the time of filing, are not recognized as prior art against the present disclosure, whether explicitly or implicitly.

[0004] Typically, a great deal of effort is spent on high-fidelity modeling and simulation of semiconductor process reactors and device features in order to better understand the physical and chemical mechanisms between the surface kinetics associated with a seed phase (e.g., in the gas phase, aqueous phase, or organosolvated phase, in the solid state) on a substrate. A sufficient understanding and predictability of such systems is very important for improving product design and optimizing process conditions, e.g., adjusting the parameters of semiconductor manufacturing tools.

Summary of the Invention

[0005] Exemplary methods, systems, and computer programs are directed to predicting the performance of semiconductor manufacturing equipment operations. The examples are merely representative of possible variations. Unless otherwise explicitly stated, the components and functions are optional and may be combined or subdivided, the operations may occur in a different order, or may be combined or subdivided. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the exemplary embodiments. However, it will be apparent to one of ordinary skill in the art that the subject matter may be practiced without these specific details.

[0006] A hybrid of a physics-based model and a data-driven model, also known as a hybrid model, provides an opportunity to construct relationships between various factors (generated from experimental or simulation datasets) and a response variable that are not directly coupled or may have non-linear dependencies. Different models (high-fidelity simulations, data-driven models, reduced-dimensional models, etc.) are used to construct these relationships, so effective hybrid models can identify relationships not previously known. These hybrid models are computationally less expensive than high-fidelity physics-based simulations and can span a wide range of spatial and temporal scales. When hybrid models are enhanced with experimental data, the predicted values of the results are improved as the uncertainties and assumptions that cause simulations to vary from experimental data are eliminated or their effects are reduced. This approach is aimed at shortening the time spent during the design, fabrication, and test phases of a project, resulting in an acceleration of the solution and a reduction in product cost.

[0007] The disclosed embodiments utilize various truncated datasets generated from physical-based simulations (e.g., high-fidelity CFD simulations, atomistic simulations), approximation methods (e.g., reduced-dimensional models, surrogate models), experimental data, and other data sources (either as individual models or stack models) to determine the collective impact on semiconductor feature processing. The data obtained from these sources is analyzed using machine learning (ML) techniques to better predict process behavior and expected outcomes. The ML techniques include any combination of statistical modeling, deep neural networks, recurrent neural networks, convolutional neural networks, kriging, dynamic mode decomposition, proper orthogonal decomposition, etc., to generate physics-constrained data-driven models.

[0008] One method includes operations for obtaining machine learning (ML) models, each model being related to predicting performance metrics regarding the operation of semiconductor manufacturing tools. Further, each ML model utilizes features that define the inputs for the ML model. The method further includes operations for receiving a process definition for manufacturing a product on a semiconductor manufacturing tool. One or more ML models are utilized to estimate the performance of the process definition used on the semiconductor manufacturing tool. Additionally, the method includes presenting, on a display, results indicating the estimated performance of manufacturing the product.

[0009] Another general aspect relates to a system including a memory containing instructions and one or more computer processors. When executed by the one or more computer processors, the instructions cause the one or more computer processors to obtain machine learning (ML) models, each model performing operations including those related to predicting performance measurement criteria for the operation of semiconductor manufacturing tools. Further, each ML model utilizes features that define inputs for the ML model. A process definition for manufacturing a product with a semiconductor manufacturing tool is received, and one or more ML models are utilized to estimate the performance of the process definition used in the semiconductor manufacturing tool. Additionally, results indicating the estimated performance of the process are presented on a display.

[0010] In yet another general aspect, a machine-readable storage medium (e.g., a non-transitory storage medium), when executed by a machine, causes the machine to obtain machine learning (ML) models, each model performing operations including those related to predicting performance measurement criteria for the operation of semiconductor manufacturing tools. Further, each ML model utilizes features that define inputs for the ML model. A process definition for manufacturing a product with a semiconductor manufacturing tool is received, and one or more ML models are utilized to estimate the performance of the process definition used in the semiconductor manufacturing tool. Additionally, results indicating the estimated performance of the process are presented on a display.

Brief Description of the Drawings

[0011] The various attached drawings merely illustrate exemplary embodiments of the present disclosure and should not be regarded as limiting its scope.

[0012]

Figure 1

[0013]

Figure 2

[0014]

Figure 3

[0015]

Figure 4A

[0016]

Figure 4B

[0017]

Figure 5

[0018]

Figure 6

[0019]

Figure 7

[0020]

Figure 8

[0021]

Figure 9

Figure 10

[0022]

Figure 11

[0023]

Figure 12

[0024]

Figure 13

DETAILED DESCRIPTION OF THE INVENTION

[0025] Depending on the application, there are several types of modeling techniques such as stress modeling, thermal modeling, computational fluid dynamics (CFD), plasma modeling, Monte Carlo simulation, etc. These models predict species concentration, temperature profile, plasma density and related distributions, flow pressure and velocity fields, and perform calculations of related dimensionless numbers, etc.

[0026] Depending on the domain of interest, the spatial scale can range from nanometers to meters, and the time interval can range from picoseconds to minutes, which means a very wide range of spatial and time scales. Since it is necessary to test multiple design conditions and map the operating range, it is very costly and may be difficult to handle in an industrial environment to model these systems using physical-based, chemical-based, or quantum-based methods. For example, the execution of an individual process may take more than a week, and in the iterative stages of design, design engineers may be limited to using only two or three such modeling operations before completing the design.

[0027] Furthermore, the fidelity of physics-based models generally depends on the ability to accurately predict or characterize input parameters or boundary conditions for the model. In many cases, such parameters cannot be measured directly and must be inferred from experimental results. Such challenges are particularly acute when attempting to model, for example, the plasma environment common in semiconductor wafer processing equipment. Such plasmas vary with time both by design and due to the accumulation or removal of products or by-products within the process equipment chamber, and including measurement probes in the plasma environment can have an undesirable effect on the characteristics of the plasma being examined. Also, as an example, the measurement of surface or interface properties such as thermal contact resistance or emissivity may be impossible or unrealistic at a resolution sufficient to fully inform a high-fidelity physics-based model.

[0028] To predict the behavior of complex systems, pure data-driven models (e.g., machine learning models) may be utilized. Data-driven models are inherently statistical and do not explicitly require knowledge of the physical system being predicted. However, such models generally exhibit improved prediction values as the amount of training data increases and require the generation of large-scale experimental or simulated datasets to achieve a useful model. In particular, when the goal of the model is to predict the behavior of a new process, the creation of such large-scale datasets may be unrealistic or impossible.

[0029] To design better and faster semiconductor devices and semiconductor manufacturing processes, there is a need for reliable methods and systems that can quickly predict the behavior of various elements in a semiconductor manufacturing process.

[0030] FIG. 1 illustrates the complexity of designing equipment and processes for semiconductor manufacturing. To manufacture a semiconductor, both the equipment and the process for making the semiconductor must be designed. Based on product specification 102, the equipment and process are designed before the product is brought into production 108.

[0031] Machine design 104 typically involves designing many of the components of semiconductor manufacturing equipment, such as the geometry of chambers, showerheads, pedestals, power supplies, etc. 110. After the components are designed, the system is tested 112 in operation and the performance is measured. Typically, there are multiple cycles to complete the component design 110 until the test results are satisfactory for the purpose of manufacturing the product.

[0032] One of the challenges for the design team is that they have to design the equipment (e.g., chamber) before the process is determined. Therefore, the equipment is designed to meet general needs rather than specific requirements for a given product. In some cases, the design process can generate feedback to the machine design 104 and the equipment is redesigned based on process tests. However, redesigning the equipment according to a new process design is costly and time-consuming.

[0033] On the process design 106 side, a recipe 118 is designed 114 that includes the definition of the process, such as workflow, fluid flow, temperature settings for the chamber, pressure inside the chamber, duration of different steps, applied radio frequency (RF) power, concentration of chemicals, electrical bias, etc. The design of the recipe is a technique where design experts define the recipe based on past experience.

[0034] After the recipe 118 is designed, the recipe 118 is tested 116 on semiconductor manufacturing equipment. However, typically multiple cycles of recipe design 114 and testing 116 are required to find a satisfactory recipe 118 for generating the product during production 108.

[0035] Both machine design 104 and process design 106 are costly and tend to take a long time, such as weeks, months, or years for each iteration. To accelerate the design, simulations may be used to perform tests without actually conducting tests on semiconductor manufacturing equipment.

[0036] For example, to predict the uniformity on a substrate, the pedestal and showerhead are modeled to predict their interaction using a physics-based model that makes predictions based on the physical aspects of the chamber elements. However, physical models are time-consuming to execute and can take hours to days to test conditions using simulations.

[0037] Physics-based simulations can involve using physical laws to predict the behavior of elements and solving conservation equations, boundaries, or mesh nodes within the system. Physics-based simulations are typically used when it is not possible to observe the evolution and results of a process, such as observing what is happening inside a chamber.

[0038] There are also behavioral models that describe the output of a manufacturing process based on analytical formulations. One exemplary simulation tool is Lam Research's SEMulator3D® which provides a voxel model of a semiconductor process based on the predicted behavior simulated by a behavioral model. The behavioral model measures the results of operation (e.g., deposition thickness on a substrate) rather than directly simulating the behavior of particles within the chamber.

[0039] Current approaches to solving issues related to uniform deposition (verified on blanket wafers) and conformal deposition (verified on patterned wafers) are highly experimental. Modeling and simulation can be useful in the current workflow for component design and prediction of filling performance, but these are often time-consuming (e.g., first-principles / physically-based computational methods are dependent on multi-physics problems and mesh size and often take days or weeks per simulation), do not always address all relevant physics (e.g., the complexity of the system and state-of-the-art technologies do not always have clearly defined physics), can be based on non-physical methods such as behavior-based simulations (e.g., not all behaviors are captured by the modeling software), and generally require experimental data for calibration.

[0040] Behavioral model simulations for feature performance prediction are employed in the process. However, these models are limited by the fact that the models are calibrated based on experimental data and adopt non-physical parameters to optimize the geometry of the features. The behavioral models cannot be used for a complete system, and the behavior can vary from structure to structure, requiring various calibrations that imply additional tests to capture data for the models.

[0041] There are also limitations in the available experimental data. Data from experiments is the ultimate ground truth of what is seen on the wafer and within the features, but these models based on the data do not enable the transfer of the learned knowledge to other systems. For example, macroscopic parameters (e.g., atmospheric pressure, wafer temperature, flow rate) do not account for the capture of fundamental flow fields, charge density, flux transport, and other microscopic parameters that define the basic learning of the system. In some cases, there is no directly observable correlation between experimental data and simulation data, and such learning cannot be extended to other designs.

[0042] In summary, current methods employ separate and different approaches (modeling, behavior-based simulation, experimentation) and do not have a way to "learn" from these separate data sources and extend "insights" to other chamber designs. Experimental data are often the only reliable results that require several costly design iterations, resulting in a long product release cycle.

[0043] FIG. 2 is an etching chamber 200 according to one embodiment. Exciting an electric field between two electrodes is one way to obtain a radio frequency (RF) gas discharge within the etching chamber. When an oscillating voltage is applied between the electrodes, the resulting discharge is called a capacitively coupled plasma (CCP) discharge.

[0044] Plasma 202 can be generated using a stable feedstock gas to obtain a wide variety of chemically reactive by-products generated by the dissociation of various molecules caused by electron-neutral collisions. The chemical aspect of etching involves the reaction of neutral gas molecules and their dissociation by-products with the molecules of the surface to be etched, as well as the generation of volatile molecules that can be evacuated. When the plasma is generated, positive ions are accelerated from the plasma across a space charge sheath that separates the plasma from the chamber walls and impinge on the wafer surface with sufficient energy to remove material from the wafer surface. This is known as ion bombardment or ion sputtering. However, some industrial plasmas do not generate ions with sufficient energy to efficiently etch the surface purely by physical means.

[0045] Controller 216 manages the operation of chamber 200 by controlling different elements within chamber 200, such as RF generator 218, gas source 222, and gas pump 220. In one embodiment, CF 4 and C-C 4 F 8Fluorocarbon gases such as are used in dielectric etching processes because of their anisotropy and selective etching capabilities, but the principles described herein can be applied to other plasma - generating gases. Fluorocarbon gases readily dissociate into smaller molecular and atomic radicals, which are chemically reactive by - products. These chemically reactive by - products etch away the dielectric material, which, in one embodiment, is SiO 2 or SiOCH.

[0046] Chamber 200 shows a processing chamber having an upper electrode 204 and a lower electrode 208. The upper electrode 204 may be grounded or coupled to an RF generator (not shown), and the lower electrode 208 is coupled to an RF generator 218 via a matching network 214. The RF generator 218 provides RF power at one, two, or three different RF frequencies. At least one of the three RF frequencies can be turned on or off according to the desired configuration of chamber 200 for a particular operation. In the embodiment shown in FIG. 2, the RF generator 218 provides frequencies of 2 MHz, 27 MHz, and 60 MHz, although other frequencies are possible.

[0047] Chamber 200 includes a gas showerhead that is on top of the upper electrode 204 and inputs gas provided by a gas source 222 into chamber 200, and a perforated confinement ring 212 that allows gas to be exhausted from chamber 200 by a gas pump 220. In some exemplary embodiments, the gas pump 220 is a turbomolecular pump, although other types of gas pumps may be utilized.

[0048] When the substrate 206 is present within chamber 200, the silicon focus ring 210 is positioned adjacent to the substrate 206 so that a uniform RF field exists at the bottom surface of the plasma 202 for uniform etching on the surface of the substrate 206. The embodiment of FIG. 2 shows a triode reactor configuration where the upper electrode 204 is surrounded by a symmetric RF ground electrode 224. The insulator 226 is a dielectric that insulates the upper electrode 204 from the ground electrode 224.

[0049] Each frequency can be selected for a specific purpose in the wafer manufacturing process. In the example of FIG. 2, RF power is provided at 2 MHz, 27 MHz, and 60 MHz, where the 2 MHz RF power provides ion energy control and the 27 MHz and 60 MHz powers provide control of the plasma density and dissociation pattern of the chemical species. This configuration, where each RF power can be turned on or off, enables specific processes that use ultra-low ion energy on the substrate or wafer, and specific processes where the ion energy must be low (less than 200 or 200 eV), such as the soft etching of low-k materials.

[0050] In another embodiment, 60 MHz RF power is used on the top electrode 204 to obtain ultra-low energy and very high density. This configuration allows for chamber cleaning with high-density plasma when the substrate 206 is not in the chamber 200, while minimizing sputtering on the electrostatic chuck (ESC) surface. When the substrate 206 is absent, the ESC surface is exposed, and thus any ion energy on the surface should be avoided, which is why the bottom 2 MHz and 27 MHz power supplies can be turned off during cleaning.

[0051] FIG. 3 shows feature modeling at multiple levels in semiconductor manufacturing according to some exemplary embodiments. The embodiments presented herein describe a method for constructing a tool to predict the behavior and performance of semiconductor manufacturing equipment.

[0052] This type of tool for predicting the behavior of physical entities is called a digital twin and is also referred to herein as a system model. A digital twin is a digital replica of a physical entity, whether living or non-living, and refers to the digital replicas of potential and actual physical assets (physical twins), processes, people, places, systems, and devices that can be used for various purposes. A digital twin emphasizes the connection between a physical model and its corresponding virtual model or virtual counterpart, and this connection can be reinforced by using sensors to generate real-time data.

[0053] A digital twin learns and updates itself from multiple sources and represents its nearly real-time status, operating conditions, or location. This learning system learns from itself (e.g., using sensor data that conveys various aspects of its operating conditions), from human experts (e.g., design engineers), from other machines, and from the environment in which it may be a part. A digital twin also integrates historical data from past machine usage and weaves it into its digital model.

[0054] In some exemplary embodiments, a digital twin for semiconductor manufacturing equipment utilizes various fragmented datasets generated from first-principles physics-based simulations, approximate empirical methods (such as reduced-dimensional models, approximate models, etc.), experimental data, and other data sources (e.g., chamber sensors). Further, the digital twin adopts ML techniques such as statistical modeling, deep neural networks, recurrent neural networks, convolutional neural networks, etc. to generate a predictive, high-fidelity, and accurate physics-constrained data-driven surrogate hybrid model.

[0055] Digital twins can be applied, for example, to analyze the impact of etching chamber design on wafer etching uniformity. For example, chamber flow simulations provide information regarding fluid flow fields, pressure fields, concentration gradients, etc. Plasma simulations provide the density and flux of charged species to the wafer, as well as plasma parameters within the process window of interest.

[0056] Experimental data provides upstream simulation data and partial sensor information that can be correlated to inputs, and includes output variables, which are on-wafer uniformity and / or overall wafer feature profile characteristics (e.g., depth, critical dimension (CD), tilt, selectivity, mask loss) on a blanket wafer. Individually, high-fidelity simulations are not expected to exactly match on-wafer experimental performance metrics, and experimental data alone does not provide chamber physical characteristics such as flow fields, plasma properties, etc. Hybrid model solutions help provide insights from these separate but related datasets and can assist in product design, engineering, and development.

[0057] The system hybrid model can include separate submodels that cover different aspects of the equipment, and these submodels can interact with each other as described below with reference to FIG. 4A.

[0058] Referring to FIG. 3, in some exemplary embodiments, the submodels include a chamber model 302, a plasma model 304, a sheath model 306, a wafer feature model 308, an atomistic model 310, and an electronic state model 312. Each model predicts its respective behavior using its respective features and training data. Further details regarding model construction are provided below with reference to FIG. 5, and further details regarding each feature are provided below with reference to FIGS. 9-10. Additionally, each of the submodels can also be divided into separate models.

[0059] The chamber model 302 is for predicting chamber-related phenomena such as chamber geometry, flow rate, thermal information, structural information, and electromagnetics. The plasma model 304 is for predicting plasma performance such as electromagnetic fields, plasma chemistry, and reaction data.

[0060] The sheath model 306 is for predicting sheath data such as RF voltage, electromagnetic fields, reaction cross sections, and reaction pathways. The wafer feature model 308 is for predicting wafer features including layout (e.g., design layout, mask layer) and chemistry (e.g., on-wafer flux, material properties).

[0061] The atomistic model 310 predicts phenomena at the atomic level such as atomic dynamics including lattice structure, species distribution, and diffusion coefficients. The electronic state model 312 is for predicting states for energy calculations such as energy states and cluster configurations.

[0062] The presented embodiments provide the advantage of determining the compositional relationships between different models generated separately for specific problems on a given system or process. These separate datasets can be combined to construct a system model that can be used for design and process optimization in semiconductor manufacturing.

[0063] These models encompass different spatial scales (e.g., on the order of meters to nanometers) and temporal scales (e.g., from transient processes such as pulse systems to steady-state conditions under equilibrium). The system model results in cost savings (e.g., use of chemicals, reduction of code iterations) as well as savings in development time (e.g., shortening the design time for new reactors and processes from 18 - 24 months to less than one year) since the design and testing of the system can be carried out much faster than previous methodologies.

[0064] Figure 4A shows the interaction between multiple machine learning models for predicting process behavior according to some exemplary embodiments. System model 402 includes multiple ML models ML1 to ML13 at different levels of prediction. Each level includes one or more ML models that can interact with each other and with ML models in other layers. In some exemplary embodiments, the models include behavior models and physics-based models.

[0065] The models include inputs 404 (e.g., pressure, temperature, flow) and generate outputs related to behavior. The output of system model 402 is system behavior 408. Inputs 404 can also include data from measurements 406 obtained from sensors. Additionally, inputs 404 can also include data from measurements 406 from the incoming substrate before application of the process being modeled. Measurements 406 provide measured values of experiments and include items such as layer thickness, resistivity, film property data. Image analysis can be used to examine the experimental results, but other types of measurements 406 can also be used.

[0066] Measurements 406 include one or more of imaging methods (e.g., scanning electron microscopy (SEM), transmission electron microscopy (TEM)), typical thickness measurements (e.g., X-ray fluorescence (XRF), ellipsometry), sheet resistance, surface resistivity, stress measurements, and other analytical methods used to determine layer thickness, composition, particle orientation, etc. These other analytical methods include one or more of X-ray diffraction (XRD), X-ray reflectivity (XRR), X-ray photoelectron spectroscopy (XPS), precession electron diffraction (PED), electron energy loss spectroscopy (EELS), energy-dispersive X-ray spectroscopy (EDS), secondary ion mass spectrometry (SIMS), etc.

[0067] In some embodiments, the raw measurement data 406 can be conditioned by a dedicated ML algorithm for each measurement source or, in some embodiments, an ML algorithm that obtains inputs from multiple measurement sources. Measurement sources in the semiconductor industry can generate an enormous amount of data, and it is often a difficult problem to distinguish useful data from individual sources. The difficulty grows exponentially when data from many sources need to be considered together. Tasking a dedicated ML algorithm to condition the incoming measurement data enables the construction of a more efficient system model 402. The output of each measurement conditioning algorithm is ROM, and each component ML in the system model can select only the output parameters relevant to its own operation.

[0068] In some exemplary embodiments, the measurements include time-series data, which includes sensor measurements obtained over time for a given parameter, such as how the pressure in a chamber changes over the course of a manufacturing process.

[0069] The inputs 404 can include controls such as pressure, flow rate, power, temperature, etc., and all of these inputs 404 affect the behavior of the chamber model 302. At the level of the chamber model 302, there are flow fields, ion densities, neutron densities, etc. Next, there is a level of the plasma model 304 that includes items such as electric fields, current densities, chemicals, etc. If there is no plasma, there are different species present in the chamber. Further, when power is supplied to the bias, there is a sheath when there is plasma deposition, and controls are used to modify the redistribution function.

[0070] Further, there is a level of the wafer feature model 308 where wafer features become observable. Typically, the models are provided for a single level, but some models may include multiple levels. Each model takes into account specific features to obtain its respective predictions. The output of a model can be used by other models operating at the same level or other levels.

[0071] During chamber design, various options can be explored (e.g., changing the hole distribution of the showerhead, the gap between the wafer and the showerhead), and the goal is to achieve specific measurement criteria regarding the wafer, such as having a uniform profile. The designer develops the chamber configuration (e.g., chamber knob, geometry) and the recipe, and the system model 402 makes predictions of system behavior 408 including measurement criteria regarding the performance of the wafer. The system model 402 can identify not only the performance regarding the wafer but also information about what is happening inside the chamber (e.g., plasma, too high or too low temperature, etc.).

[0072] If the performance is not satisfactory, the designer can change specific parameters, such as changing one or more variables (e.g., pressure, chemical substance, timing) at a time, and repeat the process to confirm the effect. However, using an ML model, this process becomes rapid (minutes or hours) without waiting weeks or months for the results. The system model 402 can evaluate the impact on performance when applying the changes to the hardware or the process recipe.

[0073] One advantage of combining models is to find the correlation between multiple models separately generated for specific problems on a given system / subsystem / process, and combine these separate datasets to construct a system model 402 that can be used for design and process optimization in the characteristics of the target.

[0074] Using such techniques, it is possible to build models from different spatial scales (from chamber to the order of m to features to the order of nm) and temporal scales (from transient processes such as pulse systems to steady-state conditions under equilibrium). This allows for a better understanding of the impact on the final desired state (e.g., process fill) based on the design conditions. Since the path from chamber-level simulations to surface dynamics on the wafer is not straightforward, this could not be done accurately in complex deposition chambers. By using the system model 402, wafers and chemicals are conserved from a process perspective, and costly design iterations are minimized.

[0075] FIG. 4B is a table 412 showing examples of modeling algorithms at different levels according to some exemplary embodiments. The different areas include etching / atomic layer etching, electroplating, plasma enhanced (PE) chemical vapor deposition (CVD) and atomic layer deposition (ALD), cleaning and stripping, physical vapor deposition, and chemical / mechanical polishing. Although some of the embodiments are described with respect to etching, the same principles may apply to other areas.

[0076] FIG. 5 shows the training and use of a machine learning program according to some exemplary embodiments. In some exemplary embodiments, a machine learning program (MLP), also referred to as a machine learning algorithm or tool, is utilized to perform operations related to searches such as job searches.

[0077] Machine learning (ML) is an application that provides a computer system with the ability to perform tasks without being explicitly programmed, by making inferences based on patterns found in the analysis of data. Machine learning explores the research and construction of algorithms, also referred to herein as tools, that can learn from existing data and make predictions about new data. Such machine learning algorithms operate by constructing an ML model 510 from exemplary training data 506 to make data-driven predictions or decisions, represented as an output or evaluation 514. Although several exemplary embodiments are presented with respect to machine learning tools, the principles presented herein may be applied to other machine learning tools.

[0078] Data representation refers to a way of organizing data for storage on a computer system, including the structure for identified features and their values. In ML, it is typical to represent data as vectors or matrices of two or more dimensions. When dealing with large amounts of data and many features, data representation is important so that training can identify correlations within the data.

[0079] There are two common modes of ML: supervised ML and unsupervised ML. Supervised ML uses prior knowledge (e.g., examples correlating inputs to outputs or results) to learn the relationship between inputs and outputs. The goal of supervised ML is to learn a function that best approximates the relationship between training inputs and outputs when given some training data, such that the ML model can implement the same relationship and generate corresponding outputs when an input is given. Unsupervised ML is the training of ML algorithms that use information that is neither classified nor labeled, enabling the algorithm to act on that information without guidance. Unsupervised ML is useful for exploratory analysis because it can automatically identify structures within the data.

[0080] Common tasks for supervised ML are classification problems and regression problems. Classification problems, also called categorization problems, aim to classify items into one of several categorical values (e.g., is this object an apple or an orange?). Regression algorithms aim to quantify several items (e.g., by providing a score as a function of several inputs). Some examples of commonly used supervised ML algorithms are logistic regression (LR), naive Bayes, random forest (RF), neural network (NN), deep neural network (DNN), matrix factorization, and support vector machine (SVM).

[0081] Some common tasks for unsupervised ML include clustering, representation learning, and density estimation. Some examples of commonly used unsupervised ML algorithms are K-means clustering, principal component analysis, and autoencoders.

[0082] Training data 506 includes examples of values for features 502. In some illustrative embodiments, the training data 506 includes labeled data having examples of values for features 502 and labels indicating results such as plasma properties, gas density, gas flow, etch rate, etch uniformity.

[0083] In some illustrative embodiments, the training data 506 is obtained by conducting experiments on semiconductor manufacturing equipment, and the data obtained from the experiments is used for ML training 508. Additionally, the training data 506 may also be obtained by conducting simulations (e.g., physics-based simulations, behavior-based simulations), and the results from the simulations (e.g., etch uniformity on a wafer, uniform deposition on a wafer) are used for ML training 508.

[0084] The machine learning algorithm uses the training data 506 to find the correlation between the identified features 502 that affect the result. The features 502 are the individual measurable properties of the observed phenomenon. The concept of features is related to the concept of explanatory variables used in statistical techniques such as linear regression. For the effective operation of ML in pattern recognition, classification, and regression, it is important to select features that are beneficial, discriminative, and independent. Features can be of different types, such as numerical features, character strings, and graphs. Further details about the features 502 used by the ML model 510 are provided below with reference to FIGS. 9 and 10.

[0085] During training 508, the ML algorithm analyzes the training data 506 based on the identified features 502 and configuration parameters 504 defined for the training 508. The result of the training 508 is an ML model 510 that can take an input to generate an evaluation.

[0086] In some embodiments, an exemplary ML model provides estimates at different levels of the semiconductor manufacturing process, as described above with reference to FIGS. 3 and 4.

[0087] Training the ML algorithm involves analyzing large amounts of data (e.g., from several gigabytes to over a terabyte) to find data correlations. The ML algorithm uses the training data 506 to find the correlation between the identified features 502 that affect the result or evaluation 514. In some exemplary embodiments, the training data 506 includes labeled data that is known data regarding one or more identified features 502 and one or more results.

[0088] The ML algorithm typically explores many possible functions and parameters before finding what it identifies as the best correlation in the data. Thus, training can require a large amount of computing resources and time.

[0089] Many ML algorithms include configuration parameters 504, and the more complex the ML algorithm, the more parameters are available to the user. Configuration parameters 504 define the variables of the ML algorithm when searching for the best ML model. Training parameters include model parameters and hyperparameters. Model parameters are learned from training data, while hyperparameters are not learned from training data but are instead provided to the ML algorithm.

[0090] Some examples of model parameters include the maximum model size, the maximum number of paths in the training data, the data shuffle type, regression coefficients, split positions of decision trees, etc. Hyperparameters can include the number of hidden layers within a neural network, the number of hidden nodes in each layer, the learning rate (and in some cases, various adaptation schemes for the learning rate), regularization parameters, the type of non-linear activation function, etc. Finding the correct (or best) set of hyperparameters can be a very time-consuming task that requires a large amount of computer resources.

[0091] When the ML model 510 is used to perform an evaluation, an input 512 is provided to the ML model 510, and the ML model 510 generates an evaluation 514 as output. For example, when analyzing the gas flow distribution for a showerhead, the output can show a time series of gas density across the entire surface of the wafer.

[0092] For example, an ML model 510 for detecting complete coverage when filling deep trenches on a wafer can provide an estimate indicating the level of coverage for the trenches, where 100% means covered and less than that means there are voids in the fill.

[0093] Feature extraction is the process of reducing the amount of resources required to describe a large dataset. When performing analysis of complex data, one of the main problems is due to the number of variables involved. Analysis using a large number of variables generally requires a large amount of memory and computing power, and there is a possibility that the classification algorithm may overfit to the training samples and generalize poorly to new samples. Feature extraction involves constructing a combination of variables to avoid these problems with large datasets while describing the data with sufficient accuracy for the desired purpose.

[0094] In some exemplary embodiments, feature extraction starts from an initial set of measurement data, constructs derived values (features) intended to be useful and non-redundant, and facilitates subsequent learning and generalization steps. Additionally, feature extraction is related to dimensionality reduction, such as reducing a large vector (possibly with very sparse data in some cases) to a smaller vector that captures the same or similar amount of information.

[0095] In some exemplary embodiments, feature extraction includes obtaining the values of features (as described in FIGS. 9 and 10) and, if necessary, converting those values into vectors or matrices. In some cases, the vectors or matrices are processed to reduce their dimensions in order to provide information to the ML model 510 in a more compact way. Additionally, data from multiple features can be combined, such as by concatenating vectors from multiple inputs into a single vector.

[0096] FIG. 6 shows a process for plasma reduced order model (ROM) simulation according to some exemplary embodiments. The exemplary embodiments are described with reference to plasma properties, but the same principles may apply to other areas such as thermal processes. In some exemplary embodiments, the goal is to use physics-based simulation, behavior-based simulation, wafer data log (WDL) (including time series information generated from sensors), and ML models to derive a reduced order model (ROM), also referred to herein as a digital twin, for a semiconductor manufacturing chamber.

[0097] A ROM is a simplification of a complex model that captures the behavior of a system, allowing engineers to quickly study the dominant effects of the system using minimal computational resources. ROMs enable engineers to achieve shorter design cycles and produce higher quality products.

[0098] Chamber knob 602 is a configurable parameter used as an input to simulation 604. Chamber knob 602 includes any of the configurable parameters for the manufacturing process, such as pressure, flow rate, transformer coupled plasma (TCP), bias, transformer coupled capacitive tuning (TCCT), chemical species information, etc.

[0099] For simulation 604, a design of experiments (DOE) matrix is created. DOE is a field of applied statistics that deals with the planning, execution, analysis, and interpretation of controlled tests to evaluate factors that control the values of parameters or groups of parameters. DOE is a powerful data collection and analysis tool that can be used in a variety of experimental situations.

[0100] In some exemplary embodiments, simulation 604 is a physics-based simulation, while in other embodiments, other types of simulations, such as behavior-based simulations, can be utilized. Typically, multiple simulations are performed, and the results are output 606, which includes physical quantity data as well as derived values obtained from simulation 604. For example, the simulation can generate predictions including flow fields, ion density, energy distribution, species distribution, and the like.

[0101] The results of simulation 604 are used as training data for ML training 508 to create ROM608. ROM608 receives the configuration of chamber knob 602 as input and generates predictions regarding the performance of a specific design and process.

[0102] For example, two designs for a showerhead are being considered. Each design can be defined, for example, by its respective computer-aided design (CAD) geometry (e.g., the number and distribution of holes in the showerhead) and gas flow. The two designs are input into ROM608, and results are obtained. Next, the performance parameters being considered for improving the showerhead are compared to determine which showerhead design is superior.

[0103] Comparing performance using ROM608 does not require actual experiments in the chamber, so the results can be obtained very quickly (e.g., within minutes or hours) without having to wait days or weeks to obtain results for each design. Since the testing of designs is very fast, design engineers can test multiple designs by fine-tuning one or more parameters at a time until a satisfactory design is found.

[0104] In addition, the ROM 608 can verify the results generated by the simulation 604 by comparing them with the predictions 610 generated by the ROM 608 when using the configuration of the same chamber knob 602 (e.g., the accuracy in formulating the predictions can be checked). In this way, the ROM 608 can be continuously enhanced until the prediction 610 substantially matches the results of the simulation 604. In general, it should be noted that the ROM 608 is much faster than the simulation 604 (e.g., a physics-based simulation) that needs to solve complex mathematical equations with many related variables. Therefore, having an accurate ROM 608 significantly accelerates the development process.

[0105] FIG. 7 shows a process for a plasma ROM using simulations and experiments according to some exemplary embodiments. The exemplary embodiments are described with reference to plasma properties, but the same principles may apply to other areas such as thermal processes. The method described above with reference to FIG. 6 can be enhanced by using actual experiments 702 conducted using semiconductor manufacturing tools. Of course, using actual experiments is more costly and requires more time for setup and execution than the simulation 604. Using experimental data improves the quality of the predicted results.

[0106] The experimental output 704 resulting from the experiment 702 may also be used as additional data for the ML training 508 to obtain the ROM 608. Additionally, the data from the sensors 706 obtained during the experiment 702 can be used for training and as an input to the ROM 608. The ROM 608 then makes a prediction 710 based on the chamber knob 602 used as the input. Therefore, the ROM 608 benefits from the training data 506 derived from the experiment 702 and the simulation 604. Further, the data from the experiment 702 is verified 807 with the output of the prediction 710.

[0107] Similar to the simulation 604, the results from the experiment can be used to verify the prediction 710 of the ROM 608. Based on the verification, the ROM 608 can be fine-tuned, such as by changing the hyperparameters for the configuration used during the ML training 508.

[0108] For example, the study of the design's impact on the non-uniformity of the film on the wafer includes simulation data with information regarding the fluid flow field, pressure field, concentration gradient, etc. The experimental data includes upstream simulation data and partial sensor information that can be correlated with the input, but mainly includes output variables, which are the non-uniformity of the film or the variation of the step coverage across the 300 mm wafer.

[0109] In some embodiments, a behavior model based on the understanding of the process can be used. Other data inputs, such as chamber sensor information and process recipe conditions, can be provided to further enhance the model. The generated model is part of an existing dataset calculated for different purposes, and the model is capable of online learning based on new experimental and simulation data. By itself, the high-fidelity simulation cannot match the experimental performance on the wafer, such as non-uniformity matching, and by itself, the experimental data cannot understand the physical characteristics of the system, such as the flow field and flux. The use of the ROM 608 aids in learning from these isolated datasets and helps design engineering for future systems.

[0110] Furthermore, for other applications, there are the design optimization of thermal-fluid systems with conjugate heat transfer as seen in high-temperature showerhead / pedestal applications, mass transfer and bias optimization in wet processing chambers, plasma-driven process systems, or other atomic layer deposition applications.

[0111] Figure 8 shows some of the inputs to an ML model according to some exemplary embodiments. Several inputs are used as data for algorithm 802 used in system modeling. The inputs include recipe information and experimental data 804 obtained from on-wafer results, experimental data 806 resulting from hardware tests, chamber sensor data 808 obtained during semiconductor manufacturing operations, chamber scale simulations 810, feature scale simulations 812, atomistic / quantum simulations 814, behavior-based models 816 at the feature scale, and the dimensionality reduction model 608.

[0112] The result is a physics-constrained ML model 818 that predicts the performance of the equipment and processes. With the physics-constrained ML model 818, designers can set the control for the chamber to obtain the desired results, which means improved predictability of wafer properties and reduced development time.

[0113] Figures 9 and 10 show tables containing some exemplary features used by the ML model. The first column indicates the scale (e.g., the level of the hierarchy described above with reference to FIG. 3), the second column indicates the observed phenomenon, the third column represents the input, and the fourth column represents the output. Each row represents a respective ML model.

[0114] The first model is for the chamber geometry with inputs including the chamber design dimensions (e.g., chamber diameter, wafer diameter). The output includes a specific chamber geometry design that can be in the form of several other geometry descriptions such as a CAD file or a mesh of the geometry shape.

[0115] The next model is for the flow within the chamber. The inputs include at least flow rate, chamber pressure, chemical substances, and flow transient phenomena. The output includes at least pressure and velocity fields, species concentration throughout the chamber, and diffusion fluxes.

[0116] The following model is for thermal research within the chamber, and the inputs include at least heat flux, heat transfer coefficient, thermal conductivity, heat capacity, and contact area. The outputs are temperature information for the entire chamber, heat flux (e.g., radiation, conduction, and convection), and heat loss.

[0117] The following model is for structural analysis of the chamber, and the inputs include at least structural load, reaction, tolerance, temperature, gradient, and pressure. The outputs include structural stress, strain, deflection, creep rate, fatigue, and elasticity / plasticity.

[0118] The final model for the chamber is for electromagnetic estimation within the chamber, and the inputs include at least voltage, current, frequency, inductance, capacitance, and impedance. The outputs include electric field, magnetic field B, induced current density, and power density.

[0119] The plasma model is for analyzing plasma behavior, and the inputs include at least electric field, magnetic field B, current density, chemical substances, reaction cross-section, reaction path, material properties (e.g., conductivity, permittivity, emission coefficient), RF frequency, RF voltage, and RF bias. The outputs include charged species density and flux, bipolar field, electron temperature, electron energy distribution function (EEDF), ion energy angular distribution (IEAD), on-wafer flux, charge density (surface and volume) sources, and species loss terms.

[0120] As described above, some outputs from one model can be used as inputs to other models. For example, the outputs from the chamber electromagnetic model can be used as inputs to the plasma behavior model (e.g., electric field, magnetic field B, current density).

[0121] The first model in Figure 10 is for analyzing the performance of the sheath, and the inputs include at least RF voltage, electric field, source term, reaction collision cross-section, and reaction path. The outputs include flux on the surface, ion energy and angular distribution, conduction and displacement current, ion transit time, and charge density.

[0122] The following model is for the layout of wafer features, and the input includes at least the design layout of the application, mask layers, and initial steps. The output includes a geometric description of the features such as a CAD file or a mesh configuration.

[0123] The following model is for the chemistry of the features, and the input includes at least on-wafer flux, material properties, reaction pathways, reaction rates, ion angle yields, etching thresholds, adhesion coefficients, and accommodation coefficients. The output includes geometric evolution and front tracking, the distribution of species within the feature, and the distribution of ion energy and angles within the feature.

[0124] The following model is for the dynamics at the atomic level, and the input includes lattice structure, material ID, and interatomic potential. The output includes lattice layout, species distribution, radial distribution function, diffusion coefficient, and reaction kinetics.

[0125] The last model in the table is for the electronic state and predicts statistics for energy calculations. The input includes at least material / species ID, cluster configuration, and electronic structure information. The output includes energy states, cluster configuration, reaction energetics, reaction pathways, bond angles, and bond lengths.

[0126] FIG. 11 shows the design of a showerhead 1102 using ML according to some exemplary embodiments. For example, the showerhead 1102 is for use in a low-fluorine tungsten reactor, but can also be used in other types of reactors.

[0127] In one case, the showerhead design did not meet the requirements regarding the fill gap, WF6 gas consumption, and Rs NU. Low gas consumption reduces the consumable cost per wafer (CoC). Furthermore, the gap from the pedestal to the showerhead could not be accurately set by the automatic gap system (AGS) at high temperatures such as 430°C. Additionally, the flatness from the pedestal to the showerhead could not be accurately set using the AGS wafer because the AGS wafer is not measured at the outer edge of the pedestal. This leads to tool-to-tool variations and impacts on the process. An improved design of the showerhead was needed.

[0128] The showerhead 1102 exhibited thermal instability. The design team tried to manually change the variables of the showerhead 1102, but the process was slow, cumbersome, and required a great deal of labor and long development time. DOE was used to learn from simulations that require long computation times. The design process was improved using simulations, but it was still time-consuming, and the results of actual experiments did not match the results from the simulations. This is why an ML model for the showerhead can speed up the design process. Also, since the ML model can be improved as additional data becomes available, the accuracy of the ML model is improved over time and can improve and accelerate the design of the showerhead 1102.

[0129] The features for the showerhead ML are defined to cover multiple facets of the showerhead 1102 (arrows in the figure showing some exemplary areas) and include one or more of the following:

[0130] - Contact conductance between plates;

[0131] - Contact conductance, gap conductance, and radiation between the manifold and the backplate;

[0132] - Contact conductance, gap conductance, and radiation between the edge ring 1110 and the manifold;

[0133] - Gap conductance and radiation between the pedestal ring and the showerhead edge;

[0134] - Contact conductance and gap conductance between the showerhead backplate and the manifold;

[0135] - Contact conductance, gap conductance, and radiation between the pedestal area and the pedestal;

[0136] - Radiation and gap conductance between the wafer and the showerhead and between the wafer and the pedestal;

[0137] - Gap conduction in the voids of the showerhead;

[0138] - Radiation to the surrounding area from the outer surface of the pedestal and the pedestal ring;

[0139] - Gap conduction and radiation between the surface of the showerhead and the edge ring;

[0140] - Radiation to the surrounding area from the outer surface of the edge ring; and

[0141] - Radiation to the surrounding area from the outer surface of the manifold.

[0142] Furthermore, zones such as the illustrated zones Z1 - Z4 are defined within the showerhead, and the parameters associated with each zone are used as characteristics such as temperature or conductivity on the zone.

[0143] Different parameters are varied as inputs and multiple experiments are conducted. The inputs can include the number of holes in the showerhead, the distribution of the holes, the different types of gases used, the gas pressure, the delivery cycle, etc. The results, along with the inputs, are used as training data for the showerhead model. The results can include the thickness of the wafer across its entire surface.

[0144] Next, the showerhead model is used to estimate each output regarding the performance of the showerhead based on various inputs. The outputs can include the uniformity measured along the wafer, as well as the measurement data obtained from sensors within the chamber.

[0145] In this way, the showerhead design team can quickly perform modeling of different showerhead configurations and determine which configuration functions best for a given product requirement.

[0146] The resulting showerhead 1102 designed with the aid of the showerhead model has a reduced internal cavity volume of the showerhead, a reduced volume between the pedestal and the showerhead, a smaller gap between the pedestal and the showerhead, and a different faceplate hole pattern. Thereby, the gas consumption is reduced while meeting the required performance parameters such as uniformity, temperature requirements, etc.

[0147] In some exemplary embodiments, the narrow measurements for the showerhead gap and the wafer are obtained using a laser attached to the showerhead facing the atmosphere side. The laser light passed through a vacuum-sealed sapphire window. Further, the laser light was reflected from the pedestal surface and its timing was used to measure the distance.

[0148] FIG. 12 is a flowchart of a method 1200 for predicting the performance of semiconductor manufacturing equipment operations according to some exemplary embodiments. Although the various operations of this flowchart are presented and described in order, those skilled in the art will understand that some or all of the operations may be performed in a different order, combined or omitted, or performed in parallel.

[0149] In operation 1202, a plurality of machine learning (ML) models are obtained, each model being related to predicting performance metrics for the operation of a semiconductor manufacturing tool, and each ML model utilizing a plurality of features that define the inputs for the ML model.

[0150] From operation 1202, method 1200 moves to operation 1204, in which one or more processors receive a process definition for manufacturing a product with a semiconductor manufacturing tool.

[0151] Furthermore, from operation 1204, method 1200 moves to operation 1206, in which one or more processors utilize one or more of the plurality of ML models from the plurality of ML models to estimate the performance of the process definition used in the semiconductor manufacturing tool.

[0152] From operation 1206, method 1200 moves to operation 1208, in which a result indicating the estimated performance of the process is presented on a display.

[0153] In one example, creating one ML model from a plurality of machine learning (ML) models involves obtaining training data for the ML model, where the training data includes providing values for the features of the ML model and training an ML algorithm to obtain the ML model.

[0154] In one example, obtaining training data for an ML model includes performing experiments with a semiconductor manufacturing tool, measuring values for the features for the experiments, and using the measured values for the training data.

[0155] In one example, obtaining training data for an ML model includes obtaining training data by performing a physics-based simulation of the operation of a semiconductor manufacturing tool.

[0156] In one example, a plurality of ML models include a chamber model, a plasma model, a sheath model, a wafer feature model, an atomistic model, and an electronic state model.

[0157] In one example, a chamber ML model is for estimating the geometry of a chamber within a semiconductor manufacturing tool using inputs including design dimensions and an output including a definition of the chamber geometry.

[0158] In one example, a plasma ML model is for analyzing plasma behavior using inputs including one or more of an electric field, magnetic field B, current density, chemical species, reaction cross-section, reaction pathway, material properties, RF frequency, RF voltage, and RF bias, and the output of the plasma ML model includes one or more of charged species densities and fluxes, bipolar fields, electron temperature, electron energy distribution function (EEDF), ion energy angle distribution (IEAD), on-wafer flux, charge density (surface and volume) sources, and species loss terms.

[0159] In one example, a sheath ML model is for analyzing sheath performance using inputs including one or more of a radio frequency (RF) voltage, electric field, source term, reaction collision cross-section, and reaction pathway, and the output of the sheath ML model includes one or more of the flux on the wafer surface, ion energy and angle distribution, conduction and displacement currents, ion transit time, and charge density.

[0160] In one example, a wafer feature ML model is for analyzing the layout of wafer features using inputs including one or more of a design layout, mask layer, and initial steps, and the output of the wafer feature ML model includes a geometric description of the wafer features.

[0161] In one example, the wafer chemical ML model is for analyzing the chemicals of wafer features using an input that includes one or more of on-wafer flux, material properties, reaction pathways, reaction rates, ion angle yields, etching thresholds, adhesion coefficients, and accommodation coefficients, and the output of the wafer chemical ML model includes one or more of geometry evolution and front tracking, the distribution of species within the wafer feature, and the distribution of ion energy and angles within the wafer feature.

[0162] FIG. 13 is a block diagram illustrating an example of a machine 1300 that can implement or control one or more exemplary processes described herein. In alternative embodiments, the machine 1300 may operate as a stand-alone device or may be connected (e.g., network-connected) to other machines. In a network deployment, the machine 1300 can operate in the capacity of a server machine, a client machine, or both in a server-client network environment. In one example, the machine 1300 can operate as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Further, although only a single machine 1300 is shown, the term “machine” should also be construed to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein via cloud computing, software as a service (SaaS), or other computer cluster configurations.

[0163] The examples described in this specification may include, or be operable by, logic, some components, or mechanisms. A circuit set is a set of circuits implemented in a tangible entity that includes hardware (e.g., simple circuits, gates, logic, etc.). The membership of a circuit set can flexibly accommodate the passage of time and the variability of the underlying hardware. A circuit set includes members that can perform certain operations during operation, either alone or in combination. In one example, the hardware of a circuit set may be designed to be invariant (e.g., hardwired) to perform a particular operation. In one example, the hardware of a circuit set may include a variably connected physical component (e.g., an execution unit, a transistor, a simple circuit) that includes a computer-readable medium physically modified (e.g., by magnetic, electrical, movable placement of immutable mass particles, etc.) to encode instructions for a particular operation. When connecting physical components, the underlying electrical properties of the hardware components are changed (e.g., from insulator to conductor, or vice versa). The instructions enable an embedded hardware (e.g., an execution unit or a loading mechanism) to create members of a circuit set within the hardware via a variable connection and execute a part of a particular operation during operation. Thus, the computer-readable medium is communicatively coupled to other components of the circuit set when the device is operating. In one example, any of the physical components may be used by multiple members of multiple circuit sets. For example, during operation, an execution unit may be used by a first circuit of a first circuit set at one point in time and reused by a second circuit within the first circuit set or by a third circuit within a second circuit set at another point in time.

[0164] A machine (e.g., a computer system) 1300 can include a hardware processor 1302 (e.g., a central processing unit (CPU), a hardware processor core, or any combination thereof), a graphics processing unit (GPU) 1303, a main memory 1304, and a static memory 1306, and some or all of which can communicate with each other via an interconnect (e.g., a bus) 1308. The machine 1300 can further include a display device 1310, an alphanumeric input device 1312 (e.g., a keyboard), and a user interface (UI) navigation device 1314 (e.g., a mouse). In one example, the display device 1310, the alphanumeric input device 1312, and the UI navigation device 1314 can be a touch screen display. The machine 1300 can further include a mass storage device (e.g., a drive unit) 1316, a signal generation device 1318 (e.g., a speaker), a network interface device 1320, and one or more sensors 1321 (such as a global positioning system (GPS) sensor, a compass, an accelerometer, or another sensor). The machine 1300 can include an output controller 1328 such as a serial (e.g., universal serial bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC)) connection to communicate with or control one or more peripheral devices (e.g., a printer, a card reader).

[0165] The mass storage device 1316 can include a machine-readable medium 1322. One or more sets of data structures or instructions 1324 (e.g., software) that implement or are utilized by any one or more of the techniques or functions described herein are stored in the machine-readable medium 1322. Also, the instructions 1324 may be present, in whole or at least in part, within the main memory 1304, within the static memory 1306, within the hardware processor 1302, or within the GPU 1303 during execution by the machine 1300. In one example, any one of the hardware processor 1302, GPU 1303, main memory 1304, static memory 1306, or mass storage device 1316, or any combination thereof, may constitute a machine-readable medium.

[0166] Although the machine-readable medium 1322 is shown as a single medium, the term "machine-readable medium" can include a single medium configured to store one or more instructions 1324, or a plurality of media (e.g., a centralized or distributed database, and / or associated caches and servers).

[0167] The term "machine-readable medium" can include any medium that can store, encode, or carry instructions 1324 for execution by a machine 1300, and that can cause the machine 1300 to perform any one or more of the techniques of this disclosure, or any medium that can store, encode, or carry a data structure used by such instructions 1324 or a data structure associated with such instructions 1324. Non-limiting examples of machine-readable media can include solid state memories, optical media, and magnetic media. In one example, a mass machine-readable medium includes a machine-readable medium 1322 having a plurality of particles with invariant (e.g., stationary) mass. Thus, a mass machine-readable medium is not a signal propagating temporarily. Specific examples of mass machine-readable media can include non-volatile memories such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0168] Further, the instructions 1324 can be transmitted or received over a communication network 1326 via a network interface device 1320 using a transmission medium.

[0169] Throughout this specification, components, operations, or structures that are described as a single instance may be implemented by a plurality of instances. Individual operations of one or more methods are illustrated and described as separate operations, but one or more of the individual operations may be performed simultaneously, and need not be performed in the order illustrated. Structures and functions presented as separate components in an exemplary configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this specification.

[0170] The illustrated embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments may be used and other embodiments may be derived from the teachings disclosed herein without departing from the scope of the present disclosure. Accordingly, this detailed description should not be construed in a limiting sense, and the scope of the various embodiments is defined only by the appended claims and all ranges of equivalents to which such claims are entitled.

[0171] As used herein, the term "or" may be construed in an inclusive or exclusive sense. Further, multiple instances may be provided for resources, operations, or structures described herein as a single instance. Additionally, the boundaries between various resources, operations, modules, engines, and data stores are somewhat arbitrary, and particular operations are presented in the context of specific exemplary configurations. Other allocations of functionality are envisioned and may be included within the scope of various embodiments of the present disclosure. In general, structures and functionality presented as separate resources in an exemplary configuration may be implemented as a combined structure or resource. Similarly, structures and functionality presented as a single resource may be implemented as separate resources. These and other variations, modifications, additions, and improvements are within the scope of the embodiments of the present disclosure as represented by the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a limiting sense.

Claims

1. 1. A method comprising: obtaining a plurality of machine learning (ML) models, each of the models relating to predicting a performance metric for operation of a semiconductor manufacturing tool, each of the ML models utilizing a plurality of features defining an input for the ML model; receiving, by one or more processors, a process definition for manufacturing a product at the semiconductor manufacturing tool; estimating, by the one or more processors, performance of the process definition for use in the semiconductor manufacturing tool using one or more ML models from the plurality of ML models; presenting on a display a result indicative of the estimate of the performance of the production of the product; and A method comprising:

2. 2. The method of claim 1 , Creating a machine learning (ML) model from the plurality of ML models includes: obtaining training data for the ML model, the training data providing values ​​of the features for the ML model; training an ML algorithm to obtain said ML model; A method comprising:

3. 3. The method of claim 2, Obtaining the training data for the ML model comprises: conducting an experiment at the semiconductor manufacturing tool; measuring values ​​of said characteristic for said experiment; using the measured values ​​for the training data; A method comprising:

4. 4. The method of claim 3, Obtaining the training data for the ML model comprises: Training a second-order ML model to produce a reduced order model (ROM) for the measurement data; using the output of the secondary ML model as additional training data; The method further comprising:

5. 3. The method of claim 2, Obtaining the training data for the ML model comprises: Obtaining the training data by performing a physics-based simulation of the operation of the semiconductor manufacturing tool. A method comprising:

6. 2. The method of claim 1 , The method, wherein the plurality of ML models include a chamber model, a process matrix model, a substrate scale model, a wafer feature model, an atomistic model, and an electronic state model.

7. 2. The method of claim 1 , A method, wherein a chamber ML model is for estimating the geometry of the chamber in the semiconductor manufacturing tool using inputs including design dimensions and outputs including a definition of a geometry of the chamber.

8. 2. The method of claim 1 , The process matrix ML model is for analyzing the behavior of the environment of a substrate during processing with inputs including one or more of electric field, magnetic field B, current density, chemicals, reaction cross sections, reaction pathways, material properties, RF frequency, RF voltage, temperature, and RF bias, and outputs of the plasma ML model include one or more of charged species density and flux, ambipolar field, electron temperature, electron energy distribution function (EEDF), ion energy angular distribution (IEAD), on-wafer flux, charge density (surface and volume) source, and species loss or creation terms.

9. 2. The method of claim 1 , The method includes a substrate-level ML model for analyzing the performance of a sheath close to the substrate using inputs including one or more of radio frequency (RF) voltage, electric field, source terms, reaction collision cross sections, solution concentrations, and reaction paths, and outputs of the sheath ML model include one or more of flux, ion energy and angular distribution, conduction and displacement currents, ion transit times, surface functionalization, and charge density on the wafer surface.

10. 2. The method of claim 1 , A method, comprising: a wafer feature ML model for analyzing a layout of a wafer feature using input including one or more of a design layout, a mask layer, and an initial step; and an output of the wafer feature ML model including a geometric description of the wafer feature.

11. 2. The method of claim 1 , A method in which a wafer chemical ML model is for analyzing the chemicals of a wafer feature using inputs including one or more of on-wafer fluxes, material properties, reaction pathways, reaction rates, ion angular yields, etch thresholds, sticking coefficients, and accommodation coefficients, and outputs of the wafer chemical ML model include one or more of geometry evolution and front tracking, distribution of species within the wafer feature, and distribution of ion energy and angles within the wafer feature.

12. 1. A system comprising: a memory containing instructions; one or more computer processors, the instructions, when executed by the one or more computer processors, providing the system with: obtaining a plurality of machine learning (ML) models, each model relating to predicting a performance metric for operation of a semiconductor manufacturing tool, each ML model utilizing a plurality of features defining an input for the ML model; receiving a process definition for manufacturing a product at the semiconductor manufacturing tool; utilizing one or more ML models from the plurality of ML models to estimate performance of the process definition for use in the semiconductor manufacturing tool; presenting on a display a result indicative of the estimate of the performance of the production of the product. one or more computer processors for performing operations including A system comprising:

13. 13. The system of claim 12, Creating a machine learning (ML) model from the plurality of ML models includes: obtaining training data for the ML model, the training data providing values ​​of the features for the ML model; training an ML algorithm to obtain said ML model; Including, the system.

14. 14. The system of claim 13, Obtaining the training data for the ML model comprises: conducting an experiment at the semiconductor manufacturing tool; measuring values ​​of said characteristic for said experiment; using the measured values ​​for the training data; and Obtaining additional training data by performing a physics-based simulation of the semiconductor manufacturing tool; Including, the system.

15. 13. The system of claim 12, A system, wherein the chamber ML model is for estimating the geometry of the chamber in the semiconductor manufacturing tool using inputs including design dimensions and outputs including a definition of the geometry of the chamber.

16. 13. The system of claim 12, The process matrix ML model is for analyzing the behavior of the environment of a substrate during processing with inputs including one or more of electric field, magnetic field B, current density, chemicals, reaction cross sections, reaction pathways, material properties, temperature, mass transport, RF frequency, RF voltage, and RF bias, and outputs of said process matrix ML model include one or more of charged species density and flux, ambipolar field, electron temperature, electron energy distribution function (EEDF), ion energy angular distribution (IEAD), on-wafer flux, charge density (surface and volume) source, and species loss or creation terms.

17. When executed by a machine, the machine: obtaining a plurality of machine learning (ML) models, each of the models relating to predicting a performance metric for operation of a semiconductor manufacturing tool, each of the ML models utilizing a plurality of features defining an input for the ML model; receiving a process definition for manufacturing a product at the semiconductor manufacturing tool; utilizing one or more ML models from the plurality of ML models to estimate performance of the process definition for use in the semiconductor manufacturing tool; presenting on a display a result indicative of the estimate of the performance of the production of the product. A machine-readable storage medium for performing operations including:

18. 20. The machine-readable storage medium of claim 17, Creating a machine learning (ML) model from the plurality of ML models includes: obtaining training data for the ML model, the training data providing values ​​of the features for the ML model; training an ML algorithm to obtain said ML model; 1. A machine-readable storage medium comprising:

19. 20. The machine-readable storage medium of claim 18, Obtaining the training data for the ML model comprises: conducting an experiment at the semiconductor manufacturing tool; measuring values ​​of said characteristic for said experiment; using the measured values ​​for the training data; and Obtaining additional training data by performing a physics-based simulation of the semiconductor manufacturing tool; 1. A machine-readable storage medium comprising:

20. 20. The machine-readable storage medium of claim 17, A machine-readable storage medium, wherein a chamber ML model is for estimating the geometry of the chamber in the semiconductor manufacturing tool with inputs including design dimensions and outputs including a definition of a geometry of the chamber.

Citation Information

Patent Citations

  • Method and process for performing machine learning on complex multivariate wafer processing equipment

    JP2019537240A

  • Virtual measuring system and method for predicting the quality of thin film transistor liquid crystal display processes

    US20120016643A1

  • Multi-model metrology

    US20150058813A1

  • Generating robust machine learning predictions for semiconductor manufacturing processes

    US20180356807A1

  • Metrology system for machine learning-based manufacturing error predictions

    WO2018204410A1