Intelligent manufacturing process optimization method and system of integrated industrial large model
By integrating intelligent manufacturing process optimization methods based on large industrial models, the problems of low efficiency and difficulty in analyzing causal relationships in high-dimensional parameter spaces in semiconductor manufacturing have been solved. This has enabled the achievement of globally optimal parameter combinations and cross-process collaboration, thereby improving the yield and production efficiency of semiconductor manufacturing.
Patent Information
- Application Number
- CN202511298226.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-11-18
AI Technical Summary
Existing semiconductor manufacturing process optimization methods are inefficient in high-dimensional parameter spaces, unable to achieve global optimal solutions, lack the ability to analyze causal relationships, have poor cross-process collaboration, and are difficult to accumulate and transfer knowledge, resulting in yield losses and low production efficiency.
The intelligent manufacturing process optimization method adopts an integrated industrial big model. Through multi-source time series data fusion, industrial big model architecture and federated learning, it generates a correlation diagram of potential defect risk areas and key parameters, generates equipment operation instructions in real time, performs cross-process collaborative scheduling and bottleneck prediction, and builds a self-evolving decision system.
It enables efficient exploration of the entire parameter space, accurately analyzes complex causal chains, improves yield and production efficiency, shortens process debugging cycle, reduces equipment downtime and energy costs, and supports knowledge transfer across factories and process nodes.
Smart Images

Figure CN120975333A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semiconductor manufacturing technology, and more specifically, to a method and system for optimizing intelligent manufacturing processes based on an integrated industrial large-scale model. Background Technology
[0002] Currently, there are three typical technical solutions in the field of semiconductor manufacturing process optimization: 1. Traditional Experimental Optimization Methods Methods such as Design of Experiments (DOE) and Response Surface Analysis (RSA) involve conducting physical experiments with pre-defined parameter gradients to construct statistical models of parameters and yield. For example, in photolithography, CD uniformity data is obtained by fixing the development time and adjusting the exposure dose (5-20 mJ / cm² gradient) to fit the optimal parameters. These methods rely on empirically pre-defined parameter ranges and are suitable for low-dimensional parameter optimization in processes above 100nm. However, in the 100-dimensional parameter space of processes below 7nm, experimental costs skyrocket and the methods cannot cover the optimal solution across the entire range.
[0003] 2. Local optimization scheme for a single AI model; Using machine learning models (such as CNN and random forest) to solve specific process problems: Defect detection: CNN was used to classify wafer SEM images, with a recognition rate of about 92%, but it could not correlate the causal relationship between defects and process parameters; Parameter prediction: The etching rate is predicted by LSTM, but the influence of plasma physics on the prediction accuracy is ignored, with an error rate of ±8%.
[0004] 3. Manufacturing Execution System (MES) and Local Decision-Making System Mainstream semiconductor factories use MES systems for production scheduling, combined with SCADA systems to collect equipment data. However, the data is stored in independent databases such as Oracle and SQL Server, forming "information silos." Decision-making relies on manual judgment based on FMEA (Failure Mode and Effects Analysis) manuals. For example, for thin film deposition defects, engineers need to consult historical cases and manually adjust the RF power, resulting in a response delay of more than 4 hours.
[0005] Parameter optimization is inefficient. Traditional DOE methods are limited by the cost of physical experiments. In the 50+ dimension parameter space of a 3nm process, if each dimension contains 10 discrete values, the number of combinations will reach 10^50. The actual sampling coverage is extremely low, covering only a very small number of parameter combinations, leading to the omission of the optimal solution. Single AI models lack global exploration capabilities. For example, CNNs can only optimize local lithography parameters and cannot coordinate with etching process parameters, resulting in a yield loss of about 7%-12%. Under nanometer-level precision requirements (such as 3nm / 5nm processes), even small fluctuations in process parameters (such as temperature, pressure, and dosage) of lithography and etching can lead to device failure. Traditional empirical trial-and-error methods are inefficient and slow in iteration.
[0006] The ability to analyze causal relationships is weak. Existing statistical models (such as regression analysis) and deep learning models (such as CNNs) can only capture the correlation between parameters and defects, but cannot uncover deep causal chains. For example, when gate oxide damage is detected, traditional models cannot distinguish whether it is caused by abnormal etching gas ratios or wafer temperature fluctuations, and the root cause localization accuracy is less than 50% (refer to SEMI standard document SEMI E178-0603). Manual reliance on experience or local small models for process adjustments cannot achieve the exploration of globally optimal parameter combinations and cross-process collaborative optimization.
[0007] Physical constraints and process coordination are lacking. Existing optimization methods do not incorporate semiconductor physics equations. For example, in thin film deposition optimization, the exponential relationship between surface diffusion coefficient and temperature (Arrhenius equation) is ignored, causing recommended parameters to exceed process physical limits, resulting in yield fluctuations of ±5% during actual execution. In terms of cross-process coordination, the scheduling rules of the MES system are based on a fixed capacity model, which cannot dynamically match the timing coupling relationship between photolithography and etching, resulting in equipment idle rates exceeding 15%.
[0008] Knowledge accumulation and transfer are difficult. Process knowledge is scattered across manuals, engineer experience, and local models. For example, the etching parameter debugging experience of TSMC's N5 process cannot be directly transferred to SMIC's 14nm production line. Federated learning technology has not been applied, and data from multiple plants cannot be collaboratively trained, resulting in a convergence cycle of up to 6 months for new production line processes. Data from production equipment (MES, SCADA), quality inspection systems, and process documents are scattered, making it difficult to establish multi-dimensional correlation analysis and limiting the depth of process optimization.
[0009] Industrial large modeling (LLM) possesses the capabilities for processing massive amounts of industrial data, fusing multimodal knowledge, and performing complex reasoning, providing a technological path to overcome the aforementioned bottlenecks. For example, the ability to generate large models can simulate combinations of process parameters, reason to uncover the root causes of defects, and analyze to predict yield trends, thereby constructing an intelligent, closed-loop process optimization system. Summary of the Invention
[0010] This invention addresses the problems of low efficiency, insufficient accuracy, and difficulty in knowledge accumulation in traditional semiconductor process optimization. It proposes an intelligent manufacturing process optimization method and system that integrates a large industrial model. This method achieves dynamic optimal combination solutions for process parameters such as lithography dose, etching rate, and film thickness; generates equipment operation instructions and anomaly intervention strategies in real time to improve the accuracy and timeliness of operations; and performs cross-process collaborative scheduling and bottleneck prediction to optimize the overall production process, ultimately improving the yield, efficiency, and flexibility of semiconductor manufacturing.
[0011] The specific implementation details of this invention are as follows: A method for optimizing intelligent manufacturing processes using an integrated industrial large-scale model, specifically including the following steps: Step S1: Call the multi-source temporal data obtained by neural symbolic structure fusion; Step S2: Input the acquired multi-source time-series data into the constructed industrial big data model architecture to generate a correlation diagram of potential defect risk areas and key parameters; the industrial big data model architecture includes a basic layer, a generation layer, an inference layer, and an analysis layer; Step S3: Based on the generated potential defect risk areas and key parameter correlation diagram, generate candidate dose compensation schemes and output the optimization candidate set; Step S4: Perform closed-loop optimization, push the optimization candidate set to the lithography machine control system for execution, and evaluate the effectiveness of the candidate dose compensation scheme; if the yield improvement meets the target, the parameter combination is automatically precipitated into the process rule engine knowledge; otherwise, backtrack the causal chain to correct the model and start a new round of optimization; at the same time, through federated learning, the optimization cases are aggregated with other wafer cases to update the model knowledge base and realize the continuous evolution of the model.
[0012] To better realize the present invention, step S1 further includes the following steps: Step S11: Acquire multi-source time-series data streams and integrate them with historical process documents and failure knowledge bases; Step S12: Call the TransE algorithm to map entities and relations to a low-dimensional vector space and calculate the distance between entity and relation vectors; Step S13: Call the generator and discriminator to construct the objective function and simulate photoresist imaging defects under different dose deviations; Step S14: Call the self-supervised learning SimCLR framework to augment the random data of the same image to generate multiple views, and calculate the contrast loss between the views; Step S15: Construct a graph attention network of wafer defect spatial distribution heatmap and process timing data, call the multi-head attention mechanism, and calculate the node feature weights; Step S16: Based on the graph attention network, predict the systematic defect chain of the acquired multi-source time-series data stream.
[0013] To better realize the present invention, step S2 further includes the following steps: Step S21: Call the pre-trained large model for vertical fine-tuning in semiconductor manufacturing, and use the quantum effect equation and the thin film growth kinetics equation as physical constraints to construct the base layer; Step S22: Invoke the diffusion model and generate an adversarial network to simulate the impact of changes in process parameters, covering extreme operating conditions to discover potential defects; Step S23: Based on the counterfactual inference parameters of the causal Bayesian network, the intervention effect is analyzed, and the graph neural network is called to parse the cross-process time constraints, construct the inference layer, and generate optimization and reordering suggestions; Step S24: Embed the physical constraint neural network into physical laws, combine real-time data to optimize the process curve, and predict process sensitivity and dynamically adjust parameters through a temporal convolutional network.
[0014] To better realize the present invention, step S3 further includes the following steps: Step S31: Input the current state parameters of semiconductor manufacturing into the large model, output candidate process schemes, and evaluate the performance benefits and physical feasibility of the schemes based on reinforcement learning combined with the physical simulation engine, and select the optimal parameter combination to deploy to the equipment control system. Step S32: Based on the constructed intelligent decision-making system, when an anomaly occurs in the wafer, the large model inference is triggered to locate the root cause and generate multi-dimensional operation instructions. Step S33: Analyze the cascading effects of changes in process steps, generate scheduling suggestions, and dynamically adjust process parameters based on digital twin simulation and physical model prediction.
[0015] To better realize the present invention, step S31 further includes the following steps: Step S311: Construct a parameter optimization framework based on an industrial large model, taking process state parameters as input, to obtain candidate dose compensation schemes; Step S312: Invoke the near-end policy optimization algorithm in conjunction with the physics simulation engine to construct the reward function; Step S313: Iteratively optimize and select the globally optimal parameter combination and deploy it to the equipment control system to improve yield.
[0016] To better realize the present invention, step S32 further includes the following steps: Step S321: Establish a large model inference mechanism under abnormal operating conditions. When the defect density on the wafer surface exceeds the threshold, the fault diagnosis process is triggered. Step S322: Call the large model to analyze multi-dimensional parameters, and call the fault tree reasoning algorithm to locate the root cause, and generate an operation instruction set including equipment calibration parameters, process restart strategy, and manual intervention priority; Step S323: When the defects on the wafer surface increase dramatically, the large model inference mechanism is triggered to quickly locate the root cause and generate multi-dimensional operation instructions.
[0017] To better realize the present invention, step S33 further includes the following steps: Step S331: Based on the obtained equipment space time, yield loss, and set weight coefficients, establish a scheduling optimization model; Step S332: Based on the scheduling optimization model, dynamically balance the Pareto front solution set of yield loss and capacity loss; Step S333: Based on the PINN prediction of wafer deformation trend, dynamically adjust the thin film deposition temperature-power window.
[0018] To better realize the present invention, step S4 further includes the following steps: Step S41: Call the decision tree algorithm to construct the reasoning path, quantify the impact of process parameters on the input process challenge through the sensitivity weight analysis equation, dynamically generate the dosage compensation scheme, and form a visual causal heat map; Step S42: Call the horizontal federated learning architecture to extract features and aggregate model parameters from the process optimization cases of multiple wafer fabs, and construct the Pareto front solution set of process parameters and defect rate; Step S43: If the process verification confirms that the parameter combination is effective, the parameter combination is transformed into decision knowledge of the process rule engine; if the process verification is a new failure case, the DistilBERT architecture is called to perform knowledge distillation, and the loss function is called to compress the model parameters.
[0019] Based on the intelligent manufacturing process optimization method of the integrated industrial big model proposed above, in order to better realize the present invention, a further intelligent manufacturing process optimization system of the integrated industrial big model is proposed to execute the intelligent manufacturing process optimization method of the integrated industrial big model mentioned above; including a multi-source data fusion and knowledge graph enhancement module, an industrial big model architecture design module, a closed-loop process optimization execution module, and a human-machine collaboration and model evolution module. The multi-source data fusion and knowledge graph enhancement module is used to call up multi-source time-series data obtained by neural symbol structure fusion. The industrial large model architecture design module is used to input the acquired multi-source time-series data into the constructed industrial large model architecture to generate potential defect risk areas and key parameter correlation diagrams; the industrial large model architecture includes a basic layer, a generation layer, an inference layer, and an analysis layer. The closed-loop process optimization execution module is used to generate candidate dose compensation schemes and output an optimization candidate set based on the generated potential defect risk areas and key parameter correlation diagrams. The human-machine collaboration and model evolution module is used for closed-loop optimization. It pushes the optimization candidate set to the lithography machine control system for execution and evaluates the effectiveness of the candidate dose compensation scheme. If the yield improvement meets the target, the parameter combination is automatically precipitated into the process rule engine knowledge. Otherwise, it backtracks the causal chain to correct the model and starts a new round of optimization. At the same time, it aggregates the optimization cases with other wafer cases through federated learning, updates the model knowledge base, and realizes the continuous evolution of the model.
[0020] The present invention has the following beneficial effects: (1) In terms of accuracy and yield, this invention achieves efficient exploration of the entire parameter space. Traditional trial-and-error methods are limited by the cost of physical experiments and the explosion of parameter combinations. However, this method simulates virtual experiments through the generation capability of industrial large models, compressing the parameter tuning cycle from the weekly level to the hourly level. At the same time, it accurately analyzes complex causal chains, reducing the error by 37% compared with traditional methods, and realizes intelligent coordination between physical constraints and process timing. The yield fluctuation control accuracy is improved from ±5% to ±1.5%. In processes below 7nm, the wafer critical dimension control accuracy reaches the sub-nanometer level (CD uniformity ±0.8nm), improving the yield by 5%-15% compared with traditional methods.
[0021] (2) In terms of efficiency and cost, this invention upgrades manual experience-based decision-making to data-model dual-driven reasoning, shortening the process debugging cycle and reducing equipment downtime and energy costs. Through process collaborative restructuring, the FabO scheduling efficiency is improved by 40%, and the overall equipment efficiency (OEE) is improved by 20%-30%.
[0022] (3) In terms of knowledge accumulation and transfer, this invention constructs a knowledge-driven self-evolving decision-making system, transforming discrete knowledge such as engineer experience, process manuals, and failure cases into a dynamic and reasonable causal knowledge graph, supporting the construction of an autonomous manufacturing system with self-diagnosis, self-decision-making, and self-evolution, shortening the learning curve for new engineers by 50%, accelerating technology transfer across factories and process nodes (such as 28nm→7nm) by 40%, and providing a core technology foundation for next-generation semiconductor manufacturing such as 6G communication and Chiplet heterogeneous integration. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the logical architecture of the modules provided by the present invention. Detailed Implementation
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments, and therefore should not be regarded as a limitation on the scope of protection. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set up," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0026] Example 1: This embodiment proposes a smart manufacturing process optimization method integrating a large industrial model, which specifically includes the following steps: Step S1: Call the multi-source temporal data obtained by neural symbolic structure fusion; Step S1 specifically includes the following steps: Step S11: Acquire multi-source time-series data streams and integrate them with historical process documents and failure knowledge bases; Step S12: Call the TransE algorithm to map entities and relations to a low-dimensional vector space and calculate the distance between entity and relation vectors; Step S13: Call the generator and discriminator to construct the objective function and simulate photoresist imaging defects under different dose deviations; Step S14: Call the self-supervised learning SimCLR framework to augment the random data of the same image to generate multiple views, and calculate the contrast loss between the views; Step S15: Construct a graph attention network of wafer defect spatial distribution heatmap and process timing data, call the multi-head attention mechanism, and calculate the node feature weights; Step S16: Based on the graph attention network, predict the systematic defect chain of the acquired multi-source time-series data stream.
[0027] Step S2: Input the acquired multi-source time-series data into the constructed industrial big data model architecture to generate a correlation diagram of potential defect risk areas and key parameters; the industrial big data model architecture includes a basic layer, a generation layer, an inference layer, and an analysis layer; Step S2 specifically includes the following steps: Step S21: Call the pre-trained large model for vertical fine-tuning in semiconductor manufacturing, and use the quantum effect equation and the thin film growth kinetics equation as physical constraints to construct the base layer; Step S22: Invoke the diffusion model and generate an adversarial network to simulate the impact of changes in process parameters, covering extreme operating conditions to discover potential defects; Step S23: Based on the counterfactual inference parameters of the causal Bayesian network, the intervention effect is analyzed, and the graph neural network is called to parse the cross-process time constraints, construct the inference layer, and generate optimization and reordering suggestions; Step S24: Embed the physical constraint neural network into physical laws, combine real-time data to optimize the process curve, and predict process sensitivity and dynamically adjust parameters through a temporal convolutional network.
[0028] Step S3: Based on the generated potential defect risk areas and key parameter correlation diagram, generate candidate dose compensation schemes and output the optimization candidate set; Step S3 specifically includes the following steps: Step S31: Input the current state parameters of semiconductor manufacturing into the large model, output candidate process schemes, and evaluate the performance benefits and physical feasibility of the schemes based on reinforcement learning combined with the physical simulation engine, and select the optimal parameter combination to deploy to the equipment control system. Step S31 specifically includes the following steps: Step S311: Construct a parameter optimization framework based on an industrial large model, taking process state parameters as input, to obtain candidate dose compensation schemes; Step S312: Invoke the near-end policy optimization algorithm in conjunction with the physics simulation engine to construct the reward function; Step S313: Iteratively optimize and select the globally optimal parameter combination and deploy it to the equipment control system to improve yield.
[0029] Step S32: Based on the constructed intelligent decision-making system, when an anomaly occurs in the wafer, the large model inference is triggered to locate the root cause and generate multi-dimensional operation instructions. Step S32 specifically includes the following steps: Step S321: Establish a large model inference mechanism under abnormal operating conditions. When the defect density on the wafer surface exceeds the threshold, the fault diagnosis process is triggered. Step S322: Call the large model to analyze multi-dimensional parameters, and call the fault tree reasoning algorithm to locate the root cause, and generate an operation instruction set including equipment calibration parameters, process restart strategy, and manual intervention priority; Step S323: When the defects on the wafer surface increase dramatically, the large model inference mechanism is triggered to quickly locate the root cause and generate multi-dimensional operation instructions.
[0030] Step S33: Analyze the cascading effects of changes in process steps, generate scheduling suggestions, and dynamically adjust process parameters based on digital twin simulation and physical model prediction.
[0031] Step S33 specifically includes the following steps: Step S331: Based on the obtained equipment space time, yield loss, and set weight coefficients, establish a scheduling optimization model; Step S332: Based on the scheduling optimization model, dynamically balance the Pareto front solution set of yield loss and capacity loss; Step S333: Based on the PINN prediction of wafer deformation trend, dynamically adjust the thin film deposition temperature-power window.
[0032] Step S4: Perform closed-loop optimization, push the optimization candidate set to the lithography machine control system for execution, and evaluate the effectiveness of the candidate dose compensation scheme; if the yield improvement meets the target, the parameter combination is automatically precipitated into the process rule engine knowledge; otherwise, backtrack the causal chain to correct the model and start a new round of optimization; at the same time, through federated learning, the optimization cases are aggregated with other wafer cases to update the model knowledge base and realize the continuous evolution of the model.
[0033] Step S4 specifically includes the following steps: Step S41: Call the decision tree algorithm to construct the reasoning path, quantify the impact of process parameters on the input process challenge through the sensitivity weight analysis equation, dynamically generate the dosage compensation scheme, and form a visual causal heat map; Step S42: Call the horizontal federated learning architecture to extract features and aggregate model parameters from the process optimization cases of multiple wafer fabs, and construct the Pareto front solution set of process parameters and defect rate; Step S43: If the process verification confirms that the parameter combination is effective, the parameter combination is transformed into decision knowledge of the process rule engine; if the process verification is a new failure case, the DistilBERT architecture is called to perform knowledge distillation, and the loss function is called to compress the model parameters.
[0034] Working Principle: In this embodiment, during a photoresist development defect analysis, 10 sets of exposure parameters were adjusted through DOE experiments, but the problem of excessive edge roughness still could not be solved. Further research revealed that traditional methods only optimized the parameters of the single photolithography machine, without considering the influence of the stress distribution of the preceding thin film deposition on the photoresist adhesion—this cross-process coupling relationship was ignored by the "locality assumption" of existing models, and the fundamental reason for this neglect is the lack of a unified modeling framework that can integrate the physical mechanisms of multiple processes.
[0035] To address the challenge of causal analysis, an attempt was made to combine Bayesian networks with semiconductor physics equations. However, the initial model crashed due to parameter dimensionality explosion (over 1000 dimensions). Further algorithmic improvements led to a "physical constraint dimensionality reduction" strategy, which, by embedding Poisson and diffusion equations, compressed the effective parameter dimension to 80 dimensions, significantly improving the speed of causal inference.
[0036] In addressing the credibility issue of virtual experiments, it was found that the parameter combinations generated by traditional GANs often violated process boundaries (such as exposure dose exceeding the hardware limits of the lithography machine). By introducing a diffusion model and loading SEMI standard process constraints (SEMI P30-0302), the physical compliance rate of the generated data was increased from 65% to 98%, ultimately enabling virtual experiments to replace 70% of silicon wafer verification.
[0037] Based on the above research, this embodiment constructs a closed-loop architecture of "generation-reasoning-analysis": multi-source data is integrated through a large industrial model, generation capabilities are used to explore global parameters, the reasoning layer analyzes causal chains, and the analysis layer ensures physical constraints, ultimately forming an optimization scheme covering the entire process chain. In a 14nm production line verification, this scheme improved the yield from 82% to 91%, validating the technical feasibility.
[0038] This embodiment addresses the problems of low efficiency, insufficient accuracy, and difficulty in knowledge accumulation in traditional semiconductor process optimization through generative simulation, causal reasoning, and multi-dimensional analysis. Specific objectives include: achieving dynamic optimal combination solutions for process parameters such as lithography dose, etching rate, and film thickness; generating equipment operation instructions and anomaly intervention strategies in real time to improve the accuracy and timeliness of operations; and performing collaborative scheduling and bottleneck prediction across processes (such as lithography → etching → CMP) to optimize the overall production process and ultimately improve the yield, efficiency, and flexibility of semiconductor manufacturing.
[0039] Example 2: This embodiment is based on the above embodiment 1, such as... Figure 1 As shown, through a closed-loop architecture of generation-inference-analysis driven by a large industrial model, the limitations of traditional semiconductor process optimization are overcome, achieving a triple paradigm shift: Efficient exploration of the entire parameter space: Traditional trial-and-error methods are limited by the cost of physical experiments and the explosion of parameter combinations (such as lithography dose, etching rate, film thickness and other hundred-dimensional parameter spaces). This method simulates virtual experiments by generating large industrial models, and generates wafer performance evolution data under tens of millions of process parameter perturbations using diffusion models or GANs. This replaces high-cost silicon wafer verification and compresses the parameter tuning cycle from the weekly level to the hour level.
[0040] Precise Analysis of Complex Causal Chains: Traditional statistical methods (such as FMEA) struggle to capture the deep causal dependencies between nanoscale process fluctuations and defects (such as LER line edge roughness and gate damage). This method integrates a causal inference engine (Graph Neural Network (GNN) + Bayesian counterfactual model) to mine a multidimensional causal graph of process parameters, defects, and electrical performance, achieving root cause localization at the 0.1% yield fluctuation level (e.g., tracing the path from abnormal etching rate to uneven plasma density to gate oxide layer damage), reducing errors by 37% compared to traditional methods.
[0041] Intelligent Collaboration of Physical Constraints and Process Timing: Traditional rule-based timing optimization, such as CMP pressure scheduling, ignores the dynamic coupling between process physical limits and wafer performance. This solution embeds semiconductor physics equations such as thin film growth kinetics and quantum tunneling effect into a Physically Constrained Neural Network (PINN), and uses a timing neural network (TCN / LSTM) to analyze the timing characteristics of multiple processes from photolithography to etching to thin film deposition. While optimizing equipment utilization, it ensures physical boundary constraints such as thin film thickness uniformity and etching selectivity, improving yield fluctuation control accuracy from ±5% to ±1.5%.
[0042] In summary, traditional MES systems rely on manual experience to iterate process rules, while this embodiment constructs a knowledge-driven, self-evolving decision-making system. Through federated learning and knowledge distillation architecture, it aggregates cross-factory process cases and failure modes, such as the experience of migrating TSMC's N3E process. Under the premise of protecting data privacy, it continuously updates the large model knowledge base, which shortens the learning curve for new engineers by 50% and supports the smooth migration of knowledge across multiple advanced process nodes.
[0043] Step S1: Multi-source data fusion and knowledge graph enhancement.
[0044] This embodiment aims to break down data silos through a specific architecture, achieving effective fusion of multi-source data and constructing a process knowledge graph. For data acquisition, it integrates multi-source high-dimensional time-series data streams, merges historical process documents and failure knowledge bases, and utilizes algorithms to mine data correlations. Addressing the scarcity of defect data, it employs generative adversarial networks to simulate defects, combined with a self-supervised learning framework to enhance image feature representation and improve model generalization ability. Simultaneously, it constructs a graph attention network to correlate the spatial distribution of wafer defects with process time-series data, predicting in advance systemic defect chains caused by equipment anomalies and reducing yield losses.
[0045] Solving the A1 data silo problem; In semiconductor smart manufacturing, the problem of data silos severely restricts process optimization.
[0046] This invention achieves effective fusion of multi-source data and knowledge graph construction through a neural symbolic architecture, meeting the needs of complex process analysis. During the data acquisition phase, it integrates real-time, multi-dimensional time-series data streams such as photoresist spin coating temperature, etching chamber RF power, wafer CD measurement, and thin film elliptic polarization spectrum, while simultaneously fusing historical process documents (e.g., photoresist Dill model parameters, thin film deposition reaction mechanisms) and a failure knowledge base. Based on the TransE algorithm, entities and relations are mapped to a low-dimensional vector space, and the distance between entity and relation vectors is calculated using the following formula: Where h, r, and t represent the head entity, relation, and tail entity vectors, respectively. The algorithm maps the correlation between etching rate anomalies and wafer edge chipping defects to a weighted edge graph. Experiments have shown that the inference accuracy can reach 85%.
[0047] Solving the problem of scarce A2 defect data To address the challenge of scarce defect data in semiconductor manufacturing, this embodiment employs a Generative Adversarial Network (GAN) to simulate photoresist imaging defects under different dosage deviations, such as line shortening and bridging defects. The generator G and discriminator D are optimized through adversarial training, with the objective function being: The objective function aims to make the generator G produce data that is as realistic as possible to pass the discriminator D, which in turn needs to distinguish between real and generated data as accurately as possible. Here, E represents the expected value, x represents the real data, and z represents the noisy data. It is the actual data distribution, that is, the actual distribution of photoresist imaging defect data; It is a noise distribution, usually a randomly set distribution, used by the generator to generate diverse simulated data.
[0048] Meanwhile, based on the self-supervised learning SimCLR framework, the feature representation of SEM images is enhanced through contrastive learning. Its core steps are: performing random data augmentation on the same image to generate multiple views, and calculating the contrast loss between the views. Where sim is the cosine similarity. This represents the feature vector used for calculation; τ is a temperature parameter used to adjust the difficulty of contrastive learning. The smaller the value, the greater the learning difficulty and the higher the feature discrimination; N is half the batch size of a single iteration, controlling the amount of data participating in contrastive learning. The indicator function has a value of 1 when k=i and 0 otherwise. The purpose of this loss function is to make similar image features closer together in the feature space and dissimilar features further apart. With only 100 training images, an F1-score of 0.89 can be achieved for LER detection, surpassing human expert performance and significantly improving the model's generalization ability.
[0049] Solution to systemic defects caused by A3 equipment malfunction To proactively mitigate systemic defects caused by equipment malfunctions, this invention constructs a graph attention network (GAT) based on a thermal map of wafer defect spatial distribution and process timing data. Through a multi-head attention mechanism, node feature weights are calculated: in, Let be the attention coefficients of nodes i and j at the k-th attention head. , For trainable parameters, Let i be the set of neighboring nodes of node i. It is the weight vector of the attention mechanism, used to learn the importance between nodes; It is a learnable weight matrix that performs a linear transformation on the node features; Represents the eigenvector; The vector concatenation function concatenates the transformed feature vectors of nodes i and j. LeakyReLU is the activation function used to introduce nonlinearity and prevent the network from getting trapped in local optima. GAT is used to process the spatial distribution heatmap of wafer defects and process timing data to analyze and predict systematic defect chains caused by equipment anomalies.
[0050] This network can predict systematic defect chains caused by equipment malfunctions (such as uneven CMP pressure → edge chipping → film step coverage failure) up to 24 hours in advance. Practical application verification has shown that it can reduce yield loss by 15%.
[0051] Step S2: Industrial large-scale model architecture design.
[0052] The foundation layer employs a pre-trained large model designed for vertical fine-tuning in semiconductor manufacturing, integrating physical equation constraints with SEMI standard process rules to construct a unified knowledge base. The layered enhancement module comprises a generation layer, an inference layer, and an analysis layer: the generation layer utilizes diffusion models and GANs to simulate the impact of process parameter variations, covering extreme operating conditions to identify potential defects; the inference layer uses causal Bayesian networks to counterfactually infer the effects of parameter interventions, employing graph neural networks to analyze cross-process timing constraints and generate optimization and reordering suggestions; the analysis layer embeds physical constraint neural networks into physical laws, combines real-time data to optimize process curves, and uses temporal convolutional networks to predict process sensitivity and dynamically adjust parameters to ensure process performance targets are met.
[0053] B11 Basic Layer Design Employing a pre-trained large-scale model designed for vertical fine-tuning in semiconductor manufacturing, and integrating constraints from physical equations such as quantum effects and thin film growth kinetics, as well as SEMI standard process rules such as photoresist baking temperature windows, a unified knowledge base of process, physics, and rules is constructed. KnowledgeBase=f(PhysicalEquations,SEMIRules,Other) Where f is the fusion function, PhysicalEquations includes quantum effect equations, thin film growth kinetic equations, etc., SEMIRules represents the standard process rules formulated by the International Organization for Semiconductor Equipment and Materials, and Other represents other necessary fusion factors.
[0054] B2 layered enhancement module design; The layered enhancement module consists of a generation layer, an inference layer, and an analysis layer: Generation Layer: Using a diffusion model, the impact of the lithography dose compensation curve on wafer CD uniformity and subsequent etching compatibility is dynamically simulated through the following process, generating a prediction set of LER evolution under a dose gradient of ±5%: LERPredictionDoseGradient=g(DiffusionModel,DoseGradient)=±5% Where g is the simulation function, DiffusionModel is the diffusion model, DoseGradient represents the dose gradient, and LERPrediction is the prediction set of line edge roughness (LER) evolution. Simultaneously, based on a generative adversarial network (GAN), the thin film deposition parameter perturbation experiment is generated using the following formula: ExperimentSet=h(GAN,PECVDParameters) In the formula, h is the generating function, GAN is the generative adversarial network, PECVDParameters represents parameters such as the flow rate ratio of plasma-enhanced chemical vapor deposition (PECVD) precursors, and ExperimentSet is the experimental set, covering extreme working condition boundaries that are difficult to reach by traditional trial and error methods.
[0055] Inference Layer: Based on a causal Bayesian network, counterfactual inference parameters are used to determine the effect of intervention. When an abnormal etch rate is detected, the large model backtracks to generate a causal path graph through the following process: CausalPath=k(CausalBayesianNetwork,EtchRateAnomaly) Where k is the causal analysis function, CausalBayesianNetwork is the causal Bayesian network, EtchRateAnomaly represents the etching rate anomaly, and CausalPath is the causal path diagram of "etching gas ratio - plasma density - wafer temperature field - gate damage". COMSOL multiphysics simulation was used to compare the etching selectivity improvement effect between the original parameters and the optimized ratio (e.g., Cl2 / HBr flow ratio +12%), with the objective being: EtchSelectivity>20:1 Meanwhile, a graph neural network (GNN) was used to analyze cross-process timing constraints, and the transmission effect of photoresist development time deviation on the chemical reaction efficiency of CMP slurry was analyzed through the following calculations: ReorderSuggestion=l(GNN,DevelopmentTimeDeviation) Where l is the analysis function, GNN is the graph neural network, DevelopmentTimeDeviation represents the photoresist development time deviation, and ReorderSuggestion is the wafer batch reordering suggestion.
[0056] Analysis layer: A Physically Constrained Neural Network (PINN) is embedded into the thin film deposition continuity equation and the law of energy conservation. Combined with real-time wafer warpage data, the PECVD deposition power ramp curve is optimized using the following formula: PowerRampCurve=m(PINN,WarpageData,PhysicalLaws) Where m is the optimization function, PINN is the physically constrained neural network, WarpageData represents wafer warpage data, and PhysicalLaws contains the thin film deposition continuity equation and the law of energy conservation, avoiding the thin film cracking threshold caused by thermal stress while satisfying the film thickness uniformity target. A temporal convolutional network (TCN) is used to predict the sensitivity of photoresist film thickness to the photolithography-etching timing window, and the developer temperature compensation curve is dynamically adjusted to ensure that CDUniformityCpk ≥ 1.5.
[0057] Step S3: Closed-loop process optimization execution.
[0058] The closed-loop process optimization execution system comprises three aspects: dynamic parameter tuning, intelligent operational decision-making, and collaborative process reconfiguration. Dynamic parameter tuning inputs the current state parameters of semiconductor manufacturing into a large model, outputs candidate process solutions, and then uses reinforcement learning combined with a physical simulation engine to evaluate the performance benefits and physical feasibility of the solutions, selecting the optimal parameter combination for deployment to the equipment control system. Intelligent operational decision-making triggers the large model to infer and locate the root cause when wafer anomalies occur, generating multi-dimensional operational instructions to accelerate anomaly response. Collaborative process reconfiguration can analyze the cascading effects of changes in process steps, generate scheduling suggestions, and dynamically adjust process parameters through digital twin pre-simulation and physical model prediction to balance yield and capacity, thereby improving process robustness.
[0059] C1 parameter dynamic tuning Construct a parameter optimization framework based on a large industrial model to incorporate process parameters such as lithography machine lens distortion coefficient and wafer edge alignment error. As input, the large model inference outputs candidate dose compensation schemes D and etching gas ratio correction vectors G, i.e., the candidate dose compensation schemes, such as a spatial gradient of +3% dose in the central region and -2% at the edges, and the etching gas ratio correction vector (such as fine-tuning the CF4 / O2 flow ratio). Then, the reward function R is constructed using the Proximal Policy Optimization (PPO) algorithm combined with the physics simulation engine (SPICE and COMSOL). in, To increase the driving current, To reduce leakage current, This represents physical feasibility constraints, such as membrane stress. α, β, and γ are weighting coefficients. These are used to quantify the electrical performance benefits (increased driving current / reduced leakage current) and physical feasibility (e.g., satisfying membrane stress < critical value) of each scheme. Through iterative optimization, the globally optimal parameter combination is selected and deployed to the equipment control system to improve yield.
[0060] C2 Operation Intelligent Decision Establish a large-scale model inference mechanism under abnormal operating conditions, such as the defect density on the wafer surface. Exceeding the threshold When this occurs, the fault diagnosis process is triggered. A large model is used to analyze multi-dimensional parameters such as CMP polishing pad pressure distribution P(x,y) and photoresist film thickness TPR, and a fault tree reasoning algorithm is used to locate the root cause. This generates a list containing equipment calibration parameters. Process restart strategy Priority of human intervention The set of operation instructions improves the speed of abnormal response.
[0061] When wafer surface defects surge, the large model inference mechanism is triggered to quickly locate the root cause, such as CMP polishing pad aging (uneven pressure → edge defects) or photoresist film thickness deviation (LER exceeding the threshold). This generates multi-dimensional operation instructions, including equipment calibration parameter files (CMP pressure distribution map correction, etc.), process restart strategies (developer batch switching verification), and manual intervention priority prompts (such as polishing pad replacement > gas purity verification), which improves the anomaly response speed by 7 times.
[0062] C3 process collaborative refactoring; Dynamic optimization of the process chain is achieved based on digital twin and physical information neural network (PINN). By analyzing the impact of disturbances such as photoresist supply delay on the etching / thin film deposition queue, wafer batch reordering suggestions are generated to avoid equipment starvation, and a scheduling optimization model is established. in, For equipment During free time, Let λ represent yield loss and λ be a weighting coefficient. By digitally twinning process chain changes, such as the insertion of additional cleaning steps, the Pareto front solution set is dynamically balanced to account for yield loss and capacity loss. Simultaneously, based on PINN-predicted wafer deformation trends, the thin film deposition temperature-power window is dynamically adjusted. In the 3D device fin structure process, film thickness uniformity is improved from ±3% using traditional methods to ±1.2%, overcoming the process robustness bottleneck constrained by physical limits, and increasing FabO scheduling efficiency by 40%.
[0063] Step S4: Human-machine collaboration and model evolution.
[0064] The human-machine collaboration and model evolution mechanism constructs a dynamic optimization system for intelligent semiconductor manufacturing through three core modules: natural language interaction, federated learning knowledge evolution, and closed-loop feedback optimization. The natural language interaction interface, utilizing voice / text input, enables rapid response to process challenges and visualization of decision tree paths, reducing debugging difficulty. Federated learning drives knowledge evolution, aggregating process optimization experience from multiple wafer fabs while ensuring data privacy, updating the defect analysis library and process optimization solution set, and promoting technology sharing. Closed-loop feedback optimization transforms validated process parameters into rule engine knowledge and uses knowledge distillation to iteratively reason the model, forming a complete chain for continuous optimization.
[0065] D1 Natural Language Interaction Interface Leveraging the natural language processing capabilities of a large language model, engineers can input lithographic linewidth roughness (LER) fluctuations and dosage deviations via voice or text input. The system utilizes a decision tree algorithm to construct inference paths and employs sensitivity weight analysis formulas. The impact of process parameters on LER is quantified, a dose compensation scheme is dynamically generated, and a visualized causal heatmap is formed by combining key points of subsequent validation. For parameter importance coefficients, This represents process variables such as photoresist concentration and exposure time.
[0066] D2 Federated Learning Drives Knowledge Evolution While complying with data privacy regulations such as GDPR, a horizontal federated learning architecture is adopted to extract features and aggregate model parameters from process optimization cases across multiple wafer fabs. Through continuous updates to the defect root cause pattern library, a Pareto front solution set for process parameters and defect rates is constructed. Let... For the Pareto solution set, where For process parameter vectors, To address the corresponding defect rate metrics, cross-node technology migration and global optimization are achieved.
[0067] D3 Closed-loop Feedback Optimization A smart closed-loop system of "verification-retention-iteration" is constructed. When process verification confirms the effectiveness of parameter combinations (such as photoresist baking temperature gradient compensation rules), it is automatically transformed into decision knowledge for the process rule engine. For new failure cases, the DistilBERT architecture is used for knowledge distillation, through a loss function... Compressed model parameters, where For cross-entropy loss, This mechanism reduces the false alarm rate to <0.03%, meeting the ASIL-D functional safety level requirements for automotive-grade chips in the ISO26262 standard.
[0068] The other parts of this embodiment are the same as those in Embodiment 1 above, so they will not be described again.
[0069] Example 3: This embodiment is based on any one of Embodiments 1-2 above, such as Figure 1 As shown, a specific embodiment will be described in detail.
[0070] Step S1: Multi-source data fusion and knowledge graph enhancement.
[0071] This embodiment uses a neural symbolic architecture to construct a process knowledge graph and employs a specific algorithm to solve defect data problems and predict equipment anomalies. The following will focus on presenting the specific process and algorithm steps of the technical implementation.
[0072] A1 Data Acquisition; It can access real-time data streams of thousands of dimensions, such as photoresist spin coating temperature, etching chamber RF power, wafer CD (critical dimension) measurement, and thin film elliptic polarization spectrum, and integrate historical process documents (such as photoresist Dill model parameters and thin film deposition reaction mechanism) and failure knowledge base to build a process knowledge graph.
[0073] The TransE algorithm maps the correlation between etch rate anomalies and wafer edge chipping defects into weighted edges in the knowledge graph, improving inference accuracy. The core calculation formula of the TransE algorithm during knowledge graph construction is as follows: in, Let r represent the embedding vector of the head entity, r represent the embedding vector of the relation, and t represent the embedding vector of the tail entity. The norm is represented. Using this algorithm, the correlation between etching rate anomalies and wafer edge chipping defects is mapped to weighted edges in the graph, thus completing the construction of a process knowledge graph.
[0074] A2 small sample enhancement; To address the scarcity of defect data, a Generative Adversarial Network (GAN) is used to simulate photoresist imaging defects under different dose deviations, such as line shortening and bridging defects. The self-supervised learning SimCLR framework is used to enhance the feature representation of SEM images.
[0075] A GAN consists of a generator G and a discriminator D, which optimize the objective function through adversarial training. in, It is the actual data distribution. This represents the noise distribution. Simultaneously, the SEM image feature representation is enhanced through the self-supervised learning SimCLR framework, with the contrastive learning loss function being: in, Represents the cosine similarity of vectors. Here, N is the temperature parameter, and N is half the batch size for a single iteration.
[0076] A3 multimodal fusion; A graph attention network (GAT) is constructed to integrate the spatial distribution thermal map of wafer defects with process timing data, predicting systematic defect chains caused by equipment malfunctions 24 hours in advance, such as uneven CMP pressure → edge chipping → thin film step coverage failure.
[0077] In GAT, compute nodes For nodes Attention coefficient The formula is: in, is the weight vector of the attention mechanism, and W is the learnable weight matrix. Let i and j be the feature vectors of nodes i and j, respectively. This represents vector concatenation, and LeakyReLU is the activation function. By processing wafer defect spatial distribution heatmaps and process timing data using GAT, we can analyze and predict systematic defect chains caused by equipment anomalies.
[0078] Calculate node feature weights using a multi-head attention mechanism: The process of calculating node feature weights after k iterations is expressed as follows: Step S2: Innovative design of industrial large-scale model architecture.
[0079] B1 base layer design; The base layer employs a pre-trained large model designed for vertical fine-tuning in semiconductor manufacturing, integrating physical constraints such as quantum effect equations and thin film growth kinetics equations, as well as SEMI standard process rules, such as photoresist baking temperature windows. This model constructs a unified knowledge base encompassing "process-physics-rules." It employs Domain-Adaptive Fine-Tuning technology to optimize the model by minimizing the following loss function: in, Ensure that the model output conforms to the laws of physics. Ensure that process parameters meet industry standards. Optimize for specific manufacturing tasks.
[0080] This model provides a solid knowledge foundation for subsequent process optimization by learning from and fine-tuning massive amounts of data in the semiconductor manufacturing field.
[0081] B2 layered enhancement module construction; (1) Generation layer The generation layer utilizes a diffusion model and a generative adversarial network (GAN) to achieve dynamic simulation of process parameters and exploration of extreme operating conditions.
[0082] The diffusion model generates a prediction set of the impact of lithography dose compensation curves on wafer CD uniformity and etching compatibility by progressively adding and removing noise through a Markov chain. The specific steps are as follows: a. Add Gaussian noise to the real data; b. Train the denoising network to predict noise; c generates prediction data through a reverse diffusion process.
[0083] Application of the diffusion model: Taking photolithography as an example, the diffusion model is used to dynamically simulate the impact of the photolithography dose compensation curve on wafer CD uniformity and subsequent etching compatibility. The specific steps are as follows: input the photolithography dose parameter range, set the dose to vary in a ±5% gradient, perform simulation calculations based on the diffusion model algorithm, output LER evolution data, and gradually generate a LER evolution prediction set to provide data support for optimizing photolithography process parameters.
[0084] GAN (Generative Adversarial Network) is used to generate experimental data on perturbations of thin film deposition parameters, covering extreme working conditions and improving the probability of defect detection.
[0085] GAN Applications: For thin film deposition processes, GAN-based experiments are used to generate perturbation experiments of thin film deposition parameters, such as variations in the flow ratio of PECVD precursors. Through adversarial training, extreme operating conditions that traditional trial-and-error methods cannot reach are covered, such as abrupt stress changes in thin films in low-pressure plasma regions. This process generates experimental data that meets the requirements through adversarial training between the generator and the discriminator. Experimental verification shows that this method increases the probability of defect detection by 4 times.
[0086] (2) Reasoning layer The inference layer uses Causal Bayesian Network and Graph Neural Network (GNN) to achieve parameter optimization and cross-process scheduling optimization.
[0087] Counterfactual inference based on causal Bayesian networks, taking an abnormal etching rate as an example, when an abnormal etching rate is detected, the large model uses the causal Bayesian network to back-generate a causal path graph of "etching gas ratio - plasma density - wafer temperature field - gate damage". The specific formula is: The causal probability relationships between factors are calculated using Bayes' theorem. An objective function is then defined. (Target value S>20:1) COMSOL multiphysics simulation was used to compare the etching selectivity improvement effect of the original parameters and the optimized ratio (such as Cl2 / HBr flow ratio +12%), so as to optimize the etching process parameters.
[0088] By utilizing graph neural networks (GNNs) to analyze cross-process timing constraints, a transmission relationship model was established between photoresist development time deviation and CMP slurry chemical reaction efficiency. Through graph structure analysis of process relationships, wafer batch reordering suggestions were generated, effectively reducing starvation waiting losses and improving capacity utilization by 15%.
[0089] Causal Bayesian networks analyze the effects of parameter interventions through counterfactual reasoning. The specific process is as follows: 1. Construct a causal path diagram of "etching gas ratio - plasma density - wafer temperature field - gate damage"; 2. Use COMSOL simulation to compare the etch selectivity before and after parameter optimization (target > 20:1).
[0090] 3GNN analyzes cross-process timing constraints and improves capacity utilization by optimizing wafer batch sequencing.
[0091] (3) Analysis layer The analysis layer uses Physically Constrained Neural Network (PINN) and Temporal Convolutional Network (TCN) to achieve real-time optimization of process parameters and timing window sensitivity analysis.
[0092] PINN embeds the thin film deposition physics equations into a neural network to optimize process parameters by minimizing the following loss function: in, The proportionality coefficients are determined empirically, data fitting ensures that the model output matches the actual data, and physical constraints guarantee compliance with physical laws.
[0093] In the actual implementation of the analysis layer, the thin film deposition continuity equation is based on the Physically Constrained Neural Network (PINN). With the law of conservation of energy Embedded in PINN and combined with real-time wafer warpage data, an objective function is constructed with the PECVD deposition power ramp curve as the optimization variable. By optimizing this function, the cracking threshold caused by thermal stress can be avoided while satisfying the film thickness uniformity target.
[0094] A temporal convolutional network (TCN) was used to predict the sensitivity of photoresist film thickness to the timing window of the photolithography-etching process, and a mapping relationship between photoresist film thickness h and developer temperature compensation curve Tcomp(t) was established. By dynamically adjusting the developer temperature compensation curve, CD uniformity Cpk ≥ 1.5 was ensured, which is 0.3 higher than that of the traditional fixed formulation.
[0095] Step S3: Closed-loop process optimization execution system; The closed-loop process optimization execution system is designed from three dimensions: dynamic parameter tuning, intelligent operational decision-making, and collaborative process reconfiguration, as detailed below: C1 parameter dynamic tuning; Real-time operating condition input. A dynamic parameter tuning framework based on a large industrial model is constructed to optimize the distortion coefficients of the lithography machine lens. wafer edge alignment error Wait for the current state parameters Input the large model. The large model generates candidate dose compensation schemes based on the training data. Such as the spatial gradient of +3% dose in the central region and -2% at the edge, and the correction vector for the etching gas ratio. Such as fine-tuning of the CF4 / O2 flow ratio.
[0096] Solution evaluation and decision-making. The Proximal Policy Optimization (PPO) algorithm is used for parameter optimization, and a reward function is constructed using a physical simulation engine (SPICE circuit parasitic parameter evaluation + COMSOL etching topography prediction). ,in The parameter combination to be optimized is as follows: In the formula, To increase the driving current, Physical feasibility constraints here to reduce leakage current. Set as This indicates that the membrane stress S should be less than the critical value. For membrane stress, This is the critical stress value. These are the weighting coefficients. An iterative optimization strategy is employed. Choose to maximize expected reward Maximize the global optimal parameter combination and deploy it to the equipment control system to quantify the electrical performance benefits (increased drive current / reduced leakage current) and physical feasibility (e.g., membrane stress < critical value) of each scheme.
[0097] C2 operation intelligent decision-making; Taking real-time anomaly response as an example, when the defect rate on the wafer surface is detected... Exceeding the threshold At this time, the large model inference mechanism is triggered. Causal analysis models are used to locate the root cause of the anomaly and construct fault diagnosis equations: in, The cause of the failure is (e.g., aging of the CMP polishing pad, or deviation in the thickness of the photoresist film). For observed anomalies (such as edge defects, LER exceeding the threshold). In order to observe anomalies Cause of the fault The posterior probability.
[0098] Based on the diagnostic results, a multi-dimensional set of operation instructions is generated. Including equipment calibration parameter files (e.g., CMP pressure distribution map correction), process restart strategy (Developer batch switching verification), priority reminder for manual intervention (e.g., grinding pad replacement > gas purity check), using a priority sorting algorithm. Complete instruction scheduling, where p is the priority weight vector.
[0099] In summary, the surge in wafer surface defects triggers a large-scale model inference mechanism, which quickly locates the root causes such as CMP polishing pad aging (uneven pressure → edge defects) or photoresist film thickness deviation (LER exceeding the threshold), and generates multi-dimensional operation instructions: equipment calibration parameter files (such as CMP pressure distribution map correction), process restart strategies (developer batch switching verification), and manual intervention priority prompts (such as polishing pad replacement > gas purity verification), thereby improving the anomaly response speed.
[0100] C3 Process Collaborative Restructuring Bottleneck prediction and scheduling. A process chain model is constructed using digital twin technology, and a graph network is used. This represents each process step and its dependencies, where node V is a process step and edge E is a material flow or process constraint. When external disturbances such as photoresist supply delays are detected, a topology sorting algorithm is used. Analyzing the cascading effects, a mixed-integer programming (MIP) model is used to solve for the wafer batch reordering scheme, aiming to minimize the production efficiency loss and cost increase caused by the disturbance: In the formula, The switching cost from process step i to j, For decision variables, a value of 1 indicates a process step. Post-execution steps If the value is 0, then it will not be executed; This is a weighting coefficient used to balance process changeover costs and delay costs; Indicates the first This model addresses the delay time caused by external disturbances in each process step. By incorporating a delay cost term while considering process changeover costs, the wafer batch reordering scheme better aligns with actual production needs and effectively addresses production plan changes caused by external disturbances.
[0101] From the implementation perspective, the above process can be summarized as follows: analyzing the cascading impact of photoresist supply delay on the etching / thin film deposition queue, generating wafer batch reordering suggestions to avoid equipment starvation; and using digital twins to simulate process chain changes, such as inserting additional cleaning steps, to dynamically balance yield loss and capacity loss with the Pareto front solution set.
[0102] Flexible process window expansion. A Physical Information Neural Network (PINN) is introduced to predict wafer deformation trends, and the model parameters are optimized by minimizing the loss function L. in, For data fitting loss, The physical constraint loss (e.g., heat conduction equation, stress balance equation) is represented by λ, and the weighting coefficient is λ. Based on the PINN prediction results, the film deposition temperature-power window is dynamically adjusted, and film thickness uniformity is optimized through an adaptive control algorithm.
[0103] Step S4: Human-machine collaboration and model evolution mechanism; The human-computer collaboration and model evolution mechanism is mainly achieved through three technical solutions: natural language interaction interface, federated learning-driven knowledge evolution, and closed-loop feedback optimization.
[0104] D1 Natural Language Interaction Interface Engineers input process challenges via voice / text, such as "How to improve the lithography yield of metal layers?", and a large model instantly generates decision tree paths (LER sensitivity weight analysis → dose compensation scheme → key points for post-process validation) and visualized causal heatmaps, significantly reducing the threshold for process debugging. The specific process is as follows: A natural language processing model based on the Transformer architecture is constructed, and semantic association weights between words in the input text (e.g., "How to improve the yield of metal layer lithography?") are calculated using a self-attention mechanism. LSTM networks are used to extract sequence features from the text representing process challenges, which are then input into a pre-trained decision tree generation model. The model employs the information gain criterion. in, Let A be the information gain of attribute A on sample set D. For dataset Information entropy For attributes No. The number of samples with each value is determined. Finally, a decision tree path is generated, and the feature importance of each decision node is calculated based on the SHAP (SHAP Additive Explanations) value, resulting in a visualized causal heatmap.
[0105] During the process parameter optimization stage, the following steps are used to generate decision tree paths and visualize feature importance: Decision tree path generation: Preprocessed semiconductor process parameter data (such as photoresist thickness, etching rate, temperature profile, etc.) are input into the decision tree model for training, and a multi-branch tree structure is constructed using the CART algorithm. After removing redundant branches through pruning strategies, 3-5 optimal decision paths are selected based on the current production conditions (such as equipment load, raw material batches). Each path corresponds to a specific combination of process parameters and the expected yield improvement.
[0106] SHAP value calculation: The SHAP (S Hapley Additive Ex Planations) framework is used to quantify the feature contribution of each decision node. Its core principle is to calculate the marginal contribution of each feature to the final decision based on game theory. For example, for the node "etching time", the SHAP value can accurately reflect its weight in the yield under different process scenarios. The calculation process supports local interpretation (single decision path) and global interpretation (the entire decision tree).
[0107] Visualized causal heatmap creation: SHAP values are mapped to color intensities in the heatmap. The horizontal axis represents decision node features, such as deposition temperature and exposure dose, while the vertical axis represents process stages, such as lithography, etching, and packaging. The heatmap visually presents the impact of different features on each process stage. For example, red areas indicate a strong positive correlation between the feature and yield improvement, while blue areas indicate potential negative impacts, providing process engineers with interpretable optimization criteria.
[0108] D2 Federated Learning Drives Knowledge Evolution Under the premise of protecting privacy (homogeneous encrypted transmission), this platform aggregates multi-wafer fab process optimization cases, continuously updates the defect root cause pattern library and process Pareto solution set, and enables cross-node technology migration. The specific technical solution is as follows: A horizontal federated learning framework is adopted, and feature alignment is performed on multi-wafer fab process optimization cases under homomorphic encryption protection. The DBSCAN clustering algorithm is used for anomaly detection and feature extraction of process parameter data. Valid case data are then aggregated using a federated averaging algorithm (FedAvg) to aggregate model parameters. in, Here are the updated parameters for the global model, and N is the number of nodes participating in the training. Let be the model parameters for the i-th node in round t. The updated defect root cause pattern library and process Pareto solution set are applied to different process optimizations through transfer learning.
[0109] D3 Closed-loop Feedback Optimization Valid parameter combinations are automatically precipitated as knowledge for the process rule engine; new failure cases are rapidly iterated through knowledge distillation to improve the inference model, reducing false alarm rates and meeting the ASIL-D functional safety requirements of automotive-grade chips. The specific technical solution is as follows: A bidirectional interaction mechanism is established between the process rule engine and the inference model. When valid parameter combinations, such as photoresist baking temperature gradient compensation rules, are generated, the TransE knowledge graph embedding algorithm is used to encode them as structured knowledge storage. For new failure cases, knowledge distillation technology is used to transfer the knowledge of the complex teacher model BERT to the lightweight student model DistilBERT. The distillation loss function is defined as: in, For cross-entropy loss, For knowledge distillation loss, This is a balancing factor. Through continuous iterative optimization, the false alarm rate is reduced to meet the ASIL-D functional safety requirements of automotive-grade chips.
[0110] The other parts of this embodiment are the same as any one of the above embodiments 1-2, so they will not be described again.
[0111] Example 4: Based on any one of Embodiments 1-3 above, this embodiment takes the optimization of photolithography process parameters as an example to explain in detail the implementation steps of this method.
[0112] First, data preparation is carried out by collecting historical lithography data (exposure dose, development time, wafer CD uniformity, defect rate) and current equipment status, including photoresist spin coating temperature, lithography machine lens distortion coefficient and other thousand-dimensional time series data, while integrating relevant historical process documents and failure knowledge base information.
[0113] Next, large-scale model inference is performed. Real-time parameters and wafer scan images are input into the large model to generate a correlation diagram of potential defect risk areas and key parameters, such as dose-line edge roughness (LER) sensitivity weights. Then, the photoresist imaging quality and subsequent etching compatibility under different dose compensation schemes (±5% gradient) are simulated, and an optimization candidate set is output (e.g., recommended dose +3% and fine-tuning developer temperature). During the inference process, a causal Bayesian network is used to analyze the causal relationship between parameter changes and defects, and a graph neural network is used to analyze the impact of time constraints across processes.
[0114] Finally, closed-loop optimization is performed, pushing the optimal parameters to the lithography machine control system for execution. Wafer CD measurement data and subsequent electrical test results are monitored online to evaluate the effectiveness of the solution. If the yield improvement meets the target, such as a 20% reduction in LER fluctuation, the parameter combination is automatically incorporated into the process rule engine knowledge; otherwise, the causal chain is backtracked to correct the model and a new round of optimization is initiated. Simultaneously, through federated learning, this optimization case is aggregated with cases from other wafer fabs to update the model knowledge base, enabling continuous model evolution.
[0115] Other process scenarios, such as the synergistic optimization of etching rate and selectivity, and the control of thin film deposition thickness uniformity, can be implemented by analogy. By dynamically adapting physical constraints and multi-objective trade-offs to a large model, the entire semiconductor manufacturing process can be comprehensively optimized.
[0116] The other parts of this embodiment are the same as any one of the embodiments 1-3 above, so they will not be described again.
[0117] Example 5: Based on any one of Embodiments 1-4 above, this embodiment proposes an intelligent manufacturing process optimization system integrating an industrial large model, used to execute the above-mentioned intelligent manufacturing process optimization method integrating an industrial large model; including a multi-source data fusion and knowledge graph enhancement module, an industrial large model architecture design module, a closed-loop process optimization execution module, and a human-machine collaboration and model evolution module; The multi-source data fusion and knowledge graph enhancement module is used to call up multi-source time-series data obtained by neural symbol structure fusion. The industrial large model architecture design module is used to input the acquired multi-source time-series data into the constructed industrial large model architecture to generate potential defect risk areas and key parameter correlation diagrams; the industrial large model architecture includes a basic layer, a generation layer, an inference layer, and an analysis layer. The closed-loop process optimization execution module is used to generate candidate dose compensation schemes and output an optimization candidate set based on the generated potential defect risk areas and key parameter correlation diagrams. The human-machine collaboration and model evolution module is used for closed-loop optimization. It pushes the optimization candidate set to the lithography machine control system for execution and evaluates the effectiveness of the candidate dose compensation scheme. If the yield improvement meets the target, the parameter combination is automatically precipitated into the process rule engine knowledge. Otherwise, it backtracks the causal chain to correct the model and starts a new round of optimization. At the same time, it aggregates the optimization cases with other wafer cases through federated learning, updates the model knowledge base, and realizes the continuous evolution of the model.
[0118] The other parts of this embodiment are the same as any one of the embodiments 1-4 above, so they will not be described again.
[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications or equivalent changes made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for optimizing intelligent manufacturing processes using an integrated large-scale industrial model, characterized in that, Specifically, the following steps are included: Step S1: Call the multi-source temporal data obtained by neural symbolic structure fusion; Step S2: Input the acquired multi-source time-series data into the constructed industrial big data model architecture to generate a correlation diagram of potential defect risk areas and key parameters; the industrial big data model architecture includes a basic layer, a generation layer, an inference layer, and an analysis layer; Step S3: Based on the generated potential defect risk areas and key parameter correlation diagram, generate candidate dose compensation schemes and output the optimization candidate set; Step S4: Perform closed-loop optimization, push the optimization candidate set to the lithography machine control system for execution, and evaluate the effectiveness of the candidate dose compensation scheme; If the yield improvement meets the target, the parameter combination will be automatically converted into process rule engine knowledge; otherwise, the causal chain will be traced back to correct the model and a new round of optimization will be launched. At the same time, through federated learning, the optimization cases will be aggregated with other wafer cases to update the model knowledge base and achieve continuous evolution of the model.
2. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S11: Acquire multi-source time-series data streams and integrate them with historical process documents and failure knowledge bases; Step S12: Call the TransE algorithm to map entities and relations to a low-dimensional vector space and calculate the distance between entity and relation vectors; Step S13: Call the generator and discriminator to construct the objective function and simulate photoresist imaging defects under different dose deviations; Step S14: Call the self-supervised learning SimCLR framework to augment the random data of the same image to generate multiple views, and calculate the contrast loss between the views; Step S15: Construct a graph attention network of wafer defect spatial distribution heatmap and process timing data, call the multi-head attention mechanism, and calculate the node feature weights; Step S16: Based on the graph attention network, predict the systematic defect chain of the acquired multi-source time-series data stream.
3. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 1, characterized in that, Step S2 specifically includes the following steps: Step S21: Call the pre-trained large model for vertical fine-tuning in semiconductor manufacturing, and use the quantum effect equation and the thin film growth kinetics equation as physical constraints to construct the base layer; Step S22: Invoke the diffusion model and generate an adversarial network to simulate the impact of changes in process parameters, covering extreme operating conditions to discover potential defects; Step S23: Based on the counterfactual inference parameters of the causal Bayesian network, the intervention effect is analyzed, and the graph neural network is called to parse the cross-process time constraints, construct the inference layer, and generate optimization and reordering suggestions; Step S24: Embed the physical constraint neural network into physical laws, combine real-time data to optimize the process curve, and predict process sensitivity and dynamically adjust parameters through a temporal convolutional network.
4. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 1, characterized in that, Step S3 specifically includes the following steps: Step S31: Input the current state parameters of semiconductor manufacturing into the large model, output candidate process schemes, and evaluate the performance benefits and physical feasibility of the schemes based on reinforcement learning combined with the physical simulation engine, and select the optimal parameter combination to deploy to the equipment control system. Step S32: Based on the constructed intelligent decision-making system, when an anomaly occurs in the wafer, the large model inference is triggered to locate the root cause and generate multi-dimensional operation instructions. Step S33: Analyze the cascading effects of changes in process steps, generate scheduling suggestions, and dynamically adjust process parameters based on digital twin simulation and physical model prediction.
5. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 3, characterized in that, Step S31 specifically includes the following steps: Step S311: Construct a parameter optimization framework based on an industrial large model, taking process state parameters as input, to obtain candidate dose compensation schemes; Step S312: Invoke the near-end policy optimization algorithm in conjunction with the physics simulation engine to construct the reward function; Step S313: Iteratively optimize and select the globally optimal parameter combination and deploy it to the equipment control system to improve yield.
6. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 3, characterized in that, Step S32 specifically includes the following steps: Step S321: Establish a large model inference mechanism under abnormal operating conditions. When the defect density on the wafer surface exceeds the threshold, the fault diagnosis process is triggered. Step S322: Call the large model to analyze multi-dimensional parameters, and call the fault tree reasoning algorithm to locate the root cause, and generate an operation instruction set including equipment calibration parameters, process restart strategy, and manual intervention priority; Step S323: When the defects on the wafer surface increase dramatically, the large model inference mechanism is triggered to quickly locate the root cause and generate multi-dimensional operation instructions.
7. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 3, characterized in that, Step S33 specifically includes the following steps: Step S331: Based on the obtained equipment space time, yield loss, and set weight coefficients, establish a scheduling optimization model; Step S332: Based on the scheduling optimization model, dynamically balance the Pareto front solution set of yield loss and capacity loss; Step S333: Based on the PINN prediction of wafer deformation trend, dynamically adjust the thin film deposition temperature-power window.
8. The intelligent manufacturing process optimization method based on an integrated industrial large-scale model according to claim 1, characterized in that, Step S4 specifically includes the following steps: Step S41: Call the decision tree algorithm to construct the reasoning path, quantify the impact of process parameters on the input process challenge through the sensitivity weight analysis equation, dynamically generate the dosage compensation scheme, and form a visual causal heat map; Step S42: Call the horizontal federated learning architecture to extract features and aggregate model parameters from the process optimization cases of multiple wafer fabs, and construct the Pareto front solution set of process parameters and defect rate; Step S43: If the process verification confirms that the parameter combination is effective, the parameter combination is transformed into decision knowledge of the process rule engine; if the process verification is a new failure case, the DistilBERT architecture is called to perform knowledge distillation, and the loss function is called to compress the model parameters.
9. A smart manufacturing process optimization system integrating a large industrial model, used to execute the smart manufacturing process optimization method integrating a large industrial model as described in claim 1; characterized in that, It includes modules for multi-source data fusion and knowledge graph enhancement, industrial large model architecture design, closed-loop process optimization and execution, and human-machine collaboration and model evolution. The multi-source data fusion and knowledge graph enhancement module is used to call up multi-source time-series data obtained by neural symbol structure fusion. The industrial large model architecture design module is used to input the acquired multi-source time-series data into the constructed industrial large model architecture to generate potential defect risk areas and key parameter correlation diagrams; the industrial large model architecture includes a basic layer, a generation layer, an inference layer, and an analysis layer. The closed-loop process optimization execution module is used to generate candidate dose compensation schemes and output an optimization candidate set based on the generated potential defect risk areas and key parameter correlation diagrams. The human-machine collaboration and model evolution module is used to perform closed-loop optimization, push the optimization candidate set to the lithography machine control system for execution, and evaluate the effectiveness of the candidate dose compensation scheme. If the yield improvement meets the target, the parameter combination will be automatically converted into process rule engine knowledge; otherwise, the causal chain will be traced back to correct the model and a new round of optimization will be launched. At the same time, through federated learning, the optimization cases will be aggregated with other wafer cases to update the model knowledge base and achieve continuous evolution of the model.
Citation Information
Cited By
Industrial decision support system and method based on large model and artificial intelligence
CN121351017A
Self-adaptive optimization and collaborative decision-making system for spin-coating process of large-diameter substrate
CN121613736A
Processing technology optimization method for semiconductor lining
CN121806752A
Layout measurement and analysis system capable of integrating CD-SEM and GDSII / OASIS by MetroAnalyzer software
CN122133486A