Method for optimizing a manufacturing process of a semiconductor device and semiconductor device

CN122797243APending Publication Date: 2026-09-22NEXCHIP SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611249888.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

优化基于理想设计版图,无法精准反映光刻、刻蚀等工艺后的实际三维形貌,忽略边缘粗糙度、侧壁倾角及关键尺寸波动,导致优化结果与实际效果偏差显著

Benefits of technology

[0035]本申请意想不到的技术效果是:基于当前关键生产工艺步骤后半导体器件的多模态检测数据,生成融合数据向量,将融合数据向量输入三维结构重建神经网络模型,由三维结构重建神经网络模型输出半导体器件的三维结构体素模型,对三维结构体素模型进行跨尺度多物理场性能仿真,得到半导体器件的电学性能预测数据,将电学性能预测数据、当前工艺配方参数及三维结构体素模型的几何特征向量输入决策神经网络模型,由决策神经网络模型基于预设的多个优化目标进行分析,输出工艺参数调整量,这样,能够显著提升半导体器件实际三维结构重建的精度与保真度,构建高可信度的器件数字孪生体,确保结构表征完整反映制造细节;对三维结构体素模型进行跨尺度多物理场性能仿真,能增强电学性能仿真预测与实际制造结果的一致性,提高性能评估的可靠性与指导价值,进而提高了工艺参数调整量的准确性,优化了制造工艺。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797243A_ABST
    Figure CN122797243A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of semiconductor device manufacturing process optimization method and semiconductor device, the method comprises: based on the multimodal detection data of semiconductor device after current key production process step, fusion data vector is generated;Fusion data vector is input three-dimensional structure reconstruction neural network model, and three-dimensional structure voxel model of semiconductor device is output by three-dimensional structure reconstruction neural network model;Cross-scale multi-physics field performance simulation is carried out to three-dimensional structure voxel model, and the electrical performance prediction data of semiconductor device is obtained;Electrical performance prediction data, current process recipe parameters and the geometric feature vector of three-dimensional structure voxel model are input decision neural network model, and decision neural network model is analyzed based on pre-set multiple optimization targets, and process parameter adjustment amount is output.The embodiment of the application realizes the leap from single-point process optimization to whole manufacturing process collaborative optimization by constructing the optimization closed loop with actual three-dimensional structure as center.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semiconductor manufacturing technology, and in particular to a method for optimizing the manufacturing process of a semiconductor device and a semiconductor device. Background Technology

[0002] As semiconductor technology nodes shrink to the nanoscale, the three-dimensional structures such as FinFET fins, GAA nanosheets, and 3D NAND stacks become increasingly complex, and their geometry has a crucial impact on device performance. However, traditional manufacturing process optimization relies on iterative cycles of "design-manufacturing-measurement-analysis" and empirical models based on design rules, facing significant challenges. First, there is a gap between the "ideal" and the "real." Optimization based on ideal design layouts cannot accurately reflect the actual three-dimensional morphology after processes such as photolithography and etching, ignoring edge roughness, sidewall tilt angles, and critical dimension fluctuations, leading to significant deviations between optimization results and actual effects. Second, the limitations of islanded optimization for single processes are obvious. Existing methods only target local tuning of individual steps, lacking global synergy between the interaction of preceding and following processes and the final electrical performance, making it difficult to achieve overall optimization and even causing subsequent problems. Third, feedback cycles are long and costs are high. Traditional verification relies on physical experiments and offline metrology, requiring a large number of wafer prototypes and destructive analysis, resulting in long iteration cycles and high costs, making it difficult to adapt to rapid market changes. In addition, there is the dilemma of multi-objective optimization. When pursuing performance, power consumption, reliability, and yield simultaneously, traditional methods often struggle to find the globally optimal window in the multidimensional parameter space, forcing compromises. Summary of the Invention

[0003] In view of the above problems, the purpose of this application is to provide a semiconductor device manufacturing process optimization method and semiconductor device, which aims to achieve a leap from single-point process optimization to full-process collaborative optimization by constructing an optimization closed loop centered on the actual three-dimensional structure, significantly shortening the process development cycle, reducing R&D costs, and improving the overall performance of the device and the production line yield.

[0004] According to a first aspect of the embodiments of this application, a method for optimizing the manufacturing process of a semiconductor device is provided, comprising:

[0005] Based on the multimodal detection data of semiconductor devices after the current key manufacturing process steps, a fused data vector is generated.

[0006] The fused data vector is input into a three-dimensional structure reconstruction neural network model, and the three-dimensional structure voxel model of the semiconductor device is output by the three-dimensional structure reconstruction neural network model.

[0007] Multi-scale multiphysics performance simulations were performed on the three-dimensional structure voxel model to obtain the predicted electrical performance data of the semiconductor device.

[0008] The electrical performance prediction data, current process formula parameters, and geometric feature vectors of the three-dimensional structure voxel model are input into the decision neural network model. The decision neural network model analyzes the data based on multiple preset optimization objectives and outputs the process parameter adjustment amount.

[0009] Optionally, the current process formula parameters include the process formula parameters of the current and one or more upstream key production process steps, and the process parameter adjustment amount includes the process parameter adjustment amount of the current and one or more upstream key production process steps.

[0010] Optionally, the manufacturing process optimization method further includes:

[0011] The current process formula parameters are optimized based on the process parameter adjustment amount to obtain the optimized process formula parameters;

[0012] The optimized process formulation parameters will be applied to the production of the next batch of semiconductor devices.

[0013] Optionally, the manufacturing process optimization method further includes:

[0014] Collect multimodal testing data and measured electrical performance data of semiconductor devices after the current key manufacturing process steps;

[0015] The three-dimensional structure reconstruction neural network model and / or the decision neural network model are optimized and updated using multimodal detection data and electrical performance measurement data after the current key production process steps.

[0016] Optionally, the three-dimensional structure reconstruction neural network model is optimized and updated using multimodal detection data and measured electrical performance data after the current key production process steps, including:

[0017] Construct a composite loss function, which includes a weighted sum of voxel-level loss, surface loss, and physical loss;

[0018] The gradient is calculated based on the composite loss function, and the parameters of the three-dimensional structure reconstruction neural network model are iteratively updated along the opposite direction of the gradient using the backpropagation algorithm to minimize the composite loss function.

[0019] The voxel-level loss uses binary cross-entropy to measure the difference between the predicted probability and the true value of each voxel; the surface loss uses chamfer distance to measure the geometric deviation between the predicted surface point set and the true surface point set; and the physical loss is used to ensure that the 3D structure reconstruction result conforms to physical laws.

[0020] Optionally, calculating the physical loss includes:

[0021] Calculate the weighted sum of the surface smoothness constraint term and the penalty term for process physical inconsistency;

[0022] The surface smoothness constraint is measured by calculating the integral of the square of the surface Laplacian operator of the three-dimensional structure voxel model, and the process physical inconsistency penalty is determined by defining different physical constraints according to the specific process type.

[0023] Optionally, the decision neural network model is optimized and updated using multimodal detection data and measured electrical performance data after the current key production process steps, including:

[0024] Construct a reinforcement learning reward function, which includes an evaluation of the optimization objective and an evaluation of the constraints. The evaluation of the optimization objective includes a weighted sum of the values ​​of each optimization objective, and the evaluation of the constraints includes the sum of the product of the violation amount and the penalty coefficient of each violated constraint.

[0025] The expected cumulative reward is calculated based on the reinforcement learning reward function, and the gradient of the expected cumulative reward with respect to the parameters of the decision neural network model is calculated using the policy gradient method.

[0026] The parameters of the decision neural network model are updated along the gradient direction using the learning rate to obtain the updated decision neural network model.

[0027] Optionally, generating a fused data vector based on the multimodal detection data of the semiconductor device after the current key manufacturing process steps includes:

[0028] Align and spatiotemporally register the multimodal inspection data of semiconductor devices from different inspection equipment after the current key production process steps to eliminate coordinate deviations between different inspection equipment;

[0029] Apply the corresponding feature extraction function to the detection data of each modality to convert the raw data into feature vectors;

[0030] The feature vectors of each modality are fused using a weighted concatenation or attention mechanism to form the fused data vector, wherein each modality detection data has a weight coefficient representing its importance.

[0031] Optionally, the step of inputting the electrical performance prediction data, current process formulation parameters, and the geometric feature vector of the three-dimensional structure voxel model into a decision neural network model, and having the decision neural network model analyze based on multiple preset optimization objectives and output process parameter adjustment amounts, includes:

[0032] A state vector is constructed, which is composed of the electrical performance prediction data, the geometric feature vector, and the current process formula parameters.

[0033] The state vector is input into the decision neural network model, which analyzes the data based on multiple preset optimization objectives and derives the adjustment amount of the process parameters in reverse.

[0034] According to a second aspect of the embodiments of this application, a semiconductor device is provided, wherein the semiconductor manufacturing process line of the semiconductor device is optimized according to the manufacturing process optimization method described above.

[0035] The unexpected technical effect of this application is as follows: Based on the multimodal detection data of semiconductor devices after the current key manufacturing process steps, a fused data vector is generated. This fused data vector is input into a three-dimensional structure reconstruction neural network model, which outputs a three-dimensional structure voxel model of the semiconductor device. Cross-scale multiphysics performance simulation is performed on the three-dimensional structure voxel model to obtain electrical performance prediction data of the semiconductor device. The electrical performance prediction data, current process formula parameters, and geometric feature vectors of the three-dimensional structure voxel model are input into a decision neural network model. The decision neural network model analyzes based on multiple preset optimization objectives and outputs process parameter adjustment amounts. In this way, the accuracy and fidelity of the actual three-dimensional structure reconstruction of semiconductor devices can be significantly improved, a highly reliable digital twin of the device can be constructed, and the structural characterization can be made to fully reflect the manufacturing details. Cross-scale multiphysics performance simulation of the three-dimensional structure voxel model can enhance the consistency between electrical performance simulation prediction and actual manufacturing results, improve the reliability and guiding value of performance evaluation, and thus improve the accuracy of process parameter adjustment amounts and optimize the manufacturing process.

[0036] Furthermore, the current process formulation parameters include the process formulation parameters of the current and one or more upstream key production process steps, and the process parameter adjustment amount includes the process parameter adjustment amount of the current and one or more upstream key production process steps. In this embodiment, the current process formulation parameters are optimized according to the process parameter adjustment amount to obtain optimized process formulation parameters, which are then applied to the production of the next batch of semiconductor devices. This breaks through the limitations of traditional isolated optimization of single process steps, achieving global collaborative optimization across process steps. Through an intelligent closed-loop feedback mechanism of "perception-analysis-decision-execution," the process window is continuously iteratively optimized, thereby significantly shortening the process development and debugging cycle, reducing R&D costs, and ultimately improving the overall performance and production yield of semiconductor devices.

[0037] Furthermore, utilizing multimodal detection data and measured electrical performance data from current key production process steps, the 3D structure reconstruction neural network model and / or decision neural network model are optimized and updated. Specifically, optimizing the 3D structure reconstruction neural network model includes constructing a composite loss function, which is a weighted sum of voxel-level loss, surface loss, and physical loss. The gradient is calculated based on the composite loss function, and the parameters of the 3D structure reconstruction neural network model are iteratively updated along the reverse direction of the gradient using a backpropagation algorithm to minimize the composite loss function. The voxel-level loss uses binary cross-entropy to measure the difference between the predicted probability and the true value of each voxel; the surface loss uses chamfer distance to measure the geometric deviation between the predicted surface point set and the true surface point set; and the physical loss is used to ensure that the 3D structure reconstruction result conforms to physical laws. This ensures that the reconstructed 3D model not only matches the observed data but also conforms to basic physical laws. This allows the model to generate reasonable structures even in regions without sufficient training data, greatly improving its generalization ability. At the same time, it can penalize unreasonable structures (such as abnormal spikes) caused by measurement noise, prompting the model to find a compromise between matching data (reducing voxel loss) and maintaining smoothness (reducing physical loss), making the reconstruction results closer to the real physical structure. Attached Figure Description

[0038] The above and other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0039] Figure 1 The diagram shown is a schematic flowchart of an exemplary semiconductor device manufacturing process optimization method according to an embodiment of this application.

[0040] Figure 2 The diagram shown is a schematic representation of an exemplary electronic device according to an embodiment of this application. Detailed Implementation

[0041] The present application will now be described in more detail with reference to the accompanying drawings. In the various drawings, the same elements are indicated by similar reference numerals. For clarity, the various parts in the drawings are not drawn to scale. Furthermore, some well-known parts may not be shown.

[0042] This application may be presented in various forms, some of which will be described below.

[0043] Figure 1 The diagram shown is a schematic flowchart of an exemplary semiconductor device manufacturing process optimization method according to an embodiment of this application. Figure 1 As shown, the manufacturing process optimization methods include:

[0044] In step S110, a fused data vector is generated based on the multimodal detection data of the semiconductor device after the current key production process step.

[0045] In some embodiments, in a semiconductor manufacturing process line, after the completion of the current critical manufacturing process step (e.g., the k-th process step, such as photolithography, etching, or deposition), at least two inspection devices (such as SEM, OCD, AFM, CDSEM, OVL, etc.) are simultaneously activated to acquire multimodal inspection data of the semiconductor device. The multimodal inspection data includes, but is not limited to, at least two of the following: scanning electron microscope (SEM) images, optical critical dimension (OCD) spectra, atomic force microscope (AFM) morphology data, critical dimension scanning electron microscope (CD-SEM) images, and overlay accuracy (OVL) measurement data. The process of generating the fused data vector further includes: first, aligning and spatiotemporally registering the multimodal inspection data of the semiconductor device from different inspection devices after the current critical manufacturing process step to eliminate coordinate deviations between different inspection devices. In some embodiments, the original data of the m-th inspection mode after the k-th critical manufacturing process step, i.e., the modal inspection data, is... Its dimensions are related to the data type (such as images, spectra). For spatial data (such as SEM, AFM), modal detection data needs to be... Aligned with the standard reference coordinate system of the chip design layout, the aligned modal detection data is In one embodiment, the alignment transformation Tm is obtained by solving for the optimal transformation matrix that minimizes the distance function, as shown below:

[0046] Tm * = argmin Dist(Tm( ,Xref))(1)

[0047] Here, the symbol Tm represents the optimal transformation matrix to be found. * It represents all possible geometric transformations, including translations in the X, Y, and Z directions, rotation angles, and possible scaling ratios. This refers to the raw data collected using the m-th type of inspection equipment after the k-th critical manufacturing process step. Xref is the standard reference coordinate system of the chip design layout. Dist is the distance function used to calculate the deviation between the transformed modal detection data and the standard reference coordinate system of the chip design layout. argmin represents finding the transformation that minimizes the value of the distance function Dist among all possible geometric transformations.

[0048] Secondly, a corresponding feature extraction function φm (e.g., convolutional neural networks for images, spectral analysis for spectra) is applied to each modality of detection data to transform the raw data into feature vectors with explicit physical meaning. In some embodiments, the feature vector for each modality of detection data is:

[0049] =φm( ), ∈R dm (2)

[0050] Where φm is the feature extraction function, and its input is the aligned modality detection data. The output is a feature vector. R dm Representing the eigenvector It is a real vector of dimension dm, where the specific value of dm depends on the design of the feature extraction function. For example, if 50 features are extracted, then dm equals 50.

[0051] Finally, a weighted concatenation or attention mechanism is used to fuse the feature vectors from different modalities into a unified, information-enhanced fused data vector, where each modality detection data has a weight coefficient representing its importance. The weights of each modality can be dynamically adjusted according to the quality of the current modality detection data (such as image contrast and spectral signal-to-noise ratio) to ensure the reliability of the fused information. In some embodiments, the fused data vector F (k) for:

[0052] (3)

[0053] Where M represents the total number of modalities participating in the fusion. Here, each modality's detection data is represented by a weighting coefficient indicating its importance; the sum of the weights for all modalities equals 1. It is a projection matrix, used to project feature vectors of different dimensions. Mapped to the same dimensional space, F (k) It is a fused data vector.

[0054] In step S120, the fused data vector is input into the three-dimensional structure reconstruction neural network model, and the three-dimensional structure voxel model of the semiconductor device is output by the three-dimensional structure reconstruction neural network model.

[0055] In some embodiments, the fused data vector F generated in step S110 (k) The input is fed into a pre-trained 3D structure reconstruction neural network model. This model employs an encoder-decoder architecture, where the encoder first processes the fused data vector F... (k)A multi-layer nonlinear transformation is performed to compress the data into a low-dimensional latent vector. This process is similar to "information condensation," extracting the core geometric and morphological features from massive amounts of multimodal detection data. Subsequently, the decoder receives this latent vector and, through layer-by-layer upsampling and feature reconstruction, finally outputs a high-fidelity voxel model V of the actual three-dimensional structure of the device, typically presented as a three-dimensional voxel mesh or surface mesh. For example, in the reconstruction of FinFET fins, the encoder analyzes and fuses the data vector F. (k) The decoder uses a combination of features related to sidewall angles, critical dimensions, and roughness to "draw" a three-dimensional fin with a specific morphology, whose edge fluctuations and dimensional deviations closely match those of the actual manufactured structure. The core innovation of this step lies in its ability to "imagine" a complete three-dimensional structure from limited two-dimensional inspection data (such as SEM images and OCD spectra), thereby avoiding the high cost and low efficiency of traditional destructive analysis (such as FIB-SEM tomography) and providing accurate geometric input for subsequent multi-scale multiphysics simulations.

[0056] In step S130, the three-dimensional structure voxel model is subjected to multi-physics performance simulation across scales to obtain the electrical performance prediction data of the semiconductor device.

[0057] In some embodiments, the high-fidelity three-dimensional structural voxel model V reconstructed in step S120 is used as the input domain for physical simulation to perform cross-scale multiphysics performance simulation, thereby predicting the key electrical performance parameters of the device under the actual morphology. In some cases, simulators of different precisions (such as TCAD and FDTD) can be called as needed to achieve cross-scale simulation from micro to macro. Specifically, the three-dimensional structural voxel model V is first converted into a finite element or finite volume simulation mesh, and the mesh is refined at material boundaries (such as the gate-oxide interface) to ensure calculation accuracy. Subsequently, corresponding material properties (such as the dielectric constant and mobility of silicon and silicon dioxide) are assigned to each mesh node, and boundary conditions for the source, drain, and gate are set (such as source grounding, drain biasing, and gate voltage scanning). Taking electrical simulation as an example, based on this, the governing equations of semiconductor device physics, namely the drift-diffusion equations, are solved, including the Poisson equation (describing the relationship between potential and charge distribution) and the current continuity equation (describing the transport and recombination of charge carriers). By iteratively solving the problem, the current flowing through the device under different gate voltages can be calculated, thereby obtaining key electrical performance parameters such as transfer characteristic curves (e.g., drive current Ion, leakage current Ioff) and threshold voltage.

[0058] In some cases, to accelerate the iterative optimization process, a pre-trained simulation proxy neural network model can be invoked. This model directly learns the mapping relationship from the geometric feature vectors of the three-dimensional structure (such as critical dimensions, sidewall angles, and roughness spectrum) to electrical performance parameters. This reduces the time for a single simulation from minutes to milliseconds while maintaining prediction accuracy, providing real-time feedback for subsequent AI collaborative optimization decisions. In some embodiments, the formula for the simulation proxy neural network model is:

[0059] E=Msurrogate(g(V);θ)(4)

[0060] Here, Msurrogate is a neural network model that defines a functional approximation relationship, θ are the model parameters of Msurrogate, and g(V) is the geometric feature vector of the three-dimensional structure voxel model V. E is the electrical performance prediction data of the semiconductor device obtained from the simulation.

[0061] In step S140, the electrical performance prediction data, the current process formula parameters, and the geometric feature vector of the three-dimensional structure voxel model are input into the decision neural network model. The decision neural network model analyzes the data based on multiple preset optimization objectives and outputs the process parameter adjustment amount.

[0062] In some embodiments, the electrical performance parameters predicted in step S130 (such as drive current Ion and leakage current Ioff), the geometric feature vectors of the three-dimensional structural voxel model reconstructed in step S2 (such as critical dimensions, sidewall angles, and roughness spectrum), and the current process formulation parameters (such as etching power, gas flow rate, and annealing temperature) are concatenated into a multidimensional state vector. The formula for calculating the state vector is:

[0063] S=[E,g(V),Pcurrent] (5)

[0064] Where S is the state vector, E is the predicted electrical performance data of the semiconductor device obtained from simulation, g(V) is the geometric feature vector of the three-dimensional structure voxel model V, and Pcurrent is the current process recipe parameter. The current process recipe parameter includes the process recipe parameters of the current and one or more upstream key production process steps.

[0065] The state vector is input into the pre-trained decision neural network model π. φ Decision neural network model π φ Guided by preset multiple optimization objectives (e.g., minimizing Ioff while ensuring Ion is not lower than a threshold, or simultaneously optimizing performance, power consumption, and yield), the process parameter adjustment amount ΔP is derived and output by analyzing the difference between the current state vector and the objective. The formula for calculating the process parameter adjustment amount ΔP is:

[0066] ΔP=π φ (S)(6)

[0067] Where S is the state vector, and φ is the decision neural network model π. φ The parameter ΔP is the process parameter adjustment amount. This process parameter adjustment amount includes the adjustment amounts for the current process and one or more upstream critical production process steps.

[0068] Understandably, the decision neural network model analyzes based on multiple preset optimization objectives, and the output process parameter adjustment amount ΔP is the adjustment amount of process parameters for the current and one or more upstream key production process steps. This decision-making process has cross-process collaborative capabilities, and can automatically learn and coordinate the parameter compensation relationship between different process steps. For example, while adjusting etching parameters to improve sidewall roughness, it can fine-tune the temperature curve of the subsequent annealing process to compensate for possible changes in contact resistance, thereby achieving a leap from single-point optimization to global optimum.

[0069] In some embodiments, the current process recipe parameter Pcurrent is optimized based on the process parameter adjustment amount ΔP to obtain the optimized process recipe parameter Pnew, which is then applied to the production of the next batch of semiconductor devices. For example, based on the process parameter adjustment amount ΔP, the system updates the current process recipe parameter Pcurrent to Pnew = Pcurrent + ΔP, obtaining the optimized process recipe parameter, and sends it to the manufacturing execution system through the advanced process control interface. The specific process equipment then executes the new recipe and applies it to the production of the next batch of semiconductor devices. It can be understood that the embodiments of this application embody a closed-loop intelligent cycle of "perception-thinking-decision-action": the current process manufacturing result (Ft) is perceived through multimodal detection data, the three-dimensional structure is reconstructed and the electrical performance is simulated, the current process recipe parameter Pcurrent is adjusted based on the state vector, and finally the optimized process recipe parameter Pnew is sent to the manufacturing execution system through the advanced process control interface for execution by the specific process equipment and applied to the production of the next batch of semiconductor devices. This approach breaks through the limitations of traditional isolated optimization of single process steps, enabling global collaborative optimization across process steps. Through an intelligent closed-loop feedback mechanism of "perception-analysis-decision-execution," the process window is continuously iterated and optimized, thereby significantly shortening the process development and debugging cycle, reducing R&D costs, and ultimately improving the overall performance and production yield of semiconductor devices.

[0070] In some embodiments, after a new batch of wafers completes its critical process steps, multimodal detection data and measured electrical performance data of the semiconductor devices following the current critical process step are collected. Using this multimodal detection data and measured electrical performance data, the 3D structure reconstruction neural network model and / or decision neural network model are optimized and updated. Through this continuous data collection and model iteration, the entire system can continuously evolve, adapt to process drift and new process conditions, and achieve long-term stable optimization results.

[0071] In some embodiments, the step of optimizing and updating the three-dimensional structure reconstruction neural network model includes: constructing a composite loss function, which comprises a weighted sum of voxel-level loss, surface loss, and physical loss; calculating the gradient based on the composite loss function; and iteratively updating the three-dimensional structure reconstruction neural network model along the reverse direction of the gradient using a backpropagation algorithm. The parameters are determined to minimize the composite loss function. The formula for calculating the composite loss function Lrecon is:

[0072] Lrecon=λ vol L BCE (V,Vture) +λ surf L Chamfer (S(V),S(Vture)) +λ phy L phy (V) (7)

[0073] Where Lrecon is the composite loss function, L BCE For voxel-level loss, L Chamfer For surface loss, L phy For physical loss, λ vol , λ surf , λ phy For the corresponding weights.

[0074] Voxel-level loss L BCE Binary cross-entropy is used to measure the difference between the predicted probability of each voxel (i.e., the 3D structural voxel model V) and the true value (i.e., the actual manufactured semiconductor device Vture). Voxel-level loss L... BCE The calculation formula is:

[0075] (8)

[0076] Among them, L BCE For voxel-level loss, N is the total number of voxels. Vtrue(i) is the actual value of the semiconductor device actually manufactured for the i-th voxel, which is either 1 (indicating the presence of material) or 0 (indicating the absence of material). The predicted probability of the three-dimensional structural voxel model V of the i-th voxel is between 0 and 1. The binary cross-entropy loss function encourages the predicted probability to be as close as possible to the true value; when the prediction is completely accurate, the voxel-level loss L... BCE The value is zero; the larger the deviation, the greater the loss at the voxel level (L). BCE The larger.

[0077] Surface loss L Chamfer The chamfer distance is used to measure the geometric deviation between the predicted surface point set S(V) and the true surface point set S(Vture). Surface loss L Chamfer The calculation formula is:

[0078] (9)

[0079] Among them, L Chamfer Let S(V) be the predicted surface point set of the 3D structural voxel model V, and S(Vture) be the actual surface point set of the manufactured semiconductor device Vture. x is a point in the predicted surface point set, and y is a point in the actual surface point set. Specifically, the chamfer distance is a bidirectional distance metric. It first calculates and averages the squared distances from each point x on the predicted surface to the nearest point y on the actual surface, then calculates and averages the squared distances from each point y on the actual surface to the nearest point x on the predicted surface, and finally adds these two averages together. Surface loss L Chamfer Ensure that the predicted surface of the three-dimensional structural voxel model V coincides geometrically with the actual surface of the actual manufactured semiconductor device Vture as much as possible.

[0080] Physical loss L phy To ensure that the reconstructed 3D structure conforms to physical laws, physical constraints are introduced to guide the model in generating more reasonable and realistic 3D structures, guaranteeing the physical reasonableness of the reconstruction results even with insufficient training data or the presence of noise. In some embodiments, the physical loss L is calculated. phy This includes: a weighted sum of the surface smoothness constraint term and the penalty term for process physical inconsistencies. Physical loss L phy The calculation formula is:

[0081] (10)

[0082] Among them, L phy (V) represents physical loss. For surface smoothness constraint terms, This is a penalty item for physical inconsistencies in the manufacturing process. The weighting coefficients for the process physics inconsistency penalty term are as follows: The surface smoothness constraint term is measured by calculating the integral of the square of the surface Laplacian operator of the 3D structure voxel model V. The Laplacian operator reflects the change in surface curvature; a larger integral value indicates a rougher surface with unreasonable burrs or depressions. Minimizing this term encourages the model to generate smooth, natural surfaces, avoiding non-physical morphologies caused by measurement noise or insufficient data. The process physics inconsistency penalty term is determined by defining different physical constraints based on the specific process type. For example, for etching processes, an important physical constraint is "no suspended structures," meaning no voids can exist beneath the material region. If such structures, which do not conform to the physical laws of etching, appear in the reconstructed model, the penalty term will increase. Similarly, for deposition processes, constraints regarding material continuity and thickness uniformity can be defined. By introducing these process-specific physical constraints, the model can learn the inherent laws of different manufacturing processes, thereby generating more accurate and reliable 3D structures.

[0083] Understandably, this application ensures that the reconstructed 3D model not only matches the observed data but also conforms to fundamental physical laws. This enables the 3D structural voxel model to generate reasonable structures even in regions lacking sufficient training data, significantly improving generalization ability. Simultaneously, it penalizes unreasonable structures caused by measurement noise (such as anomalous spikes), prompting the model to find a trade-off between matching data (reducing voxel loss) and maintaining smoothness (reducing physical loss), making the reconstruction results closer to the true physical structure.

[0084] In some embodiments, the step of optimizing and updating the decision neural network model includes: constructing a reinforcement learning reward function, the reward function including an evaluation of the optimization objective and an evaluation of the constraints, wherein the evaluation of the optimization objective includes a weighted sum of the values ​​of each optimization objective, and the evaluation of the constraints includes the sum of the products of the violation amount and the penalty coefficient of each violated constraint; calculating the expected cumulative reward based on the reinforcement learning reward function, and calculating the gradient of the expected cumulative reward relative to the parameters of the decision neural network model using the policy gradient method; updating the parameters of the decision neural network model along the gradient direction using the learning rate, to obtain the updated decision neural network model. The formula for calculating the reinforcement learning reward function r is:

[0085] (11)

[0086] Where r is the reinforcement learning reward function, To evaluate the optimization objective, For the evaluation of constraints. For the k-th optimization objective (such as maximizing the drive current (Ion) or minimizing the leakage current (Ioff), etc.), For the corresponding weighting coefficients, Let $j$ be the violation amount of the $j$-th constraint (e.g., the degree to which process parameters exceed the safe range). $j$ is only considered a constraint violation if it is committed. A penalty is only applied when the value is greater than 0. This is the penalty coefficient; the more serious the violation, the greater the penalty coefficient.

[0087] Understandably, the reward function transforms a complex multi-objective optimization problem into a unified scalar signal. By maximizing the expected cumulative reward, the decision neural network model learns, through continuous trial and adjustment, how to maximize the comprehensive performance of multiple optimization objectives such as drive current and minimize leakage current, while satisfying all constraints, and thus finds the optimal process parameter adjustment strategy to achieve the best overall device performance.

[0088] This application also provides an electronic device 200, such as... Figure 2 As shown, it includes a memory 220, a processor 210, a power supply component 230, a network interface 240, an input / output interface 250, and a program stored on the memory 220 and executable on the processor 210. When the program is executed by the processor 210, it can implement the various processes of the embodiments of the above methods and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0089] This disclosure also provides a chip that implements the steps of the method described above. The chip includes general-purpose processors (such as CPUs and GPUs), mobile device main processors (APs), programmable logic chips (such as FPGAs), and application-specific integrated circuits (such as ASICs). It can achieve the beneficial effects achievable by the method provided in this disclosure, as detailed in the preceding embodiments, and will not be repeated here.

[0090] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, this disclosure also provides a storage medium storing a computer program or instructions that, when executed by a processor, can implement the various processes of the embodiments of the above methods.

[0091] Since the instructions stored in the storage medium can execute the steps of the method provided in the embodiments of this disclosure, the beneficial effects achievable by the method provided in the embodiments of this disclosure can be realized, as detailed in the preceding embodiments, and will not be repeated here. Specific implementations of the above operations can be found in the preceding embodiments, and will not be repeated here.

[0092] Finally, it should be noted that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The embodiments described above, as per the implementation of this application, do not exhaustively describe all details, nor do they limit the application to only the specific embodiments described. Clearly, many modifications and variations can be made based on the above description. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of this application, thereby enabling those skilled in the art to make good use of this application and modifications based on it. This application is limited only by the claims and their full scope and equivalents.

Claims

1. A method for optimizing the manufacturing process of a semiconductor device, comprising: Based on the multimodal detection data of semiconductor devices after the current key manufacturing process steps, a fused data vector is generated. The fused data vector is input into a three-dimensional structure reconstruction neural network model, and the three-dimensional structure voxel model of the semiconductor device is output by the three-dimensional structure reconstruction neural network model. Multi-scale multiphysics performance simulations were performed on the three-dimensional structure voxel model to obtain the predicted electrical performance data of the semiconductor device. The electrical performance prediction data, current process formula parameters, and geometric feature vectors of the three-dimensional structure voxel model are input into the decision neural network model. The decision neural network model analyzes the data based on multiple preset optimization objectives and outputs the process parameter adjustment amount.

2. The manufacturing process optimization method according to claim 1, wherein, The current process formula parameters include the process formula parameters of the current and one or more upstream key production process steps, and the process parameter adjustment amount includes the process parameter adjustment amount of the current and one or more upstream key production process steps.

3. The manufacturing process optimization method according to claim 2, wherein, The manufacturing process optimization method further includes: The current process formula parameters are optimized based on the process parameter adjustment amount to obtain the optimized process formula parameters; The optimized process formulation parameters will be applied to the production of the next batch of semiconductor devices.

4. The manufacturing process optimization method according to claim 3, wherein, The manufacturing process optimization method further includes: Collect multimodal testing data and measured electrical performance data of semiconductor devices after the current key manufacturing process steps; The three-dimensional structure reconstruction neural network model and / or the decision neural network model are optimized and updated using multimodal detection data and electrical performance measurement data after the current key production process steps.

5. The manufacturing process optimization method according to claim 4, wherein, Using multimodal testing data and measured electrical performance data following current key production process steps, the neural network model for reconstructing the three-dimensional structure is optimized and updated, including: Construct a composite loss function, which includes a weighted sum of voxel-level loss, surface loss, and physical loss; The gradient is calculated based on the composite loss function, and the parameters of the three-dimensional structure reconstruction neural network model are iteratively updated along the opposite direction of the gradient using the backpropagation algorithm to minimize the composite loss function. The voxel-level loss uses binary cross-entropy to measure the difference between the predicted probability and the true value of each voxel; the surface loss uses chamfer distance to measure the geometric deviation between the predicted surface point set and the true surface point set; and the physical loss is used to ensure that the 3D structure reconstruction result conforms to physical laws.

6. The manufacturing process optimization method according to claim 5, wherein, Calculating the physical loss includes: Calculate the weighted sum of the surface smoothness constraint term and the penalty term for process physical inconsistency; The surface smoothness constraint is measured by calculating the integral of the square of the surface Laplacian operator of the three-dimensional structure voxel model, and the process physical inconsistency penalty is determined by defining different physical constraints according to the specific process type.

7. The manufacturing process optimization method according to claim 5, wherein, The decision neural network model is optimized and updated using multimodal detection data and measured electrical performance data following current key production process steps, including: Construct a reinforcement learning reward function, which includes an evaluation of the optimization objective and an evaluation of the constraints. The evaluation of the optimization objective includes a weighted sum of the values ​​of each optimization objective, and the evaluation of the constraints includes the sum of the product of the violation amount and the penalty coefficient of each violated constraint. The expected cumulative reward is calculated based on the reinforcement learning reward function, and the gradient of the expected cumulative reward with respect to the parameters of the decision neural network model is calculated using the policy gradient method. The parameters of the decision neural network model are updated along the gradient direction using the learning rate to obtain the updated decision neural network model.

8. The manufacturing process optimization method according to claim 1, characterized in that, The generation of a fused data vector based on the multimodal detection data of semiconductor devices after the current key manufacturing process steps includes: Align and spatiotemporally register the multimodal inspection data of semiconductor devices from different inspection equipment after the current key production process steps to eliminate coordinate deviations between different inspection equipment; Apply the corresponding feature extraction function to the detection data of each modality to convert the raw data into feature vectors; The feature vectors of each modality are fused using a weighted concatenation or attention mechanism to form the fused data vector, wherein each modality detection data has a weight coefficient representing its importance.

9. The manufacturing process optimization method according to claim 2, characterized in that, The process involves inputting the predicted electrical performance data, current process formulation parameters, and the geometric feature vectors of the three-dimensional structure voxel model into a decision neural network model. The decision neural network model then analyzes the data based on multiple preset optimization objectives and outputs adjustments to the process parameters, including: A state vector is constructed, which is composed of the electrical performance prediction data, the geometric feature vector, and the current process formula parameters. The state vector is input into the decision neural network model, which analyzes the data based on multiple preset optimization objectives and derives the adjustment amount of the process parameters in reverse.

10. A semiconductor device, wherein the semiconductor manufacturing process line of the semiconductor device is optimized according to the manufacturing process optimization method according to any one of claims 1 to 9.