A multi-physics tcad proxy modeling method
By constructing a continuous conditional diffusion model and Sobolev edge regularization, the problems of high computational cost and insufficient consistency of TCAD tools in nanoscale integrated circuit design are solved, realizing efficient and accurate simulation of multiphysics fields and supporting the rapid design and evaluation of novel devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-04-07
AI Technical Summary
Existing TCAD tools are computationally intensive, time-consuming to simulate, and difficult to converge in nanoscale integrated circuit design. Multiphysics proxy models lack consistency and cross-field coupling relationships, making it difficult to accurately capture the gradient characteristics and behavior of key junction regions in novel device structures.
A multiphysics TCAD proxy modeling method is adopted, which achieves self-consistent prediction of multiphysics and shape fidelity in high gradient regions by constructing a continuous conditional diffusion model, a conditional context attention injection mechanism, and Sobolev edge regularization.
While reducing simulation latency, it ensures the overall consistency and accuracy of multiphysics, improves the simulation efficiency and accuracy of novel devices, and supports rapid process window evaluation and full exploration of device design.
Smart Images

Figure CN121525613B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated circuits, specifically to a multiphysics TCAD proxy modeling method. Background Technology
[0002] As integrated circuit manufacturing processes continue to evolve towards the nanoscale, device geometry is shrinking dramatically, and device structures and process conditions are becoming increasingly complex. Even minute changes in parameters such as channel length, junction depth, channel doping, and gate bias can significantly alter the device's potential and carrier distribution, thus affecting threshold voltage, current-voltage characteristics, and reliability metrics. Therefore, during process development and DTCO (Design-Process Co-optimization) phases, the industry widely employs Technology Computer-Aided Design (TCAD) tools to perform extensive simulations of devices under different geometries, doping distributions, and bias conditions. This allows for the acquisition of steady-state distributions of multiple physical fields, such as potential, electron concentration, and hole concentration. Based on these fields, performance metrics are further calculated to evaluate and optimize the devices.
[0003] Traditional commercial TCAD tools solve partial differential equations self-consistently on non-uniform meshes to obtain multiphysics solutions with spatial distributions of physical quantities, providing a basis for device behavior analysis. However, their numerical solution process is computationally intensive, with simulations of a single operating point often taking minutes or even hours, and convergence issues exist under extreme structural conditions. This makes large-scale parameter scanning and rapid process window evaluation extremely time-consuming, limiting the full exploration of the device design space.
[0004] To reduce simulation overhead, various machine learning-based TCAD surrogate models have been proposed in existing technologies. Existing scalar-level surrogate models map device geometry, fabrication processes, and biases to a few metrics such as threshold voltage and leakage current for rapid sorting and feasibility screening, outputting only a small number of scalar metrics and losing the field-level information required for further physical analysis. Existing field-level surrogate models directly predict spatial distributions such as potential and carrier concentration on discrete grids, mostly using pixel-level loss functions such as pixel-wise mean square error for training. Physically, they lack cross-field consistency and explicit constraints on the coupling relationships between physical fields.
[0005] Therefore, the present invention aims to solve the following three technical problems:
[0006] (1) How to predict multiple coupled physical fields simultaneously while maintaining overall physical consistency under the premise of significantly reducing simulation latency and continuous changes in device structure and bias conditions.
[0007] (2) How to effectively suppress the numerical smoothing phenomenon of the surrogate model and improve the field-level fidelity of the junction region in the high gradient regions such as PN junction, source / drain-channel junction, and gate oxide-silicon interface.
[0008] (3) Given the availability of limited target domain TCAD data (e.g., novel device structures), how can a reasonable training mechanism enable the multiphysics proxy model to accurately capture the gradient features and cross-layer interface behavior of key junction regions, thereby achieving efficient transfer and generalization of new device types? Summary of the Invention
[0009] The purpose of this invention is to provide a multiphysics TCAD proxy modeling method to solve the problems mentioned in the background art.
[0010] To achieve the above objectives, the present invention provides the following technical solution:
[0011] A multiphysics TCAD proxy modeling method includes:
[0012] Step 1: Acquire and standardize multiphysics TCAD data;
[0013] Step 2: Construct a continuous conditional diffusion model and a conditional context attention injection mechanism;
[0014] Step 3: Perform Sobolev edge regularization and loss design for the node region.
[0015] Further, step 1 includes:
[0016] First, based on the commercial TCAD simulation tool, the structural parameters and bias conditions are scanned on the two-dimensional top-gate MOSFET reference device template, and three steady-state multiphysics fields, namely potential, electron concentration and hole concentration, are extracted. Then, through a unified post-processing process, a standardized multiphysics tensor dataset that can be used for training is formed.
[0017] Furthermore, the construction of the continuous conditional diffusion model in step 2 includes:
[0018] The complete device configuration is uniformly represented as a global continuous conditional vector c:
[0019] ,
[0020] Among them, the global continuous condition vector c is used to uniformly and continuously encode the device geometric parameters, process and doping parameters, and electrical bias conditions; where L g X represents the gate length of the device. j Indicates the source-drain junction depth, N sub N represents the substrate doping concentration. peak V represents the peak doping concentration of the source / drain.d This represents the drain bias voltage, V. g This represents the gate bias voltage; a shared condition encoder, CondEncoder(·), is introduced to project c into a global condition embedding h. c , where h c By maintaining a spatial constant for individual samples and reusing them across all network depths and resolutions, conditional semantic consistency is ensured and smooth responses are supported under continuous scanning.
[0021] ,
[0022] Among them, h c As a unified high-dimensional representation of device conditions, a channel-level conditional affine modulation mechanism is adopted. For the i-th convolutional block, a clear distinction is made between the denoising process modulation defined by the noise level of the diffusion model and the modulation defined by the unified high-dimensional representation h of device conditions. c The device conditions are defined to modulate two types of modulation sources, and gating fusion is performed at the channel level.
[0023] Furthermore, the denoising process modulation in step 2 includes:
[0024] The time step embedding e is calculated based on the current noise level σ. t (σ) and mapped to obtain the time step latent variable (γ) t ,β t ), where γ t β represents the channel-level scaling modulation factor corresponding to the current denoising stage. t This represents the channel-level translation modulation factor corresponding to the current denoising stage, used to characterize the overall response of the current denoising stage.
[0025] Further, the device condition modulation in step 2 includes:
[0026] Embed global conditions in h c The input is a shared conditional projector, which generates a spliced sequence of modulation vectors. This sequence is then segmented according to the network topology, ensuring that each convolutional block receives a device conditional latent variable (γ) matching its own channel count. c,i ,β c,i ), where γ c,i β represents the device conditional channel-level scaling and modulation factor corresponding to the i-th convolutional block. c,i This represents the device-conditional channel-level translation modulation factor corresponding to the i-th convolutional block; subsequently, two sets of trainable channel-gating hyperparameters are defined and introduced within the i-th block. , ,in, Used to adjust the injection intensity of device conditional scaling modulation in the channel dimension. This is used to adjust the injection intensity of the device conditional translation modulation in the channel dimension, and the gating coefficients are initialized to near zero in the early stage of training. The two types of modulation are fused in the channel dimension to form the final scaling and translation coefficients; using s i b represents the final channel-level scaling factor for the i-th convolutional block. i The final channel-level translation coefficients of the i-th convolutional block are calculated as follows:
[0027] ,
[0028] in, This represents element-wise multiplication. Let y be the intermediate activation feature obtained by the i-th convolutional block after the previous convolution operation and normalization. i y i The channel-normalized feature representation, which indicates that scale differences have been eliminated, then channel-level conditional affine modulation is achieved by adjusting y. i Channel-by-channel scaling and translation operations are applied, and the modulated output x is obtained through the SiLU nonlinear activation function. i ,Right now:
[0029] .
[0030] Furthermore, step 2 also includes:
[0031] In implementation, the channel-level conditional affine modulation mechanism is set between two convolutional operations within the convolutional block. That is, the features are first processed and normalized by the first convolutional layer to form the intermediate activation features, then channel-level affine modulation that integrates time step information and device condition information is injected, and finally input to the second convolutional layer to generate the feature representation for the next stage.
[0032] Furthermore, the construction of the conditional context attention injection mechanism in step 2 includes:
[0033] Two parallel attention paths are set in the attention available block of the diffusion network and fused by in-block gating, where the feature query vector is extracted from the current spatial feature map;
[0034] First, the network receives the global continuous conditional vector c and obtains the global conditional embedding h through a shared conditional encoder. c Global conditional embedding serves as the basic representation for subsequent conditional injection. It is reused and projected into a fixed number of conditional tokens at all depths and different spatial resolutions of the network in a sample-level, spatially invariant form. These tokens are shared among attention layers with different spatial resolutions.
[0035] Construct two attention paths in parallel:
[0036] (1) Self-attention path: Within the attention block, queries, keys and values are generated simultaneously from the current spatial feature map, so that the query vector performs attention aggregation on the spatial feature itself, in order to model the long-range spatial dependencies within the feature map;
[0037] (2) Conditional attention path: Using the same query vector as the self-attention path, conditional cross-attention calculation is performed on the key and value composed of conditional tokens, thereby injecting the continuous constraints of device geometric parameters and bias conditions into the spatial feature representation;
[0038] The outputs of the two attention paths are fused within each attention block through the learnable gating corresponding to that block. The gating makes the contribution of the conditional cross attention branch close to zero in the initial stage of training and is updated along with the other parameters of the model during training.
[0039] Further, step 3 includes:
[0040] This involves constructing a junction mask, calculating gradients, defining two classes of Sobolev edge regularization constraints, and implementing gating and budgeted weighting to match the noise schedule; resulting in a clean TCAD multiphysics ground truth field. ,in, This represents the dimension of multi-channel two-dimensional field data organized in tensor form, where C is the number of physical field channels, and H and W are the number of sampling points in the height and width of the discrete grid in two spatial directions, respectively; the network output denoising prediction at the current noise level σ. and define residuals To characterize the deviation between the predicted field and the true field; the key region mask M of the junction region is extracted from the peak gradient "ridge" of the doped junction region and obtained by narrowband dilation, and is used as a spatial selector fixed by sample, and its area is normalized to eliminate the scale difference between samples.
[0041] The discrete first derivatives for r and x are calculated separately, and an in-channel Sobel filter is applied in conjunction with a fixed scaling factor matched to the grid spacing to obtain the following results. , ,in, and Δr represents the discrete partial derivative operators along the first spatial direction and the second spatial direction, respectively. Δr represents the gradient vector pair of the residual field in the two-dimensional space, and Δx represents the gradient vector pair of the multi-physics truth field in the two-dimensional space. It also allows the selection of some physical field channels and the application of channel weights.
[0042] Furthermore, step 3 also includes:
[0043] Define two complementary Sobolev edge regularization loss terms: the relative gradient magnitude constraint term L H1 and normal alignment constraint terms ;
[0044] Relative gradient magnitude constraint term L H1 The format is:
[0045] in Indicates spatial averaging only over the area covered by mask M and the selected channel; normal alignment constraint term. The form is:
[0046] in, This represents the two-dimensional first-order gradient vector of the residual r. To ensure relative stability and prevent the denominator from becoming too small and causing numerical divergence, For masking key areas of the junction region, the symbol is... Indicates only on the mask Spatial averaging is performed within the coverage area and the selected physical field channel range. Represents the true value field of a multiphysics system The gradient normalization yields the cross-junction normal direction vector. The form is:
[0047] ,
[0048] Subsequently, a noise gating function g(σ) is introduced, emphasizing regularization in the low-noise stage and weakening regularization in the high-noise stage, with the following form:
[0049] ,
[0050] Among them, hyperparameters The characteristic noise threshold of the gating is represented, and the hyperparameter α is used to control the transition of the gating from weak to strong. It is a smooth mapping function; when When the value is small, g(σ) approaches 1 to increase the weight of the Sobolev marginal regularization term. When the value is large, g(σ) approaches 0 to reduce the interference of the regularization term on training;
[0051] The Sobolev edge regularization loss term is preconditionally weighted with noise scale uniformity and works in conjunction with the noise gating function g(σ).
[0052] Furthermore, step 3 also includes: employing a budgeted weighting strategy: given the total budget ratio of Sobolev edge regularization. And set the internal split ratio for the two losses, and dynamically update L. H1 and The weights are determined, and a correction factor γ(σ) is introduced to compensate for small-batch statistical fluctuations.
[0053] Compared with the prior art, the beneficial effects of the present invention are:
[0054] 1) Sobolev edge regularization and loss design for node regions:
[0055] This invention optimizes the shape fidelity of transitions in high-gradient regions (such as junctions) during multiphysics generation by constraining the first derivative of the residual. This regularization method effectively avoids physical distortion caused by gradient smoothing in traditional and diffusion models, representing a key innovation for improving model accuracy and physical consistency.
[0056] 2) Continuous Conditional Diffusion Denoising Network and Conditional Context Attention Mechanism:
[0057] This invention proposes a conditional contextual attention mechanism and a continuous conditional diffusion denoising network to ensure the consistency and accuracy of conditional information at different depths and resolutions. Through the effective injection of conditional information, this invention improves the self-consistency and responsiveness of the model during multiphysics generation.
[0058] 3) Transfer learning and model fine-tuning:
[0059] This invention introduces a transfer learning method to address the scarcity of simulation data for novel devices by fine-tuning a pre-trained model. This method, based on the model's prior knowledge of the device, rapidly adapts to other devices through a transfer learning process.
[0060] 4) Construction of the physics simulation dataset:
[0061] This invention constructs a TCAD simulation dataset based on a real simulation environment, covering multi-physics data such as electric potential and carrier concentration, providing a reliable data foundation for model training and testing, and further enhancing the applicability of the technology. Attached Figure Description
[0062] Figure 1 Flowchart of TCAD proxy modeling method for multiphysics fields.
[0063] Figure 2 This is a schematic diagram of the channel-level conditional affine modulation mechanism of the continuous conditional diffusion model constructed in step 2.
[0064] Figure 3 This is a schematic diagram of the conditional context attention injection mechanism constructed in step 2.
[0065] Figure 4 This is a flowchart of the Sobolev edge regularization and loss design for the node region in step 3.
[0066] Figure 5This diagram illustrates the comparison of prediction results of the present invention. It consists of nine sub-figures (a)–(i), where: (a) represents the TCAD result of the device potential distribution; (b) represents the potential prediction result output by the method of the present invention under the corresponding conditions; (c) represents the error distribution between the potential prediction result and the true value; (d) represents the TCAD result of the electron concentration distribution; (e) represents the corresponding electron concentration prediction result; (f) represents the electron concentration prediction error distribution; (g) represents the TCAD result of the hole concentration distribution; (h) represents the corresponding hole concentration prediction result; and (i) represents the hole concentration prediction error distribution. The above comparison can intuitively demonstrate the high-precision reconstruction capability of the present invention method for junction regions and high gradient regions in multi-physics joint modeling. Detailed Implementation
[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] Please see Figures 1-4 A multiphysics TCAD proxy modeling method, comprising:
[0069] Step 1: Acquire and standardize multiphysics TCAD data, including:
[0070] First, based on the commercial TCAD simulation tool, the structural parameters and bias conditions are scanned on the two-dimensional top-gate MOSFET reference device template, and three steady-state multiphysics fields, namely potential, electron concentration and hole concentration, are extracted. Then, through a unified post-processing process, a standardized multiphysics tensor dataset that can be used for training is formed.
[0071] Step 2: Construct a continuous conditional diffusion model and a conditional context attention injection mechanism.
[0072] Constructing a continuous conditional diffusion model includes:
[0073] First, the complete device configuration is uniformly represented as a global continuous conditional vector c:
[0074]
[0075] in, This represents a six-dimensional real vector space consisting of six real variables. The global continuous condition vector c is used to uniformly and continuously encode device geometric parameters, process and doping parameters, and electrical bias conditions; where L g X represents the gate length of the device.j Indicates the source-drain junction depth, N sub N represents the substrate doping concentration. peak V represents the peak doping concentration of the source / drain. d This represents the drain bias voltage, V. g This represents the gate bias voltage. A shared condition encoder, CondEncoder(·), is introduced to project c onto a globally conditionally embedded h. c , where h c By maintaining a spatial constant for individual samples and reusing them across all network depths and resolutions, conditional semantic consistency is ensured and smooth responses are supported under continuous scanning.
[0076]
[0077] in, Let h represent a 256-dimensional real vector space consisting of 256 real components. c As a unified high-dimensional representation of device conditions, this invention employs a channel-wise affine modulation mechanism within each convolutional block of the diffusion generation network. This is done to avoid conditional injection interfering with the temporal modeling of the diffusion denoising process, while simultaneously allowing for finely controllable adjustment of physically sensitive channels. For example... Figure 2 As shown, for the i-th convolutional block, two types of modulation sources are clearly distinguished and gated fusion is performed in the channel dimension: (1) Denoising process modulation, the time step embedding e is calculated based on the current noise level σ. t (σ) and mapped to obtain the time step latent variable (γ) t ,β t ), where γ t β represents the channel-level scaling modulation factor corresponding to the current denoising stage. t This represents the channel-level translation modulation factor corresponding to the current denoising stage, used to characterize the overall response of the current denoising stage.
[0078] (2) Device condition modulation, embedding global conditions into h c The input is a shared conditional projector, which generates a concatenated sequence of modulation vectors. This sequence is then segmented according to the network topology, ensuring that each convolutional block receives a device conditional latent variable (γ) matching its own channel count. c,i ,β c,i ), where γ c,i β represents the device conditional channel-level scaling and modulation factor corresponding to the i-th convolutional block. c,i This represents the device-conditional channel-level translation modulation factor corresponding to the i-th convolutional block. Subsequently, two sets of trainable channel-gating hyperparameters are defined and introduced within the i-th block. , Among them, channel gating hyperparameters Channel gating hyperparameters used to adjust the injection intensity of device conditional scaling modulation in the channel dimension. This is used to adjust the injection intensity of the device's conditional translation modulation in the channel dimension. The gating coefficients are initialized to near zero during the initial training phase. Using a "first stabilize and denoise, then gradually learn conditional control" approach, the two types of modulation are fused in the channel dimension into the final scaling and translation coefficients. (Using s...) i b represents the final channel-level scaling factor for the i-th convolutional block. i The final channel-level translation coefficients of the i-th convolutional block are represented as follows:
[0079]
[0080] in, This represents element-wise multiplication. Let y be the intermediate activation feature obtained by the i-th convolutional block after the previous convolution operation and normalization. i y i The channel-level conditional affine modulation, which represents the channel normalized feature representation after eliminating scale differences, is then applied to y. i Channel-by-channel scaling and translation operations are applied, and the modulated output x is obtained through the SiLU (Sigmoid Linear Unit) nonlinear activation function. i ,Right now:
[0081]
[0082] In implementation, the channel-level conditional affine modulation mechanism is positioned between two convolutional layers within a convolutional block. Specifically, the features are first processed and normalized by the first convolutional layer to form the intermediate activation features, denoted as y. i Then, channel-level affine modulation, which integrates time step information and device condition information, is injected and finally input into the second convolutional layer to generate the feature representation for the next stage.
[0083] The construction of a conditional context attention injection mechanism includes:
[0084] like Figure 3 As shown, in order to enable continuous conditions to participate in feature modeling in the form of attention during the diffusion denoising process, this invention sets up two parallel attention paths in the attention available block of the diffusion network and fuses them by in-block gating, wherein the feature query vector is extracted from the current spatial feature map.
[0085] First, the network receives the globally continuous conditional vector c constructed at point 3 and obtains the global conditional embedding h through a shared conditional encoder. cThe global conditional embedding is reused at all network depths and different spatial resolutions in a sample-level, spatially invariant form; simultaneously, the global conditional embedding h c As the basis for subsequent conditional injection, the projection is a fixed number of conditional tokens that depend only on the global conditional embedding h. c Furthermore, it is shared between attention layers with different spatial resolutions, thereby avoiding the problem of inconsistent representation of conditional information at different network depths.
[0086] Construct two attention paths in parallel:
[0087] (1) Self-attention path: Within the attention available block, the query vector Q is extracted from the current spatial feature map, and the key and value of the path are derived from the same spatial feature map. This allows the query vector Q to perform attention aggregation on its own spatial features within the block, so as to establish long-range dependencies within the feature map and maintain the ability of the diffusion denoising backbone to represent the spatial structure.
[0088] (2) Conditional attention path: embed global conditions into h c The projection is a fixed number of condition tokens; based on this, using the same query vector Q as the self-attention path, conditional cross-attention computation is performed on the keys and values composed of condition tokens, so that spatial features can retrieve and align the conditional context, thereby explicitly injecting continuous constraints of geometry and bias into the attention mechanism.
[0089] The outputs of the two attention paths are fused within each attention block through the learnable gating corresponding to that block. The gating makes the contribution of the conditional cross attention branch close to zero in the initial stage of training, and updates it together with the other parameters of the model during training.
[0090] Step 3, perform Sobolev regularization and loss design for the node region, including:
[0091] like Figure 4 As shown, to suppress excessive smoothing in the knot region and improve shape fidelity in high-gradient regions, this invention introduces knot region Sobolev edge regularization. At the loss level, a consistency constraint is applied to the first-order residual derivative in key regions of the knot region. Sobolev refers to applying an L2 penalty to the first-order residual derivative, corresponding to H... 1 The seminorm energy constraint is used to couple adjacent pixels in high gradient regions, thereby promoting sharp, physically reliable transitions across the junction region (Sobolev refers to the L2 penalty of the first residual derivative).
[0092] Step 3 specifically includes: constructing a junction mask, calculating gradients, defining two types of Sobolev edge regularization constraints, and gating and budgeted weighting to match the noise schedule; making the clean TCAD multiphysics ground truth field... ,in Let c represent the dimension of the multi-channel two-dimensional field data organized in tensor form, where c is the number of physical field channels (corresponding to field components of different physical quantities), and H and W are the number of sampling points in the height and width of the discrete grid in two spatial directions, respectively. The network outputs a denoising prediction at noise level σ. and define residuals To characterize the deviation between the predicted field and the true field; the key region mask M of the junction is extracted from the peak gradient "ridge" of the doped junction and obtained by narrowband dilation, and is used as a spatial selector fixed by sample, and its area is normalized to eliminate the scale difference between samples.
[0093] The discrete first derivatives for r and x are calculated separately, and an in-channel Sobel filter is applied in conjunction with a fixed scaling factor matched to the grid spacing to obtain the following results. , ,in and Let represent the discrete partial derivative operators along the first and second spatial directions, respectively. Δr represents the gradient vector pair of the residual field in two-dimensional space (used to characterize the rate of change of the error in space), and Δx represents the gradient vector pair of the true field in two-dimensional space (used to characterize the rate of change of the physical field in space). Simultaneously, it is allowed to select some physical field channels and apply channel weights. Based on this, the following two complementary Sobolev edge regularization loss terms are defined:
[0094] (1) Relative gradient magnitude constraint term L H1
[0095] This item is used to constrain the relative error of the residual gradient magnitude in the critical region of the junction, focusing on penalizing the interface smoothing phenomenon caused by "gradient magnitude mismatch". Its form is as follows:
[0096] in This indicates that the spatial average is applied only to the area covered by mask M and the selected channel.
[0097] (2) Normal alignment constraint terms
[0098] This term is used to constrain the consistency of the residual gradient in the cross-junction region direction. In the high-gradient junction region, the cross-junction normal can be considered to be consistent with (or approximately consistent with) the local electric field direction. Therefore, the cross-junction normal n is first constructed based on the true gradient. x :
[0099]
[0100] Subsequently, the residual gradient is calculated in the normal direction n. x Normalization penalty is applied to the projection on the surface:
[0101] in, Let represent the two-dimensional first-order gradient vector of the residual field r, and let represent the gradient vector of the multiphysics truth field. The gradient normalization yields the cross-junction normal direction vector. This represents the projected component of the residual gradient in the normal direction. To ensure relative stability and prevent the denominator from becoming too small and causing numerical divergence, For masking key areas of the junction region, the symbol is... Indicates only on the mask Spatial averaging is performed within the coverage area and the selected physical field channel range, thereby ensuring the consistency of the residual gradient across the junction region (normal direction) and suppressing excessive smoothing of the junction interface. To match the above regularization with the diffusion noise schedule, this invention introduces a noise gating function. The regularization is emphasized in the low-noise stage and weakened in the high-noise stage, and its form is as follows:
[0102]
[0103] Where σ represents the current noise scale (noise level) during the diffusion training / denoising process, and the hyperparameters are... The characteristic noise threshold of the gating is represented, and the hyperparameter α is used to control the transition of the gating from weak to strong. It is a smoothing mapping function used to implement gated switches. When σ is small (low noise, close to the stage of reconstructing fine structure), The value approaches 1 to increase the weight of the Sobolev edge regularization term. When σ is large (high noise, still coarse structure stage), Approaching 0 reduces the interference of regularization terms on training, achieving a "low-noise reinforcement, high-noise suppression" regularization injection strategy that matches the noise schedule.
[0104] The Sobolev edge regularization loss term is preconditionally weighted with noise scale uniformity and then combined with the noise gating function. Together, they achieve stable injection and controllable intensity distribution of edge regularization at different noise stages of diffusion training.
[0105] Furthermore, to control the overall regularization strength and stabilize training, step 3 also employs a budgeted weighting strategy: given the proportion of the Sobolev total budget... And set the internal split ratio for the two losses, and dynamically update L. H1 and The weights are determined, and a correction factor is introduced. This strategy compensates for small-batch statistical fluctuations. It only applies to the loss-weighted level and does not change the feedforward network structure or sampling process.
[0106] The dataset used in the following examples is as follows:
[0107] Parallel parametric sweep simulations were performed in parallel on a computing cluster using commercial TCAD tools. The sweep variables were six continuous design / operational parameters, including structural and bias parameters: gate length, junction depth, substrate doping, source / drain peak doping, drain bias, and gate bias. The ranges of each parameter are shown in Table 1. A total of 10,000 different device instances were obtained within these ranges and divided into training, validation, and test sets in a 7:2:1 ratio. Standard physical models were used in the simulations, including Fermi-Dirac statistics, Old-Slotboom bandgap narrowing correction, high-field mobility saturation model, and Shockley-Read-Hall composite model. For the steady-state converged solution of each simulation, only the semiconductor region was retained, and a two-dimensional potential field was extracted from the converged steady-state results. electron density Hole density As a multiphysics supervision target, after exporting the data, each physics field is resampled to a unified 128×128 structured grid and normalized to obtain a standardized multichannel tensor representation used for training.
[0108]
[0109] Table 1 Examples of design parameters generated from the TCAD dataset
[0110] Example 1:
[0111] This embodiment was performed in an Ubuntu server environment, with hardware including an NVIDIA A100 (80GB VRAM) GPU and an Intel Xeon Gold 6348 CPU. The diffusion model was pre-trained using the EDM training settings, with parameter updates performed using the AdamW optimizer. The learning rate was set to 6e-4, the batch size to 80, and approximately 2×10⁻⁶ cycles were pre-trained. 6 Each noisy physics sample took approximately 8 hours to complete. Different comparison methods maintained consistency in data partitioning, loss function, learning rate, and model capacity. Evaluation metrics included logarithmic mean squared error and relative error, supplemented by pixel residual map comparisons.
[0112] like Figure 5As shown, this embodiment compares the method of the present invention with two representative TCAD field surrogate models: Lee-UNet with a pixel regression paradigm, and NDD based on the Fourier Neural Operator. Table 2 shows the comparison results between the TCAD alternative model described in this example and existing methods. The results show that, under the same model learning settings, our model outperforms other models in both mean squared error and relative error. For example, in the potential prediction task, the MSE is reduced from 3.95 × 10⁻⁶. -5 Reduced to 4.87×10 -6 The relative error decreased from 2.15% to 0.72%; in the electron density prediction task, the MSE decreased from 1.39 × 10⁻⁶. -4 Reduced to 1.62×10 -5 The relative error decreased from 2.10% to 0.72%; the MSE for hole density decreased from 1.15 × 10⁻⁶. -4 Reduced to 2.49×10 -5 The relative error decreased from 1.16% to 0.58%. Furthermore, the present invention achieves a single inference time of approximately 4 seconds, while commercial TCAD tools require approximately 120–300 seconds per steady-state simulation and may not converge under certain configurations, indicating that at least a 30-fold speedup can be achieved in convergent cases.
[0113]
[0114] Table 2 Performance comparison between the method of the present invention and existing methods
[0115] Example 2:
[0116] This embodiment uses an SOI-FET device as the representative target device for transfer learning fine-tuning. Its structure includes a buried oxide layer that isolates the thin silicon channel from the substrate, thereby altering the material interface and electric field distribution. To characterize the corresponding transport behavior, a more advanced carrier transport and quantum correction model is used in the simulation. Candidate designs and bias configurations are extracted from the same scan range as the source domain. Data is still divided into training / validation / test sets in a 7:2:1 ratio, and the input continues to use the aforementioned six-dimensional continuous conditional vector. Data-efficient transfer learning is employed: pre-training is completed on the MOSFET dataset, followed by fine-tuning on a small sample dataset of SOI-FET. During fine-tuning, the weights are initialized from the pre-trained weights, and the same training objective as the source domain is used, allowing the continuous conditional diffusion model and Sobolev-edge regularization to be optimized together. Simultaneously, to adapt to new high-gradient interfaces in the SOI structure (such as the gate oxide-silicon boundary), the key interface region mask M is recalculated based on the SOI doping distribution.
[0117] Based on the same training budget (approximately 2 × 10) 6(Using a sample of noisy physical fields), the results compare the performance of the scratch-learning method with the fine-tuning method based on MOSFET pre-training. The results show that the fine-tuning method significantly outperforms the scratch-learning method in all three physical fields. Specifically, in the ϕ field, the MSE decreases from 7.40e-05 to 1.72e-05 after fine-tuning, and the relative error decreases from 2.82% to 1.36%; in the n field, the MSE decreases from 5.31e-04 to 8.58e-05, and the relative error decreases from 2.51% to 1.01%; in the p field, the MSE decreases from 2.26e-04 to 3.85e-05, and the relative error decreases from 3.04% to 1.26%. These results demonstrate that the FUSE-TCAD model pre-trained on MOSFETs can effectively provide prior knowledge for new device structures and higher-order physical problems, achieving efficient adaptation and optimization even with limited data in the target domain.
[0118]
[0119] Table 3. Error Comparison between Direct Training and Transfer Learning Fine-tuning
[0120] These experiments demonstrate that the model of this invention possesses strong field-level prior information, significantly improving prediction accuracy on novel devices such as SOI-FETs. Furthermore, due to the complex quantum correction model of SOI-FETs, although the absolute error of the fine-tuned model in the target domain is slightly higher than that in the source domain, it still exhibits a significant performance improvement. This fine-tuning-based transfer learning method provides an efficient and data-intensive solution for modeling complex devices.
[0121] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multiphysics TCAD proxy modeling method, characterized in that, include: Step 1: Acquire and standardize multiphysics TCAD data; Step 2: Construct a continuous conditional diffusion model and a conditional context attention injection mechanism; Constructing a continuous conditional diffusion model includes: The complete device configuration is uniformly represented as a global continuous conditional vector c: , Among them, the global continuous condition vector c is used to uniformly and continuously encode the device geometric parameters, process and doping parameters, and electrical bias conditions; where L g X represents the gate length of the device. j Indicates the source-drain junction depth, N sub N represents the substrate doping concentration. peak V represents the peak doping concentration of the source / drain. d This represents the drain bias voltage, V. g This represents the gate bias voltage; a shared condition encoder, CondEncoder(·), is introduced to project c into a global condition embedding h. c , where h c By maintaining a spatial constant for individual samples and reusing them across all network depths and resolutions, conditional semantic consistency is ensured and smooth responses are supported under continuous scanning. , Among them, h c As a unified high-dimensional representation of device conditions, a channel-level conditional affine modulation mechanism is adopted. For the i-th convolutional block, a clear distinction is made between the denoising process modulation defined by the noise level of the diffusion model and the modulation defined by the unified high-dimensional representation h of device conditions. c The device conditions are defined to modulate two types of modulation sources, and gated fusion is performed at the channel level; The construction of a conditional context attention injection mechanism includes: Two parallel attention paths are set in the attention available block of the diffusion network and fused by in-block gating, where the feature query vector is extracted from the current spatial feature map; First, the network receives the global continuous conditional vector c and obtains the global conditional embedding h through a shared conditional encoder. c Global conditional embedding serves as the basic representation for subsequent conditional injection. It is reused and projected into a fixed number of conditional tokens at all depths and different spatial resolutions of the network in a sample-level, spatially invariant form. These tokens are shared among attention layers with different spatial resolutions. Construct two attention paths in parallel: (1) Self-attention path: Within the attention block, queries, keys and values are generated simultaneously from the current spatial feature map, so that the query vector performs attention aggregation on the spatial feature itself, in order to model the long-range spatial dependencies within the feature map; (2) Conditional attention path: Using the same query vector as the self-attention path, conditional cross-attention calculation is performed on the key and value composed of conditional tokens, thereby injecting the continuous constraints of device geometric parameters and bias conditions into the spatial feature representation; The outputs of the two attention paths are fused within each attention block through the learnable gating corresponding to that block. The gating makes the contribution of the conditional cross attention branch close to zero in the initial stage of training and is updated along with the other parameters of the model during training. Step 3: Perform Sobolev edge regularization and loss design for the junction region. Sobolev refers to applying L2 penalty to the first-order residual derivative.
2. The multiphysics TCAD proxy modeling method according to claim 1, characterized in that, Step 1 includes: First, based on the commercial TCAD simulation tool, the structural parameters and bias conditions are scanned on the two-dimensional top-gate MOSFET reference device template, and three steady-state multiphysics fields, namely potential, electron concentration and hole concentration, are extracted. Then, through a unified post-processing process, a standardized multiphysics tensor dataset that can be used for training is formed.
3. The multiphysics TCAD proxy modeling method according to claim 1, characterized in that, The noise reduction process modulation in step 2 includes: The time step embedding e is calculated based on the current noise level σ. t (σ) and mapped to obtain the time step latent variable (γ) t ,β t ), where γ t β represents the channel-level scaling modulation factor corresponding to the current denoising stage. t This represents the channel-level shift modulation factor corresponding to the current denoising stage, used to characterize the overall response of the current denoising stage.
4. The multiphysics TCAD proxy modeling method according to claim 3, characterized in that, The device condition modulation in step 2 includes: Embed global conditions in h c The input is a shared conditional projector, which generates a spliced sequence of modulation vectors. This sequence is then segmented according to the network topology, ensuring that each convolutional block receives a device conditional latent variable (γ) matching its own channel count. c,i ,β c,i ), where γ c,i β represents the device conditional channel-level scaling and modulation factor corresponding to the i-th convolutional block. c,i This represents the device-conditional channel-level translation modulation factor corresponding to the i-th convolutional block; subsequently, two sets of trainable channel-gating hyperparameters are defined and introduced within the i-th block. , ,in, Used to adjust the injection intensity of device conditional scaling modulation in the channel dimension. This is used to adjust the injection intensity of the device conditional translation modulation in the channel dimension, and the gating coefficients are initialized to near zero in the early stage of training. The two types of modulation are fused in the channel dimension to form the final scaling and translation coefficients; using s i b represents the final channel-level scaling factor for the i-th convolutional block. i The final channel-level translation coefficients of the i-th convolutional block are calculated as follows: , in, This represents element-wise multiplication. Let y be the intermediate activation feature obtained by the i-th convolutional block after the previous convolution operation and normalization. i y i The channel-normalized feature representation, which indicates that scale differences have been eliminated, then channel-level conditional affine modulation is achieved by adjusting y. i Channel-by-channel scaling and translation operations are applied, and the modulated output x is obtained through the SiLU nonlinear activation function. i ,Right now: 。 5. The multiphysics TCAD proxy modeling method according to claim 4, characterized in that, Step 2 also includes: In implementation, the channel-level conditional affine modulation mechanism is set between two convolutional operations within the convolutional block. That is, the features are first processed and normalized by the first convolutional layer to form the intermediate activation features, then channel-level affine modulation that integrates time step information and device condition information is injected, and finally input to the second convolutional layer to generate the feature representation for the next stage.
6. The multiphysics TCAD proxy modeling method according to claim 1, characterized in that, Step 3 includes: This involves constructing a junction mask, calculating gradients, defining two classes of Sobolev edge regularization constraints, and implementing gating and budgeted weighting to match the noise schedule; resulting in a clean TCAD multiphysics ground truth field. ,in, This represents the dimension of multi-channel two-dimensional field data organized in tensor form, where C is the number of physical field channels, and H and W are the number of sampling points in the height and width of the discrete grid in two spatial directions, respectively; the network output denoising prediction at the current noise level σ. and define residuals To characterize the deviation between the predicted field and the true field; the key region mask M of the junction region is extracted from the peak gradient ridge of the doped junction region and obtained by narrow band expansion, and is used as a spatial selector fixed by sample, and its area is normalized to eliminate the scale difference between samples. The discrete first derivatives for r and x are calculated separately, and an in-channel Sobel filter is applied in conjunction with a fixed scaling factor matched to the grid spacing to obtain the following results. , ,in, and Let represent the discrete partial derivative operators along the first spatial direction and the second spatial direction, respectively. This represents a pair of gradient vectors of the residual field in two-dimensional space. It represents the gradient vector pair of the multi-physics truth field in two-dimensional space, and allows the selection of some physics channels and the application of channel weights.
7. The multiphysics TCAD proxy modeling method according to claim 6, characterized in that, Step 3 also includes: Define two complementary Sobolev edge regularization loss terms: the relative gradient magnitude constraint term L H1 and normal alignment constraint terms ; Relative gradient magnitude constraint term L H1 The format is: in Indicates spatial averaging only over the area covered by mask M and the selected channel; normal alignment constraint term. The form is: in, To ensure relative stability and prevent the denominator from becoming too small and causing numerical divergence, Masking of key areas in the junction region, Represents the true value field of a multiphysics system The gradient normalization yields the cross-junction normal direction vector. The form is: , Subsequently, a noise gating function g(σ) is introduced, emphasizing regularization in the low-noise stage and weakening regularization in the high-noise stage, with the following form: , Among them, hyperparameters The characteristic noise threshold of the gating is represented, and the hyperparameter α is used to control the transition of the gating from weak to strong. It is a smooth mapping function; when When the value is small, g(σ) approaches 1 to increase the weight of the Sobolev marginal regularization term. When the value is large, g(σ) approaches 0 to reduce the interference of the regularization term on training; The Sobolev edge regularization loss term is preconditionally weighted with noise scale uniformity and works in conjunction with the noise gating function g(σ).
8. The multiphysics TCAD proxy modeling method according to claim 7, characterized in that, Step 3 further includes: employing a budgeted weighting strategy: given the total budget ratio of Sobolev edge regularization. And set the internal split ratio for the two losses, and dynamically update L. H1 and The weights are determined, and a correction factor γ(σ) is introduced to compensate for small-batch statistical fluctuations.
Citation Information
Patent Citations
TCAD simulation calibration method of SOI field effect transistor
CN102184879A
Training method of flow field reconstruction model
CN117094220A