Semiconductor device simulation solving method based on physical information neural network
By adopting a spatiotemporal decoupling format in physical information neural networks, decoupling into spatial modal subnets and temporal dynamic subnets, the problems of low computing efficiency, poor convergence and insufficient interpretation in the prior art are solved, and a more efficient and interpretable semiconductor device simulation is achieved.
Patent Information
- Application Number
- CN202510197042.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing physical information neural networks have low computational efficiency and poor convergence in semiconductor device simulation, and the network architecture is insufficiently interpretable, making it difficult to meet the complex needs of TCAD simulation.
The space-time decoupling format is adopted to decouple the neural network into a spatial modal subnet and a temporal dynamic subnet. The spatial modal subnet is used to learn orthogonal spatial modality, and the temporal dynamic subnet is used to learn the dynamic evolution of time coefficients, and the solution of partial differential equations is reconstructed through the POD decomposition format.
It improves the interpretability and convergence efficiency of the network, obtains higher simulation accuracy with fewer iterations, and enhances the practical value and reliability of TCAD simulation.
Smart Images

Figure CN120124461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer science and artificial intelligence, and specifically refers to a method for simulating and solving semiconductor devices based on a physics-informed neural network. Background Art
[0002] Semiconductor devices are the cornerstone of modern information technology and play a crucial role in devices such as microprocessors, memories, and power devices. As the size of semiconductor devices continues to shrink and the process nodes continue to advance, the complexity of design and manufacturing also increases. This requires TCAD (Technology Computer-Aided Design) simulation tools to be able to more accurately predict device performance and optimize designs. However, traditional TCAD simulation methods are facing bottlenecks when dealing with the increasingly complex semiconductor design requirements. Traditional methods such as the finite difference method and the finite element method based on grid division, although effective in the early semiconductor device design, also have some long-standing limitations. With the reduction in the size and increase in the complexity of integrated circuits, these methods are increasingly difficult to meet the simulation requirements.
[0003] To address the limitations of traditional simulation methods, in recent years, the academic and industrial communities have actively explored applying artificial intelligence and machine learning technologies to semiconductor device simulation. Based on large-scale numerical simulation data, neural networks can learn the complex non-linear mapping relationship between the electrical and physical characteristics of devices and the simulation output, quickly capture the physical characteristics of semiconductor devices, and efficiently predict simulation results, providing a new TCAD simulation technology route for semiconductor design and avoiding the cumbersome iterative solution steps in traditional numerical methods. Based on traditional TCAD simulation methods, researchers have tried to use deep learning technologies for acceleration. For example, in a study, a convolutional neural network was applied to predict the initial potential field of a metal-oxide-semiconductor field-effect transistor under given bias conditions, and the predicted initial potential field was used for numerical solution to finally accelerate the TCAD simulation. In addition, in a recent study, a technology called a physics-informed neural network was proposed. This method combines physical prior knowledge and machine learning technologies, making the model prediction results consistent with physical models, thereby improving the accuracy of the model and reducing the amount of data required for training. In addition, another network architecture, the deep operator network, was designed to learn and approximate the infinite-dimensional partial differential equation solution operator mapping. By learning the solution operator mapping under different input conditions, it finally shows good generalization ability. In the latest study, the physics-informed neural network and the deep operator network were combined and applied to data-efficient neural network modeling and real-time physical analysis in the spatial domain of TCAD simulation, and preliminary results were achieved in the TCAD simulation of three-dimensional nanowire field-effect transistors.
[0004] Although the physics-informed neural network framework has demonstrated excellent effectiveness in the numerical simulation of partial differential equations, a large number of studies have also shown that such networks have the disadvantages of low computational efficiency and poor convergence, and the interpretability of the network architecture is still insufficient. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a network architecture with stronger interpretability and higher convergence efficiency to enhance its practical value and reliability in TCAD simulation, based on the semiconductor device simulation and solution method of the physics-informed neural network, aiming at the deficiencies proposed in the above background technology.
[0006] To solve the above technical problem, the technical solution provided by the present invention is: a semiconductor device simulation and solution method based on the physics-informed neural network, which includes the following steps:
[0007] S1. Selection of the spatio-temporal decoupling format;
[0008] S2. Design of the spatial modal sub-network;
[0009] S3. Design of the time dynamic sub-network;
[0010] S4. Integrate the sub-networks and give the loss function;
[0011] S5. Apply the spatio-temporal decoupled physics-informed neural network to the actual partial differential equation mathematical model and conduct simulation experiments.
[0012] Further, the design of the spatial modal sub-network includes the following steps:
[0013] S2.1. Basis function embedding layer;
[0014] S2.2. Rank-preserving mapper;
[0015] S2.3. Basis function weight generator based on the Stiefel manifold.
[0016] Further, in order to input a set of basis functions to the neural network, a basis function embedding layer needs to be designed for the basis function embedding layer. Let the input of the spatial modal sub-network be x ∈ Ω, where Ω is the spatial domain where x is located, then the representation of the basis function embedding layer is:
[0017] Φ(x) = [φ 1 (x) φ 2 (x) … φ i (x)].
[0018] Further, the specific definition of the rank-preserving linear layer of the rank-preserving mapper is shown in the following formula:
[0019]
[0020] Combining the rank-preserving linear layer and the Tanh activation function, taking a two-layer rank-preserving mapper as an example, the network can be expressed as:
[0021]
[0022] In this invention patent, the rank-preserving mapper acts on the output of S2.1. Therefore, the basis function after passing through the rank-preserving mapper is denoted as
[0023]
[0024] Furthermore, the basis function weight generator based on the Stiefel manifold includes the following steps:
[0025] Step 1: Calculate the inner product matrix G of linearly independent basis functions on x ∈ Ω:
[0026]
[0027] Step 2: Calculate the Cholesky decomposition of the inner product matrix G:
[0028]
[0029] Step 3: Generate a matrix belonging to the Stiefel manifold according to the Stiefel manifold parameterized weight Θ
[0030]
[0031] Step 4: Calculate the basis function weight W φ :
[0032]
[0033] Step 5: The spatial mode sub-network outputs orthogonal spatial modes:
[0034]
[0035] Furthermore, the designed time dynamics sub-network includes the following steps:
[0036] Step 1: In the time dynamics sub-network, first use a delay layer to expand the time input. The role of the delay layer is to learn a delay embedding, extending a single time input to multiple equally spaced discrete time steps to extrapolate the time series:
[0037] Delay(t) = [t, t + Δt,..., t + kΔt],
[0038] Step 2: Use a multi-layer perceptron with a Tanh activation function as the observation space mapper, denoted as MLP 1 , so as to learn an initial temporal dynamics g in the observation space from the extrapolated time series 0 (t):
[0039] g 0 (t) = MLP 1 (Delay(t)),
[0040] Step 3: Therefore, the temporal dynamics subnetwork uses another multi-layer perceptron with a Tanh activation function, denoted as MLP 2 , to learn a time-varying Koopman matrix
[0041]
[0042] Step 4: Use the third multi-layer perceptron MLP with a Tanh activation function 3 as the Krylov subspace mapper to learn a set of coordinates [c 1 , c 2 ,..., c l in the Krylov subspace, where l is a hyperparameter for adjusting the dimension of the Krylov subspace.
[0043] [c 1 , c 2 ,..., c l = LayerNorm(MLP 3 (Delay(t))),
[0044] Step 5: The output of the temporal dynamics subnetwork is as follows:
[0045]
[0046] Furthermore, the comprehensive subnetwork and the given loss function include the following steps:
[0047] Step 1: According to the POD decomposition format, use the spatial mode and temporal dynamics to reconstruct the output of the spatio-temporal decoupled physical information neural network:
[0048] NN(x, t) = NN Ψ (x)NN a (t),
[0049] Step 2: The final network will learn the optimization of a mean square error loss, and the definition of this loss function is as follows:
[0050] MSE = MSE 0 +MSEBC +MSE PDE 。
[0051] After adopting the above method, the present invention has the following advantages: The method of the present invention extends the idea of the POD decomposition method to the design of the physics-informed neural network, decouples the neural network into a spatial modal sub-network and a temporal dynamics sub-network. The spatial modal sub-network is used to learn a set of orthogonal spatial modes, while the temporal dynamics sub-network is used to learn the dynamic evolution of the corresponding temporal coefficients. In the design process of the spatial modal sub-network, a variety of new network modules such as the basis function embedding layer, the rank-preserving mapper, and the basis function weight generator based on the Stiefel manifold are proposed. In the temporal dynamics sub-network, a network structure based on the Koopman operator and the Krylov subspace is used. Finally, the solution of the partial differential equation is reconstructed using the POD decomposition format. The spatio-temporal decoupled physics-informed neural network proposed in this invention realizes the learning of a set of orthogonal modal decompositions of the solution of the partial differential equation under unsupervised conditions, thereby ensuring the interpretability of the network and obtaining higher simulation accuracy with fewer iterations. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a schematic flow chart of a semiconductor device simulation and solution method based on a physics-informed neural network.
[0053] Figure 2 is a schematic structural diagram of a semiconductor device simulation and solution method based on a physics-informed neural network. DETAILED DESCRIPTION OF THE INVENTION
[0054] Although the physics-informed neural network framework in the prior art has shown excellent effectiveness in the numerical simulation of partial differential equations, a large number of studies have also shown that such networks have the disadvantages of low computational efficiency and poor convergence, and the interpretability of the network architecture is still insufficient. Therefore, this invention aims to start from the theoretical basis and combine neural network technology to realize a network architecture with stronger interpretability and higher convergence efficiency, so as to enhance its practical value and reliability in TCAD simulation.
[0055] The following further details the present invention with reference to the accompanying drawings.
[0056] Combined with the attached Figure 1-2 , in order to achieve the above object, the present invention adopts the following steps:
[0057] S1. Selection of spatio-temporal decoupling format
[0058] In the technical solution described in this patent, it is assumed that the general format of the mathematical model of the partial differential equation to be solved is where the solution of the partial differential equation is represents a certain linear or non - linear operator, and \(u\) t is the partial derivative of \(u(x,t)\) with respect to the time variable , \(u\) x and \(u\) xx are the partial derivative and the second - order partial derivative of \(u(x,t)\) with respect to the space variable respectively, and so on. \(\beta\) is the parameter included in the mathematical model of the partial differential equation.
[0059] For the solution of the partial differential equation, its complexity usually comes from the coupling of space and time. For most of the dynamic systems described by partial differential equations, proper orthogonal decomposition (POD) is an effective and natural way of spatio - temporal decoupling, and its decomposition format is as follows:
[0060]
[0061] where \(r\) represents the rank of the POD decomposition, the spatial mode \(\psi\) i (x) describes the main spatial characteristics of the solution, which is usually related to the physical structure of the problem, such as wave patterns, symmetries, etc. The time coefficient \(a\) i (t) correspondingly describes the dynamic changes of these spatial characteristics over time. And the spatial modes should satisfy orthogonality, that is, \(\langle\psi\) i , \(\psi\) j \rangle = 0\), \(i\neq j\), \(\langle\psi\) i , \(\psi\) j \rangle\) represents the inner product of \(\psi\) i and \(\psi\) j .
[0062] S2. Design the spatial mode sub - network
[0063] In this patent, the neural network for solving the mathematical model of the partial differential equation is decoupled into a spatial mode sub - network and a time - dynamics sub - network. The spatial mode sub - network is used to learn a set of orthogonal spatial modes, while the time - dynamics sub - network is used to learn the dynamic evolution of the corresponding time coefficients.
[0064] S2.1. Basis function embedding layer
[0065] In order to learn a set of orthogonal spatial modes, this invention patent proposes a scheme for generating orthogonal modes using linearly independent basis functions. First, in order to input a set of basis functions to the neural network, a basis function embedding layer needs to be designed. Let the input of the spatial mode sub - network be \(x\in\Omega\), where \(\Omega\) is the spatial domain where \(x\) is located, then the representation of the basis function embedding layer is:
[0066] \(\varPhi(x)=[\varphi\) 1 (x)\ \varphi\) 2 (x)\ \cdots\ \varphi\) i (x)]
[0067] Among them, {φ 1 , φ 2 , …, φ i} is a set of linearly independent basis functions. In this invention patent, the basis functions are selected as:
[0068]
[0069] Among them, α i is an adjustable parameter. As long as α i ≠α j , the basis functions Φ i (x) and Φ j (x) are linearly independent. Usually, α i can be selected as the grid vertex positions after equally dividing the spatial domain Ω.
[0070] S2.2, Rank-Preserving Mapper
[0071] Since the basis function embedding layer is an unlearnable layer, in order to improve the diversity of basis function selection, reduce the subjectivity of manual selection of basis functions, and thus enhance the generalization ability of the spatial modal sub-network. This invention patent proposes a learnable rank-preserving mapper, which ensures that the input data will not introduce linear correlation after passing through the rank-preserving mapper while introducing a learnable network structure, that is, the rank-preserving mapper will not change the rank of the input data. This is crucial for the design of the subsequent Stiefel manifold generator.
[0072] The difference between the rank-preserving mapper and the multi-layer perceptron used in traditional physics-informed neural networks is that the rank-preserving mapper uses a special rank-preserving linear layer to replace the traditional linear layer. The specific definition of the rank-preserving linear layer is shown in Equation (1):
[0073]
[0074] Among them, is the network input with dimension m, is the learnable weight vector in the rank-preserving linear layer, is the learnable diagonal weight matrix, is the identity matrix with dimension m. Using the rank-preserving linear layer can ensure that the input x will only undergo one orthogonal transformation and scale scaling, without introducing linear correlation. At the same time, the rank-preserving mapper uses the Tanh function as the non-linear activation function. Since the Tanh function is a bijective function, the Tanh function is invertible, and thus it is also a rank-preserving operation.
[0075] Combining the rank-preserving linear layer and the Tanh activation function, taking a two-layer rank-preserving mapper as an example, this network can be expressed as Equation (2):
[0076]
[0077] In this invention patent, the rank-preserving mapper acts on the output of S2.1, so the basis functions after passing through the rank-preserving mapper are denoted as
[0078]
[0079] S2.3, Basis function weight generator based on Stiefel manifold
[0080] After S2.1 and S2.2, the network obtains a set of linearly independent basis functions. In order to further construct a set of orthogonal spatial modes, a set of weights of the basis functions needs to be learned where n is the dimension of the solution u of the partial differential equation mathematical model, and k is the number of basis functions. Then the spatial mode constructed according to the basis function is:
[0081]
[0082] Since the spatial mode z must satisfy orthogonality, the spatial mode z belongs to the Stiefel manifold, that is
[0083] In this invention patent, a parameterization method of the Stiefel manifold is proposed, as shown in Equation (3). Thus, the basis function weight generator is constructed as follows:
[0084] Step 1: Calculate the inner product matrix G of the linearly independent basis functions on x ∈ Ω:
[0085]
[0086] Step 2: Calculate the Cholesky decomposition of the inner product matrix G:
[0087]
[0088] where U is an upper triangular matrix.
[0089] Step 3: Generate a matrix belonging to the Stiefel manifold according to the Stiefel manifold parameterized weight Θ
[0090]
[0091] where MatExp is the matrix exponential function, and LayerNorm is the layer normalization operation. The LayerNorm operation is an optional item to improve the convergence speed and numerical stability of the model.
[0092] Step 4: Calculate the basis function weights W φ :
[0093]
[0094] where denotes extracting the submatrix composed of the first r columns.
[0095] Step 5: The spatial modal subnetwork outputs orthogonal spatial modes:
[0096]
[0097] S3. Design the time dynamics subnetwork
[0098] In this patent, the time dynamics subnetwork is used to learn the time dynamics characteristics of the solution of the partial differential equation mathematical model, that is, it describes the evolution of orthogonal spatial modes over time. According to the Koopman operator theory, for the time dynamics a(t), there exists an observation space g(a(t)) and a Koopman operator such that And the Koopman operator can be approximated by a Koopman matrix. For the non-autonomous partial differential equation mathematical model, simply learning a time-varying Koopman matrix through a neural network has poor effects. Therefore, this invention patent designs a new learning strategy: using the approximate Koopman matrix to generate a Krylov sequence and performing time dynamics learning in the new Krylov subspace to capture richer dynamic characteristics.
[0099] In the time dynamics subnetwork, first, a delay layer is used to expand the time input. The role of the delay layer is to learn a delay embedding, expanding a single time input into multiple equally spaced discrete time steps to extrapolate the time series:
[0100] Delay(t) = [t, t + Δt,..., t + kΔt],
[0101] where the time increment Δt is a learnable weight, and the delay times k is a hyperparameter.
[0102] Subsequently, a multi-layer perceptron with a Tanh activation function is used as the observation space mapper, denoted as MLP 1 , so as to learn an initial time dynamics g 0 (t) in the observation space according to the extrapolated time series:
[0103] g 0 (t) = MLP 1 (Delay(t)),
[0104] Therefore, the time-dynamic sub-network uses another multi-layer perceptron with a Tanh activation function, denoted as MLP 2 , to learn a time-varying Koopman matrix
[0105]
[0106] where Reshape is a reshaping operation used to convert the vector output of MLP 2 into a Koopman matrix
[0107] Subsequently, a third multi-layer perceptron MLP with a Tanh activation function is used 3 as a Krylov subspace mapper to learn a set of coordinates [c 1 , c 2 ,..., c l in a Krylov subspace, where l is a hyperparameter that adjusts the dimension of the Krylov subspace
[0108] [c 1 , c 2 ,..., c l = LayerNorm(MLP 3 (Delay(t))),
[0109] Finally, the output of the time-dynamic sub-network is
[0110]
[0111] where MLP 4 acts as an inverse observer to map the time dynamics in the observation space back to the original solution space
[0112] S4. Integrate the sub-networks and present the loss function
[0113] According to the POD decomposition format, the output of the spatio-temporal decoupled physical information neural network is reconstructed using spatial modes and time dynamics
[0114] NN(x, t) = NN Ψ (x)NN a (t)
[0115] Finally, the network will learn the optimization of a mean squared error loss, and the definition of this loss function is as follows
[0116] MSE = MSE 0 + MSE BC + MSE PDE
[0117] where
[0118]
[0119] Mean square error representing the initial conditions, N 0 is the number of training samples for the initial conditions;
[0120]
[0121] Mean square error representing the boundary conditions, N BC is the number of training samples for the boundary conditions;
[0122]
[0123] Mean square error representing the partial differential equation, N PDE is the number of unsupervised training samples; NN t is the partial derivative of NN(x, t) with respect to the time variable of, NN x and NN xx are the partial derivative and the second - order partial derivative of NN(x, t) with respect to the spatial variable respectively, and so on.
[0124] S5. Apply the spatio - temporal decoupled physics - informed neural network to the actual partial differential equation mathematical model and conduct simulation experiments
[0125] Combining the steps described above, a spatio - temporal decoupled physics - informed neural network is formed. Under the given partial differential equation mathematical model, initial conditions, and boundary conditions, apply the neural network described above to the simulation solution of the mathematical model for experiments, verify the feasibility and superiority of the proposed algorithm, and collect and analyze the experimental data.
[0126] The content of the present invention will be further described below in conjunction with an implementation example. This implementation example is based on the benchmark partial differential equation mathematical model - the nonlinear Schrödinger equation and its publicly available dataset, and simulation experiments are carried out. The nonlinear Schrödinger equation is as follows:
[0127]
[0128] where the solution of the nonlinear Schrödinger equation is h, h t is the partial derivative of h with respect to time, h xx is the second - order partial derivative of h with respect to the position x. For this implementation example, the initial conditions of the nonlinear Schrödinger equation are set as:
[0129] h(x, 0) = 2sech(x),
[0130] The boundary conditions are set to two periodic boundary conditions as follows:
[0131] h(t, -5) = h(t, 5),
[0132] h x (t, -5) = h x (t, 5),
[0133] In this embodiment, 50 initial condition training samples, 50 boundary condition training samples, and 20,000 grid node positions are collected as the input for unsupervised learning. In terms of network settings, for the spatial modal sub-network, the number of output modes is set to 10, the number of basis functions in the basis function embedding layer is set to 30, the network depth of the rank-preserving mapper is set to 6, and each layer has 30 hidden neurons. For the Stiefel manifold parameterized weight Θ, it is correspondingly set to a 30×30 weight matrix. For the time dynamics sub-network, the delay times of the delay layer are set to 10, the dimension of the Koopman matrix is 10×10, the dimension of the Krylov subspace is 15, and the observation space mapper, Koopman matrix mapper, Krylov subspace mapper, and inverse observer are set to multi-layer perceptrons with 3 hidden layers and 20 hidden neurons. Finally, this embodiment uses the Adam optimizer with a learning rate of 0.001, and the number of iterations for network training is set to 10,000 times.
[0134] S1. Selection of the spatio-temporal decoupling format
[0135] In the technical solution described in this patent, it is assumed that the general format of the partial differential equation mathematical model to be solved is where the solution of the partial differential equation is represents a certain linear or non-linear operator, u t is the partial derivative of u(x, t) with respect to the time variable , u x and u xx are the partial derivative and the second-order partial derivative of u(x, t) with respect to the spatial variable respectively, and so on. β is the parameter included in the partial differential equation mathematical model.
[0136] For the solution of the partial differential equation, its complexity usually comes from the coupling of space and time. For most dynamic systems described by partial differential equations, proper orthogonal decomposition (POD) is an effective and natural spatio-temporal decoupling method, and its decomposition format is as follows:
[0137]
[0138] where r represents the rank of the POD decomposition, and the spatial mode ψ i(x) describes the main spatial features of understanding, usually related to the physical structure of the problem, such as wave patterns, symmetries, etc. The time coefficient a i (t) then correspondingly describes the dynamic changes of these spatial features over time. Moreover, the spatial modes should satisfy orthogonality, that is, <ψ i , ψ j > = 0, i ≠ j, <ψ i , ψ j > represents the inner product of ψ i and ψ j .
[0139] S2. Design of the spatial mode sub-network
[0140] In this patent, the neural network for solving the partial differential equation mathematical model is decoupled into a spatial mode sub-network and a time dynamics sub-network. The spatial mode sub-network is used to learn a set of orthogonal spatial modes, while the time dynamics sub-network is used to learn the dynamic evolution of the corresponding time coefficients.
[0141] S2.1. Basis function embedding layer
[0142] To learn a set of orthogonal spatial modes, this invention patent proposes a scheme for generating orthogonal modes using linearly independent basis functions. First, to input a set of basis functions into the neural network, a basis function embedding layer needs to be designed. Let the input of the spatial mode sub-network be x ∈ Ω, where Ω is the spatial domain where x is located, then the representation of the basis function embedding layer is:
[0143] Φ(x) = [φ 1 (x) φ 2 (x) … φ i (x)]
[0144] where {φ 1 , φ 2 , …, φ i} is a set of linearly independent basis functions. In this invention patent, the basis functions are selected as:
[0145]
[0146] where α i is an adjustable parameter. As long as α i ≠ α j , then the basis functions φ i (x) and φ j (x) are linearly independent. Usually, α i can be selected as the grid vertex positions after equally dividing the spatial domain Ω.
[0147] S2.2. Rank-preserving mapper
[0148] Since the basis function embedding layer is a non-learnable layer, in order to improve the diversity of basis function selection and reduce the subjectivity of manual selection of basis functions, thereby enhancing the generalization ability of the spatial modal sub-network. This invention patent proposes a learnable rank-preserving mapper, which ensures that the input data will not introduce linear correlation after passing through the rank-preserving mapper while introducing a learnable network structure, that is, the rank-preserving mapper will not change the rank of the input data. This is crucial for the subsequent design of the Stiefel manifold generator.
[0149] The difference between the rank-preserving mapper and the multi-layer perceptron used in traditional physics-informed neural networks is that the rank-preserving mapper uses a special rank-preserving linear layer to replace the traditional linear layer. The specific definition of the rank-preserving linear layer is shown in Equation (1):
[0150]
[0151] where, is the network input with dimension m, is the learnable weight vector in the rank-preserving linear layer, is the learnable diagonal weight matrix, is the identity matrix with dimension m. Using the rank-preserving linear layer, it can be ensured that the input x will only undergo one orthogonal transformation and scale scaling, and will not introduce linear correlation. At the same time, the rank-preserving mapper uses the Tanh function as the non-linear activation function. Since the Tanh function is a bijective function, the Tanh function is invertible, so it is also a rank-preserving operation.
[0152] Combining the rank-preserving linear layer and the Tanh activation function, taking a two-layer rank-preserving mapper as an example, the network can be expressed as Equation (2):
[0153]
[0154] In this invention patent, the rank-preserving mapper acts on the output of S2.1, so the basis function after passing through the rank-preserving mapper is denoted as
[0155]
[0156] S2.3. Basis Function Weight Generator Based on Stiefel Manifold
[0157] After S2.1 and S2.2, the network obtains a set of linearly independent basis functions. In order to further construct a set of orthogonal spatial modes, a set of weights of the basis functions needs to be learned where n is the dimension of the solution u of the partial differential equation mathematical model, and k is the number of basis functions. Then the spatial mode constructed according to the basis function is:
[0158]
[0159] Since the spatial mode z must satisfy orthogonality, the spatial mode z belongs to the Stiefel manifold, that is
[0160] In this invention patent, a parameterization method of the Stiefel manifold is proposed, as shown in Equation (3). Thus, a basis function weight generator as shown below is constructed:
[0161] Step 1: Calculate linearly independent basis functions The inner product matrix G on x ∈ Ω:
[0162]
[0163] Step 2: Calculate the Cholesky decomposition of the inner product matrix G:
[0164]
[0165] where U is an upper triangular matrix.
[0166] Step 3: Generate a matrix belonging to the Stiefel manifold according to the Stiefel manifold parameterized weight W s ,
[0167]
[0168] where MatExp is the matrix exponential function and LayerNorm is the layer normalization operation. The LayerNorm operation is an optional item to improve the convergence speed and numerical stability of the model.
[0169] Step 4: Calculate the basis function weight W φ :
[0170]
[0171] where denotes taking out the submatrix composed of the first r columns.
[0172] Step 5: The spatial mode subnetwork outputs orthogonal spatial modes:
[0173]
[0174] S3. Design the time dynamic subnetwork
[0175] In this patent, the time-dynamic sub-network is used to learn the time-dynamic characteristics of the solution of the partial differential equation mathematical model, that is, it describes the evolution of orthogonal spatial modes over time. According to the Koopman operator theory, for the time dynamics a(t), there exists an observation space g(a(t)) and a Koopman operator such that The Koopman operator can be approximated by a Koopman matrix. For a non-autonomous partial differential equation mathematical model, simply learning a time-varying Koopman matrix through a neural network has poor effects. Therefore, this invention patent designs a new learning strategy: using the approximate Koopman matrix to generate a Krylov sequence and performing time dynamics learning in the new Krylov subspace to capture richer dynamic characteristics.
[0176] In the time-dynamic sub-network, first, a delay layer is used to expand the time input. The role of the delay layer is to learn a delay embedding, extending a single time input to multiple equally spaced discrete time steps to extrapolate the time series:
[0177] Delay(t) = [t, t + Δt,..., t + kΔt],
[0178] where the time increment Δt is a learnable weight, and the delay times k is a hyperparameter.
[0179] Subsequently, a multi-layer perceptron with a Tanh activation function is used as the observation space mapper, denoted as MLP 1 , so as to learn an initial time dynamics g 0 (t) in the observation space according to the extrapolated time series:
[0180] g 0 (t) = MLP 1 (Delay(t)),
[0181] Therefore, the time-dynamic sub-network uses another multi-layer perceptron with a Tanh activation function, denoted as MLP 2 , to learn a time-varying Koopman matrix
[0182]
[0183] where Reshape is a deformation operation used to convert the vector output of MLP 2 into a Koopman matrix.
[0184] Immediately afterwards, a third multi-layer perceptron MLP with a Tanh activation function is used 3As a Krylov subspace mapper, it is used to learn the coordinates [c 1 , c 2 ,..., c l in a set of Krylov subspaces, where l is a hyperparameter for adjusting the dimension of the Krylov subspace.
[0185] [c 1 , c 2 ,..., c l = LayerNorm(MLP 3 (Delay(t))),
[0186] Finally, the output of the time dynamics subnetwork is:
[0187]
[0188] where MLP 4 acts as an inverse observer to map the time dynamics in the observation space back to the original solution space.
[0189] S4. Integrate the subnetwork and give the loss function
[0190] According to the POD decomposition format, use the spatial modes and time dynamics to reconstruct the output of the spatio-temporal decoupled physical information neural network:
[0191] NN(x, t) = NN Ψ (x)NN a (t)
[0192] Finally, the network will learn the optimization of a mean square error loss, and the definition of this loss function is as follows:
[0193] MSE = MSE 0 + MSE BC + MSE PDE
[0194] where
[0195]
[0196] represents the mean square error of the initial conditions, and N 0 is the number of training samples of the initial conditions;
[0197]
[0198] represents the mean square error of the boundary conditions, and N BC is the number of training samples of the boundary conditions;
[0199]
[0200] Represents the mean square error of the partial differential equation, N PDE is the number of unsupervised training samples; NN t is the partial derivative of NN(x, t) with respect to the time variable and NN x and NN xx are the partial derivative and the second-order partial derivative of NN(x, t) with respect to the spatial variable respectively, and so on.
[0201] S5. Apply the spatio-temporal decoupled physics-informed neural network to the actual partial differential equation mathematical model and conduct simulation experiments
[0202] Combining the steps described above, a spatio-temporal decoupled physics-informed neural network is formed. Under the given partial differential equation mathematical model, initial conditions, and boundary conditions, apply the neural network described above to the simulation solution of this mathematical model for experiments to verify the feasibility and superiority of the proposed algorithm, and collect and analyze the experimental data. In the simulation experiment of this implementation example, the neural network model is implemented using PyTorch and trained on a single NVIDIA A100 GPU. The experimental results are obtained using the software developed by ourselves.
[0203] The above describes the present invention and its implementation manners. This description is not restrictive, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and without departing from the purpose of the present invention, they design similar structural forms and embodiments to this technical solution without creative efforts, which should all fall within the protection scope of the present invention.
Claims
1. A semiconductor device simulation solution method based on physical information neural network, characterized in that: It includes the following steps: S1. Selection of spatiotemporal decoupling format; S2. Design spatial modal sub-network; S3. Design time dynamic subnetwork; S4. Integrate the sub-network and give the loss function; S5. Apply the space-time decoupled physical information neural network to the actual partial differential equation mathematical model and conduct simulation experiments.
2. The semiconductor device simulation solution method based on physical information neural network according to claim 1 is characterized in that: The design space modal subnetwork comprises the following steps: S2.1, basis function embedding layer; S2.2, rank-preserving mapper; S2.
3. Basis function weight generator based on Stiefel manifold.
3. The semiconductor device simulation solution method based on physical information neural network according to claim 2 is characterized in that: In order to input a set of basis functions into the neural network, a basis function embedding layer needs to be designed. Assuming that the input of the spatial modal subnetwork is x∈Ω, Ω is the spatial domain where x is located, the basis function embedding layer is expressed as: Φ(x)=[φ1(x) φ2(x) … φ i (x)].
4. The semiconductor device simulation solution method based on physical information neural network according to claim 3 is characterized in that: The specific definition of the rank-preserving linear layer of the rank-preserving mapper is shown in the following formula: Combining the rank-preserving linear layer and the Tanh activation function, taking a two-layer rank-preserving mapper as an example, the network can be expressed as: In the present invention, the rank-preserving mapper acts on the output of S2.1, so the basis function after the rank-preserving mapper is recorded as 5. The semiconductor device simulation solution method based on physical information neural network according to claim 4 is characterized in that: The Stiefel manifold-based basis function weight generator comprises the following steps: Step 1: Compute linearly independent basis functions The inner product matrix G on x∈Ω: Step 2: Calculate the Cholesky decomposition of the inner product matrix G: Step 3: Parameterize the weight Θ according to the Stiefel manifold and generate a matrix belonging to the Stiefel manifold Step 4: Calculate the basis function weight W φ : Step 5: The spatial modality subnetwork outputs orthogonal spatial modes:
6. The semiconductor device simulation solution method based on physical information neural network according to claim 1, characterized in that: The design of the time dynamic subnetwork comprises the following steps: Step 1: In the time dynamic subnetwork, first use the delay layer to expand the time input. The role of the delay layer is to learn a delay embedding, which expands a single time input to multiple equidistant time steps to extrapolate the time series: Delay(t)=[t, t+Δt,..., t+kΔt]; Step 2: Use a multi-layer perceptron with Tanh activation function as the observation space mapper, denoted as MLP1, to learn an initial time dynamics g0(t) in the observation space based on the extrapolated time series: g0(t)=MLP1(Delay(t)); Step 3: Therefore, the time-dynamic subnetwork uses another multi-layer perceptron with Tanh activation function, denoted as MLP2, to learn a time-varying Koopman matrix Step 4: Use the third multi-layer perceptron MLP3 with Tanh activation function as a Krylov subspace mapper to learn a set of coordinates [c1, c2, ..., c l ], l is a hyperparameter for adjusting the dimension of the Krylov subspace: [c1,c2,...,c l ]=LayerNorm(MLP3(Delay(t))); Step 5: The output of the time dynamic subnetwork is:
7. The semiconductor device simulation solution method based on physical information neural network according to claim 1 is characterized in that: The comprehensive sub-network and loss function include the following steps: Step 1: According to the POD decomposition format, use spatial modalities and temporal dynamics to reconstruct the output of the spatiotemporal decoupled physical information neural network: NN(x,t)=NN Ψ (x)NN a (t); Step 2: Finally, the network will learn the optimization of a mean square error loss, and the loss function is defined as follows: MSE=MSE0+MSE BC +MSE PDE 。