Method for improving physical information neural network training based on loss curved surface
By analyzing and improving the residual loss surface of PINNs and weighted optimization combined with boundary and initial conditional losses, the problem of poor convergence of PINNs during the optimization process is solved, and higher prediction accuracy and training stability are achieved.
Patent Information
- Application Number
- CN202510124318.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-16
AI Technical Summary
PINNs face complex loss surfaces during optimization, resulting in convergence problems during training, affecting the final performance of the model.
By conducting in-depth analysis of the loss surface of residual loss, an improved PINNs training method is designed, using a fully connected neural network to train boundary and initial conditional losses, and combined with residual loss, it is weighted to optimize the prediction accuracy of the model.
This method significantly improves the prediction accuracy of the model, reduces the relative L2 error, and improves the stability and convergence speed of training.
Smart Images

Figure CN120012862A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and the direction of deep learning, and invents a method for training improved physical information neural networks (PINNs) based on loss surfaces. Background Art
[0002] In recent years, with the rapid development of deep learning technology, the application of deep learning-based physics-informed neural networks (PINNs) in solving partial differential equations (PDEs) has gradually attracted widespread attention. PINNs use the powerful expressive power of neural networks to minimize the weighted sum of residual loss and boundary / initial condition loss through training, thereby effectively solving complex physical problems. The success of this method not only improves the modeling capabilities of various physical phenomena, but also provides new tools for solving important problems in many scientific fields such as fluid mechanics, medical imaging, and climate simulation. The solution to these practical problems not only helps to promote the progress of basic research, but also has a profound impact on industrial applications and social development.
[0003] Although PINNs have made significant progress in many fields, they still face many challenges in the optimization process. First, the training loss of PINNs usually contains multiple loss terms, which makes the loss surface extremely complex and brings difficulties to the optimization process. Understanding the structure of these loss surfaces is of great significance for identifying and solving the root causes of PINNs optimization problems. Once the characteristics and behaviors of the loss surface are deeply understood, more effective optimization strategies can be developed to improve the stability and convergence speed of training. Summary of the invention
[0004] The problem to be solved by the present invention is that although PINNs have significant advantages, in practice, researchers have found that their loss functions, especially the loss surface of the residual loss, are often complex, with multiple local minima and non-smooth characteristics. This complexity can lead to convergence problems in the training process, affecting the final performance of the model. Since traditional training methods perform poorly when faced with complex loss surfaces, it is extremely urgent to explore new training strategies to improve the convergence speed of PINNs and the accuracy of the solution. Based on an in-depth analysis of the loss surface of the residual loss, the present invention designs a method for improving PINNs training, which greatly improves the prediction accuracy of the model. The following are several important concepts related to PINNs and their mathematical expressions:
[0005] (1) Partial differential equations (PDEs)
[0006] Partial differential equations are equations that contain partial derivatives of unknown functions. They are widely used in science and engineering and are of the following form:
[0007] F(u(x,t))=0,x∈Ω,t∈[0,T];
[0008]
[0009] I(u(x,0))=0,x∈Ω,
[0010] Where u(x,t) is the solution of the partial differential equation, that is, the value to be solved, F is the differential operator, B is the boundary condition operator, and I is the initial condition operator. Ω is the spatial domain, is the boundary, and [0,T] is the time domain.
[0011] The partial differential equation used in the experiment of the present invention is the convection equation. The convection equation is mainly used to simulate fluid flow, heat transfer and biological processes. The form of the convection equation is as follows:
[0012]
[0013] u(0,t)=u(2π,t);
[0014] u(x,0)=sin(x);
[0015] Wherein, β is the convection coefficient, and the experimental value of the present invention is β=40.
[0016] (2) Physical Information Neural Networks (PINNs)
[0017] To solve partial differential equations, PINNs train a neural network to approximate u(x,t). res :=Ω×(0,T)Collection N res Internal points, in the boundary domain Collection N b boundary points, in the initial domain ζ i :=Ω Collection N i PINNs use the following loss function to enforce the PDEs constraints as well as the boundary and initial conditions:
[0018] L total (θ) = L res (θ)+L b (θ)+L i (θ),
[0019] Among them, θ is the weight parameter of the neural network. PINNs uses a fully connected neural network, so θ is the connection weight between different layers of the fully connected neural network. is the residual loss corresponding to the partial differential equation, is the boundary condition loss corresponding to the boundary condition, is the initial condition loss corresponding to the initial condition.
[0020] (3) Relative L2 error
[0021] The prediction accuracy of the experiment of the present invention is expressed as a relative L2 error, which is in the following form:
[0022]
[0023] in is the predicted solution of the neural network, and u(x,t) is the true solution.
[0024] The present invention is generally divided into three parts:
[0025] (1) First, for a certain partial differential equation, its boundaries and initial conditions can be changed. This is equivalent to the fact that during the training process of the neural network, the loss surface of the residual loss of the partial differential equation is fixed. By giving different constraints on boundaries and initial conditions, the neural network will be able to choose from the global minimum points of the loss surfaces of different residual losses.
[0026] So the residual loss L res Different global minimum points in the loss surface of (θ) correspond to the solutions of partial differential equations constrained by different boundary / initial conditions, and can form connected regions with low loss values in the loss surface.
[0027] (2) Use a fully connected neural network to calculate the boundary and initial condition loss L b (θ)+L i (θ) is trained. After more than 10,000 rounds of training, the neural network weight parameters at this time are recorded as model * .
[0028] (3) Loss of boundary and initial conditions L b (θ)+L i (θ) and residual loss L res (θ) is weighted and summed to obtain the total loss L total (θ) = L res (θ)+λ(L b (θ)+L i (θ)), using the same neural network from model * The total training loss L total (θ). where by giving L b (θ)+Li (θ) a large weight λ (λ is usually 1≤λ≤10 5 ,) fix the boundary and initial conditions, so that the residual loss L res (θ) enters the global minimum point of the residual loss corresponding to the current boundary and initial conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a complete flow chart of the present invention.
[0030] Figure 2 Form a connected region with low loss value for the global minimum point of the residual loss.
[0031] Figure 3 is a graph of the true solution to the convection equation.
[0032] Figure 4 This is an image of the predicted solution of the present invention.
[0033] Figure 5 This is an image of the predicted solution of a traditional physical information neural network. DETAILED DESCRIPTION
[0034] The present invention proposes a method for improving physical information neural network training based on loss surface. The experimental environment resources used are as follows. It runs under Windows 10 system and is implemented using python3.8.3 and pytorch1.7.0 framework. In order to accelerate the optimization of deep neural network, the present invention uses NVIDIAGTX 3080 (16G) GPU with CUDA version 11.4.
[0035] Step 1: First, for a certain partial differential equation, there is a unique solution under given boundaries and initial conditions, but its boundaries and initial conditions can be changed. This is equivalent to the fact that in the training process of the neural network, the loss surface of the residual loss of the partial differential equation is fixed. By giving different constraints on boundaries and initial conditions, the neural network will be able to choose from the global minimum points of the loss surfaces of different residual losses.
[0036] So the residual loss L res Different global minimum points in the loss surface of (θ) correspond to the solutions of partial differential equations constrained by different boundary / initial conditions, and can form connected regions with low loss values in the loss surface.
[0037] Without loss of generality, consider a residual loss L res (θ) has a global minimum with a loss of 0 at the weight parameter θ of the neural network, that is:
[0038] Lres (θ)=0,
[0039] Assume that δθ represents a weight parameter perturbation vector and ‖δθ‖≤10 -5 , that is, its norm is less than or equal to 10 -5 If , then:
[0040]
[0041] Among them, θ+δθ refers to the neural network weight parameter after perturbation from the neural network weight parameter θ along the weight parameter perturbation vector δθ. is the residual loss L res (θ) The first derivative of the weight parameter θ.
[0042] for It is always possible to find a weight parameter perturbation vector δθ perpendicular to it, such that Then, we have:
[0043] L res (θ+δθ)=L res (θ)=0.
[0044] That is, from the global minimum point of the neural network weight parameter θ, another global minimum point with a residual loss of 0 can be found along the direction of the weight parameter perturbation vector δθ. Therefore, it is theoretically proved that different global minimum points of residual loss are connected areas with low loss values in the neural network weight parameter space.
[0045] In order to study the connectivity between two different global minima of residual loss, let θ1 and θ2 represent the same neural network with different initialization methods, and minimize the residual loss L res (θ) The two sets of neural network weight parameters are obtained. Then, the connected piecewise linear curve is constructed as follows:
[0046]
[0047] 0≤t≤1,
[0048] 0≤i≤n,
[0049] Where t is the position parameter on the curve. n is the number of inflection points of the piecewise linear curve, and in the experiment, n=3. i is the position number of the inflection point or endpoint on the curve. w i and w i+1 is the neural network weight parameter at the inflection point or endpoint of the piecewise linear curve.
[0050] Then, θ1 and θ2 are the two endpoints of the piecewise linear curve, that is: Φ θ(0) = w0 = θ1, Φ θ (1) = w n+1 =θ2.
[0051] So Φ θ (t) is expressed as the neural network weight parameter at each position on the piecewise linear curve, so the residual loss at each position on the piecewise linear curve is:
[0052] L res (Φ θ (t)),
[0053] Then for the above residual loss L res (Φ θ (t)) finds the definite integral on the interval [0,1] of t, that is:
[0054] l(θ)=∫0 1 L res (Φ θ (t))dt,
[0055] Further, the residual loss L is obtained res (Φ θ (t)) is the expectation of a uniform distribution on the interval [0,1] with respect to t, that is:
[0056] l(θ)=E t~U(0,1) L res (Φ θ (t)),
[0057] Among them, U(0,1) is the uniform distribution on the interval [0,1], t~U(0,1) means t obeys the uniform distribution on the interval [0,1], and E represents expectation.
[0058] Then, the neural network is used to minimize the expected loss l(θ) with respect to the neural network weight parameters θ.
[0059] like Figure 2 , which is the global minimum point of the residual loss of the convection equation, will form a connected area with a low loss value. The two endpoints in the figure (the black dashed lines in each figure) represent the two global minimum points (the neural network weight parameters θ1 and θ2 mentioned above). It can be observed that these two global minimum points can be connected by a linear line segment (black solid line), and their loss is very low. It is worth noting that the connected area revealed above is different from the pattern connectivity in traditional deep learning, which mainly comes from the inherent invariance of neural networks, that is, in traditional deep learning, the output of the neural network will not change under certain weight parameter transformations.
[0060] The connected regions of low loss values formed by the global minimum of the residual loss are unfavorable to first-order optimizers (such as Adam). Despite the existence of boundary and initial condition losses, first-order optimizers can easily get stuck in these nearly flat regions and fail to find the true global minimum of the residual loss corresponding to the given boundary and initial conditions.
[0061] Step 2: Use a neural network to calculate the boundary and initial condition loss L b (θ)+L i (θ) is trained. After 10,000 rounds of training, the current neural network weight parameters are recorded as model * , that is, both boundary and initial condition losses have been successfully optimized.
[0062] Step 3: Loss L for boundary and initial conditions b (θ)+L i (θ) and residual loss L res (θ) is weighted and summed to obtain the total loss L total (θ) = L res (θ)+λ(L b (θ)+L i (θ)), using the same neural network from the weight parameters model * The total training loss L total (θ). In which, by giving L b (θ)+L i (θ) A large weight λ (values: 1, 10, 100, 1000) fixes the boundary / initial condition, so that the residual loss L res (θ) is optimized to the corresponding global minimum point.
[0063] The above four experiments with different λ values are carried out, and the best result is selected. As the value of λ increases, the control over the boundary and initial condition losses becomes stronger, which leads to the optimization of the fixed residual loss.
[0064] Step 4: Algorithm experiment process and result evaluation
[0065] The partial differential equation used in the experiment of the present invention is the convection equation, and the equation form has been described above.
[0066] The neural network used in the experiment of the present invention is a fully connected neural network with 5 layers, 100 neurons in each layer, and tanh as the activation function. The optimizer is Adam. The number of sampling points for training is the internal point N res =2000, boundary point N b =200, initial point N i =256.
[0067] For boundary and initial condition loss L b (θ)+L i (θ), the randomly initialized fully connected neural network is optimized and trained using the Adam optimizer. The learning rate is set to 0.001 and the weight decay is set to 0.0005. After 10,000 rounds of training, the network weight parameters at this time are recorded as model * , the training loss is R(θ * )=0.000074.
[0068] According to step 3, the boundary and initial condition loss L b (θ)+L i (θ) and residual loss L res (θ) is weighted, and the total loss is L total (θ) = L res (θ)+λ(L b (θ)+L i (θ)), the λ in this experiment is set to 1, 10, 100, and 1000, which represent the different degrees of fixed boundary and initial conditions. total (θ) from the network weight parameter model * The Adam optimizer was used for optimization training at the beginning, and the learning rate and weight decay remained unchanged. After 20,000 rounds of training, the loss and relative L2 error were as follows:
[0069] λ <![CDATA[Loss L total (θ)]]> <![CDATA[Relative L2 error]]> 1 1.8E-3 6.6E-2 10 1.2E-3 4.1E-2 100 2.1E-2 1.3E-1 1000 7.1E-1 7.6E-1
[0070] It can be found that when λ is 10, the loss L total (θ) and the relative L2 error are both the lowest, so the relative L2 error at this time is recorded as the result of the present invention. Figure 3 is a graph of the true solution to the convection equation, Figure 4 is the predicted solution image of the method of the present invention, Figure 5 This is the image of the predicted solution of the traditional physical information neural network. It can be seen that the image of the method of the present invention is closer to the image of the real solution.
[0071] Through the experimental process and result evaluation, it can be proved that the method designed in the present invention can obtain the global minimum point of the residual loss with lower relative L2 error in the process of optimizing the physical information neural network, thereby reducing the prediction error of the physical information neural network.
Claims
1. A method for improving physical information neural network training based on loss surface, characterized in that: (1) Partial differential equations (PDEs) Partial differential equations are equations that contain partial derivatives of unknown functions. They are widely used in science and engineering and are of the following form: F(u(x,t))=0,x∈Ω,t∈[0,T]; I(u(x,0))=0,x∈Ω, Where u(x, t) is the solution of the partial differential equation, that is, the value to be solved, F is the differential operator, B is the boundary condition operator, I is the initial condition operator; Ω is the spatial domain, is the boundary, [0,T] is the time domain; The partial differential equation used is the convection equation; the convection equation has the following form: u(0,t)=u(2π,t); u(x,0)=sin(x); Where, β is the convection coefficient, β = 40; (2) Physical Information Neural Networks (PINNs) PINNs trains a neural network to approximate u(x,t); res :=Ω×(0,T)Collection N res Internal points, in the boundary domain Collection N b boundary points, in the initial domain ζ i :=Ω Collection N i initial points; PINNs use the following loss function to enforce PDEs constraints as well as boundary and initial conditions: L total (θ)=L res (θ)+L b (θ)+L i (i), Among them, θ is the weight parameter of the neural network. PINNs uses a fully connected neural network, so θ is the connection weight between different layers of the fully connected neural network. is the residual loss corresponding to the partial differential equation, is the boundary condition loss corresponding to the boundary condition, is the initial condition loss corresponding to the initial condition; (3) Relative L2 error The prediction accuracy of the experiment is expressed in terms of relative L2 error, which is in the following form: in is the predicted solution of the neural network, and u(x,t) is the true solution; Step 1: Residual Loss L res Different global minimum points in the loss surface of (θ) correspond to the solutions of partial differential equations constrained by different boundary / initial conditions, and can form connected regions with low loss values in the loss surface; Consider a residual loss L res (θ) has a global minimum with a loss of 0 at the weight parameter θ of the neural network, that is: L res (θ)=0, Let δθ represent a weight parameter perturbation vector, and ‖δθ‖≤10 -5 , that is, its norm is less than or equal to 10 -5 If , then: Among them, θ+δθ refers to the neural network weight parameter after perturbation from the neural network weight parameter θ along the weight parameter perturbation vector δθ. is the residual loss L res (θ) the first derivative of the weight parameter θ; for It is always possible to find a weight parameter perturbation vector δθ perpendicular to it, such that Then, we have: L res (θ+δθ)=L res (θ)=0; That is, from the global minimum point of the neural network weight parameter θ, find another global minimum point with a residual loss of 0 along the direction of the weight parameter perturbation vector δθ; Let θ1, θ2 represent the same neural network with different initialization methods, and minimize the residual loss L res (θ) The two sets of neural network weight parameters are obtained; then, the connected piecewise linear curve is constructed in the following form: 0≤t≤1, 0≤i≤n, Where t is the position parameter on the curve; n is the number of inflection points of the piecewise linear curve, n = 3; i is the position number of the inflection point or endpoint on the curve; w i and w i+1 is the neural network weight parameter at the inflection point or endpoint of the piecewise linear curve; Then, θ1 and θ2 are the two endpoints of the piecewise linear curve, that is: Φ θ (0) = w0 = θ1, Φ θ (1) = w n+1 =θ2; So Φ θ (t) is expressed as the neural network weight parameter at each position on the piecewise linear curve, so the residual loss at each position on the piecewise linear curve is: L res (F θ (t)), Then for the above residual loss L res (Φ θ (t)) finds the definite integral on the interval [0,1] of t, that is: l(θ)=∫0 1 L res (Φ θ (t))dt, The residual loss L is obtained res (Φ θ (t)) is the expectation of a uniform distribution on the interval [0,1] with respect to t, that is: l(θ)=E t~U(0,1) L res (F θ (t)), Among them, U(0,1) is a uniform distribution on the interval [0,1], t~U(0,1) means t obeys a uniform distribution on the interval [0,1], and E represents expectation; Then, use the neural network to minimize the expected loss l(θ) with respect to the neural network weight parameters θ; Step 2: Use a neural network to calculate the boundary and initial condition loss L b (θ)+L i (θ) is trained. After more than 10,000 rounds of training, the current neural network weight parameters are recorded as model * , that is, both boundary and initial condition losses have been successfully optimized; Step 3: Loss L for boundary and initial conditions b (θ)+L i (θ) and residual loss L res (θ) is weighted and summed to obtain the total loss L total (θ) = L res (θ)+λ(L b (θ)+L i (θ)), using the same neural network from the weight parameters model * The total training loss L total (θ); where, by giving L b (θ)+L i (θ) A weight λ fixes the boundary / initial condition, so that the residual loss L res (θ) is optimized to the corresponding global minimum point; Conduct the above experiments with different λ values and select the best one for the results; λ is 1≤λ≤10 5 ; The selected neural network is a fully connected neural network with 5 layers, 100 neurons in each layer, tanh activation function, and Adam optimizer. For boundary and initial condition loss L b (θ)+L i (θ), the randomly initialized fully connected neural network is optimized and trained using the Adam optimizer. The learning rate is set to 0.001 and the weight decay is set to 0.0005. After training for more than 10,000 rounds, the network weight parameters at this time are recorded as model * .
Citation Information
Cited By
PINN proxy model construction method based on imprecise sparse sampling information
CN121480260A