Learning device, simulation device, learning method, simulation method, and program
The learning device addresses the challenge of simulating unknown physical phenomena by estimating solution operator parameters with an objective function that includes energy conservation, achieving high-precision simulations without relying on known differential equations.
Patent Information
- Application Number
- JP2023189489
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing operator learning methods for simulating physical phenomena assume known differential equations, leading to decreased simulation accuracy for unknown physical phenomena.
A learning device that uses data composed of input and output functions representing physical phenomena, estimates parameters of a solution operator to minimize an objective function that includes a loss term for reproducibility and a regularization term for energy conservation.
Enables high-precision simulation of physical phenomena even when the differential equation is unknown, preventing overfitting by utilizing the law of conservation of energy as prior knowledge.
Smart Images

Figure 2025077362000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a learning device, a simulation device, a learning method, a simulation method, and a program.
Background Art
[0002] With the development and spread of sensing technology, various information in the real world is being accumulated as data. Attention has been focused on realizing the discovery of new natural laws, future prediction, etc. by analyzing physical phenomena by utilizing such data.
[0003] A pair of an observation time and a physical quantity is regarded as one sample, and its time series is regarded as observation data. At this time, operator learning has been proposed as a method for efficiently simulating physical phenomena by using machine learning technology (Non-Patent Documents 1 and 2). In addition, operator learning that takes into account differential equations describing physical phenomena has also been proposed in addition to the observation data (for example, Non-Patent Documents 3 and 4).
[0004] In the operator learning proposed in Non-Patent Documents 1 and 2, by learning the mapping (operator) from an input function (for example, initial conditions, boundary conditions, external force, etc.) to an output function (solution of a differential equation) by a neural network, it is possible to efficiently simulate the solution when an arbitrary input function is given. However, the operator learning proposed in Non-Patent Documents 1 and 2 assumes that a large amount of observation data represented by a pair of an input function and an output function is available. For this reason, when a large amount of observation data cannot be used, overfitting to the observation data may occur, and the simulation accuracy for an unknown input function may decrease. On the other hand, in the operator learning proposed in Non-Patent Documents 3 and 4, by adding a loss related to the differential equation describing the physical phenomenon, even when a large amount of observation data cannot be used, a more accurate simulation is realized.
Prior Art Documents
Non-Patent Documents
[0005] [Non - Patent Document 1] Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. liu, K. Bhattacharya, A. Stuart, and A. Anandkumar. Fourier neural operator for parametric partial differential equations. In International Conference on Learning Representations, 2021. [Non - Patent Document 2] L. Lu, P. Jin, G. Pang, Z. Zhang, and G. E. Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3(3):218 - 229, 2021. [Non - Patent Document 3] Z. Li, H. Zheng, N. Kovachki, D. Jin, H. Chen, B. Liu, K. Azizzadenesheli, and A. Anandkumar. Physics - informed neural operator for learning partial differential equations. In arXiv, 2023. [Non - Patent Document 4] S. Wang, H. Wang, and P. Perdikaris. Learning the solution operator of parametric partial differential equations with physics - informed DeepONets. Science Advances, 7(40):eabi8605, 2021.
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the operator learning proposed in Non-Patent Documents 3 and 4, it is assumed that the differential equation is known. Therefore, the accuracy of the simulation decreases for unknown physical phenomena.
[0007] The present disclosure has been made in view of the above points, and provides a technique capable of realizing a high-precision simulation of physical phenomena.
Means for Solving the Problems
[0008] A learning device according to an aspect of the present disclosure uses learning data composed of a pair of an input function representing conditions of a physical phenomenon in a predetermined physical system and an output function representing a solution of a differential equation expressing the physical phenomenon, and at least estimates parameters of a solution operator representing a mapping from the input function to the output function so as to minimize a predetermined objective function. The objective function includes a loss for evaluating the reproducibility of the solution operator with respect to the learning data, and a regularization term for making the value of the output function predicted by the solution operator satisfy the energy conservation law when the input function is given.
Effects of the Invention
[0009] A high-precision simulation of physical phenomena can be realized.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Embodiment for Carrying Out the Invention
[0011] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0012] <Theoretical Configuration> Hereinafter, the theoretical configuration of the method proposed in this embodiment (hereinafter also referred to as the "proposed method") will be described. This proposed method is applicable to all physical phenomena. By introducing the theory of Hamiltonian mechanics, the physical law (law of conservation of energy) is utilized as prior knowledge for operator learning. As a result, in the proposed method, even when the differential equation describing the physical phenomenon is unknown, it is possible to learn and predict (simulate) the physical phenomenon with high accuracy from only the observed data while preventing overfitting.
[0013] T = R ≧0 is the time domain, and X ⊂ R D is the D-dimensional space domain. Here, R is the set of all real numbers, and in particular, R ≧0 is the set of all non-negative real numbers. At this time, consider a physical system defined in the spacetime domain Y = T × X.
[0014] Let two function spaces be A and U, the input function be a ∈ A, and the output function be u ∈ U. As the input function a, various things such as initial conditions, boundary conditions, external forces, etc. in a certain physical system are assumed. On the other hand, the output function u is the solution to the physical system (that is, the solution to the differential equation describing a certain physical phenomenon in the physical system). Hereinafter, the input function and the output function will be collectively referred to as the "input-output function". Note that the input function a is defined on the spacetime domain Y and is a function that outputs a vector representing initial conditions, boundary conditions, external forces, etc. in a certain physical system. Similarly, the output function u is defined on the spacetime domain Y and is a function that outputs a vector representing the solution to the physical system. However, the outputs of the input function a and the output function u may not be vectors, but may be, for example, scalars, matrices, tensors, etc.
[0015] The purpose of the proposed method is to learn a mapping from an input function to an output function (hereinafter, this mapping is referred to as a "solution operator") G: A → U from observed data. The solution operator satisfies the relationship u = G(a).
[0016] The proposed method has a "parameter estimation phase" for estimating various parameters and a "simulation phase" for simulating physical phenomena using the various parameters. Note that the parameter estimation phase may be referred to as a learning phase or the like, for example. Also, the simulation phase may be referred to as a test phase, a prediction phase, or the like, for example.
[0017] In the parameter estimation phase, it is assumed that I pairs of input-output functions {(a i , u i ) | i = 1, ···, I} are given as learning data. Let the independent variable of the input-output function be y ∈ Y, and y will be called a query. At this time, consider the situation where the values of the input-output function are observed at each of J discrete points (query set) Y' = {y j | j = 1, ···, J} ⊂ Y. That is, for each pair of input-output functions (a i , u i ) and each query y j ∈ Y', pairs of values of the input-output function (a i (y j ), u i (y j )) are obtained. Since the values of the input-output functions are observed for the I pairs of input-output functions {(a i , u i ) | i = 1, ···, I} given as learning data, the learning data may be called "observed data". Also, in this embodiment, for simplicity of notation, it is assumed that the values of all input-output functions are observed in the same query set, but this is just an example and is not limited to this. For example, different query sets may be used for each pair of input-output functions.
[0018] In the simulation phase, when an unknown input function (hereinafter also referred to as "test data") a that is not included in the learning data is given, an approximate solution u * is output using the learned solution operator. * =G(a * ). As a result, a highly accurate solution of the differential equation describing a certain physical phenomenon is obtained, and a high-precision simulation of the physical phenomenon is realized.
[0019] Hereinafter, a simulation device 10 that realizes learning of the solution operator and high-precision simulation of a physical phenomenon by the proposed method will be described.
[0020] Here, as an example, hereinafter, it is assumed that the solution operator G is realized by a neural network, and its parameter is θ. Also, when explicitly indicating the parameter θ, the solution operator G will be denoted as "G θ ". As an implementation of the neural network that realizes the solution operator G, any implementation can be used. For example, a standard multi-layer perceptron may be used, or more advanced methods such as Deep Operator Network (Non-Patent Document 2) or Fourier Neural Operator (Non-Patent Document 1) may be used. However, the fact that the solution operator G is realized by a neural network is just an example, and the solution operator G may be realized by a machine learning model other than a neural network.
[0021] Also, hereinafter, an energy function H(u):R N →R that expresses a physical law (law of conservation of energy) is introduced, and this energy function H is also assumed to be realized by a neural network, and its parameter is φ. Here, N is the dimension number of the output of the output function u. Similar to the solution operator G, when explicitly indicating the parameter φ, the energy function H will also be denoted as "H φ ". The energy function H φThe implementation of the neural network that realizes this can also use any implementation, similar to the solution operator G. For example, a standard multi-layer perceptron may be used, or other implementations may be used.
[0022] <An example of the hardware configuration of the simulation device 10> An example of the hardware configuration of the simulation device 10 according to this embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing an example of the hardware configuration of the simulation device 10 according to this embodiment.
[0023] As shown in FIG. 1, the simulation device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. These hardware components are communicably connected to each other via a bus 109.
[0024] The input device 101 is, for example, a keyboard, a mouse, a touch panel, physical buttons, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the simulation device 10 may not have at least one of the input device 101 and the display device 102, for example.
[0025] The external I / F 103 is an interface with an external device such as a recording medium 103a. The simulation device 10 can read and write to the recording medium 103a via the external I / F 103. Examples of the recording medium 103a include a flexible disk, a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), a USB (Universal Serial Bus) memory card, etc.
[0026] The communication I / F 104 is an interface for the simulation device 10 to connect to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily holds programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can hold programs and data even when the power is turned off. The auxiliary storage device 107 is a storage device (storage device) such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, for example. The processor 108 is an arithmetic device such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit), for example.
[0027] The simulation device 10 according to this embodiment can realize various processes to be described later by having the hardware configuration shown in FIG. 1. Note that the hardware configuration shown in FIG. 1 is an example, and the hardware configuration of the simulation device 10 is not limited to this. For example, the simulation device 10 may have a plurality of auxiliary storage devices 107 and a plurality of processors 108, may not have a part of the illustrated hardware, or may have various hardware other than the illustrated hardware.
[0028] <An example of the functional configuration of the simulation device 10> An example of the functional configuration of the simulation device 10 according to this embodiment will be described with reference to FIG. 2. FIG. 2 is a diagram showing an example of the functional configuration of the simulation device 10 according to this embodiment.
[0029] As shown in FIG. 2, the simulation apparatus 10 according to the present embodiment includes a parameter estimation unit 201, a reception unit 202, a simulation unit 203, and an output unit 204. Each of these units is realized, for example, by a process executed by one or more programs installed in the simulation apparatus 10 on a processor 108 or the like. Further, the simulation apparatus 10 according to the present embodiment includes a learning data storage unit 205, a first model parameter storage unit 206, a second model parameter storage unit 207, a regularization parameter storage unit 208, and a test data storage unit 209. Each of these storage units is realized, for example, by a storage area such as an auxiliary storage device 107. Note that, for example, at least one of these storage units may be realized by a storage area of a storage device or the like that is communicably connected to the simulation apparatus 10.
[0030] In the parameter estimation phase, the parameter estimation unit 201 uses the learning data stored in the learning data storage unit 205 to estimate the parameter θ of the solution operator G θ and the parameter φ of the energy function H φ and the regularization parameter λ. The regularization parameter λ is a parameter representing the strength of the influence of the regularization term included in the objective function for estimating the parameters θ and φ.
[0031] In the simulation phase, the reception unit 202 receives the designation of the test data a * , its simulation coordinates x ∈ X, and its simulation time [t s , t e (where t s ≦ t e and t s , t e ∈ T). Note that the simulator coordinates x ∈ X are not limited to one, and the reception unit 202 may receive a plurality of simulation coordinates.
[0032] In the simulation phase, the simulation unit 203 uses the test data a corresponding to the designation received by the reception unit 202 *After obtaining it from the test data storage unit 209, u is calculated using the parameter θ stored in the first model parameter storage unit 206. * (y)=G θ (a * (y)). Here, y = (x, t) (where t ∈ [t s , t e ). As a result, the simulation result of the time evolution of a certain physical phenomenon at the simulation coordinate x is calculated.
[0033] In the simulation phase, the output unit 204 outputs the simulation result {u * (y)|y = (x, t) (where t ∈ [t s , t e} to a predetermined output destination. Note that the output destination may be any output destination, for example, a display device 102 such as a display, an auxiliary storage device 107, other devices or other equipment connected via a communication network, and the like.
[0034] The learning data storage unit 205 stores the given learning data {(a i , u i )|i = 1, ···, I}. Note that u i may contain some noise (observation noise).
[0035] The first model parameter storage unit 206 stores the parameter θ estimated by the parameter estimator 201.
[0036] The second model parameter storage unit 207 stores the parameter φ estimated by the parameter estimator 201.
[0037] The regularization parameter storage unit 208 stores the regularization parameter λ estimated by the parameter estimator 201.
[0038] The test data storage unit 209 stores I * pieces of test data {a i *|i = 1, ···, I * stores it.
[0039] Note that in the example shown in FIG. 2, the functional configuration in the case where the parameter estimation phase and the simulation phase are executed by the same simulation device 10 is shown, but this is just an example and is not limited thereto. For example, the parameter estimation phase and the simulation phase may be executed by different devices. At this time, for example, the device that executes the parameter estimation phase may be called a "parameter estimation device", a "learning device", or the like.
[0040] <An example of the flow of processing executed by the simulation device 10> Hereinafter, an example of the flow of processing executed by the simulation device 10 according to the present embodiment will be described with reference to FIG. 3. FIG. 3 is a flowchart showing an example of the flow of processing executed by the simulation device according to the present embodiment. Note that steps S101 to S102 in FIG. 3 are the processing executed in the parameter estimation phase, and steps S103 to S105 are the processing executed in the simulation phase.
[0041] The parameter estimation unit 201 uses the learning data stored in the learning data storage unit 205 to estimate the parameter θ of the solution operator G θ and the parameter φ of the energy function H φ and the regularization parameter λ (step S101). The parameter estimation unit 201 estimates the parameter θ and the parameter φ by an arbitrary continuous optimization method so as to minimize the following objective function L(θ, φ).
[0042] L(θ, φ) = J(θ) + λΩ(θ, φ) Here, λ > 0 is the regularization parameter. J(θ) is a loss for evaluating how well the solution operator G θ can reproduce the learning data, and is represented as follows, for example, using the mean squared error.
Equation
[0043] Ω(θ, φ) is a regularization term based on the theory of Hamiltonian mechanics and is represented, for example, as follows.
Equation
[0044] By introducing the above regularization term Ω(θ, φ), it is possible to bias the partial derivative with respect to time of the solution operator G θ (a i )(y k ) to be equal to the symplectic gradient. This can encourage the simulation result by the solution operator G θ to satisfy the energy conservation law.
[0045] Also, the above S ∈ R N×Nis a strain symmetric matrix and is set according to the class of the physical system targeted by the simulation device 10. For example, when targeting a Hamiltonian system including a physical system such as a simple pendulum, S is represented as follows. [Number] However, I is the identity matrix and O is the zero matrix. Also, for example, when targeting a Hamiltonian partial differential equation system such as the KdV equation representing shallow water waves (assuming periodic boundary conditions), S is represented as follows. [Number] However, Δx is the volume of the grid cell used for spatial discretization. In addition, an appropriate S can be set according to the class of the physical system targeted (see, for example, Reference 1). When targeting a Hamiltonian partial differential equation system, G θ (a) not only, but also the partial derivatives ∂ x u, ∂ xx u, ··· with respect to space need to be input to the energy function H φ . These partial derivatives are obtained by automatically differentiating G θ (a) with respect to space.
[0046] The regularization term Ω(θ, φ) shown in the above number 2 can be calculated for any query y k that is not necessarily included in the query set Y'. For example, for each update of the parameters θ and φ, K queries may be sampled from a uniform distribution so as to satisfy {y k | k = 1, ···, K} ⊂ Y.
[0047] Also, the parameter estimator 201 can, for example, use a part of the training data as evaluation data and the rest as training data, and then estimate the parameters θ and φ respectively using each of the candidates of the value of the regularization parameter λ prepared in advance and the training data, and then select the value that gives the best simulation accuracy for the evaluation data to determine (estimate) the value of the regularization parameter λ.
[0048] The parameter estimation unit 201 stores the parameter θ, the parameter φ, and the regularization parameter λ estimated in the above step S101 in the first model parameter storage unit 206, the second model parameter storage unit 207, and the regularization parameter storage unit 208, respectively (step S102).
[0049] The reception unit 202 receives the designation i ∈ {1, ···, I *} of the test data a * , the simulation coordinates x ∈ X, and the simulation time [t s , t e (where t s ≤ t e and t s , t e ∈ T) (step S103).
[0050] The simulation unit 203 uses the test data a * = a i * specified in the above step S101, the simulation coordinates x ∈ X, and the simulation time [t s , t e to calculate the simulation result u θ = G * (a θ )(step S104). That is, the simulation unit 203 calculates u * (y) = G s (a e )(y)) for each time t ∈ [t * , t θ and the simulation coordinates x ∈ X, where y = (t, x). Here, θ is the parameter stored in the first model parameter storage unit 206 (i.e., the parameter of the learned solution operator G). * (y)) (where y = (t, x)). Here, θ is the parameter stored in the first model parameter storage unit 206 (i.e., the parameter of the learned solution operator G).
[0051] The output unit 204 outputs the simulation result {u * (y)|y = (x, t) (where t ∈ [t s , te Output it to a predetermined output destination that has been determined in advance (step S105).
[0052] <Summary> As described above, the simulation device 10 according to the present embodiment introduces the theory of Hamiltonian mechanics and utilizes physical laws (the law of conservation of energy) as prior knowledge for operator learning, thereby preventing overfitting while learning the parameters of the solution operator from only the observed data. As a result, the simulation device 10 according to the present embodiment can simulate the time evolution of the physical phenomenon at a predetermined coordinate with high accuracy even when the differential equation describing the physical phenomenon is unknown.
[0053] The present invention is not limited to the specifically disclosed above embodiments, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the description of the claims.
[0054] [References] Reference 1: E. Celledoni, V. Grimm, R. McLachlan, D. McLaren, D. O'Neale, B. Owren, and G. Quispel. Preserving energy resp. dissipation in numerical PDEs using the "average vector field" method. Journal of Computational Physics, 231(20):6770 - 6789, 2012.
Explanation of Signs
[0055] 10 Simulation device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary memory device 108 Processor 109 Bus 201 Parameter estimation unit 202 Reception unit 203 Simulation unit 204 Output unit 205 Learning data storage unit 206 First model parameter storage unit 207 Second model parameter storage unit 208 Regularization parameter storage unit 209 Test data storage unit
Claims
1. a parameter estimation unit that estimates at least parameters of a solution operator that represents a mapping from an input function that represents a condition of a physical phenomenon in a predetermined physical system to an output function that represents a solution of a differential equation that represents the physical phenomenon, so as to minimize a predetermined objective function, The learning device, wherein the objective function includes a loss for evaluating the reproducibility of the solution operator on the learning data, and a regularization term for ensuring that the value of the output function predicted by the solution operator when the input function is given satisfies the law of conservation of energy.
2. The regularization term is a term for inputting the value of the output function and outputting a real number, and for making a symplectic gradient, which is expressed by the product of the gradient of the energy function expressing the law of conservation of energy and a predetermined skew-symmetric matrix, equal to a time derivative of the solution operator; The parameter estimation unit The learning device according to claim 1 , further comprising: an energy function estimation unit for estimating parameters of the energy function.
3. The loss is 3. The learning device according to claim 2, wherein the term is for making equal a value of the output function at a predetermined discrete point and a value at the discrete point of an output function predicted by the solution operator from the input function corresponding to the output function.
4. the objective function includes the loss and a product of a regularization parameter and the regularization term; The parameter estimation unit The learning device according to claim 1 , further comprising: a regularization parameter estimation unit configured to estimate the regularization parameter.
5. a parameter estimation unit that uses learning data consisting of a pair of an input function that represents a condition of a physical phenomenon in a predetermined physical system and an output function that represents a solution of a differential equation that represents the physical phenomenon, and estimates at least parameters of a solution operator that represents a mapping from the input function to the output function so as to minimize a predetermined objective function; a simulation unit that calculates a value of an output function predicted from the unknown input function by the solution operator using the parameters estimated by the parameter estimation unit, an unknown input function, a simulation coordinate, and a simulation time; having The simulation device, wherein the objective function includes a loss for evaluating the reproducibility of the solution operator for the learning data, and a regularization term for making the value of the output function predicted by the solution operator when the input function is given satisfy the law of conservation of energy.
6. a parameter estimation procedure for estimating at least parameters of a solution operator that represents a mapping from an input function that represents a condition of a physical phenomenon in a predetermined physical system to an output function that represents a solution of a differential equation that represents the physical phenomenon, so as to minimize a predetermined objective function, by a computer; A learning method, wherein the objective function includes a loss for evaluating the reproducibility of the solution operator on the learning data, and a regularization term for ensuring that the value of the output function predicted by the solution operator when the input function is given satisfies the law of conservation of energy.
7. a parameter estimation step of estimating at least parameters of a solution operator that represents a mapping from an input function that represents a condition of a physical phenomenon in a predetermined physical system to an output function that represents a solution of a differential equation that represents the physical phenomenon, so as to minimize a predetermined objective function; a simulation step of calculating a value of an output function predicted from the unknown input function by the solution operator using the parameters estimated in the parameter estimation step, an unknown input function, simulation coordinates, and a simulation time; The computer executes The simulation method, wherein the objective function includes a loss for evaluating the reproducibility of the solution operator for the training data, and a regularization term for ensuring that the value of the output function predicted by the solution operator when the input function is given satisfies the law of conservation of energy.
8. A program for causing a computer to function as the learning device according to claim 1 or the simulation device according to claim 5.