Device and method

By employing linear algebraic methods and Fourier functions to approximate Koopman operators, the method constructs efficient neural networks that achieve universal approximation with reduced calculation time and computational costs.

WO2026033850A1PCT designated stage Publication Date: 2026-02-12NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/028809
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing ODE-Nets perform actual calculations using numerical solutions of differential equations, leading to increased calculation time as the value of time increases, and require a 2d-dimensional neural network for universal approximation, which is impractical for finite parameter settings.

Method used

A method is proposed to construct a d+1-dimensional neural network based on multiple differential equations using linear algebraic formulations, approximating the generator with Fourier functions and Koopman operators, allowing for efficient calculation without numerical solutions.

Benefits of technology

This approach enables the construction of neural networks that achieve universal approximation with constant calculation time, reducing computational costs and improving efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024028809_12022026_PF_FP_ABST
    Figure JP2024028809_12022026_PF_FP_ABST
Patent Text Reader

Abstract

A device according to one embodiment disclosed herein includes a network configuration unit that uses the generator of the Koopman operator to represent a first neural network defined by the flow of a plurality of differential equations and a function included in a function space defined by a plurality of Fourier functions, and approximates the generator by the Fourier function to configure a second neural network that approximates the first neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Apparatus and method

[0001] The present disclosure relates to an apparatus and a method.

[0002] Neural networks are applied in a wide range of fields, and research into them is being actively conducted. In particular, ODE-Net (Ordinary Differential Equation Network) has been proposed as a method for constructing neural networks based on differential equations (Non-Patent Document 1), and is used, for example, for data generation by density estimation and analysis of time-series data.

[0003] Chen, RTQ, Rubanova, Y., Bettencourt, J., and Duvenaud, DK Neural ordinary differential equations, NeurIPS 2018.

[0004] However, in the existing ODE-Net, actual calculations are performed using a numerical solution of differential equations, which poses a problem that the calculation time increases as the value of time t increases.

[0005] The present disclosure has been made in view of the above points, and aims to configure an efficient neural network based on differential equations.

[0006] An apparatus according to one aspect of the present disclosure includes a network configuration unit that represents a first neural network defined by a flow of multiple differential equations and a function included in a function space defined by multiple Fourier functions using a generator of a Koopman operator, and configures a second neural network that approximates the first neural network by approximating the generator using a Fourier function.

[0007] Efficient neural networks based on differential equations can be constructed.

[0008] FIG. 1 is a diagram showing an example of the hardware configuration of a neural network device according to an embodiment of the present invention; FIG. 2 is a diagram showing an example of the functional configuration of a neural network device according to an embodiment of the present invention; FIG. 3 is a flowchart showing an example of the operation of a neural network device according to an embodiment of the present invention; FIG. 4 is a diagram showing an example of training data; and FIG. 5 is a diagram showing an example of evaluation results.

[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0010] <Problems with existing ODE-Net> Existing ODE-Nets perform actual calculations using a numerical solution of differential equations, so 1 , ..., t J The problem is that the calculation time increases as the value of .

[0011] In addition, a classical ODE-Net is a neural network constructed based on a single differential equation. In order for the neural network to exhibit the property of being able to approximate any function (universal approximation), it is necessary to construct a 2d-dimensional neural network, where d is the number of dimensions of the actual input and output (Reference 1).

[0012] There are some results that show that universal approximation can be achieved by constructing a neural network based on multiple differential equations (Reference 2), but this assumes an ideal case where the number of parameters is infinite, and it is unclear how to actually perform the calculations on a computer.

[0013] Therefore, in the following, we propose a method for constructing and calculating a d+1-dimensional neural network based on multiple differential equations and satisfying universal approximation by performing linear algebraic formulation and calculation without using a numerical solution. This proposed method is based on the t 1 , ..., t J This method solves the problem that the calculation time increases as the value of increases, and allows the calculation time to be evaluated at a constant value.

[0014] <Proposed Method> The proposed method will be explained below.

[0015] <1. Representation of ODE-Net using Koopman operator and its generator> Consider the following J differential equations.

[0016] Note that t is a variable representing time.

[0017] In this case, let T be a torus, d be a predetermined integer of 1 or more, and g j : R x T d →T d Let g be the flow, that is, j (t, x j (0)) = x j (t) is a mapping that satisfies (R). Note that R represents the set of all real numbers. Hereafter, we will not distinguish between mappings and functions, and will also refer to mappings as functions.

[0018] Also, q n (x) = e inx (n∈Z d , i is the imaginary unit) is the Fourier function, 2 (T d ) is a subspace of {q n |n∈Z d Let V be the subspace spanned by \{0}}, and Z represent the set of all integers.

[0019] Furthermore, L 2 (T d ) the orthogonal complement of V in ⊥ Let v be a non-zero (real-valued) function of . Then, consider a neural network f shown in the following equation (1).

[0020] In addition, J in the neural network f represents the number of layers.

[0021] Hereinafter, we consider expressing the neural network f shown in the above formula (1) using the Koopman operator.

[0022] u∈L 2 (T d ) for the linear operator (Koopman operator) K j t is determined as follows:

[0023] Also, {Kj t} t∈R The generator (Koopman generator) L j is determined as follows:

[0024] However, g j,k is the function g j The kth component of f j,k is the function f j is the k-th component of

[0025] In this case, the Koopman operator K j t is expressed as follows:

[0026] Therefore, the neural network f shown in the above formula (1) is expressed as the following formula (2).

[0027]

[0028] ≪2. Approximation of the generator of the Koopman operator using Fourier functions≫ Fourier function q n (x) = e inx (n∈Z d ) to obtain the Koopman generator L j Consider approximating the finite index set N ⊂ Z d \{(n 1 , ..., n d ) | n 1 , ..., n d ≠0}, the linear operator Q N : C |N| →L 2 (T d ) is defined as follows:

[0029] Note that C represents the set of all complex numbers.

[0030] The above linear operator Q N Using Q N Q N * L j Q N Q N * Thus, the Koopman generator L j is approximated. N* is a linear operator Q N represents the adjoint operator of

[0031] In this case, the representation matrix of the Koopman generator approximated above is Q N * L j Q N and this will be written below.

[0032] However, since symbols with "~" directly above them cannot be expressed in the text of the specification, they are expressed in the text of the specification with a superscript immediately before them. For example, in the text of the specification, the above expression matrix Q N * L j Q N of" ~ L j "I will write it as ".

[0033] A finite index set M ⊂ Z d For f j,k =Σ m∈M a m j,k q m (where |M|<<|N|), ~ L j The (n, l) component of is as follows:

[0034] where n∈N, l∈Z d and 〈・,・〉 is L 2 (T d ) on the inner product defined by k is the k-th component of l.

[0035] Here, h 1 , h 2 is T d In the above function, the bar (-) represents the complex conjugate.

[0036] a n-l j,k A Toeplitz matrix with (n, l) elements is j,k , il k The diagonal matrix D has the following as its l-th diagonal element: k Then, ~ Lj can be expressed as follows:

[0037] Note that n-l∈M, |M|<<|N|, so A j,k Note that is a sparse matrix.

[0038] In general, f j,k Let be the following:

[0039] At this time, ~ L j The (n, l) component of is as follows:

[0040] For this reason, a n-l,s j,k A Toeplitz matrix with (n, l) elements is j,k,s , il k The diagonal matrix D has the following as its l-th diagonal element: k Then, ~ L j can be expressed as follows:

[0041] However, since f is a real-valued function, the following must also be a real-valued function.

[0042] Therefore, the following should be set:

[0043]

[0044] ≪3. Learning the neural network f≫ Let us consider learning the neural network f shown in the above formula (1). The generator L introduced in the above "2. Approximation of the generator of the Koopman operator using Fourier functions" j Approximation Q of N ~ L j Q N * Using the above equation (2), a neural network that approximates the neural network f is calculated. ~ f is constructed as shown in the following equation (3).

[0045] Here, each ~ Lj is expressed by the above formula 14, and A j,k,s is a sparse matrix, D k is a diagonal matrix. Therefore, ~ L j The product of x and a vector can be calculated with low computational cost. In addition, the following calculations can be performed with low computational cost by using, for example, the Krylov subspace method (Reference 3).

[0046] As a result, a neural network can be created using only algebraic calculations. ~ Therefore, it is possible to calculate f using the given training data. ~ Parameter a of f n-l,s j,k By learning ~ It is possible to learn f.

[0047] ≪4. Universal Approximation≫ Any matrix can be expressed as a product of Toeplitz matrices (Theorem 2 in Reference 4). Also, for any n∈Z\{0}, D k There exists a k for which the n-th component of is non-zero. Using these, it can be seen that the following space coincides with the space M(|N|, C) of all |N|×|N| matrices.

[0048] M(|N|,C) is the Lie algebra for the Lie group GL(|N|,C) (all n × n regular matrices). The Lie group GL(|N|,C) can be expressed as follows (Colorary 3.47 in Reference 5):

[0049] Therefore, by choosing appropriate S and J, any regular matrix can be expressed in the following form:

[0050] GL(|N|, C) is C |N| \{0} is transitive, so any vector c∈C |N| For \{0}, the following holds:

[0051] Therefore, u = Σ n∈N cn q n Any non-zero function expressed in the form ~ It is possible to express it in terms of L 2 (T d ) is a subspace of {q n |n∈Z d Let V be the subspace spanned by \{0}}. The set of Fourier functions {q n |n∈Z d The space spanned by} is L 2 (T d ) is dense among them, so if |N| is increased, L 2 (T d ) the orthogonal complement of V in ⊥ It is possible to approximate the function of with any accuracy. ⊥ Since the function corresponds to a function without a constant component, any function can be expressed by shifting the output result.

[0052] ≪5. Construction of a neural network with d-dimensional output dimensions≫ In the above sections "1. Representation of ODE-Net using Koopman operator and its generator" to "4. Universal approximation", ~ It has been assumed that the number of input dimensions of f is d and the number of output dimensions is 1, but hereinafter, it will be considered that the number of output dimensions is expanded to d.

[0053] A d-dimensional function is expressed as u = [u 1 ,...u d ], and then the following T d+1 Define the above complex-valued function.

[0054] At this time, as mentioned in "4. Universal Approximation" above, each u i (x) can be approximated with any precision by T d+1 Since the above neural network exists, ~ f i Then, the complex value function shown in the above formula 23 can be approximated with any precision by T d+1 There is a neural network shown in the following equation (4) above.

[0055] Hereinafter, the neural network shown in Equation (3) or Equation (4) will be generated by the above proposed method. ~ f and this neural network ~ After learning f, the trained neural network ~ The neural network device 10 that performs inference using f will be described below. ~ The phase of constructing f is called the "network construction phase", and the neural network ~ The phase of learning f is called the "learning phase", and the trained neural network ~ The phase in which inference is performed using f is called the "inference phase."

[0056] <Example of Hardware Configuration of Neural Network Device 10> An example of the hardware configuration of the neural network device 10 according to this embodiment is shown in Fig. 1. As shown in Fig. 1, the neural network device 10 according to this embodiment has an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.

[0057] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the neural network device 10 does not necessarily have to have at least one of the input device 101 and the display device 102, for example.

[0058] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0059] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is a variety of arithmetic devices such as a central processing unit (CPU) or a graphic processing unit (GPU).

[0060] 1 is merely an example, and the hardware configuration of the neural network device 10 is not limited to this. For example, the neural network device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various types of hardware other than the hardware shown in the figure.

[0061] <Example of Functional Configuration of Neural Network Device 10> An example of the functional configuration of the neural network device 10 according to this embodiment is shown in Fig. 2. As shown in Fig. 2, the neural network device 10 according to this embodiment has a network configuration unit 201, a learning unit 202, and an inference unit 203. Each of these units is realized, for example, by a process in which one or more programs installed in the neural network device 10 are executed by the processor 108 or the like. The neural network device 10 according to this embodiment also has a memory unit 204. The memory unit 204 is realized, for example, by a memory area of ​​the auxiliary storage device 107 or the like. Note that the memory unit 204 may also be realized, for example, by a memory area of ​​a storage device (e.g., a storage device included in a database server) communicatively connected to the neural network device 10.

[0062] The network configuration unit 201 generates a neural network shown in equation (3) or equation (4) by the above-mentioned proposed method. ~Construct f.

[0063] The learning unit 202 uses the learning data stored in the storage unit 204 to generate the neural network constructed by the network construction unit 201. ~ Learn f.

[0064] The inference unit 203 uses the test data stored in the storage unit 204 to generate a trained neural network. ~ Inference is made using f.

[0065] The storage unit 204 stores various data (e.g., neural network ~ f parameters, training data, test data, etc.) are stored.

[0066] <Example of Operation of Neural Network Device 10> An example of operation of the neural network device 10 according to this embodiment will be described with reference to Fig. 3. Here, step S101 is processing in the network configuration phase, steps S102 and S103 are processing in the learning phase, and steps S104 and S105 are processing in the inference phase. Note that the network configuration phase is executed before the learning phase, and the learning phase is executed before the inference phase.

[0067] The network configuration unit 201 generates a neural network shown in equation (3) or equation (4) by the above-mentioned proposed method. ~ f is constructed (step S101).

[0068] The learning unit 202 acquires the learning data stored in the storage unit 204 (step S102).

[0069] The learning unit 202 uses the learning data acquired in step S102 to generate the neural network constructed in step S101. ~ In other words, the learning unit 202 learns f by using a neural network to optimize (minimize or maximize) a preset objective function (e.g., a loss function, a negative log-likelihood function, etc.). ~ Update the parameters of f, which will ~As mentioned in "3. Learning of neural network f" of the proposed method, ~ Learning f can be done efficiently with low computational cost.

[0070] The inference unit 203 acquires the test data stored in the storage unit 204 (step S104).

[0071] The inference unit 203 uses the test data acquired in step S104 to generate a trained neural network ~ Inference is performed using f (step S105), thereby obtaining an inference result for the test data.

[0072] <Application Example> Hereinafter, as an application example of the neural network device 10 according to this embodiment, an application example to density estimation using normalizing flow will be described.

[0073] Let p be the density function of a two-dimensional normal distribution with mean 0 and variance 1. x = [x 1 , x 2 ]for v 1 (x) = x 1 , v 2 (x) = x 2 In addition, the following f 1 and f 2 For f = [f 1 , f 2 ]

[0074] At this time, f -1 can be expressed as follows:

[0075] When f is set as a mapping that maps points that follow a normal distribution to points that follow the distribution of data, the density function p of the distribution of data can be expressed by the following equation (5).

[0076] Here, the following is f -1 is the Jacobian matrix of

[0077] Based on the above formula (5), given training data x 1 , ..., xK By minimizing the negative log-likelihood shown in the following equation (6), 1 , ..., L J We learn the parameters that determine the approximation of

[0078] Note that the application to density estimation using normalized flow is just one example, and various other applications are possible, such as analysis of time-series data.

[0079] <Evaluation> The results of an experiment conducted to evaluate the neural network device 10 according to this embodiment will be described below. In this experiment, 1 , w 2 ) | -10≦w 1 ≦10, −10≦w 2 ≦10}, J=3. In the first layer (j=1), S=2, M 1 = M 2 = {(w 1 , w 2 ) | -5≦w 1 ≦5, −5≦w 2 ≦5}, in the second layer (j=2), S=2, M 1 = M 2 = {(w 1 , w 2 ) | -3≦w 1 ≦3, −3≦w 2 ≦3}, in the third layer (j=3), S=1, M 1 = {(w 1 , w 2 ) | -5≦w 1 ≦5, −5≦w 2 ≦5}.

[0080] In this case, the density function was estimated by minimizing the negative log-likelihood shown in equation (6) above, using the 1,000 data points shown in Figure 4 as training data. The likelihood during testing was also calculated using 100 test data sets created in the same way as the training data. The results are shown in Figure 5. As shown in Figure 5, it can be seen that the likelihood during testing is higher than during training.

[0081] <Summary> As described above, the neural network device 10 according to this embodiment can configure a neural network that satisfies universal approximation by a linear algebraic method based on a plurality of differential equations. As a result, in the neural network device 10 according to this embodiment, t in the above equation (1) can be 1 , ..., t J This solves the problem of the conventional ODE-Net that the calculation time increases as the value of θ increases, and makes it possible to evaluate the calculation time at a constant value.

[0082] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.

[0083] [References] Reference 1: Zhang, H., Gao, X., Unterman, J., and Arodz, T., Approximation capabilities of neural ODEs and invertible residual networks, NeurIPS 2020. Reference 2: Teshima, T., Tojo, K., Ikeda, M., Ishikawa, I., and Oono, K., Universal approximation property of neural ordinary differential equations, NeurIPS 2020 Workshop on Differential Geometry meets Deep Learning. Reference 3: Hashimoto, Y. and Nodera, T., Inexact shift-invert Arnoldi method for evolution equations, ANZIAM Journal, Vol. 58, E1-E27, 2016. Reference 4: Ye, K., Lim, LH. Every matrix is ​​a product of Toeplitz matrices. Found Comput Math, Vol. 16, 577-598, 2016. Reference 5: Hall, Brian C., Lie Groups, Lie Algebras, and Representations: An Elementary Introduction, Graduate Texts in Mathematics, vol. 222 (2nd ed.), Springer, 2015.

[0084] REFERENCE SIGNS LIST 10 Neural network device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Network configuration unit 202 Learning unit 203 Inference unit 204 Storage unit

Claims

1. A device having a network configuration unit that represents a first neural network defined by a flow of multiple differential equations and a function included in a function space defined by multiple Fourier functions using a generator of a Koopman operator, and configures a second neural network that approximates the first neural network by approximating the generator using a Fourier function.

2. The apparatus according to claim 1, further comprising a training unit for training said second neural network using training data.

3. The apparatus according to claim 2, further comprising an inference unit that performs inference using test data with the second neural network trained by the training unit.

4. A method executed by a computer, which comprises a network construction procedure for expressing a first neural network defined by a flow of multiple differential equations and a function included in a function space defined by multiple Fourier functions using a generator of a Koopman operator, and approximating the generator by a Fourier function to construct a second neural network that approximates the first neural network.