Information processing device, information processing method, and program
By employing iterative processes with forward and reverse-mode automatic differentiation and convergence checks, the method addresses the high computational cost of calculating derivatives in high-dimensional parameter spaces, achieving reduced costs and maintained accuracy.
Patent Information
- Application Number
- PCT/JP2024/045384
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-10
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-17
AI Technical Summary
Existing methods for calculating the derivative of an objective function with respect to high-dimensional parameters using automatic differentiation are computationally costly.
A method involving iterative processes with a combination of forward and reverse-mode automatic differentiation, including convergence checks, is employed to reduce computational cost while maintaining accuracy in calculating the derivative.
The method effectively reduces the computational cost of calculating the derivative of the objective function with respect to high-dimensional parameters while preserving high accuracy.
Smart Images

Figure JP2024045384_17072025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present invention relates to an information processing device, an information processing method, and a program.
[0002] There are known techniques for analyzing, for example, the optimization of parameters that describe a target system such as a material system. For example, Patent Document 1 discloses a technique for designing materials based on quantum mechanical calculations.
[0003] Japanese Patent Application Publication No. 2016-69302
[0004] As one of the above-described analysis methods, a method using automatic differentiation is also known. By using automatic differentiation, it is possible to improve the accuracy of the analysis, and it is desirable to be able to perform automatic differentiation at a lower calculation cost. One aspect of the present invention aims to realize a technology that can reduce the calculation cost in a calculation method using automatic differentiation.
[0005] In order to solve the above problem, an information processing device according to one aspect of the present invention comprises an acquisition unit that acquires initial values for calculating the value of a target function that directly or indirectly includes one or more parameters as arguments, and a control unit that calculates a derivative of the target function, wherein the calculation process of the derivative by the control unit includes a first process and a second process, the calculation process of the derivative by the control unit includes a first iterative process, the first process includes a part of the first iterative process, and the second process includes a part other than the part of the first iterative process, the calculation process of the derivative by the control unit includes a third process that includes a calculation process of automatic differentiation for a process that does not include the first process but includes the second process, the first process includes a first iterative process of the first iterative process, and the first iterative process is an iterative process that finds a solution to a self-consistent equation.
[0006] In order to solve the above problem, an information processing method according to one aspect of the present invention includes an acquisition step of acquiring initial values for calculating a value of an object function that directly or indirectly includes one or more parameters as arguments, and a calculation step of calculating a derivative of the object function, wherein the calculation process of the derivative by the calculation step includes a first process and a second process, wherein the calculation process of the derivative by the calculation step includes a first iterative process, wherein the first process includes a part of the first iterative process, and the second process includes a part other than the part of the first iterative process, and the calculation process of the derivative by the calculation step includes a third process including a calculation process of automatic differentiation for a process that does not include the first process but includes the second process, wherein the first process includes a first iterative process of the first iterative process, and the first iterative process is an iterative process that finds a solution to a self-consistent equation.
[0007] The information processing device according to each aspect of the present invention may be realized by a computer. In this case, the program that realizes the information processing device on a computer by causing the computer to operate as each part (software element) of the information processing device, and the computer-readable recording medium on which the program is recorded, also fall within the scope of the present invention.
[0008] According to one aspect of the present invention, it is possible to reduce calculation costs in a calculation method using automatic differentiation.
[0009] FIG. 1 is a block diagram showing the configuration of an information processing device according to an embodiment of the present invention. FIG. 2 is a flow chart showing the flow of information processing according to an embodiment of the present invention. FIG. 3 is a flow chart showing the flow of information processing according to an embodiment of the present invention. FIG. 4 is a flow chart showing the flow of information processing based on an embodiment of the present invention. FIG. 5 is a flow chart showing the flow of information processing according to an embodiment of the present invention. FIG. 6 is a block diagram showing the configuration of an information processing device according to an embodiment of the present invention. FIG. 7 is a flow chart showing the flow of information processing according to an embodiment of the present invention.
[0010] (Outline of Information Processing According to Each Present Embodiment) First, an outline of information processing according to each embodiment described in this specification will be described. The information processing device and information processing method according to each embodiment described in this specification can be used, for example, to design (also referred to as reverse design) a system or material having desired performance or physical properties. In a typical simulation, for example, parameters p are set, such as atomic arrangement, semiconductor impurity distribution, and manufacturing conditions, and an objective function related to performance or physical properties, such as superconducting transition temperature, voltage-current characteristics, and defect density, determined according to the parameters is calculated. However, the information processing device and information processing method according to each embodiment described in this specification can be suitably used to solve the problem of designing (reverse designing) parameters p, such as atomic arrangement, semiconductor impurity distribution, and manufacturing conditions, so as to achieve a desired superconducting transition temperature, desired voltage-current characteristics, and desired defect density. In another example, the information processing device and information processing method can be used to estimate parameters describing a target system from limited measurement results related to the target system.
[0011] In other words, the information processing device and information processing method according to each embodiment described in this specification can be suitably applied to the problem of finding a parameter p that maximizes or minimizes an objective function related to performance or physical properties in the inverse design described above. To solve such a problem, it is useful to know the derivative of the objective function with respect to p. However, when p is high-dimensional, there is a problem in that finding all the derivatives of the objective function with respect to p requires a high computational cost. The information processing device and information processing method according to each embodiment described in this specification can reduce the computational cost when calculating the derivative of the objective function with respect to p using an iterative method.
[0012] More specifically, in the information processing device and information processing method of the present disclosure, at least one parameter p a (where a is an index that lists the at least one parameter); Q is a subset of the set P; ・ At least one state variable x b Consider a state variable X, which is a set of P and Q (where b is an index that lists at least one parameter), and a function F(Q, X). Hereinafter, P or Q will be simply referred to as a parameter. X will also be simply referred to as a state variable. Also, X, which is the value of the state variable X determined based on the value of the parameter P, ANS Consider (P). Below, X ANS (P) can also be written simply as X(P). If the set Q is not an empty set, the function F is a function of the parameters p a This corresponds to the case where such a function F depends explicitly on one or more parameters (p a On the other hand, if the set Q is an empty set, the function F is a function that directly includes the parameter p a This corresponds to a case where the function F does not explicitly depend on any of the parameters p a ) as an argument. Also, F(Q, X) is a function that b For example, F depends on at least one of x b It may be any one of them itself.
[0013] In this way, the target function (F) can be expressed as a function having as arguments a subset (Q) of the one or more parameters and one or more state variables (X(P)) having the one or more parameters as arguments. b does not have to be about a physical state.
[0014] In the information processing device and information processing method according to the present specification, as an example, F(Q, X ANS (P)) each p a Derivative with respect to ∂F(Q,X ANS (P)) / ∂p a Numerically calculate the value of F(Q, X ANS The specific form of ∂F(Q, X) depends on the target of the information processing device and the information processing method. ANS (P)) / ∂p a contains contributions where the objective function F depends on the parameters P via one or more state variables X.
[0015] X ANS Although (P) may also depend on parameters not included in P, in the following, for the sake of simplicity, we will use X to parameters not included in P. ANS In addition, the function F may also depend on parameters not included in P, but in the following, for the sake of simplicity, the dependence of the function F on parameters not included in P may not be written explicitly.
[0016] X ANS (P) is determined based on the target of the information processing device and information processing method. ANS (P) may be an implicit function of P, and X ANS All you need to know is the iterative calculation method to find the numerical solution of (P). For example, if you have a suitable initial condition X 0 Starting from X i =φ(P, X i-1 ) by repeating the calculation of X i The convergence value of X ANSIt is only necessary to know the function φ that is expected to give a numerical solution to (P). Also, the function φ does not have to be fixed, but can be a function φ that changes depending on the iteration. i And X i =φ i (P, X i-1 ) by repeating the calculation of X i The convergence value of X ANS It may be a function that is expected to be solved numerically in (P). For example, φ i is a function V that depends on i and not on P itself. i And lim i→∞ V i (P) = P, a function V i The value of V i (P) as a parameter i = φ(V i (P), X i-1 ) by X i For example, the solution X of X=φ(P,X) can be obtained by using the annealing method. ANS (P) may be a numerical solution. i Each state variable x included in b The value of x (i,b) It is written as follows.
[0017] φ is always X ANS It is not necessary to obtain a numerical solution for (P), for example, if a certain initial condition X 0 For the repeated calculation, X ANS It is also possible that a numerical solution of (P) cannot be obtained. For example, φ is a function that X = φ(P, X) has multiple solutions and X ANS (P) is one of the multiple solutions, and has a certain initial condition X 0 When the repeated calculation is performed starting from X = φ(P, X), X ANS It is also possible to obtain a numerical solution corresponding to a solution different from (P). In that case, for example, by starting from multiple initial values and performing calculations, X ANS A numerical solution of (P) is obtained.
[0018] [First Embodiment] Next, one embodiment of the present invention will be described in detail.
[0019] 1 is a block diagram showing the configuration of an information processing device 1 according to this embodiment. As shown in FIG. 1, the information processing device 1 includes a control unit 2, a storage unit 3, and an input / output unit 4.
[0020] (Storage Unit 3) The storage unit 3 stores various data and information referenced by the control unit 2 and various data and information derived by the control unit 2. As an example, as shown in Fig. 1, the storage unit 3 stores a parameter P, calculation process information CP, a threshold Cth, and the number of iterations m.
[0021] Here, the parameter P is, for example, at least one of the parameters p a (where a is an index assigned to the at least one parameter). The calculation process information CP is, for example, information related to various calculation processes executed by the control unit 2, and includes, for example, information used in automatic differentiation, which will be described later. The threshold value Cth is, for example, a threshold value referenced in the determination process that determines whether or not the convergence condition is satisfied. The number of iterations m is, for example, a value that specifies the number of iterations of the target calculation by the control unit 2. The various data and information stored in the memory unit 3 will be described in more detail later.
[0022] (Input / Output Unit 4) The input / output unit 4 is configured to include at least one of input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel, for example. Alternatively, the input / output unit 4 may be configured to have input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel connected to it. In this configuration, the input / output unit 4 accepts various types of information input to the information processing device 1 from the connected input devices. Furthermore, the input / output unit 4 outputs various types of information to connected output devices under the control of the control unit 2. An example of the input / output unit 4 is an interface such as a USB (Universal Serial Bus).
[0023] (Control Unit 2) As shown in FIG. 1, the control unit 2 has an initial value determination unit 21, a state update unit 22, an end determination unit 23, a function value calculation unit 24, a differential value calculation unit 26, and an end determination unit 28.
[0024] The initial value determination unit 21 determines the initial value X of the state variable X. 0 The initial value determination unit 21 determines, for example, the initial value X 0 may be set to a predetermined value, or may be set to an initial value X using a random number. 0 The initial value determination unit 21 may determine the initial value X 0 may be set to a value input from the outside via the input / output unit 4. The initial value determination unit 21 determines one or more parameters (p a ) as an argument of the target function (F) directly or indirectly. 0 ) may be expressed as an acquisition unit that acquires the
[0025] The state update unit 22 updates the state variable X. As an example, the state update unit 22 updates X i =φ(P, X i-1 ) to obtain the value of the new state variable X, X i The details of the processing by the state update unit 22 will be described later.
[0026] The end determination unit 23 determines whether or not to end the iteration of the target calculation. Details of the processing by the end determination unit 23 will be described later.
[0027] The function value calculation unit 24 calculates the value of the target function. As an example, the function value calculation unit 24 calculates the value F(Q, X n+m The processing by the function value calculation unit 24 will be described in detail later.
[0028] The differential value calculation unit 26 functions as a calculation unit that calculates the differential coefficient of the target function (F). As an example, as shown in Fig. 1, the differential value calculation unit 26 has a backpropagation calculation unit 252 and a calculation process recording unit 253. Details of the processing by the differential value calculation unit 26 will be described later.
[0029] The completion determination unit 28 determines whether the target process has been executed a predetermined number of times. Details of the process by the completion determination unit 28 will be described later.
[0030] <Processing Flow by Information Processing Device 1> The processing flow by the information processing device 1 will be described below. The information processing device 1 generally executes the following steps: A convergence solution calculation step (also referred to as step S1); and A processing step S2 (also referred to as step S2). Note that steps S1 and S2 as a whole may be referred to as a differential coefficient calculation process. Here, the differential coefficient calculation process includes, for example, an iterative process (also referred to as a first iterative process). Step S1 is also referred to as a first process, and as described below, includes, for example, a part of the iterative process. Meanwhile, steps S211 and S212 of step S2, which will be described below, are also referred to as a second process, and as described below, include a part other than the part of the iterative process. Step S213 of step S2, which will be described below, may be included in the second process. Furthermore, as described below, the differential coefficient calculation process includes an automatic differentiation calculation process for a process that does not include the first process but includes the second process. The "calculation process of automatic differentiation for the process that does not include the first process but includes the second process" is also referred to as the third process.
[0031] As will be described later, the iterative process (first iterative process) includes updating one or more state variables (X). In other words, the iterative process (first iterative process) obtains a numerical solution for the one or more state variables (X) using a parameter P as an argument. The third process includes calculating a differential coefficient of the object function (F) using automatic differentiation. Furthermore, in the process according to this embodiment, the third process includes calculating a differential coefficient of the object function (F) using reverse mode automatic differentiation, as will be described later.
[0032] (Step S1) The convergence solution calculation step (step S1) includes, for example, steps S11 to S13 as shown in FIG.
[0033] (Step S11) In step S11, the initial value determination unit 21 determines the initial value X of the state variable X. 0Define the initial value X 0 may be input from the outside via the input / output unit 4.
[0034] (Step S12) Next, in step S12, the state update unit 22 updates the state variable X. That is, X i =φ(P, X i-1 ) to obtain the value of the new state variable X, X i Here, i corresponds to the number of times step S12 has been executed. For example, when step S12 is executed for the first time, the state update unit 22 calculates X 1 =φ(P, X 0 ) by X 1 Ask for.
[0035] (Step S13) Next, in step S13, the end determination unit 23 determines whether to end the iteration of the calculation in step S12. In this embodiment, the end determination unit 23 determines whether to end the iteration of the calculation in step S12. i In other words, the termination determination unit 23 executes a convergence determination process for at least one of the one or more state variables. i If X does not satisfy the convergence condition, step S12 is executed again. i If the convergence condition is satisfied, step S1 is completed and step S2 is executed. The convergence condition used in step S13 will be described later.
[0036] From then on, X i When it is determined that the convergence condition is satisfied, i is also written as n.
[0037] (Note: Convergence determination in step S13) X i The method by which the termination determination unit 23 determines whether X satisfies the convergence condition is not particularly limited, and may be a method according to the purpose of the information processing method. i-1 and X i The value C(X i-1 , X i ) and the threshold Cth, C(X i-1 , X i ) < Cth is satisfied, then Xi If it does not, X i does not satisfy the convergence condition. 0 , X 1 , ..., X i-2 The function C may depend on one or more of X i Depends only on X 0 , X 1 , ..., X i-1 The values of these state variables are added to the calculation process information CP stored in the storage unit 3 during the calculation process, as needed.
[0038] The end determination unit 23 determines, for example, the X newly obtained in step S12. i and X obtained in the previous step S12. i-1 For example, the termination determination unit 23 determines whether or not the convergence has occurred by using a predetermined threshold Cth. i -X i-1 |<Cth when X i If it does not satisfy the convergence condition, X i It is determined that does not satisfy the convergence condition. i -X i-1 |=(Σ b (x (i,b) -x (i-1,b) ) 2 ) 0.5 is.
[0039] Function C is X 0 , X 1 , ..., X i-1 If it does not depend on C(X i ) < Cth is satisfied, then X i If it does not satisfy the convergence condition, X i does not satisfy the convergence condition. For example, such a convergence determination may be performed when a state variable X that makes the value of C(X) 0 is obtained by an iterative method.
[0040] (Step S2) Step S1 is executed as described above, and after step S1 is completed, step S2 (processing step S2) is performed. FIG. 4 is a flowchart showing step S2. As shown in FIG. 4, step S2 includes steps S211, S212, S213, and S214. Note that all or part of steps S211 to S213 constitute an example of the second process described above. Step S214 can also be regarded as an example of the third process described above.
[0041] (Step S211) In step S211, the state update unit 22 n+k =φ(P, X n+k-1 ) by X n+k Here, k corresponds to the number of times step S211 has been executed. For example, when step S211 is executed for the first time, the state update unit 22 calculates X n+1 =φ(P, X n ) by X n+1 However, X n is the X obtained in step S1 n is.
[0042] In step S211, the calculation process recording unit 253 of the differential value calculation unit 26 adds information necessary for automatic differentiation in the reverse mode to the calculation process information CP for the calculation performed by the state update unit 22 in step S211. The information necessary for automatic differentiation in the reverse mode may be recorded in the form of a calculation graph, for example, or in the form of a calculation graph for each X n+k is recorded as checkpoint information, and when backpropagation calculation is performed in step S214 described later, X n+kIn other words, in the third process described above, information on the calculation process of the second process recorded at the checkpoint is used. As an example, of the information included in the calculation process information CP, information on the calculation process of the second process recorded at the checkpoint may be used in the third process described above. Note that in each embodiment of the present application, "performing automatic differentiation in reverse mode" refers to performing a calculation of backpropagation of differential values, and does not include recording of the calculation process necessary for the original forward calculation and the calculation of backpropagation of differential values.
[0043] (Step S212) In step S212, the completion determination unit 28 determines whether step S211 has been executed a predetermined number of times. If step S211 has been executed a predetermined number of times, step S213 is executed. If step S211 has not been executed a predetermined number of times, step S211 is executed again. If the predetermined number of times step S211 is executed is m, X is obtained by executing step S211 m times. n+m The way in which m is determined is, for example, as follows.
[0044] The number of iterations m may be determined based on consideration of the target of the information processing method, for example, before the execution of step S1. For example, i If the convergence of X is known, m may be determined based on the convergence. i When it is known that converges faster than linear convergence, for example, when it is known that converges quadratically, m may be set to 1. This can also be expressed as the application of automatic differentiation may be one of multiple iterations in the first iterative process.
[0045] The number of iterations m may be determined based on a test calculation before the execution of step S1. aWhen calculating the derivative of F with respect to P at multiple values of P, a preferred value for the number of iterations m in step S2 may be determined for a representative value of P, and the preferred value for the number of iterations m may be used in the calculations at multiple values of P.
[0046] The number of iterations m is, for example, X obtained in step S1. i For example, in step S1, X i Using a parameter λ whose value is less than 1, |X i -X i-1 |<λ|X i-1 -X i-2 If the convergence is such that |, then λ m (λ to the mth power) is small enough, or λ m In this case, for example, the control unit 2 includes an iteration number determination unit (not shown), which determines the number of iterations m. i where |X i -X i-1 |<λ|X i-1 -X i-2 The fact that the convergence is such that | is the case when i is close to n i This may be determined from the behavior of the
[0047] (Step S213) In step S213, the function value calculation unit 24 calculates the value F(Q, X n+m In step S213, the calculation process recording unit 253 of the differential value calculation unit 26 records information necessary for performing automatic differentiation in the reverse mode regarding the calculation performed by the function value calculation unit 24 in step S213. This information is included in the calculation process information CP recorded in the storage unit 3.
[0048] In step S214, the backpropagation calculation unit 252 of the differential value calculation unit 26 calculates X by an iterative method based on the calculation process information CP recorded in the storage unit 3. n From X n+m Find the value of the function F(Q, X n+m) is calculated. (n,m) (P)) / ∂p a Here, in order to clarify the dependency on P, we newly add Y (n,m) (P) was introduced. (i,0) (P) is Y (i,0) =X i (P), which is a constant that does not depend on P. 、 Y (i,k) (P) is Y (i,0) It is a state variable obtained by applying φ k times starting from Y. For example, Y (i,1) (P) = φ(P, Y (i,0) ) For a parameter value P' different from P, Y (i,1) (P') = φ(P', Y (i,0) ) but in this case too, Y on the right side (i,0) Is Y (i,0) =X i (P). Y (i,k) (P) = X i+k (P), but for a parameter value P' different from P and k>0, Y (i,k) (P')≠X i+k (P'). As can be seen from the above description, the first process includes the first iteration of the process that is repeated in the first iterative process.
[0049] In this way, in step S2, X is calculated by the iterative method performed in steps S211 to S213. n From X n+m and then calculate the value of the function F(Q, X n+m ) is calculated by applying reverse mode automatic differentiation. a For ∂F(Q,X ANS (P)) / ∂p a As a numerical value of ∂F(Q, Y (n,m) (P)) / ∂p aAs can be seen from the above description, the target to which automatic differentiation is applied at least partially includes a calculation for determining the state variable X through a calculation that depends on a parameter, and the differential coefficient determined by automatic differentiation includes a contribution of the target function F that depends on the parameter P via one or more state variables X. Furthermore, as can be seen from the above description, automatic differentiation is applied to a process that includes step S211, which is a process that is repeated at least once, but does not include step S12, which is a process that is repeated at least once.
[0050] In this embodiment, the calculation of step S1 is performed before step S2, so that, for example, the number of iterations m in step S2 can be reduced, and each p a This can reduce the calculation cost of automatic differentiation for finding the derivative with respect to . This effect is significant when the ratio n / m of the number of iterations n in step S1 to the number of iterations m in step S2 is large.
[0051] In step S2, automatic differentiation in reverse mode is used, so that p a Even if the number of Q,Y is large, the computational cost can be reduced. (n,m) (P)) / ∂p a can be calculated.
[0052] As described above, in the information processing device 1 and the information processing method according to the present embodiment, one or more parameters (p a ) as an argument directly or indirectly, obtains an initial value for calculating the value of the object function (F), and calculates a differential coefficient of the object function (F), and the calculation process of the differential coefficient includes a first process (e.g., step S1) and a second process (e.g., steps S211 and S212). Also, the calculation process of the differential coefficient includes an iterative process, the first process includes a part of the iterative process, and the second process includes a part other than the part of the iterative process, and the calculation process of the differential coefficient includes a third process which is a calculation process of automatic differentiation for a process which does not include the first process but includes the second process. Therefore, according to the above configuration, the calculation process of each parameter (pa In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0053] Furthermore, as described above, the second process includes a process of calculating the differential coefficient of the target function (F) using automatic differentiation of backpropagation. This makes it possible to more effectively reduce the calculation cost. (Notes to the First Embodiment) The processes described in the first embodiment and the processes described in the following embodiments can also be considered to be processes based on the comparative example described below. Figure 5 is a flowchart showing an information processing method according to the comparative example.
[0054] The information processing method according to the comparative example includes steps S1010, S1020, and S1030.
[0055] (Step S1010) First, step S1010 is performed. Fig. 6 is a flowchart showing the flow of each process included in step S1010. As shown in Fig. 6, step S1010 includes steps S1011 to S1013.
[0056] (Step S1011) In step S1011, the initial value X of the state variable X is 0 is determined.
[0057] (Step S1012) Next, in step S1012, X i =φ(P, X i-1 ) to obtain the new value of the state variable X i Here, i corresponds to the number of times step S1012 is executed. For example, when step S1012 is executed for the first time, X 1 =φ(P, X 0 ) by X 1 is required.
[0058] (Step S1013) Next, in step S1013, X i It is determined whether X satisfies the convergence condition. iIf it is determined that does not satisfy the convergence condition, step S1013 is executed again. i If it is determined that satisfies the convergence condition, step S1012 is executed.
[0059] The convergence condition in step S1013 is, for example, X i and X i-1 This is similar to what will be explained in other embodiments, for example.
[0060] In step S1013, X i If it is determined that X satisfies the convergence condition, ANS As a numerical solution of (P), n Here, n is the number of X i is i when it is determined that the convergence condition is satisfied. 0 Starting from X obtained by applying φ n times n depends on P, so to make the dependency on P explicit, we use X n Also written as (P).
[0061] (Step S1020) X i If it is determined that F(Q, X) satisfies the convergence condition, then in step S1020, n ) is calculated.
[0062] (Step S1030) Next, in step S1030, back propagation of the differential value of F is calculated, which corresponds to the calculations in steps S1010 and S1020. As a result, ∂F(Q, X n (P)) / ∂p a is obtained.
[0063] In this way, by applying automatic differentiation to the calculations in steps S1010 and S1020, each p included in P a F(Q,X ANS (P)) ANS (P)) / ∂p a As a numerical solution, ∂F(Q,X n (P)) / ∂p aIn the above explanation, ∂F(Q,X n (P)) / ∂p a However, forward mode automatic differentiation may also be used.
[0064] When forward-mode automatic differentiation is used, for example, the calculations of steps S1010 and S1020 can be performed once using an extended double number with multiple non-real parts to obtain the desired value for each p a The derivative with respect to ∂F(Q, X ANS (P)) / ∂p a may be calculated.
[0065] When reverse mode automatic differentiation is used, the back propagation of the differential coefficients is calculated, and each p a For the differential coefficient ∂F(Q, X ANS (P)) / ∂p a can be calculated together.
[0066] [Embodiment 2] Next, another embodiment of the present invention, embodiment 2, will be described in detail. For convenience of explanation, members having the same functions as those described in the above embodiment will be denoted by the same reference numerals, and their description will not be repeated.
[0067] The information processing device according to this embodiment includes a differential value calculation unit 25 instead of the differential value calculation unit 26 included in the information processing device 1 according to the above-described first embodiment. Other configurations of the information processing device according to this embodiment are similar to those of the information processing device 1 according to the above-described first embodiment.
[0068] (Differential Value Calculation Unit 25) The differential value calculation unit 25 includes, for example, a forward propagation calculation unit 251. The forward propagation calculation unit 251 of the differential value calculation unit 25 performs calculation of forward propagation of differential values, which corresponds to the calculation performed by the state update unit 22. Details of the processing by the differential value calculation unit 25 will be described later.
[0069] <Processing flow by the information processing device according to this embodiment> The processing flow by the information processing device according to this embodiment will be described below. In summary, the information processing device according to this embodiment executes the following steps, similar to the information processing device 1 according to the first embodiment: A convergence solution calculation step (step S1, also referred to as the first processing) A processing S2b step (step S2b)
[0070] Step S2b includes steps S201, S202, and S203, as described below. Some of the processes performed in steps S201 and S202 are also referred to as second processes. The processes performed by the forward propagation calculation unit 251 in steps S201 and S203 are an example of processes included in third processes.
[0071] (Step S1) Step S1 according to this embodiment is the same as step S1 executed by the information processing device 1 according to the first embodiment, and therefore a description thereof will be omitted.
[0072] (Step S2b) After step S1 is completed, step S2b is performed. Step S2b according to this embodiment includes, as an example, steps S201, S202, and S203, which will be described below.
[0073] (Step S201) In step S201, the state update unit 22 n+k =φ(P, X n+k-1 ) by X n+k Here, k corresponds to the number of times step S201 has been executed. For example, when step S201 is executed for the first time, the state update unit 22 calculates X n+1 =φ(P, X n ) by X n+1 However, X n is the X obtained in step S1 n is.
[0074] In step S201, the forward propagation calculation unit 251 of the differential value calculation unit 25 performs forward propagation (forward mode) calculation of differential values, which corresponds to the calculation performed by the state update unit 22 in step S201.
[0075] (Step S202) Subsequently, in step S202, the end determination unit 28 determines whether step S201 has been executed a predetermined number of times. If step S201 has been executed a predetermined number of times, step S203 is executed. If step S201 has not been executed a predetermined number of times, step S201 is executed again. If the predetermined number of times step S201 is executed is m, X n+m is obtained.
[0076] (Step S203) Subsequently, in step S203, the function value calculation unit 24 calculates the value F(Q, X) of the function F. n+m ) is calculated.
[0077] In step S203, the forward propagation calculation unit 251 of the differential value calculation unit 25 performs calculation of forward propagation of differential values, which corresponds to the calculation performed by the function value calculation unit 24 in step S203.
[0078] In step S2b, an iterative method is performed to obtain X n From X n+m The objective function value F(Q, X n+m ) is calculated by applying forward mode automatic differentiation to each p a On the other hand, F's p a The derivative with respect to ∂F(Q, X ANS (P)) / ∂p a As a numerical value of ∂F(Q, Y (n,m) (P)) / ∂p a Here, Y (n,m) (P) is the Y introduced in embodiment 1. (n,m) Same as (P).
[0079] After step S2b is completed, the information processing device according to this embodiment transmits, for example, via the input / output unit 4, a ∂F(Q, Y (n,m) (P)) / ∂p a is output to the outside.
[0080] When automatic differentiation in forward mode is used as in this embodiment, each p aTo find the derivative with respect to, for example, p a In this embodiment, the calculation cost is approximately proportional to the number of a Automatic differentiation to obtain the derivatives with respect to is not applied to step S1. By performing step S1, for example, the number of iterations m in step S2b can be reduced, and each p of F can be a The calculation cost of automatic differentiation for determining the differential coefficient with respect to can be reduced compared to the comparative example. This effect is significant when the ratio n / m of the number of iterations n in step S1 to the number of iterations m in step S2b is large. Note that the method for setting the number of iterations m is the same as the process described in the first embodiment, so a description thereof will be omitted here.
[0081] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0082] Next, another embodiment of the present invention, that is, embodiment 3, will be described in detail. For the sake of convenience, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0083] The information processing device according to this embodiment includes a termination determination unit 29 instead of the termination determination unit 28 included in the information processing device according to the above-described embodiment 2. Furthermore, the storage unit 3 of the information processing device according to this embodiment stores a threshold value Cth2 instead of the number of iterations m in embodiment 2. Other configurations of the information processing device according to this embodiment are the same as those of the information processing device according to the above-described embodiment 2.
[0084] (End Determination Unit 29) The end determination unit 29 determines whether the target process satisfies a convergence condition. Details of the process by the end determination unit 29 will be described later.
[0085] <Processing Flow by Information Processing Device According to this Embodiment> The processing flow by the information processing device according to this embodiment will be described below. In outline, the information processing device according to this embodiment, like the information processing device 2 according to the second embodiment, executes: a convergence solution calculation step (step S1, also referred to as the first processing); and a processing S2c step (also referred to as step S2c). Step S2c includes steps S201, S203, and S205, as described below. Some of the processing performed in step S2c is also referred to as the second processing. The processing performed by the forward propagation calculation unit 251 in steps S201 and S203 is an example of processing included in the third processing.
[0086] (Step S1) Step S1 according to this embodiment is the same as step S1 executed by the information processing device 1 according to the first embodiment, and therefore a description thereof will be omitted.
[0087] (Step S2c) After step S1 is completed, step S2c is performed. Step S2c according to this embodiment includes, as an example, steps S201, S203, and S205, which will be described below.
[0088] (Step S201) In step S201, the state update unit 22 n+k =φ(P, X n+k-1 ) by X n+k Here, k corresponds to the number of times step S201 has been executed. For example, when step S201 is executed for the first time, the state update unit 22 calculates X n+1 =φ(P, X n ) by X n+1 However, X n is the X obtained in step S1 n In step S201, the forward propagation calculation unit 251 of the differential value calculation unit 25 performs a calculation of forward propagation of differential values, which corresponds to the calculation performed by the state update unit 22 in step S201.
[0089] (Step S203) Subsequently, in step S203, the function value calculation unit 24 calculates the value F(Q, X) of the function F. n+k ) is calculated.
[0090] In step S203, the forward propagation calculation unit 251 of the differential value calculation unit 25 performs calculation of forward propagation of differential values, which corresponds to the calculation performed by the function value calculation unit 24 in step S203.
[0091] Step S201 is repeatedly executed to n Based on X n+k In step S203, the value F(Q, X n+k ) is calculated, the parameter p a The derivatives ∂F(Q, Y (n,k) (P)) / ∂p a is obtained. Here, Y (n,k) is the same as Y introduced in the first embodiment.
[0092] (Step S205) Next, in step S205, the end determination unit 29 calculates ∂F(Q, Y (n,k) (P)) / ∂p a In other words, the termination determination unit 29 executes a convergence determination process for the differential coefficient of the target function (F). (n,k) (P)) / ∂p a If it is determined that ∂F(Q, Y) satisfies the convergence condition, step S2c is completed. In this case, if it is determined that ∂F(Q, Y) satisfies the convergence condition, (n,k) (P)) / ∂p a is ∂F(Q, X ANS (P)) / ∂p a This is the numerical solution of ∂F(Q, Y (n,k) (P)) / ∂p a If it is determined that the convergence condition is not satisfied, step S201 is executed again.
[0093] In step S205, ∂F(Q, Y (n,k) (P)) / ∂p a The method by which the termination determination unit 29 determines whether g has converged may be determined appropriately depending on the purpose. For example, (k,a) = ∂F(Q, Y (n,k) (P)) / ∂p a As such, (Σa (g (k,a) -g (k-1,a) ) 2 ) 0.5 <Cth2 is satisfied, then ∂F(Q, Y (n,k) (P)) / ∂p a is determined to be converged, otherwise, ∂F(Q, Y (n,k) (P)) / ∂p a is determined to not have converged.
[0094] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0095] Next, another embodiment of the present invention, embodiment 4, will be described in detail. For convenience of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0096] The information processing device according to this embodiment includes a termination determination unit 29 instead of the termination determination unit 28 included in the information processing device 1 according to the above-described embodiment 1. Furthermore, the storage unit 3 of the information processing device according to this embodiment stores a threshold value Cth2 instead of the number of iterations m in embodiment 1. Other configurations of the information processing device according to this embodiment are similar to those of the information processing device 1 according to the above-described embodiment 1.
[0097] (End Determination Unit 29) The end determination unit 29 determines whether the target process satisfies a convergence condition. Details of the process by the end determination unit 29 will be described later.
[0098] <Processing flow by the information processing device according to this embodiment> The processing flow by the information processing device according to this embodiment will be described below. In summary, the information processing device according to this embodiment executes the following steps, similar to the information processing device 1 according to the first embodiment: a convergence solution calculation step (step S1, also referred to as the first process); and a process S2d step (step S2d).
[0099] Step S2d includes steps S211, S213, S214, and S205, as described below. Steps S211 and S213 are also referred to as a second process. Steps S214 and S205 are an example of a process included in a third process.
[0100] (Step S1) Step S1 according to this embodiment is the same as step S1 executed by the information processing device 1 according to the first embodiment, and therefore a description thereof will be omitted.
[0101] (Step S2d) After step S1 is completed, step S2d is performed. Step S2d according to this embodiment includes, as an example, steps S211, S213, S214, and S205, which will be described below.
[0102] (Step S211) In step S211, the state update unit 22 n+k =φ(P, X n+k-1 ) by X n+k In step S211, the calculation process recording unit 253 of the differential value calculation unit 26 adds information required for automatic differentiation in the reverse mode to the calculation process information CP, regarding the calculation performed by the state update unit 22 in step S211.
[0103] (Step S213) Subsequently, in step S213, the function value calculation unit 24 calculates the value F(Q, X) of the function F. n+k In step S213, the calculation process recording unit 253 of the differential value calculation unit 26 adds information required for automatic differentiation in the reverse mode to the calculation process information CP regarding the calculation performed by the function value calculation unit 24 in step S213.
[0104] (Step S214) Subsequently, in step S214, the back propagation calculation unit 252 of the differential value calculation unit 26 performs back propagation calculation based on the calculation process information CP recorded in the storage unit 3. As a result, ∂F(Q, Y (n,k) (P)) / ∂p a For example, the back propagation calculation unit 252 calculates X n Starting from X by iterative calculationn+k Then, F(Q, X n+k ) and perform the backpropagation calculation corresponding to the entire calculation.
[0105] (Step S205) Subsequently, in step S205, the end determination unit 29 determines whether ∂F(Q,Y (n,k) (P)) / ∂p a In other words, the termination determination unit 29 executes a convergence determination process for the differential coefficient of the target function (F). (n,k) (P)) / ∂p a If it is determined that ∂F(Q,Y) satisfies the convergence condition, step S2d is completed. In this case, if it is determined that ∂F(Q,Y) satisfies the convergence condition, (n,k) (P)) / ∂p a is ∂F(Q,X ANS (P)) / ∂p a This is the numerical solution of ∂F(Q, Y (n,k) (P)) / ∂p a If it is determined that ∂F(Q,Y (n,k) (P)) / ∂p a The method by which the termination determination unit 29 determines whether satisfies the convergence condition is the same as the method described in the third embodiment, for example.
[0106] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0107] As described above, the third process includes a process of calculating the differential coefficient of the objective function (F) using automatic differentiation in reverse mode, which can further reduce the calculation cost.
[0108] Next, another embodiment of the present invention, embodiment 5, will be described in detail. For the sake of convenience, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0109] The information processing device according to this embodiment is similar to the information processing device according to the second embodiment described above.
[0110] <Processing flow by the information processing device according to this embodiment> The processing flow by the information processing device according to this embodiment will be described below. In summary, the information processing device according to this embodiment executes the following steps, similar to the information processing device 2 according to the second embodiment: a convergence solution calculation step (also referred to as step S1); and a processing S2e step (also referred to as step S2e).
[0111] (Step S1) Step S1 according to this embodiment is the same as step S1 executed by the information processing device 1 according to the first embodiment, and therefore a description thereof will be omitted.
[0112] (Step S2e) Step S2e is performed after step S1 is completed. Step S2e according to this embodiment includes, as an example, step S201e, step S202, and step S203, which will be described below.
[0113] (Step S201e) In step S201e, the state update unit 22 n-m+k =φ(P, X n-m+k-1 ) by X n-m+k Here, k corresponds to the number of times step S201e has been executed. For example, when step S201e is executed for the first time, the state update unit 22 calculates X n-m+1 =φ(P, X n-m ) by X n-m+1 However, X n-m is the X obtained in step S1 n-m and X when i=n-m i Furthermore, m is the number of iterations m stored in the storage unit 3.
[0114] The value of n is not known until step S1 is completed, and the value of nm is also not known. Therefore, for example, in step S1, each X obtained each time step S12 is executed is i is added to the calculation process information CP, and in step S201e, X stored in the calculation process information CP is n-mAlternatively, before executing step S2e, step S1 is executed again from the beginning to the middle, and X n-m In this case, the memory unit 3 stores each X i You don't need to memorize it.
[0115] In step S201e, the forward propagation calculation unit 251 of the differential value calculation unit 25 performs calculation of forward propagation of differential values, which corresponds to the calculation performed by the state update unit 22 in step S201e.
[0116] (Step S202) Subsequently, in step S202, the end determination unit 28 determines whether step S201e has been executed a predetermined number of times. If step S201e has been executed a predetermined number of times, step S203 is executed. If step S201e has not been executed a predetermined number of times, step S201e is executed again. If the predetermined number of times step S201e is executed is m, X is executed again by executing step S201e m times. n is obtained.
[0117] (Step S203) Subsequently, in step S203, the function value calculation unit 24 calculates the value F(Q, X) of the function F. n In step S203, the forward propagation calculation unit 251 of the differential value calculation unit 25 performs forward propagation calculation of the differential value, which corresponds to the calculation performed by the function value calculation unit 24 in step S203.
[0118] From the above, each p included in P a On the other hand, F's p a The derivative with respect to ∂F(Q, X ANS (P)) / ∂p a As a numerical value of ∂F(Q, Y (n-m,m) (P)) / ∂p a The number of iterations m in step S2e is determined in the same manner as in the first embodiment, for example. The number of iterations m is determined so as to satisfy m<n. As a result, ∂F(Q,Y (n-m,m) (P)) / ∂p aThe cost of calculating X can be reduced. 0 Starting from and repeatedly applying φ to X n-m The process up to obtaining X is also called the first process. n-m From X n The process up to obtaining ( ) or part of the process in step S201e is also called a second process. The process performed by the forward propagation calculation unit 251 in step S201e is an example of a process included in a third process.
[0119] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0120] Next, another embodiment of the present invention, embodiment 6, will be described in detail. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0121] The information processing device according to this embodiment is similar to the information processing device 1 according to the first embodiment described above.
[0122] <Processing flow by the information processing device according to this embodiment> The processing flow by the information processing device according to this embodiment will be described below. In summary, the information processing device according to this embodiment executes the following steps, similar to the information processing device 1 according to the first embodiment: a convergence solution calculation step (also referred to as step S3); and a processing S4 step (also referred to as step S4).
[0123] (Step S3) Step S3 includes, for example, steps S31, S32, and S33 as shown in FIG.
[0124] (Step S31) In step S31, the initial value determination unit 21 determines the initial value X of the state variable X. 0 The following is established.
[0125] (Step S32) Next, in step S32, the state update unit 22 updates the state variable X. That is, X i =φ(P, X i-1 ) to obtain the value of the new state variable X, X i Here, i corresponds to the number of times step S32 has been executed. For example, when step S32 is executed for the first time, the state update unit 22 calculates X 1 =φ(P, X 0 ) by X 1 This step S32 can be expressed as constituting at least a part of the process of finding a numerical solution for the state variable X using the parameter P as an argument through an iterative process (first iterative process).
[0126] In step S32, the calculation process recording unit 253 records information necessary for automatic differentiation in the reverse mode regarding the calculation performed by the state update unit 22 in step S32. The information necessary for automatic differentiation in the reverse mode may be recorded in the form of, for example, a calculation graph, or i is recorded as checkpoint information, and when backpropagation calculation is performed in step S414 described later, X i A computation graph may be constructed using
[0127] (Step S33) Next, in step S33, the termination determination unit 23 determines whether to terminate the iteration of the calculation in step S32. The method for determining in step S33 whether to terminate the iteration of the calculation in step S32 by the termination determination unit 23 may be the same as in the first embodiment.
[0128] (Step S4) Step S4 according to the present embodiment includes, as an example, step S413 and step S414, which will be described below.
[0129] (Step S413) In step S413, the function value calculation unit 24 calculates the value F(Q, X nIn step S413, the calculation process recording unit 253 of the differential value calculation unit 26 adds information necessary for performing automatic differentiation in the reverse mode regarding the calculation performed by the function value calculation unit 24 in step S213 to the calculation process information CP.
[0130] (Step S414) Subsequently, in step S414, the back propagation calculation unit 252 of the differential value calculation unit 26 performs back propagation calculations corresponding to the calculations in step S32 from the n-m+1th to the nth calculations and the calculations in step S413 based on the calculation process information CP recorded in the storage unit 3. As a result, ∂F(Q,Y (n-m,m) (P)) / ∂p a is obtained.
[0131] The value m in step S4 is determined in the same manner as the value m in step S2e in the fifth embodiment. The number of iterations m is determined so as to satisfy m<n. This allows the value ∂F(Q,Y (n-m,m) (P)) / ∂p a In addition, by using reverse mode automatic differentiation, the cost of calculating p a Even if there are a large number of p a For ∂F(Q, Y (n-m,m) (P)) / ∂p a In the process performed in step S3, X 0 Starting from and repeatedly applying φ to X n-m The process up to obtaining X is also called the first process. n-m From X n The process up to obtaining is also called the second process. Step S413 may be included in the second process. Step S414 is also called the third process.
[0132] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0133] As described above, the third process includes a process of calculating the differential coefficient of the objective function (F) using automatic differentiation in reverse mode, which can further reduce the calculation cost.
[0134] Next, another embodiment of the present invention, embodiment 7, will be described in detail. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0135] The information processing device according to this embodiment further includes a termination determination unit 29 in addition to the components included in the information processing device according to the second embodiment. Furthermore, the storage unit 3 of the information processing device according to this embodiment stores a threshold value Cth2 instead of the number of iterations m. Other components of the information processing device according to this embodiment are the same as those of the information processing device according to the second embodiment.
[0136] <Processing flow by the information processing device according to this embodiment> The processing flow by the information processing device according to this embodiment will be described below. In summary, the information processing device according to this embodiment executes the following steps, similar to the information processing device according to the second embodiment: A convergence solution calculation step (also referred to as step S1) A processing S2g step (also referred to as step S2g).
[0137] (Step S1) Step S1 according to this embodiment is the same as step S1 executed by the information processing device 1 according to the first embodiment, and therefore a description thereof will be omitted.
[0138] (Step S2g) After step S1 is completed, step S2g is performed. Step S2g according to this embodiment includes steps S200g, S201g, S202, S203, and S205g, which will be described below.
[0139] (Step S200g) In step S200g, an initial state for calculating the derivative is set. When step S200g is executed for the kth time, X n-k is set as the initial state for the calculation of the derivatives.
[0140] (Step S201g) In step S201g, the state update unit 22 n-k+s =φ(P, X n-k+s-1 ) to obtain the new value of the state variable X, X n-k+s where s corresponds to the number of times step S201g has been executed since the last execution of step S200g.
[0141] In step S201g, the forward propagation calculation unit 251 performs calculations of forward propagation of differential values, which correspond to the calculations performed by the state update unit 22 in step S201g.
[0142] (Step S202) Subsequently, in step S202, the completion determination unit 28 determines whether step S201g has been executed k times after step S200g. If step S201g has been executed k times after step S200g, step S203 is executed. If step S201g has not been executed k times after step S200g, step S201g is executed again. By executing step S201g k times after step S200g, X n-k From X n is recalculated and the corresponding forward propagation calculation is performed.
[0143] (Step S203) Subsequently, in step S203, the function value calculation unit 24 calculates the value F(Q, X) of the function F. n In step S203, the forward propagation calculation unit 251 performs a calculation of forward propagation of the differential value, which corresponds to the calculation performed by the function value calculation unit 24 in step S203. a The derivative with respect to ∂F(Q, X ANS (P)) / ∂p a As a candidate for the numerical value of ∂F(Q, Y (n-k,k) (P)) / ∂p a is obtained.
[0144] (Step S205g) Next, in step S205g, the end determination unit 29 determines whether ∂F(Q,Y (n-k,k) (P)) / ∂p aIn other words, the termination determination unit 29 executes a convergence determination process for the derivative of the target function (F). (n-k,k) (P)) / ∂p a If it is determined that ∂F(Q,Y) satisfies the convergence condition, step S2g is completed. In this case, ∂F(Q,Y) satisfies the convergence condition. (n-k,k) (P)) / ∂p a is ∂F(Q,X ANS (P)) / ∂p a This is the numerical solution of ∂F(Q, Y (n-k,k) (P)) / ∂p a If it is determined that does not satisfy the convergence condition, step S200g is executed again.
[0145] In step S205g, ∂F(Q,Y (n-k,k) (P)) / ∂p a The method for determining whether g satisfies the convergence condition may be determined appropriately depending on the purpose. For example, (n-k,a) = ∂F(Q, Y (n-k,k) (P)) / ∂p a As such, (Σ a (g (n-k,k) -g(n-k+1,k+1)) 2 ) 0.5 <Cth2 is satisfied, then ∂F(Q, Y (n-k,k) (P)) / ∂p a is determined to be converged, otherwise, ∂F(Q,Y (n-k,k) (P)) / ∂p a is determined to not have converged.
[0146] In step S205g, ∂F(Q, Y (n-k,k) (P)) / ∂p a When it is determined that the convergence condition is satisfied, k is set to m. It is expected that m is smaller than n. When m is smaller than n, the calculation cost can be reduced compared to the comparative example, and ∂F(Q,X ANS (P)) / ∂p a In the process performed in step S1, the numerical solution of X 0 Starting from and repeatedly applying φ to X n-mThe process up to obtaining X is also called the first process. n-m From X n The process up to obtaining ( ) or part of the process of step S2g is also called a second process. The process performed by the forward propagation calculation unit 251 in step S2g is an example of a process included in a third process.
[0147] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0148] Next, another embodiment of the present invention, embodiment 8, will be described in detail. For ease of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0149] 8 is a block diagram showing the configuration of an information processing device 1h according to this embodiment. The configuration of the information processing device 1h is the same as the configuration of the information processing device according to the fourth embodiment.
[0150] <Processing flow by the information processing device according to this embodiment> The processing flow by the information processing device according to this embodiment will be described below. Fig. 9 is a flowchart showing an information processing method of this embodiment. As shown in Fig. 9, the information processing device according to this embodiment generally executes the following steps, similar to the information processing device according to the first embodiment: A convergence solution calculation step (also referred to as step S3) A processing S4h step (also referred to as step S4h)
[0151] (Step S3) Step S3 is the same as step S3 in the sixth embodiment, and therefore a description thereof will be omitted here.
[0152] (Step S4h) As shown in FIG. 10, step S4h includes step S413, step S414h, and step S415.
[0153] (Step S413) Step S413 is the same as step S413 in the sixth embodiment.
[0154] (Step S414h) Subsequently, in step S414h, the backpropagation calculation unit 252 performs backpropagation calculation on the differential value. The backpropagation calculation is performed based on the calculation process information added to the calculation process information CP in steps S3 and S413.
[0155] When step S414h is executed for the kth time, X n-k Starting from, by applying φ k times, we get X n Find F(Q, X n ) is calculated, the backpropagation calculation is performed to obtain ∂F(Q,X ANS (P)) / ∂p a As a candidate for the numerical value of ∂F(Q, Y (n-k,k) (P)) / ∂p a It is to be noted that X n-k Starting from, by applying φ k times, we get X n is part of the calculation performed in step S3. As can be seen from the above description, the first process includes the first iteration of the process performed in the first iterative process.
[0156] When step S414h is executed the kth time, the calculation result when step S414f was executed the (k-1)th time can be used, and it is sufficient to perform calculations only for the part that requires additional calculations from the calculation when step S414f was executed the (k-1)th time. Note that, as can be seen from the above explanation, the target to which automatic differentiation is applied at least partially includes a calculation for determining a state variable X through a calculation that depends on a parameter, and the differential coefficient determined by automatic differentiation includes a contribution of the target function F that depends on the parameter P via one or more state variables X.
[0157] (Step S415) Subsequently, in step S415, the end determination unit 29 determines whether ∂F(Q,Y (n-k,k) (P)) / ∂p asatisfies the convergence condition. In other words, the termination determination unit 29 executes a convergence determination process for the derivative of the target function (F). The method of determining whether the convergence condition is satisfied in step S415 is the same as, for example, step S205g in the seventh embodiment.
[0158] In step S415, ∂F(Q,Y (n-k,k) (P)) / ∂p a If it is determined that ∂F(Q,Y (n-k,k) (P)) / ∂p a If it is determined that the convergence condition is not satisfied, S414h is executed again.
[0159] In step S415, ∂F(Q,Y (n-k,k) (P)) / ∂p a When it is determined that the convergence condition is satisfied, k is set to m. It is expected that m is smaller than n. When m is smaller than n, the calculation cost can be reduced compared to the comparative example, and ∂F(Q,X ANS (P)) / ∂p a In the process performed in step S3, the numerical solution of X 0 Starting from and repeatedly applying φ to X n-m The process up to obtaining X is also called the first process. n-m From X n The process up to obtaining F is also called second process. Step S413 may be included in the second process. Steps S414h and S415 are also called third process. As described above, in this embodiment, the third process includes a calculation process of automatic differentiation for a process including the second process and a convergence determination process for the differential coefficient of the object function F.
[0160] The information processing device and information processing method according to this embodiment also allow each parameter (p a In the calculation for determining the differential coefficient with respect to , the calculation cost can be reduced while maintaining high accuracy of the differential calculation by using automatic differentiation.
[0161] As described above, the third process includes a process of calculating the differential coefficient of the objective function (F) using automatic differentiation in reverse mode. This makes it possible to more effectively reduce the calculation cost. [Embodiment 9] In the first to eighth embodiments, X i =φ(P, X i-1 ) by X i-1 From X i To accelerate or improve convergence, known methods may be employed, for example, mixing may be used to find new state variables.
[0162] For example, when simple mixing is used, it is as follows: X i '=φ(P, X i-1 ) and X i ', each state variable in x' b In this case, for each b, x is calculated using a predetermined α. (i,b) = αx' (i,b) +(1-α)x (i-1,b) Therefore, X i Each state variable x included in (i,b) Here, α is a weighting coefficient, for example, between 0 and 1. In other words, in the information processing method described in each embodiment, the state variable X i-1 From X i '=φ(P, X i-1 ) to derive a new state variable, and the new state variable X i ' and the (original) state variable X i-1 The state variable X is calculated by the average or weighted average of i In this embodiment, Anderson mixing, Broyden mixing, or the like may be used instead of simple mixing.
[0163] [Embodiment 10] In the first to ninth embodiments, the information processing device a Although the first derivative of F with respect to ∂ is calculated, the second or higher derivative may be calculated. In other words, the second process described in each embodiment may include a process of calculating the second or higher derivative with respect to the objective function (F) using automatic differentiation. For example, with respect to each a0 and each a1, ∂2 F / ∂p a0 ∂p a1 Here, a0 and a1 are indices that list the parameters included in P.
[0164] [Embodiment 11] In the first to ninth embodiments, the information processing device a In this embodiment, the application of the information processing methods of the first to tenth embodiments will be described.
[0165] <Application 1> The information processing methods of the first to tenth embodiments are applied to, for example, first-principles electronic state calculations, in which an electronic state is calculated based on a Hamiltonian expressed by the following formula: In the above formula (1), H represents the first-principles Hamiltonian, Ψ represents the wave function, T represents the kinetic energy, V represents the potential between the atom and the electron, U represents the electron-electron interaction, and E represents the energy. i , r j is the electron position. The dependence of the Hamiltonian on the atomic position is V(r i ) is included in the functional form of
[0166] In first-principles electronic state calculations, for example, the Hartree-Fock approximation or density functional theory is used to repeatedly update the electronic state and calculate the electronic state at the ground state or a finite temperature. In this case, the state variable X is, for example, a set of variables that specify the electronic state. The state variable X may include a variable that specifies the atomic position. φ is a function that updates the electronic state to find the electronic state at the ground state or a finite temperature in the Hartree-Fock approximation or density functional theory.
[0167] Although the formula (1) does not include a term corresponding to an external field, for example, any one or more of a magnetic field, an electric field, an electromagnetic wave, a sound wave, and pressure may be applied as the external field. a is, for example, a parameter describing the applied external field. a may be, for example, the temperature. Also, the atomic position may be positioned as a parameter instead of the state variable X.
[0168] The function F corresponds to, for example, a superconducting transition temperature, a superconducting gap, a thermoelectric efficiency, a solar power generation efficiency, or a linear or nonlinear response (e.g., a Hall coefficient) to an external field (e.g., a magnetic field, an electric field, an electromagnetic wave, a sound wave, or pressure). The storage unit 3 stores an equation for calculating the value of the function F based on the electronic state.
[0169] <Application 2> Instead of the first-principles Hamiltonian shown in the above (Equation 1), for example, the Hamiltonian of the tight-binding model shown below may be considered. On the right side of the above formula (2), the first term represents electron hopping, the second term represents electron-electron interaction, the third term represents phonon dispersion, and the fourth term represents electron-phonon interaction. The formula (2) includes a case where a quantum spin system is described as an example where all terms on the right side except the second term are zero.
[0170] To calculate the ground state or the state at finite temperature in the Hamiltonian shown in (Equation 2), the electronic state is iteratively updated using, for example, the Hartree approximation or the Hartree-Fock approximation method, and the electronic state at the ground state or the finite temperature is calculated. In this case, the state variable X is, for example, a set of variables that specify the electronic state. φ is a function that updates the electronic state based on the Hartree approximation or the Hartree-Fock approximation method.
[0171] In order to calculate the ground state or the state at a finite temperature in the Hamiltonian shown in (Equation 2), for example, the Lanczos method is used to repeatedly update the electronic state to calculate the electronic state at the ground state or the finite temperature. In this case, the state variable X is, for example, a set of variables that specify the electronic state. φ is, for example, a function that updates the electronic state based on the Lanczos method.
[0172] For example, each of t, V, ω, and g (subscripts omitted) in the formula (2) or some of them is p a An example of the function F is the same as that described in <Application 1>.
[0173] When calculating finite temperatures based on the Lanczos method, X is calculated by iterative calculation starting from a certain initial condition. n Do this for multiple initial conditions, and calculate each X n The expected value of the physical quantity of the physical system under consideration is calculated by averaging the physical quantities calculated based on the function F. n It may be calculated based on the following.
[0174] <Application 3> The following formula shows the Hamiltonian that indicates the interaction between classical spins. On the right side of the above formula (3), S i , S j is a vector representing classical spin. i For example, the variable specified in polar coordinates is x b For example, each S i Given that Σ j≠i J ij S j Each new S i The operation of determining φ corresponds to φ. For example, J ij Each of these is p a In addition to the term shown in (Equation 3), an external magnetic field may be considered, and p a may be a parameter representing, for example, an external magnetic field. F corresponds, for example, to the magnetization or magnetic susceptibility.
[0175] <Application 4> φ may be used to solve a self-consistent equation in a weakly coupled analysis that analyzes multiple systems that interact with each other. The multiple systems that are the targets of the weakly coupled analysis may represent physically different objects, or may represent different properties of the same physical object.
[0176] The weakly coupled analysis, for example, combines analysis of a structure and analysis of a fluid. A fluid analysis is performed under a given state of the structure to determine a new state of the fluid, and a new state of the structure is determined under the new state of the fluid. In this way, the function that determines the new state of the fluid and the new state of the structure corresponds to φ. By repeating this process, the states of the structure and fluid that interact with each other can be determined. The parameter p a is a parameter that represents the property of the structure, for example, its strength. b are variables that represent the state of the structure and the state of the fluid. The function F corresponds to, for example, the displacement of the structure.
[0177] [Embodiment 12] Next, another embodiment of the present invention, embodiment 12, will be described in detail. For convenience of explanation, the same reference numerals will be used to designate components having the same functions as those described in the above embodiment, and the description thereof will not be repeated.
[0178] 11 is a diagram showing an information processing device 100 according to this embodiment. As shown in FIG. 11, the information processing device 100 includes an information processing device 1, a control unit 5, a storage unit 6, and an input / output unit 7.
[0179] (Control unit 5) The control unit 5 has an initial value setting unit 51, a parameter updating unit 52, and an end determination unit 53. Here, as an example, the initial value setting unit 51 sets an initial value of a parameter P. As an example, the parameter updating unit 52 updates the parameter P. As an example, the end determination unit 53 determines whether to end the update of the parameter P. Specific processes performed by each unit included in the control unit 5 will be described later.
[0180] (Storage Unit 6, Input / Output Unit 7) The storage unit 6 stores, for example, data and information similar to those stored in the storage unit 3 of the information processing device 1. The input / output unit 7 has, for example, a configuration similar to that of the input / output unit 4 of the information processing device 1.
[0181] The information processing device 100 may include the information processing device described in any one of the second to tenth embodiments instead of the information processing device 1.
[0182] <Processing flow by the information processing device according to this embodiment> In the first to eleventh embodiments, each p a The differential value of the function F with respect to each p a The parameter P is optimized using the value of the derivative of the function F with respect to F. For example, the parameter P is determined so as to maximize or minimize the function F. FIG. 12 is a flowchart showing the information processing method of this embodiment. As shown in FIG. 12, the information processing method according to this embodiment includes steps S5 to S8. The iterative process consisting of steps S6 to S8 shown in FIG. 12 is also referred to as the second iterative process. The first iterative process described above corresponds to the process of step S6.
[0183] (Step S5) First, in step S5, the initial value setting unit 51 sets the initial value P of the parameter P. 0 Set the parameter P 0 may be given randomly, or may be input from the outside via an input / output unit.
[0184] (Step S6) Next, in step S6, the information processing device 1 a Calculate the value of the derivative of the function F with respect to
[0185] (Step S7) Next, in step S7, the parameter update unit 52 updates each p a The parameter P is updated based on the value of the derivative of the function F with respect to the parameter P. The parameter update unit 52 updates the parameter P by, for example, the steepest descent method. a In step S7, the parameter update unit 52 may calculate the Hessian of the function F with respect to the parameter P by the Newton method. The value of the parameter P obtained when step S7 is executed for the i-th time may be referred to as P i Also, P i The values of the parameters included in p (i,a)In step S7, the quasi-Newton method or the conjugate gradient method may be used. The process in step S7 can also be expressed as a process of updating one or more parameters using the differential coefficient of the target function F.
[0186] (Step S8) Next, in step S8, the end determination unit 53 determines whether to end the update of the parameter P. The conditions for determining whether to end the update of the parameter P include, for example, whether step S7 has been performed a predetermined number of times or whether the parameter P i This is a condition on whether or not (Σ a (p (i,a) -p (i-1,a) ) 2 ) 0.5 <Cth3 (where Cth3 is a threshold value) is satisfied, the parameter P i is determined to satisfy the convergence condition, otherwise the parameter P i It is determined that the parameter P does not satisfy the convergence condition. i If the parameter P satisfies the convergence condition, the process ends, otherwise the process returns to step S6. i satisfies the convergence condition, it may be determined whether the norm of the differential value of the function F obtained in step S6 is smaller than a predetermined threshold value.
[0187] As described above, the information processing device of this embodiment calculates the one or more parameters (p a) is included. The information processing device of this embodiment optimizes the one or more parameters so as to maximize or minimize (minimize) the value of the object function F through processing by the parameter updating unit 52 and the termination determination unit 53. The information processing device of this embodiment may, for example, optimize the one or more parameters so that the value of the object function F becomes a target value. Furthermore, the information processing device of this embodiment may, for example, optimize the one or more parameters using the differential coefficient of the object function so as to maximize or minimize (minimize) the value of a function G other than the object function, or so that the value of function G becomes a target value. The differential value of function G is calculated, for example, using the differential value of function F calculated by the method described in any of embodiments 1 to 11. Function G is, for example, the sum of function F and a function H other than function F. In this case, the differential value of function G is calculated as the sum of the differential value of function F and the differential value of function H. The derivative of the function H may be calculated by the method described in any one of the first to eleventh embodiments, or may be calculated by another method.
[0188] [Example of implementation by software] The functions of the information processing device (hereinafter referred to as "device") according to each embodiment can be realized by a program for causing a computer to function as the device, and a program for causing a computer to function as each control block of the device (particularly each part included in the control unit).
[0189] In this case, the device includes a computer having at least one control device (e.g., a processor) and at least one storage device (e.g., a memory) as hardware for executing the program. The functions described in each of the above embodiments are realized by executing the program using the control device and storage device.
[0190] The program may be non-transitory and may be recorded on one or more computer-readable recording media. The recording media may or may not be included in the device. In the latter case, the program may be supplied to the device via any wired or wireless transmission medium.
[0191] Furthermore, some or all of the functions of the control blocks can be realized by logic circuits. For example, an integrated circuit in which a logic circuit that functions as each of the control blocks is formed is also included in the scope of the present invention. In addition, the functions of the control blocks can also be realized by, for example, a quantum computer.
[0192] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0193] 1, 1h, 100... Information processing device 2, 2d, 5... Control unit 21... Initial state determination unit (acquisition unit) 22... State update unit 23... End determination unit 24... Function value calculation unit 26... Differential value calculation unit (calculation unit) 28... End determination unit
Claims
1. An acquisition unit that acquires an initial value regarding calculation of a value of a target function that directly or indirectly includes one or more parameters as arguments, and a control unit that calculates a derivative coefficient of the target function, wherein the calculation process of the derivative coefficient by the control unit includes a first process and a second process, the calculation process of the derivative coefficient by the control unit includes a first iterative process, the first process includes a part of the first iterative process, the second process includes a part other than the part of the first iterative process, the calculation process of the derivative coefficient by the control unit includes a third process that includes a calculation process of automatic differentiation for a process that does not include the first process and includes the second process, the first process includes the first time of the processes repeated in the first iterative process, and the first iterative process is an iterative process for obtaining a solution of a self-collision-free equation. An information processing apparatus.
2. The target function takes as arguments a subset of the one or more parameters and one or more state variables that take the one or more parameters as arguments, the iterative process includes an update process of the one or more state variables, and the third process includes a process of calculating a derivative coefficient of the target function using automatic differentiation. The information processing apparatus according to claim 1.
3. The control unit obtains a numerical solution of the one or more state variables that take the one or more parameters as arguments by the first iterative process, and the third process includes a process of calculating, using automatic differentiation, a derivative coefficient regarding the one or more parameters of the target function, the derivative coefficient including a contribution in which the target function depends on the one or more parameters via the one or more state variables. The information processing apparatus according to claim 2.
4. The information processing apparatus according to claim 1, further comprising a parameter update unit that updates the one or more parameters using the derivative coefficient of the target function, wherein the control unit performs a second iterative process, and in the second iterative process, a process including the first iterative process and a process of updating the one or more parameters using the derivative coefficient of the target function is repeated.
5. The objective function takes as arguments a subset of the one or more parameters and one or more state variables that take the one or more parameters as arguments. The first iterative process includes an update process for the one or more state variables. The control unit obtains a numerical solution of the one or more state variables that take the one or more parameters as arguments by the first iterative process. The third process includes a process of calculating, using automatic differentiation, a derivative coefficient with respect to the one or more parameters of the objective function, the derivative coefficient including a contribution in which the objective function depends on the one or more parameters via the one or more state variables. The automatic differentiation is applied to one or more of the processes that are repeated a plurality of times in the first iterative process and do not include at least one time. The information processing apparatus according to claim 4.
6. The third process includes a process of calculating a derivative coefficient of the objective function using reverse-mode automatic differentiation. The information processing apparatus according to any one of claims 1 to 5.
7. At least any one of the first process, the second process, and the third process includes at least any one of a convergence determination process for at least any one of the one or more state variables and a convergence determination process for a derivative coefficient of the objective function. The information processing apparatus according to claim 6.
8. In the third process, information on the calculation process of the second process recorded at a checkpoint is used. The information processing apparatus according to claim 6.
9. An acquisition step of acquiring an initial value regarding calculation of a value of a target function that directly or indirectly includes one or more parameters as arguments, and a calculation step of calculating a derivative coefficient of the target function, wherein the calculation process of the derivative coefficient by the calculation step includes a first process and a second process, the calculation process of the derivative coefficient by the calculation step includes a first iterative process, the first process includes a part of the first iterative process, the second process includes a part other than the part of the first iterative process, the calculation process of the derivative coefficient by the calculation step includes a third process including a calculation process of automatic differentiation for a process including the second process and not including the first process, the first process includes the first time among the processes repeated in the first iterative process, and the first iterative process is an iterative process for obtaining a solution of a self-collision-free equation. An information processing method.
10. A program for causing a computer to function as the information processing apparatus according to claim 1, the program for causing a computer to function as the acquisition unit and the control unit.
Citation Information
Patent Citations
Computation method of external field, designing method of substance, program, and recording medium
JP2016069302A
Information processing device, information processing method, and program
WO2023162493A1