Learning device, control device, learning method, and computer program

The learning device addresses long development times by determining sparse optimization problems and updating parameters to mimic optimal control, reducing lead times and costs in system development.

JP7853875B2Active Publication Date: 2026-04-30KK TOYOTA CHUO KENKYUSHO +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
KK TOYOTA CHUO KENKYUSHO
Filing Date
2022-09-15
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing systems require running a simulator each time a fit is needed, leading to long development lead times and increased costs in system development.

Method used

A learning device that determines sparse optimization problems using a nonlinear model, calculates optimal control, and updates parameters using derivative values to mimic optimal control, allowing for quicker system development.

Benefits of technology

Significantly reduces system development lead time and costs by learning parameters that can mimic optimal control, even in systems with constraints, using well-known gradient methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007853875000035
    Figure 0007853875000035
  • Figure 0007853875000036
    Figure 0007853875000036
  • Figure 0007853875000037
    Figure 0007853875000037
Patent Text Reader

Abstract

To provide a technique for reducing system development lead time.SOLUTION: A learning device includes: a determination unit that determines a sparse optimization problem which uses a nonlinear model approximately representing output time-series from input time-series and is determined by a predetermined parameter; a calculation unit that solves the determined sparse optimization problem to calculate optimal control which corresponds to the parameter and is the input time-series evaluated to be optimal under an evaluation function; and an update unit that update the parameter using a differential value by the parameter for the optimal control. The learning device repeatedly performs determination of the sparse optimization problem by the determination unit, calculation of the optimal control by the calculation unit, and updating of the parameter by the update unit, so as to learn the parameter for determining the sparse optimization problem by which the optimal control with a minimum error from teacher data can be obtained.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a technique for learning parameters that determine sparse optimization problems. [Background technology]

[0002] A control device (controller) that controls the state of a controlled object is known. For example, Patent Document 1 describes a control parameter adaptation system that can automatically estimate the optimal adaptation value even in a control system where it is difficult to determine the target value in advance. In the system described in Patent Document 1, the control state of the air-fuel ratio is evaluated based on the output of an exhaust gas analyzer and an oxygen sensor, and the indicated value of the target air-fuel ratio is adjusted based on the evaluation result to generate an adaptation map that is adapted so that the indicated value of the target air-fuel ratio becomes the optimal value. After the adaptation map is generated, the indicated value of the target air-fuel ratio is determined according to the adaptation map. [Prior art documents] [Patent Documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2010-086405 [Overview of the project] [Problems that the invention aims to solve]

[0004] However, the technology described in Patent Document 1 requires the simulator to be run each time a fit is needed to generate a fit map, which results in a long time required to generate the fit map and a long system development lead time. This problem is not limited to cases where an internal combustion engine is the target of control, but is common to all cases where a system that can be represented by a simulator is the target of control.

[0005] This invention was made to solve at least some of the problems described above, and aims to provide a technology that contributes to shortening the development lead time of a system. [Means for solving the problem]

[0006] The present invention has been made to solve at least some of the above-mentioned problems and can be realized in the following forms.

[0007] (1) According to one embodiment of the present invention, a learning device is provided. This learning device comprises: a determination unit that determines a sparse optimization problem using a nonlinear model that approximately represents an output time series from an input time series, and which is determined by predetermined parameters; a calculation unit that calculates an optimal control corresponding to the parameters, which is the input time series evaluated as optimal under a certain evaluation function, by solving the determined sparse optimization problem; and an update unit that updates the parameters using the derivative values ​​of the optimal control with respect to the parameters. The device learns the parameters for determining the sparse optimization problem that yields the optimal control with a small error from the training data by repeatedly performing the determination of the sparse optimization problem by the determination unit, the calculation of the optimal control by the calculation unit, and the update of the parameters by the update unit.

[0008] With this configuration, the update unit updates the parameters using the derivative values ​​of the parameters relative to the optimal control, allowing it to learn parameters that can mimic the optimal control using well-known gradient methods. By using the parameters learned in this way in the development of the system including the controlled object, the system development lead time can be significantly reduced compared to the conventional configuration that generates a fit map by running the simulator each time a fit is needed. This can reduce system development costs.

[0009] (2) In the learning device of the above form, the nonlinear model may be expressed as a differentiable function in which each component of the return value is designed to be convex and monotonically nondecreasing with respect to the argument. Generally, in order to simulate optimal control using well-known gradient methods, it is necessary to differentiate the optimal input (i.e., optimal control) of the optimal control problem with respect to the parameters used to obtain that optimal control. Here, if the evaluation function is differentiable, the derivative value with respect to the parameters of the optimal control can be obtained. However, in sparse optimization problems, the evaluation function is not differentiable, so it has not been possible to obtain this derivative value using conventional methods. In this respect, with this configuration, the nonlinear model representing the plant is a differentiable function in which each component of the return value is designed to be convex with respect to the argument and monotonically non-decreasing, making it possible to obtain the derivative value with respect to the parameters of the optimal control.

[0010] (3) In the learning device of the above form, the sparse optimization problem may be expressed as the sum of the input time series U of the nonlinear model, the output time series Y determined by the input time series U using the nonlinear model, the parameter θ, a differentiable scalar-valued function V with these as arguments, and the L1 regularization term of the input time series U. In this configuration, the sparse optimization problem is represented as the sum of an input time series U of a nonlinear model, a differentiable scalar-valued function V whose arguments are the output time series Y determined by the input time series using the nonlinear model and a parameter θ, and the L1 regularization term of the input time series U. Therefore, the matrix determined from the sparse optimization problem is always made invertible, and the implicit differentiation theorem (arXiv:2106.04350, a theorem stating that if a matrix determined from an equation is always invertible, then implicit differentiation is possible) can be used to obtain the parameter-dependent derivative of the optimal control.

[0011] (4) In the learning device of the above form, the sparse optimization problem may further have constraints that limit at least one of an upper limit and a lower limit on the input time series U. For example, in the case of fuel injection in an internal combustion engine, the controlled object may have constraints such as upper and lower limits on the input time series U due to the operational requirements of the controlled object. With this configuration, the sparse optimization problem also has a constraint that limits at least one of the upper and lower limits on the input time series U. Therefore, even for controlled objects with such constraints, it is possible to learn parameters that can mimic optimal control.

[0012] (5) In the learning device of the above form, the update unit calculates the differential value J defined by the following equation for updating the parameter, In the above equation, I is the identity matrix, γ is a positive constant, and D is the function l(U,θ)=V(U,F(U,-U,θ),θ), where Nr-dimensional vector v=U-γ∇ U l(U,θ ) for the i-th component d i Select and generate the Nr-dimensional vector d, and arrange them diagonally. It is a matrix, where q is the i-th component of the function l, given the Nr-dimensional vector v. i Select The generated Nr-dimensional vector is where H is the Hessian matrix of the function l with respect to U, U is the input time series U, θ is the parameter θ, and V may be the scalar-valued function V. With this configuration, the update unit can calculate the derivative value J for parameter updates.

[0013] (6) In the learning device of the above form, the sparse optimization problem further has at least one of an equality constraint on the input time series and an inequality constraint on a function that takes the input time series and the output time series as arguments, and the evaluation function is the input time series and the output time series The function may be designed to be convex with respect to the time series, monotonically non-decreasing with respect to the output time series, and strongly convex with respect to the input time series. For example, when setting the fuel injection amount in an internal combustion engine to a predetermined value or suppressing the noise (dB) below a predetermined value, there may be cases where constraints expressed by equations or inequalities are provided for the control target with respect to the input time series and the output time series during operation of the control target. According to this configuration, since the sparse optimization problem further has at least one of an equality constraint and an inequality constraint, it is possible to learn parameters that can imitate optimal control even for a control target having such constraints.

[0014] (7) In the learning device of the above aspect, when the evaluation function includes a non-differentiable function, the calculation unit may solve the sparse optimization problem using the alternating direction multiplier method as an optimization algorithm. According to this configuration, the update rule of the input time series as a variable includes solving an unconstrained sparse optimization problem, and the solution of the evaluation function can be obtained using a well-known sparse optimization algorithm (for example, the proximal gradient algorithm).

[0015] (8) In the learning device of the above aspect, the update unit calculates the differential value J defined by the following equation for updating the parameter. In the above equation, I is the identity matrix, γ and ρ are positive constants, and U * is the input time series U is the solution of the non-linear model for U, the functions V and g are differentiable scalar-valued functions taking the input time series U, the output time series Y, and the parameter θ as arguments, and when the function F is a function that returns the output time series Y with the input time series U determined by the non-linear model and the parameter θ as arguments, D is set as the function V~(U,θ)=V(U,F(U,θ),θ), and for the Nr-dimensional vector v = U * -γ∇ U V~(U * ,θ), the Nr-dimensional vector d generated by selecting the i-th component d i is a diagonal matrix arranged diagonally, and q is the Nr-dimensional vector v of the function V~, and the Nr-dimensional vector i generated by selecting the i-th component q The sparse optimization problem is n e It has the above linearly independent constraints, When the aforementioned equality constraint is expressed as AU+b=0 for the input time series U, n e ×Nr It is a dimensional matrix, where H is the Hessian matrix of the function V~ for U, and Pi is the i-th component of the function g~ (U,θ) = V(U,g(U,θ),θ). i ~no U The Hessian matrix E is such that the inequality constraint g~(U,θ)≦0 is expressed as g~(U,θ)≦0 and the dual variable corresponding to g~(U,θ)≦0 is η. * In that case, η * It is a diagonal matrix with the diagonal elements being ηi, where ηi is the i-th component of the dual variable corresponding to g~(U,θ)≦0, and G is g~ (U * A diagonal matrix with diagonal elements ,θ) and the inverse matrix R -1 in R 11 inv ,R 13 inv ,R 14 inv is matrix R -1 Each component is partitioned into blocks in the same way as matrix R. In that case, each row and column may be a corresponding submatrix. With this configuration, the update unit can calculate the derivative value J for parameter updates.

[0016] (9) In the learning device of the above form, the nonlinear model may be expressed as a linear function of the input time series, where each component of the output time series is linear with respect to the input time series. In this configuration, each component of the output time series is expressed as a linear function with respect to the input time series. In other words, since the nonlinear model is expressed as a linear function, the range of optimization problems that can be handled is broadened.

[0017] (10) In the learning device of the above form, the sparse optimization problem further has at least one of an equality constraint on the input time series and an inequality constraint on a function that takes the input time series and the output time series as arguments, and the evaluation function may be a function designed to be convex with respect to the input time series and the output time series, and to be strongly convex with respect to the input time series. For example, when setting the fuel injection amount in an internal combustion engine to a predetermined amount, or when trying to keep the noise level (dB) below a predetermined value, the controlled object may have constraints that can be expressed as equations or inequalities with respect to the input time series or output time series, depending on the operation of the controlled object. With this configuration, since the sparse optimization problem also has at least one of the equation constraint and the inequality constraint, it is possible to learn parameters that can mimic optimal control even for controlled objects with such constraints.

[0018] (11) In the learning device of the above form, the calculation unit may solve the sparse optimization problem using the alternating direction multiplier method as the optimization algorithm when the evaluation function includes a function that is not differentiable. In this configuration, the update rule for the input time series, which is a variable, involves solving an unconstrained sparse optimization problem, and the solution to the evaluation function can be obtained using a well-known sparse optimization algorithm (e.g., the nearest gradient algorithm).

[0019] (12) In the learning device of the above form, the update unit calculates the differential value J defined by the following equation for updating the parameter, TIFF0007853875000003.tif82170 In the above formula, I is the identity matrix, γ and ρ are positive constants, and U * The input time series This is the solution to the nonlinear model for U, where functions V and g are the input time series U and the output time series If the function F is a differentiable scalar-valued function that takes a time series Y and the parameter θ as arguments, and the function F is a function that returns an output time series Y with an input time series U determined by a nonlinear model and the parameter θ as arguments, then D is defined as the function V~(U,θ)=V(U,F(U,θ),θ), and the Nr-dimensional vector v=U * -γ∇ U V~(U * For θ), the i-th component d i This is a diagonal matrix formed by arranging Nr-dimensional vectors d, which are generated by selecting a specific element, diagonally, where q is the i-th component of the function V~ for the Nr-dimensional vector v. i Select and generate an Nr-dimensional vector The sparse optimization problem is n e It has the above linearly independent constraints, When the aforementioned equality constraint is expressed as AU+b=0 for the input time series U, n e ×Nr It is a dimensional matrix, where H is the Hessian matrix of the function V~ for U, and Pi is the i-th component of the function g~ (U,θ) = V(U,g(U,θ),θ). i ~no U The Hessian matrix E is such that the inequality constraint g~(U,θ)≦0 is expressed as g~(U,θ)≦0 and the dual variable corresponding to g~(U,θ)≦0 is η. * In that case, η * It is a diagonal matrix with the diagonal elements being ηi, where ηi is the i-th component of the dual variable corresponding to g~(U,θ)≦0, and G is g~(U * A diagonal matrix with diagonal elements ,θ) and the inverse matrix R -1 in R 11 inv ,R 13 inv ,R 14 inv is matrix R -1 Each component is partitioned into blocks in the same way as matrix R. In that case, each row and column may be a corresponding submatrix. With this configuration, the update unit can calculate the derivative value J for parameter updates.

[0020] (13) According to one embodiment of the present invention, a control device is provided. This control device comprises an acquisition unit that acquires the parameters learned by the learning device of the above embodiment, and a control unit that calculates the optimal control by solving the sparse optimization problem determined by the acquired parameters. With this configuration, the control unit can quickly and easily calculate optimal control by utilizing parameters learned by the learning device. Therefore, compared to conventional configurations that generate a calibration map by running the simulator each time calibration is required, the development lead time for the system including the controlled object can be significantly reduced, and system development costs can be lowered.

[0021] Furthermore, the present invention can be realized in various forms, for example, a learning device for learning parameters that determine a sparse optimization problem, a control device or ECU (Engine Control Unit) for controlling a system that can be represented by a nonlinear model that approximately represents the output time series from the input time series, a system including a control device and a controlled object, and a control device with a built-in control This can be implemented in the form of the target, a control method for these devices and systems, a computer program executed in these devices and systems, a server device for distributing the computer program, or a non-temporary storage medium storing the computer program. [Brief explanation of the drawing]

[0022] [Figure 1] This is an explanatory diagram illustrating the configuration of a control system. [Figure 2] This diagram illustrates the overview of the learning process. [Figure 3] This is a flowchart illustrating an example of the learning process. [Figure 4] This graph shows the transition of the imitation error obtained as a result of the experiment. [Figure 5] This graph shows the parameter transitions obtained as a result of the experiment. [Figure 6] This is an explanatory diagram illustrating the configuration of the control system in the second embodiment. [Figure 7] This flowchart shows an example of the learning process in the second embodiment. [Figure 8] This is an explanatory diagram illustrating the configuration of the control system in the third embodiment. [Figure 9] This is a flowchart showing an example of the learning process in the third embodiment. [Figure 10] This graph shows the transition of the imitation error obtained as a result of the experiment. [Figure 11] This is an explanatory diagram illustrating the configuration of the control system according to the fourth embodiment. [Figure 12] This is a flowchart showing an example of the learning process in the fourth embodiment. [Figure 13] This graph shows the transition of the imitation error obtained as a result of the experiment. [Modes for carrying out the invention]

[0023] <First Embodiment> Figure 1 is an explanatory diagram illustrating the configuration of control system 1. Control system 1 is a system for calculating the optimal control with small error from training data for a controlled object represented by a nonlinear model that approximately represents the output time series from the input time series. Here, "optimal control" means the input time series that is evaluated as optimal under a certain evaluation function. The input time series means the time series pattern of the input to the controlled object. Details will be described later. Control system 1 as one embodiment of the present invention comprises a learning device 100 and a control device 200.

[0024] The learning device 100 is a device that learns the parameter θ that determines the sparse optimization problem for obtaining optimal control. Here, a sparse optimization problem refers to mathematical optimization in which, in an optimization problem with vectors or matrices as decision variables, the goal is to represent the solution using only some information, such as increasing the number of zero elements in the optimal solution vector or decreasing the rank of the matrix.

[0025] The learning device 100 is, for example, a personal computer (PC) or It can be implemented as an in-vehicle ECU (Electronic Control Unit). The learning device 100 is C The learning device 100 comprises a PU 110, a memory unit 120, a ROM / RAM 130, a communication unit 140, and an input / output unit 150. Each part of the learning device 100 is interconnected by a bus. The CPU 110 includes a determination unit 111, a calculation unit 112, and an update unit 113. These functional units execute the learning process described later by loading the computer program stored in the ROM 130 into the RAM 130 and executing it.

[0026] Figure 2 is a diagram illustrating the overview of the learning process. The decision unit 111 determines a sparse optimization problem 20 using a nonlinear model 10 during the learning process. The nonlinear model 10 is a model of the controlled object and is constructed by pre-learning the input-output relationship of the controlled object. The nonlinear model 10 is a recurrent neural network model that approximately represents the output time series Y from the input time series U of the controlled object, and is expressed by the following equation (1). As is clear from equation (1), the nonlinear model 10 calculates the output Y for the input U under a certain parameter θ. The nonlinear model 10 can be considered a type of simulator, and it is also possible to obtain information (derivative value) of the change in output Y in response to a change in input U from the nonlinear model 10. The nonlinear model 10 is a well-known machine learning model and can represent the input-output relationship of the controlled object system with arbitrary accuracy. The nonlinear model 10 can be implemented, for example, as a model of an internal combustion engine. When the nonlinear model 10 is used as a model for an internal combustion engine, the fuel injection amount can be exemplified as the input time series U, and the torque of the internal combustion engine can be exemplified as the output time series Y.

[0027]

number

[0028] The variables in equation (1) are as shown below in a1 to a5. Note that r, m, n, and p are arbitrary natural numbers, and k=1,...,N are indices representing time. (a1)U: Input time series for nonlinear model 10, where U=[u1,...,u N Each input u is represented by an r-dimensional vector. Therefore, the input time series U is an Nr-dimensional vector. (a2) Y: Output time series from nonlinear model 10, where Y=[y1,···,y N ]. Each output y is represented by an m-dimensional vector. Therefore, the output time series Y is an Nm-dimensional vector. It is a kuthol. (a3)x: This is the internal state time series of the nonlinear model 10, where x = [x1,···,x N+1 Each internal state x is represented by an n-dimensional vector. Therefore, the internal state time series x is an Nn+N-dimensional vector. (a4) The input time series with the sign reversed (hereinafter also called the "reverse input time series"), where U~=[u~1,···,u~ N ]=[-u1,···,-u N ] (a5)θ: A parameter that determines equation (1), more specifically, a parameter that determines the function f and function h in equation (1), and is represented by a p-dimensional vector.

[0029] Here, functions f and h are differentiable functions whose return components are designed to be convex and monotonically non-decreasing with respect to the arguments x, u, and u~. In other words, functions f and h are designed so that, with respect to some of the arguments x, u, and u~, which are part of the arguments x, u, and u~, z is grouped together as z=(x,u,u~) and is convex and monotonically non-decreasing with respect to z. This nonlinear model 10 defines a function Y=F(U,U~,θ) with input time series U, inverted input time series U~, and parameter θ as arguments, respectively, where each component of the output time series is a function that is convex and monotonically non-decreasing with respect to the input time series U and the inverted input time series U~.

[0030] The calculation unit 112 solves the sparse optimization problem 20 determined by the decision unit 111 during the learning process, thereby determining the optimal control U corresponding to the parameter θ. * Calculate (θ). Note: Optimal control U * (θ) is the solution to the sparse optimization problem 20. The update unit 113 is a study In the training process, the obtained optimal control U * (θ) and training data U expert The imitation error is 30 The parameter θ is updated so that it becomes smaller. At this time, the update unit 113 performs optimal control U * The parameter θ is updated using the derivative value 40 of (θ) with respect to the parameter θ. In the learning process, the decision unit 111 determines the sparse optimization problem 20, and the calculation unit 112 performs optimal control U * The process involves repeatedly calculating (θ) and updating the parameter θ using the update unit 113. This minimizes the imitation error 30 (Figure 2: Imitation error V im ). This means that in the learning process, the training data U expert Optimal control U with small error * Sparse optimization problem where (θ) is obtained The parameter θ that determines 20 can be learned. Details will be described later.

[0031] The memory unit 120 is a storage medium composed of a hard disk, flash memory, memory card, etc. The memory unit 120 stores the parameters θ121 calculated and updated by the learning process. In addition, the memory unit 120 has the above-mentioned nonlinear model 10 and the constant γ used in the learning process stored in advance (not shown).

[0032] The communication unit 140 controls communication between the learning device 100 and other devices via a communication interface. Other devices include, for example, the control device 200 described later and other information processing devices. The input / output unit 150 is a variety of interfaces used for inputting and outputting information between the learning device 100 and the user. The input / output unit 150 can include, for example, a touch panel, keyboard, mouse, operation buttons, microphone as an input unit, and a touch panel, monitor, speaker, LED (Light Emitting Diode) indicator as an output unit.

[0033] The control device 200 acquires the parameters θ learned by the learning device 100 and calculates the optimal control with small error from the training data (i.e., the input time series U evaluated as optimal under a certain evaluation function) by solving a sparse optimization problem determined by the parameters θ. The input time series U calculated by the control device 200 is a sparse input time series. Here, a sparse input time series means an input time series with high sparsity (in other words, a large number of zero elements).

[0034] The control device 200 can be implemented, for example, as a personal computer or an in-vehicle ECU. The control device 200 includes a CPU 210, a storage unit 220, a ROM / RAM 230, a communication unit 240, and an input / output unit 250. The various parts of the control device 200 are interconnected by a bus. The CPU 210 includes an acquisition unit 211 and a control unit 212. These functional units are realized by loading computer programs stored in the ROM 230 into the RAM 230 and executing them.

[0035] The acquisition unit 211 acquires the parameter θ learned by the learning process from the learning device 100 and stores it as parameter θ221 in the storage unit 220. The control unit 212 solves the sparse optimization problem determined by parameter θ221 to obtain the optimal control U with small error from the training data. * (θ), in other words, evaluated as optimal under a certain evaluation function. The control unit 212 calculates the input time series U. The control unit 212 stores the calculated input time series U as input time series U222 in the storage unit 220. This input time series U222 is used to control the controlled object.

[0036] Figure 3 is a flowchart of an example of the learning process. The learning process is the process of learning the parameters θ that determine the sparse optimization problem for obtaining optimal control. The learning process can be started at any trigger. For example, the learning process may be started simultaneously with the power being turned on to the learning device 100, or it may be started in response to an instruction given from the input / output unit 150. In the following explanation, we consider the problem of determining the parameters θ that minimize the imitation error 30 expressed by the following equation (2). That is, the imitation error 30 expressed by equation (2) is the optimal control U * (θ) And, sparse teacher input U expert This is the squared error. Furthermore, in the following explanation, the gradient descent method will be used as an example of a method for determining the parameter θ using the derivative value.

[0037]

number

[0038] In step S1, the decision unit 111 of the learning device 100 obtains the nonlinear model 10 (Figure 2) and a positive constant γ from the storage unit 120. The nonlinear model 10 is a model that approximates the output time series Y from the input time series U, as explained in equation (1). The positive constant γ is a constant with the smallest possible value and can be determined arbitrarily.

[0039] In step S2, the determination unit 111 of the learning device 100 initializes the parameter θ. The initial value of the parameter θ may be determined arbitrarily. The determination unit 111 also initializes the index k, which represents time, by substituting 0 for index k.

[0040] In step S3, the decision unit 111 of the learning device 100 determines a sparse optimization problem 20, which is expressed by the following equation (3), from the current parameter θ. The sparse optimization problem 20 is also called the "optimal control problem". Step S3 corresponds to the "decision process", and in step S3, the decision unit 111 performs the "decision function".

[0041]

number

[0042] The variables in equation (3) are as shown below, b1 to b4. (b1) V: A differentiable scalar-valued function with input time series U, output time series Y, and parameter θ as arguments. The function V is convex with respect to both the input time series U and the output time series Y, and is strongly convex with respect to the input time series U, and simply convex with respect to the output time series Y. It is a non-decreasing function. (b2)|U|1: This is the L1 regularization term (L1 norm) of the input time series U. Note that in term b2... The two pillars are represented by a single vertical line. (b3)F: A function whose arguments are the input time series U, the inverted input time series U~, and the parameter θ, as defined by equation (1). (b4)λ: A positive scalar-valued function that takes parameter θ as an argument. Regarding term b1, a differentiable scalar-valued function f(x) defined in a domain X is said to be strongly convex with respect to the variable x if there exists a positive constant ε>0 such that for all x∈X, the Hessian matrix H of f in terms of x is positive semi-definite, where H-εI is positive semi-definite. Alternatively, there exists a positive constant ε>0 such that for all x∈X, all eigenvalues ​​of the Hessian matrix H of f in terms of x are greater than or equal to ε.

[0043] In step S4, the calculation unit 112 of the learning device 100 solves the sparse optimization problem 20 determined in step S3, thereby achieving optimal control U * Calculate (θ). Here, The output unit 112 can solve the sparse optimization problem 20 using the nearest gradient algorithm or the ADMM algorithm. Since the sparse optimization problem 20 is a strictly convex problem, the solution is uniquely determined regardless of the initial value of the parameter θ. Step S4 corresponds to the "calculation process," in which the calculation unit 112 performs the "calculation function."

[0044] In step S5, the update unit 113 of the learning device 100 performs the optimal control U calculated in step S4. * Obtain the derivative J of (θ) with respect to the parameter θ. Specifically, first update Section 113 takes the positive constant γ obtained in step S1. Then, functions α and β that take a scalar ξ as input and return a set of scalar values ​​are defined as shown in equation (4) below. Note that in equation (4), the function value λ(θ) is written as "λ" for simplification of notation.

[0045]

number

[0046] Subsequently, the update unit 113 sets the function l(U,θ)=V(U,F(U,-U,θ),θ), and the Nr-dimensional vector v=U-γ∇ U For l(U,θ), the i is given by equation (5) below. ingredient d i ,q i Select and generate Nr-dimensional vectors d and q. Here, the diagonal matrix obtained by arranging vectors d diagonally is denoted as D. Since the solution to the sparse optimization problem 20 shown in equation (3) depends on the parameter θ, the solution to the sparse optimization problem 20 is U * It can be written as (θ). Solution U of sparse optimization problem 20 * (θ) is the optimal control U * This is synonymous with (θ).

[0047]

number

[0048] Subsequently, the update unit 113 uses the following equation (6) to obtain the solution (i.e., optimal control) of the sparse optimization problem 20, U * We calculate the derivative J of (θ) with respect to the parameter θ. In (6), I is the identity matrix, D is a diagonal matrix with vectors d arranged diagonally, and H is the Hessian matrix relating the input time series U of the function l.

[0049]

number

[0050] Subsequently, the update unit 113 uses the calculated derivative value J to calculate the derivative value g according to the following equation (7). As is clear from equation (7), the derivative value g is equal to the imitation error 30 (i.e., the imitation error V) explained in equation (2). im This is the derivative of ). The derivative g is also called "gradient information". Note that in equation (7), U expert This is the training data, which is the input time series U actually obtained from the controlled object in prior experiments, etc. In equation (7), for the sake of simplicity of notation, U * ( θ) is "U * It states, V im (θ) is "V im It states, "."

[0051]

number

[0052] In step S6, the update unit 113 of the learning device 100 determines whether or not to terminate the parameter update process θ. Specifically, the update unit 113 determines to terminate the process if predetermined convergence criteria are met, and to not terminate (continue) the process if the convergence criteria are not met. The convergence criteria include the imitation error V imConditions can be set using the value of (θ), the norm value of the derivative g, and the index value of k. If the process is not terminated (step S6: NO), the update unit 113 increments the index k and then transitions the process to step S7. On the other hand, if the process is terminated (step S6: YES), the update unit 113 stores the latest parameter θ value in the storage unit 120 and terminates the process.

[0053] In step S7, the update unit 113 of the learning device 100 updates the parameter θ using the derivative value g calculated in step S5 according to the following equation (8). In equation (8), α is a positive constant and is also called the "learning rate". After that, the update unit 113 transitions the process to step S3 and repeats the process described above. Steps S5 to S7 correspond to the "update process", and the "update function" is executed by the update unit 113 in steps S5 to S7.

[0054]

number

[0055] As described above, according to the learning device 100 of the first embodiment, the update unit 113 performs optimal control U * The parameter θ is updated using the derivative J with respect to the parameter θ (Figure 3: Steps S5-S7). Therefore, the learning device 100 can learn parameters θ that can mimic optimal control using well-known gradient methods. By utilizing the parameters θ thus learned in the development of a system including the controlled object, the system development lead time can be significantly shortened and system development costs reduced compared to the conventional configuration that generates a fitting map by running the simulator each time a fit is needed. Furthermore, according to the learning device 100 of the first embodiment, the evaluation function of a non-differentiable optimization problem can be automatically tuned using well-known gradient methods with a teacher input that is considered to have good performance. The properties of the nonlinear model 10's functions f and h, as well as the evaluation function of the optimization problem (convex, strongly convex, monotonically non-decreasing), are such that gradient information for the optimization problem can be constructed.

[0056] Generally, in order to simulate optimal control using well-known gradient methods, the optimal input U of the optimal control problem is required. * (θ) (i.e., optimal control U) * (θ)) is the optimal control U * To obtain (θ) It is necessary to differentiate with respect to the parameter θ used. Here, if the evaluation function is differentiable, the derivative value with respect to the optimal control parameter can be obtained, but in the sparse optimization problem 20, the evaluation function is not differentiable, so it was not possible to obtain this derivative value by the usual method. In this respect, according to the learning device 100 of the first embodiment, as shown in equation (1), the nonlinear model 10 representing the plant is a differentiable function in which each component of the return value is designed to be convex and monotonically non-decreasing with respect to the argument, so the optimal control U * (θ) It is possible to obtain the derivative value J with respect to the parameter θ.

[0057] Furthermore, in the learning device 100 of this embodiment, as shown in equation (3), the sparse optimization problem 20 is expressed as the sum of the input time series U of the nonlinear model 10, a differentiable scalar-valued function V whose arguments are the output time series Y determined by the input time series U using the nonlinear model 10 and a parameter θ, and the L1 regularization term of the input time series U. For this reason, the matrix determined from the sparse optimization problem 20 is always made invertible, and the implicit differentiation theorem (arXiv:2106.04350, a theorem stating that implicit differentiation is possible if a matrix determined from an equation is always invertible) is used to determine the optimal control U * The derivative of (θ) with respect to the parameter θ. It is possible to obtain J. However, although the paper arXiv:1810.13400 describes a method of imitation learning in control problems, it implicitly assumes that the evaluation function of the optimal control problem is differentiable, and does not consider cases where the evaluation function is not differentiable, such as the sparse optimization problem in this embodiment. In other words, the paper arXiv:2105.01637 is incomplete as it does not clarify the method for constructing the derivative value at points where the function is not differentiable. Furthermore, although the paper arXiv:2105.01637 describes differentiating the optimal solution of an optimization problem with an evaluation function that is not differentiable, such as a sparse optimization problem, by the parameters used to obtain that optimal solution, it does not consider at all how to generate the derivative value at points where the function is not differentiable.

[0058] Furthermore, according to the control device 200 of this embodiment, by utilizing the parameters learned by the learning device 100, optimal control U can be performed quickly and easily. * (θ) can be calculated. Compared to conventional configurations that generate a calibration map by running the simulator each time calibration is required, this significantly reduces the development lead time for the system including the controlled object, and also lowers system development costs.

[0059] Numerical experiments were conducted to confirm the effectiveness of the above embodiment. In these experiments, two parameters Qu and Qy included in the evaluation function of the sparse optimization problem 20 were identified as parameters θ according to the method described above. The sparse teacher input was defined as the two parameters Qu and Qy included in the evaluation function of the sparse optimization problem 20, respectively. _expert ,Qy _expert The input obtained as the optimal solution to the optimization problem was given. The other parameters included in the evaluation function were the same as those used when generating the training input. In this experiment, the derivative value J was obtained based on the method described above, and the parameter θ (i.e., parameters Qu, Qy) was updated in the direction of the descent of the imitation error 30 using the gradient method.

[0060] FIG. 4 is a graph showing the transition of the imitation error 30 obtained as a result of the experiment. In FIG. 4, the logarithmic value of the imitation error 30 (that is, the imitation error V im ) is plotted on the vertical axis, and the step size (that is, the index k) is plotted on the horizontal axis. As is clear from FIG. 4, according to the learning device 100 of the present embodiment, the larger the step size, in other words, the larger the value of the index k, the smaller the imitation error 30, in other words, the optimal control U * (θ ) and the difference from the teacher data U expert become smaller. In the example of FIG. 4, finally, the logarithmic value of the imitation error 30 is about 10 -7 .

[0061] FIG. 5 is a graph showing the transition of the parameters Qu and Qy obtained as a result of the experiment. FIG. 5(A) is a graph showing the change in the parameter Qu. FIG. 5(B) is a graph showing the change in the parameter Qy. In FIGS. 5(A) and 5(B), the parameter values are plotted on the vertical axis, and the step size (that is, the index k) is plotted on the horizontal axis. As is clear from FIGS. 5(A) and 5(B), according to the learning device 100 of the present embodiment, the larger the step size, in other words, the larger the value of the index k, the parameter Qu _learner (solid line) that is updated approaches the parameter Qu _expert (dashed line) used by the teacher input, and the parameter Qy (solid line) that is updated approaches the parameter Qy _learner (dashed line) used by the teacher input. In both graphs, since the solid line and the dashed line overlap at a point where the step size is about 700, it can be seen that the estimation of the parameters Qu and Qy is completed.

[0062] Thus, from the results of FIGS. 4 and 5 as well, according to the learning device 100 of the present embodiment, it can be seen that the parameters θ for determining the sparse optimization problem 20 for obtaining the optimal control U * (θ) with a small error from the teacher data can be learned.​​​​

[0063] <Second Embodiment> Figure 6 is an explanatory diagram illustrating the configuration of the control system 1A of the second embodiment. In the control system 1A of the second embodiment, parameters θ can be learned for a control target having predetermined constraints, similar to the first embodiment. In the control system 1A of the second embodiment, the learning device 100 is replaced with a learning device 100A in the configuration of the first embodiment.

[0064] The learning device 100A includes a determination unit 111A in place of the determination unit 111, and an update unit 113A in place of the update unit 113. The processing content of the determination unit 111A and the update unit 113A in the learning process differs from that of the first embodiment. Details will be described later. The storage unit 120 of the learning device 100A also has constraints 122 stored in advance. The constraints 122 are conditions for limiting at least one of an upper limit and a lower limit for the input time series U to the nonlinear model 10. In this embodiment, an example is given in which both the upper limit and the lower limit of the input time series U are limited by the constraints 122.

[0065] Figure 7 is a flowchart showing an example of the learning process in the second embodiment. The difference from the first embodiment described in Figure 3 is that step S1A is executed instead of step S1, step S3A is executed instead of step S3, and step S5A is executed instead of step S5.

[0066] In step S1A, the decision unit 111A of the learning device 100A obtains the nonlinear model 10 (Figure 2), the positive constant γ, and the constraint 122 from the memory unit 120.

[0067] In step S3A, the decision unit 111A of the learning device 100A determines a sparse optimization problem 20, expressed by the following equation (9), from the current parameter θ. Equation (9) is a constraint 122 on the input time series U (i.e., U∈U boxExcept for the points having (0), it is the same as the formula (3) described in the first embodiment. The constraint 122 is defined by the following formula (10). The formula (10) means, in other words, that the admissible set U of the sparse optimization problem 20 box has an upper limit value u i  ̄ and a lower limit value u i _ for the i-th element.

[0068] <9000580> [Number] [Number]

[0069] In step S5A, the update unit 113A of the learning device 100A obtains the derivative value J of the optimal control U * (θ) with respect to the parameter θ. Specifically, first , the update unit 113A takes the positive constant γ obtained in step S1A. Then, functions α i , β i that return a scalar value set with a scalar ξ as an input are defined as the following formula (11). In the formula (11), for simplicity of notation, the function value λ(θ) is denoted as "λ".

[0070] [Number]

[0071] After that, the update unit 113A sets the function l(U, θ) = V(U, F(U, -U, θ), θ), and for the Nr-dimensional vector v = U - γ∇ U l(U, θ), as in the following formula (12) selects the i-th component d i , q i to generate the Nr-dimensional vectors d and q. Here, the diagonal matrix with the vector d arranged diagonally is denoted as D. Since the solution of the sparse optimization problem 20 shown in the formula (9) depends on the parameter θ, the solution of the sparse optimization problem 20 can be described as U * (θ). Then After that, the update unit 113A uses equation (6) described in the first embodiment to obtain the solution (i.e., optimal control) of the sparse optimization problem 20, which is U * Calculate the derivative J of (θ) with respect to the parameter θ. To release.

[0072]

number

[0073] Thus, the sparse optimization problem 20 is constrained by the input time series U (i.e., U∈U) box ) may have the same effect as the first embodiment described above. For example, in the case of a controlled object, such as the fuel injection amount in an internal combustion engine, constraints 122 such as upper and lower limits on the input time series U may be imposed on the controlled object for operational reasons. In this regard, according to the learning device 100A of the second embodiment, the sparse optimization problem 20 shown in equation (9) further includes an upper limit u on the input time series U. i  ̄ and lower limit u i Constraint U that restricts at least one of the following: box Because of this, it is possible to learn parameters θ that allow for the imitation of optimal control even for controlled objects with such constraints 122.

[0074] <Third Embodiment> Figure 8 is an explanatory diagram illustrating the configuration of the control system 1B of the third embodiment. In the control system 1B of the third embodiment, it is possible to learn the same parameter θ as in the second embodiment for a control target having predetermined constraints different from those of the second embodiment. In the control system 1B of the third embodiment, the learning device 100B is provided in place of the learning device 100A in the configuration of the second embodiment.

[0075] The learning device 100B includes a determination unit 111B instead of a determination unit 111A, a calculation unit 112B instead of a calculation unit 112, and an update unit 113B instead of an update unit 113A. The processing content of the determination unit 111B, the calculation unit 112B, and the update unit 113B in the learning process differs from that of the second embodiment. Details will be described later. The storage unit 120 of the learning device 100B also has constraints 122B pre-stored. The constraints 122B have at least one of an equality constraint on the input time series U and an inequality constraint on a function that takes the input time series U and the output time series Y as arguments to the nonlinear model 10. In this embodiment, the case in which both equality constraints and inequality constraints are present (accompany each other) is illustrated.

[0076] Figure 9 is a flowchart showing an example of the learning process in the third embodiment. The difference from the second embodiment described in Figure 7 is that step S1B is executed instead of step S1A, step S3B is executed instead of step S3A, step S4B is executed instead of step S4, and step S5B is executed instead of step S5A. The third embodiment will describe a process that differs from the first and second embodiments.

[0077] In step S1B, the decision unit 111B of the learning device 100B obtains the nonlinear model 10 (Figure 2), the positive constant γ, the constraint 122B, and an additional positive constant ρ from the memory unit 120.

[0078] In step S3B, the decision unit 111B of the learning device 100B determines a sparse optimization problem 20 (Figure 2) represented by the following equation (13) from the current parameter θ. The sparse optimization problem 20 represented by equation (13) is accompanied by the constraints AU+b=0 representing the equality constraint and g(U,Y,θ)≦0 representing the inequality constraint, compared with the sparse optimization problem 20 represented by equation (9) of the second embodiment, and U∈U box The difference is that it does not involve the constraint of equality. Equality constraints and inequality constraints are stored as constraint 122B.

[0079]

number

[0080] The optimization problem expressed by equation (13) is n e It is possible to accompany individual linearly independent equality constraints. The equality constraint is n e ×Nr-dimensional matrix A and n e It is determined by the dimensional vector b. A concrete example of an equality constraint is the condition that the total amount of fuel injected in one engine cycle must be constant.

[0081] Furthermore, the optimization problem represented by equation (13) is n i It is possible to include individual inequality constraints. The inequality constraint is n i It is determined by the dimensional vector-valued function g. That is, by equation (13) The optimization problem of the third embodiment described has both equality constraints and inequality constraints.

[0082] In equation (13), the function V is a differentiable scalar-valued function with input time series U, output time series Y, and parameter θ as arguments. Function V is convex with respect to the input time series U and output time series Y. Furthermore, function V is strongly convex with respect to the input time series U. Also, function V is monotonically non-decreasing with respect to the output time series Y.

[0083] Function F, as shown in equation (13), is a function that takes an input time series U defined by a nonlinear model and a parameter θ as arguments and returns an output time series Y. Function λ is a positive scalar-valued function and is differentiable. Note that function F in the third embodiment is substantially the same as function F shown in equation (3) of the first embodiment. In the first embodiment, the components of the return value of the nonlinear model are designed to be monotonically non-decreasing and convex with respect to the input time series U and the inverted input time series U~. Substituting -U into the inverted input time series U~ from this function F gives the function F~(U,θ)=F(U,-U,θ), which satisfies the design conditions for function F in the third embodiment. Conversely, from function F in the third embodiment, a function F~ can be designed that satisfies the design conditions for function F in the first embodiment and satisfies the equation F(U,θ)=F~(U,-U,θ) when -U is substituted into the argument U~ of function F~.

[0084] The function g is a differentiable scalar-valued function that takes an input time series U, an output time series Y, and a parameter θ as arguments. The function g is a convex function with respect to the input time series U and the output time series Y. The method for designing such a function g is to construct n from the function g, F. i Dimensional vector value The function g~(U,θ)=g(U,F(U,θ),θ) has the effect of being a convex function with respect to the input time series U. Note that the function g in the third embodiment and the fourth embodiment described later is different from the derivative g shown in equation (7) of the first embodiment. A concrete example of an inequality constraint is the control such that the engine noise (dB) is below a certain threshold.

[0085] In step S4B, the calculation unit 112B of the learning device 100B solves the sparse optimization problem 20 determined in step S3B, thereby achieving optimal control U * Calculate (θ). As an optimization algorithm for solving the optimization problem represented by equation (13), the ADMM (Alternating Direction Method of Multiplier) represented by equation (14) is known. In the third embodiment, when the evaluation function includes a non-differentiable function as represented by equation (13), the calculation unit 112B solves the sparse optimization problem 20 using the Alternating Direction Method of Multiplier as the optimization algorithm.

number

[0086] In equation (14), V~(U) is the function V~(U)=V(U,F(U,θ),θ) obtained by substituting the nonlinear model function Y=F(U,θ) into the evaluation function V defined in equation (13). Here, V~ and λ also depend on θ, but since the dependency is clear, θ is omitted from the argument to avoid complexity in notation.

[0087] In ADMM, the positive constant ρ obtained in step S1B and the initial value U 0 ,V 0 ,μ 0 suitable For this purpose, we solve the update rule for the variables U, V, and μ. Initial value U 0 The input time series This is the initial value of U. Initial value V 0 Unlike the function V represented by equation (13), the input time series This is the initial value of the Nr-dimensional vector, which is a variable of the same dimension as U. Note that only the variable V shown in equation (14) differs from the function V expressed by the other equations in the third embodiment. Initial value μ 0 teeth As shown in equation (14), this is the initial value of the variable μ, which is an Nr-dimensional vector defined by the input time series U, which is an Nr-dimensional vector, and the variable V.

[0088] The update rule for the variable U involves solving the unconstrained sparse optimization problem 20, and the solution can be found using well-known sparse optimization algorithms such as the nearest gradient algorithm. In the update rule for the variable V, the set U represents the set of all input time series that satisfy the equality and inequality constraints, and is a convex set under the design method of equation (13). In other words, the optimization problem containing the update rule for the variable V is a constrained convex optimization problem, and the solution can be found using optimization algorithms such as SQP (sequential quadratic programming), and can also be found using general-purpose solvers. Note that the update rule is also called the learning rule. ||UV in equation (14) k +μ k ||2 is UV k +μ k This is the L2 norm.

[0089] Here, we define the function V~(U,θ)=V(U,F(U,θ),θ) from the functions V and F defined in equation (13), and the function g~(U,θ)=g(U,F(U,θ),θ) from the functions g and F. Thus, using the variable U~, equation (13) can be rewritten as equation (15).

[0090]

number

[0091] The solution to the optimization problem represented by equation (15) is U * Corresponding to the equality constraint UU~=0 The value of the dual variable is y * In this case, the positive constant ρ obtained in step S1B is used. And, μ * =y * Let / ρ. Here, U * and μ * Against this, we consider the convex optimization problem represented by equation (16). Note that the variable U~ defined in the third embodiment is different from the inverted input time series of the input time series U defined in equation (3) of the first embodiment.

[0092]

number

[0093] In step S5B, the update unit 113B of the learning device 100B performs the optimal control U calculated in step S4B. * We obtain the derivative J of the function with respect to the parameter θ. Specifically, first, New section 113B gives the dual variable corresponding to the inequality constraint g ~ (U, θ) ≤ 0 in equations (15) and (16) as η * Let's assume that. For convenience, η * Let E be a diagonal matrix with diagonal elements and G be a diagonal matrix with diagonal elements g~(U,θ). Furthermore, using the positive constant γ obtained in step S1B, functions α and β that take a scalar ξ as input and return a set of scalar values ​​are defined as shown in equation (4) of the first embodiment.

[0094] Subsequently, the update unit 113B generates an Nr-dimensional vector v=y * / ρU * -γ∇V~(U * For θ), the i-th component d i ,q i to d i ∈α(v i ),q i ∈β(v i Select as shown above to generate Nr-dimensional vectors d and q. The update unit 113B takes a diagonal matrix D formed by arranging vector d diagonally. The solution to the optimization problem shown in equation (13) depends on the parameter θ, and U * (θ) and table It is revealed. The update unit 113B is U * The derivative J of (θ) is found using equation (17). Oh, in equation (17), matrix H is the Hessian matrix of function V~ with respect to U. Matrix Pi is the i-th component of function g~ i This is the Hessian matrix for U.

[0095]

number

[0096] The matrix R in equation (17) is represented by a block partition. The inverse matrix of matrix R is matrix R -1 In R 11 inv ,R 13 inv ,R 14 inv is matrix R -1 Each component is the same as matrix R When a matrix is ​​partitioned into blocks, each row and column corresponds to a submatrix. For example, the top-left matrix of matrix R is an Nr×Nr-dimensional matrix, so R 11 inv is matrix R -1 The upper left Nr×Nr dimension This represents a matrix. In step S5B, once the differential value J as gradient information is obtained, the processing from step S6 onwards is performed in the same manner as in the first embodiment.

[0097] Thus, the sparse optimization problem 20 may have at least one of an equality constraint and an inequality constraint, such as the constraint 122B on the input time series U, as shown in equation (13). Furthermore, in the evaluation function for the sparse optimization problem 20, at least one of the following constraints is set as constraint 122B: an equality constraint on the input time series U and an inequality constraint on the function that takes the input time series U and the output time series Y as arguments to the nonlinear model 10. Also, the function V, which is a differentiable scalar-valued function that takes the input time series U, output time series Y, and parameter θ as arguments, as defined in the evaluation function, is a convex function with respect to the input time series U and the output time series Y. Furthermore, the function V is a strongly convex function with respect to the input time series U and a monotonically non-decreasing function with respect to the output time series Y. The learning device 100B of the third embodiment described above can also achieve the same effects as the first and second embodiments described above. For example, in the case of a controlled object, such as the fuel injection amount in an internal combustion engine, an equality constraint may be set so that the input time series U in one cycle remains constant for the operation of the controlled object. Furthermore, in the case of noise in an internal combustion engine, the controlled object may be subject to an inequality constraint such that the function g of the input time series U, the output time series Y, and the parameter θ must be less than or equal to a predetermined value, depending on the operation of the controlled object. In this regard, according to the learning device 100B of the third embodiment, the sparse optimization problem 20 shown in equation (13) further has at least one of an equality constraint and an inequality constraint, so that even for a controlled object with such constraints, it is possible to learn parameters that can mimic optimal control. As a result, a general optimization problem can be treated as a sparse optimization problem 20.

[0098] Furthermore, in the learning device 100B of the third embodiment, as shown in equation (13), if the evaluation function includes a non-differentiable function, the calculation unit 112B may solve the sparse optimization problem 20 using the alternating direction multiplier method as the optimization algorithm. Therefore, the update rule for the input time series U, which is a variable, includes solving the unconstrained sparse optimization problem 20, and the solution to the evaluation function can be obtained using a well-known sparse optimization algorithm (e.g., the nearest gradient algorithm).

[0099] Furthermore, in the learning device 100B of the third embodiment, the update unit 113B of the learning device 100B performs the optimal control U calculated in step S4B. * Obtain the derivative J of with respect to the parameter θ. In other words, the update unit 113B can calculate the derivative value J for parameter updating. Here, the implicit differentiation theorem, which states that the equation is constructed of non-differentiable functions, is used to obtain the derivative of the optimal solution of the sparse optimization problem 20 with respect to the parameters. According to this theorem, implicit differentiation is possible if a certain matrix determined from the equation is always invertible. Since the matrix R in equation (17) is always invertible, implicit differentiation is possible. In the third embodiment, by using a model in which each component of the output time series Y is convex with respect to the input time series U as a nonlinear model representing the plant, and by formulating the sparse optimization problem 20 as a convex problem, the matrix R is always made invertible under appropriate assumptions, and the derivative value of the optimal solution with respect to the parameters is obtained using implicit differentiation.

[0100] Numerical experiments were conducted to confirm the effectiveness of the third embodiment described above. In the experiment, a 2-input, 2-output ICRNN model was used as the plant model. An optimization problem with an L1 normalization term added was modeled using the sum of the squares of the inputs at each time step and the sum of the values ​​obtained by applying a soft plus function to the outputs at each time step as the evaluation function. 100 initial states of the plant model were randomly generated, and the solutions to the sparse optimization problem 20 for each generated initial state were prepared as training input. Using the four weight parameters included in the evaluation function of the sparse optimization problem 20 as parameter θ, the differential value J expressed by equation (17) was obtained based on the method of the third embodiment. The parameter θ was updated in the direction of the descent of the imitation error 30 (Figure 2) by stochastic gradient descent using mini-batch learning.

[0101] Figure 10 is a graph showing the transition LN1 of the imitation error 30 obtained as a result of the experiment. In Figure 10, as in Figure 4, the common logarithm of the imitation error 30 is plotted on the vertical axis and the step size is plotted on the horizontal axis. As is clear from Figure 10, when the step size is large The smaller the imitation error of 30, the smaller the imitation error. Specifically, 10 1 The imitation error of approximately 30 was, 10 -3 It has decreased to that point.

[0102] <Fourth Embodiment> Figure 11 is an explanatory diagram illustrating the configuration of the control system 1C of the fourth embodiment. In the control system 1C of the fourth embodiment, the same parameter θ can be learned as in the third embodiment for a controlled object having the same constraints 122B as in the third embodiment. In the control system 1C of the fourth embodiment, the learning device 100C is provided in place of the learning device 100B in the configuration of the third embodiment.

[0103] The learning device 100C includes a determination unit 111C instead of a determination unit 111B, and an update unit 113C instead of an update unit 113B. The processing content of the determination unit 111C and the update unit 113C in the learning process differs from that of the third embodiment. Details will be described later.

[0104] Figure 12 is a flowchart showing an example of the learning process in the fourth embodiment. The difference from the third embodiment described in Figure 9 is that step S3C is executed instead of step S3B.

[0105] In step S3C, the decision unit 111C of the learning device 100C determines a sparse optimization problem 20 (Figure 2) expressed by the following equation (18) from the current parameter θ.

number

[0106] The sparse optimization problem 20 represented by equation (18) differs from the sparse optimization problem 20 represented by equation (13) in the third embodiment in that the output time series Y, which is the return value of the input time series U and parameter θ, is set differently. Specifically, in the fourth embodiment, it is represented as Y = L(θ)U + q(θ) as shown in equation (18). Similar to the third embodiment, the function V is a convex function with respect to both the input time series U and the output time series Y, and is strongly convex with respect to the input time series U. The functions L and q are a matrix-valued function and a vector-valued function, respectively, with parameter θ as an argument, and are defined as a linear model that returns the output time series Y based on the input time series U. In other words, in the sparse optimization problem 20 of the fourth embodiment, there is no condition that the function V defined in the sparse optimization problem 20 of the third embodiment is monotonically non-decreasing with respect to the output time series Y, and each component of the output time series Y is expressed as a linear function with respect to the input time series U.

[0107] In the fourth embodiment, the sparse optimization problem 20 represented by equation (18) is processed in the same way as in the third embodiment from step S4 onward, so the explanation from step S4 onward is omitted.

[0108] Thus, in the sparse optimization problem 20 using a nonlinear model, each component of the output time series Y is expressed as a linear function with respect to the input time series U. In other words, in the fourth embodiment, the nonlinear model is expressed as a linear function, which broadens the range of optimization problems that can be handled.

[0109] Furthermore, in the evaluation function for the sparse optimization problem 20 of the fourth embodiment, similar to the third embodiment, at least one of the following is set as constraint 122B: an equality constraint on the input time series U and an inequality constraint on a function that takes the input time series U and the output time series Y as arguments to the nonlinear model 10. Also, as shown in equation (18), the function V takes the input time series U and the output time series Y as arguments. The function is convex with respect to the force time series Y and strongly convex with respect to the input time series U. That is, according to the learning device 100C of the fourth embodiment, the sparse optimization problem 20 shown in equation (18) is accompanied by equality constraints and inequality constraints, so even for a controlled object with such constraints 122B, it is possible to learn parameters θ that can mimic optimal control. As a result, general optimization problems can be treated as sparse optimization problems 20.

[0110] Numerical experiments were conducted to confirm the effectiveness of the fourth embodiment described above. A linear time-invariant system with two inputs and two outputs was used as the plant model. An optimization problem with an L1 normalization term added was modeled using the weighted sum of the sum of squares of the inputs at each time step and the sum of squares of the outputs at the final time step as the evaluation function. 100 initial states of the plant model were randomly generated, and the solutions to the sparse optimization problem 20 for each generated initial state were prepared as training input. Using the four weight parameters included in the evaluation function of the sparse optimization problem 20 as parameter θ, the differential value J expressed by equation (17) was obtained based on the method of the fourth embodiment. The parameter θ was updated in the direction of the descent of the imitation error 30 (Figure 2) by stochastic gradient descent using mini-batch learning.

[0111] Figure 13 is a graph showing the transition LN2 of the imitation error 30 obtained as a result of the experiment. In Figure 13, as in Figure 10, the common logarithm of the imitation error 30 is plotted on the vertical axis and the step size is plotted on the horizontal axis. As is clear from Figure 13, the larger the step size, the smaller the imitation error 30. Specifically, the imitation error 30 eventually becomes 10 -3 It has decreased to a certain extent.

[0112] <Modified form of this embodiment> The present invention is not limited to the embodiments described above, and can be implemented in various forms without departing from its spirit. For example, a part of the configuration implemented by hardware may be replaced with software, or conversely, a part of the configuration implemented by software may be replaced with hardware. In addition, the following modifications are also possible, for example.

[0113] [Example 1] The above embodiment shows an example of the configuration of control systems 1, 1A, 1B, and 1C. However, the configuration of control systems 1, 1A, 1B, and 1C can be modified in various ways. For example, the learning device 100, the control device 200, and the controlled object of control system 1 may be connected to each other via a network for communication and may be located in physically separate locations. For example, control system 1 may have a different configuration from the learning device 100 and control device 200 exemplified. For example, at least a part of the nonlinear model 10, the positive constants γ and ρ, and the constraints 122 may not be pre-stored in the memory unit 120, but may be provided from the outside via the input / output unit 150 or the communication unit 140. For example, the learning device 100 and the control device 200 may be configured as a single device. For example, control system 1 may consist only of the learning device 100, or it may consist only of the control device 200.

[0114] [Differentiation 2] The above embodiment shows an example of a learning process. However, the procedure of the learning process can be modified in various ways, and the processing content in each step may be added, omitted, or changed, and the execution order of each step may also be changed.

[0115] For example, the diagonal matrix D described in equation (6) is a sparse matrix with most elements being zero, so this property can be used to reduce the computational complexity of the inverse matrix operation. Specifically, if we express the diagonal matrix D by separating the subscripts that are zero and the subscripts that are not zero, we get the following equation (19). In equation (19), D~ represents all diagonal elements This is a positive diagonal matrix. Using this subscript division, matrix H is divided into blocks and expressed as shown in equation (20). In this case, the inverse matrix of matrix P can be written as shown in equation (21). Equation (21) allows us to reduce the size of the matrix from which to take the inverse matrix, thereby reducing the computational complexity.

[0116]

number

number

number

[0117] For example, in the regularization term λ(θ)|U|1 explained in equations (3) and (9), the count λ Change this for each component i and Σ i ∈ [1, ···, Nr] λ i (θ)|U i We can also use |. In this way, the proximity operator can be explicitly written out, and therefore the differential value J can be obtained.

[0118] In the third embodiment described above, ICRNN was given as an example of a nonlinear model, but any nonlinear model other than ICRNN may be used, as long as each component of the output time series Y is a differentiable function that is convex with respect to the input time series U.

[0119] In the third and fourth embodiments described above, the case in which constraint 122B includes both an equality constraint and an inequality constraint was explained. However, in the sparse optimization problem 20 of other embodiments, only one of the equality constraint or the inequality constraint may be included. In this case, if there is no equality constraint, the derivative J can be calculated by removing the 5th row and 5th column of matrix R in equation (17). Similarly, if there is no inequality constraint, the derivative J can be calculated by removing the 4th row and 4th column of matrix R in equation (17).

[0120] In the above embodiment, the normalization term that determines sparsity is introduced in the form λ(θ)|U|1. However, if the proximity operator of the normalization term can be explicitly written, then in equation (17) of the third embodiment, the i-th component d of the Nr-dimensional vectors d,q i ,q i The third and fourth embodiments can also be implemented by replacing the set functions α and β that define the function with the subgradient of the nearest operator. For example, by changing the coefficient λ for each component i, Σ i ∈ [1 ,··· ,Nr] λ i ( θ)|U i If we set it to |, then change the λ included in the definitions of the set-valued functions α and β for each subscript i. That's all you need to do.

[0121] The fourth row and third and fourth column block of matrix R in equation (17) of the third embodiment described above is as shown in equation (22).

number

[0122] The part represented by equation (22) is the equality of equation (23) included in the complementarity condition of the inequality constraint (U * ,η * It can be obtained by differentiating with respect to ).

number

[0123] The complementarity condition can be replaced by the equality condition expressed in equation (24) using the function φ defined with the complementarity function Ψ.

number

[0124] Therefore, the block shown in equation (22) of the matrix R represented by equation (17) may be replaced using a complementarity function as shown in equation (25).

number

[0125] The embodiments of this specification have been described above based on the embodiments and modifications described above. The embodiments described above are for the purpose of facilitating understanding of this specification and do not limit it. This specification may be modified and improved without departing from its spirit and the scope of the claims, and equivalents thereof are included in this specification. Furthermore, any technical features that are not described as essential in this specification may be deleted as appropriate.

[0126] The present invention can also be realized in the following forms. [Application Example 1] A learning device, A sparse optimization problem using a nonlinear model that approximately represents the output time series from the input time series, comprising a determination unit that determines the sparse optimization problem determined by predetermined parameters, A calculation unit calculates an optimal control that corresponds to the parameters and is the input time series evaluated as optimal under a certain evaluation function by solving the determined sparse optimization problem, An update unit updates the parameters using the derivative value of the parameters with respect to the optimal control, Equipped with, A learning device that learns the parameters for determining the sparse optimization problem that yields the optimal control with a small error from the training data, by repeatedly performing the following: determination of the sparse optimization problem by the determination unit, calculation of the optimal control by the calculation unit, and updating of the parameters by the update unit. [Application Example 2] The learning device described in Application Example 1, The aforementioned nonlinear model is a learning device in which each component of the return value is expressed as a differentiable function designed to be convex and monotonically non-decreasing with respect to the argument. [Application Example 3] A learning device as described in Application Example 1 or Application Example 2, The aforementioned sparse optimization problem is, A differentiable scalar-valued function V takes as arguments the input time series U of the nonlinear model, the output time series Y determined by the input time series U using the nonlinear model, and the parameter θ. The L1 regularization term of the aforementioned input time series U, A learning device that can be expressed as a sum of [numbers]. [Application Example 4] A learning device described in any one of the three application examples, The aforementioned sparse optimization problem is further, A learning device having constraints that limit at least one of an upper limit and a lower limit for the input time series U. [Application Example 5] A learning device described in any one of Application Examples 1 to 4, The update unit calculates the differential value J defined by the following equation for updating the parameter, In the above formula, TIFF0007853875000029.tif12170, I is the identity matrix, γ is a positive constant, Let D be the function l(U,θ)=V(U,F(U,-U,θ),θ), and the Nr-dimensional vector v=U-γ∇ U For l(U,θ), the i-th component d i This is a diagonal matrix formed by arranging the Nr-dimensional vectors d, which were generated by selecting the appropriate elements, diagonally. q is the i-th component of the function l, given the Nr-dimensional vector v. i Select and generate This is an Nr-dimensional vector, H is the Hessian matrix of the function l with respect to U, A learning device in which U is the input time series U, θ is the parameter θ, and V is the scalar-valued function V. [Application Example 6] A learning device described in any one of Application Examples 1 to 5, The aforementioned sparse optimization problem is further, having at least one of an equation constraint on the input time series and an inequality constraint on a function including the input time series and the output time series as arguments, the evaluation function is designed to be convex with respect to the input time series and the output time series, and designed to be monotonically non-decreasing with respect to the output time series, and is a function designed to be strongly convex with respect to the input time series. [Application Example 7] A learning device according to any one of Application Examples 1 to 6, wherein when the evaluation function includes a non-differentiable function, the calculation unit solves the sparse optimization problem using the alternating direction multiplier method as an optimization algorithm. [Application Example 8] A learning device according to any one of Application Examples 1 to 7, wherein the update unit calculates the differential value J defined by the following equation for updating the parameter, In the above equation, I is the identity matrix, γ and ρ are positive constants, U * is the solution of the non-linear model for the input time series U, when the function V and the function g are differentiable scalar-valued functions taking the input time series U, the output time series Y, and the parameter θ as arguments, and the function F is a function that returns the output time series Y with the input time series U determined by the non-linear model and the parameter θ as arguments, let D be the function V~(U,θ)=V(U,F(U,θ),θ), and let the Nr-dimensional vector v = U * -γ∇ U V~(U * ,θ), and let the Nr-dimensional vector d generated by selecting the i-th component d i be a diagonal matrix with the vectors d arranged diagonally, let q be the Nr-dimensional vector v of the function V~, and let the i-th component q i be selected to The generated Nr-dimensional vector is A is a sparse optimization problem n e It has the above linearly independent constraints, and the When the equality constraint is expressed as AU+b=0 for the input time series U, n e ×Nr dimension It is a matrix, H is the Hessian matrix of the function V~ for U, Let Pi be the i-th component of the function g~(U,θ)=V(U,g(U,θ),θ), where g~ is the function g~ i ~ is the Hessian matrix of U, E is such that the inequality constraint is expressed as g~(U,θ)≦0 and the dual variable corresponding to g~(U,θ)≦0 is η * In that case, η * It is a diagonal matrix with the diagonal elements as follows: ηi is the i-th component of the dual variable corresponding to g~(U,θ)≦0, G is g~(U * It is a diagonal matrix with diagonal elements ,θ) The inverse matrix of matrix R is matrix R -1 In R 11 inv ,R 13 inv ,R 14 inv is matrix R -1 A learning device where each component of is divided into blocks in the same way as matrix R, and each row and column corresponds to a submatrix. [Application Example 9] A learning device described in any one of Application Examples 1 to 8, The aforementioned nonlinear model is a learning device in which each component of the output time series is expressed as a linear function with respect to the input time series. [Application Example 10] A learning device according to any one of the application examples 1 to 9, The aforementioned sparse optimization problem is further, The system has at least one of the following: an equality constraint on the input time series and an inequality constraint on a function that takes the input time series and the output time series as arguments. The aforementioned evaluation function is, Designed to be convex with respect to the input time series and the output time series, Furthermore, the learning device is a function designed to be strongly convex with respect to the input time series. [Application Example 11] A learning device described in any one of Application Examples 1 to 10, The calculation unit is a learning device that solves the sparse optimization problem using the alternating direction multiplier method as an optimization algorithm when the evaluation function includes a function that is not differentiable. [Application Example 12] A learning device according to any one of Application Examples 1 to 11, The update unit calculates the differential value J defined by the following equation for updating the parameter, In the above formula, TIFF0007853875000031.tif84170, I is the identity matrix, γ and ρ are positive constants, U * This is the solution to the nonlinear model for the input time series U, When functions V and g are differentiable scalar-valued functions that take an input time series U, an output time series Y, and the parameter θ as arguments, and function F is a function that returns an output time series Y that takes an input time series U and the parameter θ determined by a nonlinear model as arguments, Let D be a function V~(U,θ)=V(U,F(U,θ),θ), and let v=U be an Nr-dimensional vector. * -γ∇ U V~(U * For θ), the i-th component d i This is a diagonal matrix formed by arranging the Nr-dimensional vectors d, which were generated by selecting the appropriate elements, diagonally. q is the i-th component of the function V~, given the Nr-dimensional vector v. i Select The generated Nr-dimensional vector is A is a sparse optimization problem n e It has the above linearly independent constraints, and the When the equality constraint is expressed as AU+b=0 for the input time series U, n e ×Nr dimension It is a matrix, H is the Hessian matrix of the function V~ for U, Let Pi be the i-th component of the function g~(U,θ)=V(U,g(U,θ),θ), where g~ is the function g~ i ~ is the Hessian matrix of U, E is such that the inequality constraint is expressed as g~(U,θ)≦0 and the dual variable corresponding to g~(U,θ)≦0 is η * In that case, η * It is a diagonal matrix with the diagonal elements as follows: ηi is the i-th component of the dual variable corresponding to g~(U,θ)≦0, G is g~(U * It is a diagonal matrix with diagonal elements ,θ) The inverse matrix of matrix R is matrix R -1 In R 11 inv ,R 13 inv ,R 14 inv is matrix R -1 A learning device where each component of is divided into blocks in the same way as matrix R, and each row and column corresponds to a submatrix. [Application Example 13] A control device, An acquisition unit that acquires the parameters learned by the learning device described in any one of Application Examples 1 to 12, A control device comprising: a control unit that calculates the optimal control by solving the sparse optimization problem determined by the acquired parameters. [Application Example 14] A learning method, where the information processing device A sparse optimization problem using a nonlinear model that approximately represents the output time series from the input time series, comprising a decision step of determining a sparse optimization problem determined by predetermined parameters, By solving the determined sparse optimization problem, a calculation step of calculating an optimal control corresponding to the parameter, which is the optimal control of the input time series evaluated as optimal under a certain evaluation function, is performed. An update step of updating the parameter using a derivative value with respect to the parameter for the optimal control. Execute, By repeating the determination of the sparse optimization problem, the calculation of the optimal control, and the update of the parameter, the parameter for determining the sparse optimization problem that obtains the optimal control with a small error from the teacher data is learned. A learning method. [Application Example 15] A computer program that causes an information processing device to A determination function for determining a sparse optimization problem using a non-linear model that approximately represents an output time series from an input time series, the sparse optimization problem being determined by a predetermined parameter. By solving the determined sparse optimization problem, a calculation function for calculating an optimal control corresponding to the parameter, which is the optimal control of the input time series evaluated as optimal under a certain evaluation function, is performed. An update function for updating the parameter using a derivative value with respect to the parameter for the optimal control. Cause to execute, By repeating the determination of the sparse optimization problem, the calculation of the optimal control, and the update of the parameter, the parameter for determining the sparse optimization problem that obtains the optimal control with a small error from the teacher data is learned. A computer program.

Explanation of Signs

[0127] 1, 1A, 1B, 1C... Control system 10... Non-linear model 20... Sparse optimization problem 30... Imitation error 40... Derivative value 100, 100A, 100B, 100C... Learning device 110... CPU 111,111A,111B,111C...Decision section 112,112B…Calculation part 113,113A,113B,113C…Update section 120...Storage section 121... Parameter θ 122,122B…constraint 130...ROM / RAM 140... Communications Department 150…Input / output section 200... Control device 210…CPU 211…Acquisition Department 212... Control Unit 220...Storage section 221... Parameter θ 230...ROM / RAM 240... Communications Department 250…Input / output section LN1, LN2… Transition of imitation errors

Claims

1. A learning device, A sparse optimization problem using a nonlinear model that approximately represents the output time series from the input time series, comprising a determination unit that determines the sparse optimization problem determined by predetermined parameters, A calculation unit calculates an optimal control that corresponds to the parameters and is the input time series evaluated as optimal under a certain evaluation function by solving the determined sparse optimization problem, An update unit updates the parameters using the derivative value of the parameters with respect to the optimal control, Equipped with, A learning device that learns the parameters for determining the sparse optimization problem that yields the optimal control with a small error from the training data, by repeatedly performing the following: determination of the sparse optimization problem by the determination unit, calculation of the optimal control by the calculation unit, and updating of the parameters by the update unit.

2. A learning device according to claim 1, The aforementioned nonlinear model is a learning device in which each component of the return value is expressed as a differentiable function designed to be convex and monotonically non-decreasing with respect to the argument.

3. A learning device according to claim 2, The aforementioned sparse optimization problem is, A differentiable scalar-valued function V takes as arguments the input time series U of the nonlinear model, the output time series Y determined by the input time series U using the nonlinear model, and the parameter θ. The L1 regularization term of the aforementioned input time series U, A learning device that can be expressed as a sum of [numbers].

4. A learning device according to claim 3, The aforementioned sparse optimization problem is further, A learning device having constraints that limit at least one of an upper limit and a lower limit for the input time series U.

5. A learning device according to claim 4, The update unit calculates the differential value J defined by the following equation for updating the parameter: In the above formula, I is the identity matrix, γ is a positive constant, D is defined as the function l(U, θ) = V(U, F(U, -U, θ), θ), and the Nr-dimensional vector v = U - γ∇ U For l(U, θ), the i-th component d i This is a diagonal matrix formed by arranging the Nr-dimensional vectors d, which were selected and generated, diagonally. q is the i-th component of the function l, relative to the Nr-dimensional vector v. i Select and generate This is an Nr-dimensional vector, H is the Hessian matrix of the function l with respect to U, U is the input time series U, θ is the parameter θ, and V is the scalar value relation A learning device with a voltage of several volts.

6. A learning device according to claim 1, The aforementioned sparse optimization problem is further, The system has at least one of the following: an equality constraint on the input time series and an inequality constraint on a function that takes the input time series and the output time series as arguments. The aforementioned evaluation function is, Designed to be convex with respect to the input time series and the output time series, Furthermore, it is designed to be monotonically non-decreasing with respect to the output time series, Furthermore, the learning device is a function designed to be strongly convex with respect to the input time series.

7. A learning device according to claim 6, The calculation unit is a learning device that solves the sparse optimization problem using the alternating direction multiplier method as an optimization algorithm when the evaluation function includes a function that is not differentiable.

8. A learning device according to claim 7, The update unit calculates the differential value J defined by the following equation for updating the parameter: In the above formula, I is the identity matrix, γ and ρ are positive constants, U * This is the solution to the nonlinear model for the input time series U, When functions V and g are differentiable scalar-valued functions that take an input time series U, an output time series Y, and the parameter θ as arguments, and function F is a function that returns an output time series Y that takes an input time series U determined by a nonlinear model and the parameter θ as arguments, Let D be a function V ~ (U, θ) = V(U, F(U, θ), θ), and then an Nr-dimensional vector v = U * -γ∇ U V~(U * For θ, the i-th component d i This is a diagonal matrix formed by arranging the Nr-dimensional vectors d, which were selected and generated, diagonally. q selects the i-th component q for the Nr-dimensional vector v of the function V~ i and This is the generated Nr-dimensional vector, A is a sparse optimization problem n e It has the above linearly independent constraints, and the When the equality constraint is expressed as AU + b = 0 for the input time series U, n e ×Nr dimension It is a matrix, H is the Hessian matrix of the function V to U, Let Pi be the i-th component of the function g ~ (U, θ) = V(U, g(U, θ), θ), where g ~ is the i-th component of the function g ~ i This is the Hessian matrix of U in ~, E is the dual variable corresponding to the inequality constraint g ~ (U, θ) ≤ 0, where η * In that case, η * It is a diagonal matrix with the diagonal elements as follows: ηi is the i-th component of the dual variable corresponding to g ~ (U, θ) ≤ 0, G is g~ (U * It is a diagonal matrix with diagonal elements θ, The inverse matrix of matrix R is matrix R -1 In R 11 inv , R 13 inv , R 14 inv is matrix R -1 A learning device where each component of is divided into blocks in the same way as matrix R, and each row and column corresponds to a submatrix.

9. A learning device according to claim 1, The aforementioned nonlinear model is a learning device in which each component of the output time series is expressed as a linear function with respect to the input time series.

10. A learning device according to claim 9, The aforementioned sparse optimization problem is further, The system has at least one of the following: an equality constraint on the input time series and an inequality constraint on a function that takes the input time series and the output time series as arguments. The aforementioned evaluation function is, Designed to be convex with respect to the input time series and the output time series, Furthermore, the learning device is a function designed to be strongly convex with respect to the input time series.

11. A learning device according to claim 10, The calculation unit is a learning device that solves the sparse optimization problem using the alternating direction multiplier method as an optimization algorithm when the evaluation function includes a function that is not differentiable.

12. A learning device according to claim 11, The update unit calculates the differential value J defined by the following equation for updating the parameter: In the above formula, I is the identity matrix, γ and ρ are positive constants, U * This is the solution to the nonlinear model for the input time series U, When functions V and g are differentiable scalar-valued functions that take an input time series U, an output time series Y, and the parameter θ as arguments, and function F is a function that returns an output time series Y that takes an input time series U determined by a nonlinear model and the parameter θ as arguments, Let D be a function V ~ (U, θ) = V(U, F(U, θ), θ), and then an Nr-dimensional vector v = U * -γ∇ U V~(U * For θ, the i-th component d i This is a diagonal matrix formed by arranging the Nr-dimensional vectors d, which were selected and generated, diagonally. q is the i-th component of the function V~, with respect to the Nr-dimensional vector v. i Select This is the generated Nr-dimensional vector, A is a sparse optimization problem n e It has the above linearly independent constraints, and the When the equality constraint is expressed as AU + b = 0 for the input time series U, n e ×Nr dimension It is a matrix, H is the Hessian matrix of the function V to U, Let Pi be the i-th component of the function g ~ (U, θ) = V(U, g(U, θ), θ), where g ~ is the i-th component of the function g ~ i This is the Hessian matrix of U in ~, E is the dual variable corresponding to the inequality constraint g ~ (U, θ) ≤ 0, where η * In that case, η * It is a diagonal matrix with the diagonal elements as follows: ηi is the i-th component of the dual variable corresponding to g ~ (U, θ) ≤ 0, G is g~ (U * It is a diagonal matrix with diagonal elements θ, The inverse matrix of matrix R is matrix R -1 In R 11 inv , R 13 inv , R 14 inv is matrix R -1 A learning device where each component of is divided into blocks in the same way as matrix R, and each row and column corresponds to a submatrix.

13. A control device, An acquisition unit that acquires the parameters learned by the learning device according to any one of claims 1 to 12, A control device comprising: a control unit that calculates the optimal control by solving the sparse optimization problem determined by the acquired parameters.

14. A learning method, where the information processing device A sparse optimization problem using a nonlinear model that approximately represents the output time series from the input time series, comprising a decision step of determining a sparse optimization problem determined by predetermined parameters, A calculation step of calculating an optimal control that corresponds to the parameters and is the input time series evaluated as optimal under a certain evaluation function by solving the determined sparse optimization problem, An update step in which the parameters are updated using the derivative value of the optimal control with respect to the parameters, Execute, A learning method for learning the parameters for determining the sparse optimization problem that yields the optimal control with a small error from the training data, by repeatedly performing the steps of determining the sparse optimization problem, calculating the optimal control, and updating the parameters.

15. A computer program, for use in an information processing device. A sparse optimization problem using a nonlinear model that approximately represents the output time series from the input time series, comprising a decision function that determines the sparse optimization problem determined by predetermined parameters, A calculation function that calculates an optimal control corresponding to the parameters, which is the input time series evaluated as optimal under a certain evaluation function, by solving the determined sparse optimization problem, The parameters are updated using the derivative of the optimal control with respect to the parameters. Update function, Make it run, A computer program that learns the parameters for determining the sparse optimization problem that yields the optimal control with a small error from the training data, by repeatedly performing the steps of determining the sparse optimization problem, calculating the optimal control, and updating the parameters.

Citation Information

Patent Citations

  • System for adapting control parameter

    JP2010086405A

  • Information processing apparatus, information processing method, and program

    JP2021174040A

  • Efficient convolutional sparse coding

    US20160335224A1