Control device, learning device, control method, learning method, and program

JPWO2024095651A5Pending Publication Date: 2025-07-11
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024554316
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-05-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

There is no general method for obtaining the Lyapunov function, which is necessary for ensuring control stability, placing a heavy burden on operators who need to manually discover it, and making it difficult to achieve control stability without prior knowledge of the Lyapunov function.

Method used

A control device and method that uses an objective function indicating stability conditions based on a Lyapunov function, with a function obtaining means to search for a control rule that maximizes stability evaluation, allowing for control commands to be issued and executed without manually discovering the Lyapunov function, utilizing a learning approach that includes actual and latent state data acquisition, transition model learning, and objective function optimization.

Benefits of technology

Enables control stability to be achieved without the need for manual Lyapunov function discovery, by searching for control rules that optimize stability evaluation using objective functions, reducing the computational load and improving control precision.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This control device searches for functions included in an objective function that indicates stability conditions based on a Lyapunov function, so that the evaluation of stability based on the objective function is as good as possible. The control device searches for a control rule for a control target so that the evaluation of stability based on the objective function is as good as possible, and determines a control command for the control target on the basis of the found control rule. The control device controls the control target on the basis of the control rule.
Need to check novelty before this filing date? Find Prior Art

Description

Control device, learning device, control method, learning method, and recording medium

[0001] The present disclosure relates to a control device, a learning device, a control method, a learning method, and a recording medium.

[0002] One method for achieving control stability is to use a Lyapunov function. For example, Patent Document 1 describes that if a Lyapunov function can be found, the stability of a nonlinear model can be guaranteed.

[0003] Japanese Patent Application Publication No. 2021-189934

[0004] A general method for obtaining a Lyapunov function is not known. In this respect, obtaining a Lyapunov function in order to obtain control stability places a heavy burden on the operator who obtains the Lyapunov function. It is preferable to obtain control stability without the need to manually find the Lyapunov function in advance.

[0005] An example of an object of the present disclosure is to provide a control device, a learning device, a control method, a learning method, and a recording medium that can solve the above-mentioned problems.

[0006] According to a first aspect of the present disclosure, a control device includes a function acquisition means that uses an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation using the objective function as much as possible, a behavior decision means that searches for control rules for a control object so as to improve the stability evaluation using the objective function as much as possible, and determines a control command for the control object based on the obtained control rules, and a control execution means that controls the control object based on the control command.

[0007] According to a second aspect of the present disclosure, a learning device includes a function acquisition means that uses an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation using the objective function as much as possible, and a behavior decision means that searches for control rules for a controlled object so as to improve the stability evaluation using the objective function as much as possible, and decides a control command for the controlled object based on the obtained control rules.

[0008] According to a third aspect of the present disclosure, a control method includes a computer using an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation using the objective function as much as possible, searching for control rules for a controlled object so as to improve the stability evaluation using the objective function as much as possible, determining a control command for the controlled object based on the obtained control rules, and controlling the controlled object based on the control command.

[0009] According to a fourth aspect of the present disclosure, a learning method includes a computer using an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation using the objective function as much as possible, searching for control rules for a controlled object so as to improve the stability evaluation using the objective function as much as possible, and determining a control command for the controlled object based on the obtained control rules.

[0010] According to a fifth aspect of the present disclosure, a recording medium stores a program for causing a computer to perform the following operations: using an objective function indicating stability conditions using a Lyapunov function, searching for functions included in the objective function so as to improve the stability evaluation using the objective function; searching for control rules for a controlled object so as to improve the stability evaluation using the objective function; determining control commands for the controlled object based on the obtained control rules; and controlling the controlled object based on the control commands.

[0011] According to a sixth aspect of the present disclosure, a recording medium stores a program for causing a computer to perform the following operations: using an objective function indicating stability conditions using a Lyapunov function, searching for functions included in the objective function so as to improve the stability evaluation using the objective function; searching for control rules for a controlled object so as to improve the stability evaluation using the objective function; and determining a control command for the controlled object based on the obtained control rules.

[0012] According to the present disclosure, it is expected that control stability can be obtained without the need to manually find the Lyapunov function in advance.

[0013] 1 is a diagram illustrating an example of a configuration of a control system according to some embodiments of the present disclosure. FIG. 2 is a diagram illustrating an example of a configuration of a control device according to some embodiments of the present disclosure. FIG. 3 is a diagram illustrating an example of a data flow when an actual state data acquisition unit according to some embodiments of the present disclosure acquires actual state data under control of a control object. FIG. 4 is a diagram illustrating an example of a data flow when a latent state data acquisition unit according to some embodiments of the present disclosure learns a mapping from an actual state to a latent state. FIG. 5 is a diagram illustrating an example of a data flow when a behavior decision unit according to some embodiments of the present disclosure learns a control rule for a control object in a latent state. FIG. 6 is a diagram illustrating an example of a data flow in a control device during additional learning and inference according to some embodiments of the present disclosure. FIG. 7 is a diagram illustrating an example of a processing procedure in which a control device according to some embodiments of the present disclosure learns a control rule for a control object. FIG. 8 is a diagram illustrating another example of a data flow in a control device during operation of a control system according to some embodiments of the present disclosure. FIG. 9 is a diagram illustrating an example of a configuration of a control device dedicated to operation when no Lyapunov reward is calculated according to some embodiments of the present disclosure. FIG. 10 is a diagram illustrating another example of a configuration of a control device according to some embodiments of the present disclosure. FIG. 11 is a diagram illustrating an example of a processing procedure in a control method according to some embodiments of the present disclosure. FIG. 1 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment.

[0014] Embodiments of the present disclosure will be described below, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all combinations of features described in the embodiments are necessarily essential to the solution of the invention. FIG. 1 is a diagram showing an example of the configuration of a control system according to some embodiments of the present disclosure. In the configuration shown in FIG. 1, the control system 1 includes a control device 100 and a controlled object 900.

[0015] The control system 1 is a system that causes a controlled object 900 to perform stable operation. The stable operation of the controlled object 900 here means that the controlled object 900 operates so that a value related to the operation of the controlled object 900 is maintained at a constant target value or a value close to the constant target value.

[0016] The controlled object 900 is not limited to a specific one, and can be various objects that can be controlled and can be expected to operate stably. For example, if the controlled object 900 is an air conditioning unit, an example of stable operation would be to adjust the air temperature to a set temperature, such as a room temperature, and operate to maintain the set temperature. If the controlled object 900 is a cruise control system for an automobile, an example of stable operation would be to drive the automobile at a constant speed. If the controlled object 900 is a power generation plant, an example of stable operation would be to adjust the generated power to a set power and operate to maintain the set power.

[0017] The control device 100 learns control rules for the control object 900 to cause the control object 900 to operate stably, and controls the control object 900. Specifically, the control device 100 searches for control rules for the control object 900 in an optimization problem using an objective function that indicates stability conditions using a Lyapunov function, so as to improve the evaluation of stability using the objective function as much as possible. Learning the control rules is also referred to as learning control.

[0018] The control device 100 is an example of a learning device. The control device 100 or a part thereof may be configured using a computer such as a personal computer, a workstation, or a programmable logic controller (PLC).

[0019] The learning of the control rules for the control target 900 performed by the control device 100 can be considered as a type of reinforcement learning. Reinforcement learning here is machine learning that learns a policy, which is an action rule of an agent that takes an action in a certain environment, based on a state in the environment and a reward that represents an evaluation of the state or the action.

[0020] The control device 100 corresponds to an example of an agent, and the control of the control target 900 performed by the control device 100 corresponds to an example of an action. The control rule by which the control device 100 determines a control command for the control target 900 corresponds to an example of a measure. Determining a control command is also referred to as determining control.

[0021] The control device 100 may learn a control rule for the control object 900 based on a state observed in the actual operating environment of the control object 900. In this case, the operating environment of the control object 900 corresponds to an example of an environment in reinforcement learning. The state observed in the operating environment of the control object 900 corresponds to an example of a state related to the control object 900. Here, the operating environment of the control object 900 includes the control object 900 itself. The state related to the control object 900 may be a state of the control object 900, or a state observed outside the control object 900, or may include both of these.

[0022] Alternatively, as will be described later, the control device 100 may learn the control rule for the control target 900 based on a state obtained by converting the observed state, rather than on the state itself observed in the actual operating environment of the control target 900. In this case, a virtual environment linked to the actual operating environment of the control target 900 corresponds to an example of an environment in reinforcement learning. The state observed in the virtual environment corresponds to an example of a state related to the control target 900.

[0023] Hereinafter, the operating environment of the actual control target 900 will also be referred to as the "real environment" or "real space." A virtual environment linked to the real environment will also be referred to as the "latent environment" or "latent space." A state in the real environment will also be referred to as the "real state." Data indicating the real state will also be referred to as real state data. A variable representing the real state will also be referred to as a "state variable," and the real state or state variable will be represented by x. The real state (value of state variable x) or real state data at time t will be expressed as x. t It is expressed as:

[0024] The state in the latent environment is also called a "latent state." Data indicating the latent state is also called latent state data. A variable representing the latent state is also called a "latent variable," and the latent state or latent variable is represented by z. The latent state (value of latent variable z) or latent state data at time t is represented by z. t The control command for the controlled object 900 is represented by u. The control command at time t is represented by u. t A control command for the control object 900 is also called an "action." A control rule for determining a control command for the control object 900 is also called a "measure."

[0025] Furthermore, hereinafter, time will be expressed in terms of time steps each having a time width Δt, and will be expressed as time step 0, time step 1, time step 2, ..., time step t, time step t+1, .... Time step 0, time step 1, time step 2, ..., time step t, time step t+1, ... will also be expressed as time 0, time 1, time 2, ..., time t, time t+1, .... The length of the time width Δt may be constant or may differ for each time step.

[0026] The stability conditions based on the Lyapunov function used by the control device 100 to learn the control rule for the controlled object 900 will be described. Here, t is an independent variable that takes a real value, and z is a dependent variable that takes the value of an n-dimensional real vector. Also, the function f(z) is a function that maps an n-dimensional real vector to a real number, and the ordinary differential equation shown in equation (1) is an autonomous system with an equilibrium point at the origin z = 0.

[0027]

[0028] For example, the condition for global asymptotic stability is that there exists a function V that satisfies the following equations (2) to (4) over the entire domain of z, and the function V is called a Lyapunov function. The first condition for global asymptotic stability is expressed as equation (2).

[0029]

[0030] The second condition for global asymptotic stability is expressed as equation (3).

[0031]

[0032] A function that satisfies equations (2) and (3) is called a Lyapunov candidate function. The third condition for global asymptotic stability is expressed as equation (4).

[0033]

[0034] When a function V exists that satisfies equations (2) to (4), the equilibrium point at the origin z = 0 is globally asymptotically stable, and all solution trajectories converge to the origin z = 0. In the control of the control object 900, t can be considered to represent time, and z can be considered to represent a latent variable or latent state. Furthermore, the ordinary differential equation shown in equation (1) can be considered to represent the transition of the latent state z under the control of the control object 900.

[0035] When a function V that satisfies equations (2) to (4) exists, the control device 100 can control the control object 900 so that the latent state z converges to the origin z = 0. The equilibrium point can be moved from the origin z = 0 to any point by translating the coordinates. Therefore, when the control device 100 can perform control that satisfies the condition of global asymptotic stability, it can control the control object 900 using any value, not limited to 0, as the target value.

[0036] Furthermore, the condition that a function V exists that satisfies equations (2) to (4) in a certain neighborhood B of the origin z = 0 corresponds to the condition of asymptotic stability, and in this case, the function V is also called a Lyapunov function. When a function V that satisfies the condition of asymptotic stability exists, the control device 100 can control the control object 900 so that the latent state z converges to the origin z = 0 when the time series of the latent state remains in neighborhood B. In this case, too, the control device 100 can control the control object 900 using any value, not limited to 0, as a target value.

[0037] Fig. 2 is a diagram showing an example of the configuration of the control device 100. In the configuration shown in Fig. 2, the control device 100 includes a communication unit 110, a display unit 120, an operation input unit 130, a storage unit 170, and a processing unit 180. The processing unit 180 includes an actual state data acquisition unit 181, a potential state data acquisition unit 182, a transition model acquisition unit 185, a function acquisition unit 191, an evaluation value calculation unit 192, an action determination unit 193, and a control execution unit 194. The transition model acquisition unit 185 includes a vector field calculation unit 186 and a numerical integration unit 187.

[0038] The communication unit 110 communicates with other devices. For example, the communication unit 110 may receive real-state data indicating the real state from a state observation sensor installed in the real environment. The communication unit 110 may also transmit a control command to the control target 900.

[0039] The display unit 120 has a display screen such as a liquid crystal panel or an LED (Light Emitting Diode) panel, and displays various images. For example, the display unit 120 may display various information such as a target value for control of the control target 900, actual state data, or an objective function value. The operation input unit 130 includes input devices such as a keyboard and a mouse, and accepts user operations. For example, the operation input unit 130 may accept a user operation for setting or changing a target value for control of the control target 900.

[0040] The storage unit 170 stores various data. The storage unit 170 is configured using a storage device provided in the control device 100. The processing unit 180 controls each unit of the control device 100 to perform various processes. The functions of the processing unit 180 are performed, for example, by a CPU (Central Processing Unit) provided in the control device 100 reading and executing a program from the storage unit 170.

[0041] The actual state data acquiring unit 181 acquires actual state data when the control device 100 controls the control target 900. The actual state data acquiring unit 181 corresponds to an example of actual state data acquiring means.

[0042] 3 is a diagram showing an example of a data flow when the actual state data acquisition unit 181 acquires actual state data under control of the control target 900. In the example of FIG. 3, the control device 100 receives a control command u t The determined control command u t to the control target 900 to control the control target 900. In addition, the actual state data acquisition unit 181 transmits the actual state data x t For example, the actual state data acquisition unit 181 extracts actual state data from data received by the communication unit 110 from sensors installed in the actual environment.

[0043] When the actual state data acquisition unit 181 acquires the actual state data, the control device 100 issues a control command u to the control target 900. t Alternatively, the action decision unit 193 may randomly decide the control command u from among the candidates of the control command. t Alternatively, the actual state data acquisition unit 181 may determine the control command u t may be determined.

[0044] Actual state data x obtained by the actual state data acquisition unit 181 t and the control command u tBy repeating the determination and transmission of the control command, the actual state data acquisition unit 181 acquires time series data of the control command for the control object 900 and time series data of the actual state when the control object 900 follows the control. The combination of the time series data of the control command and the time series data of the actual state corresponds to an example of training data that indicates, by actual state data, the state transition under control of the control object 900. The training data that indicates, by actual state data, the state transition under control of the control object 900 is also referred to as training data based on the actual state.

[0045] Alternatively, the actual state data acquisition unit 181 acquires the actual state data x t and control command u t and the actual state x t Next state x t+1 Actual state data x indicating t+1 Triple data (x t , u t , x t+1 ) Dataset D x The data set D may be acquired. x corresponds to an example of training data showing state transitions under control of the control object 900 using actual state data.

[0046] The latent state data acquisition unit 182 learns how to convert actual state data into actual state data and converts the data based on the learning results. The latent state data acquisition unit 182 is an example of an actual state data acquisition means. In particular, the latent state data acquisition unit 182 learns a diffeomorphism from the actual state to the latent state. For example, the latent state data acquisition unit 182 includes a neural network and learns the diffeomorphism using a normalizing flow.

[0047] 4 is a diagram showing an example of the flow of data when the latent state data acquisition unit 182 learns the mapping from the actual state to the latent state. In the example of FIG. 4, the latent state data acquisition unit 182 acquires actual state data x t The latent state data z t diffeomorphism z = g(x) that transforms -1(z) is learned. The inverse of a diffeomorphism is also a diffeomorphism. Inverse mapping x = g -1 (z) is the latent state data z t Actual state data x t This corresponds to a diffeomorphism that transforms

[0048] The latent state data acquisition unit 182 performs learning using the Normalizing Flow, thereby obtaining a diffeomorphism and its inverse mapping. However, the method by which the latent state data acquisition unit 182 performs learning is not limited to a specific method as long as it is a method that can obtain a diffeomorphism and its inverse mapping.

[0049] However, no general method has been found to obtain Lyapunov functions or Lyapunov candidate functions. Therefore, obtaining Lyapunov functions is generally not easy. On the other hand, diffeomorphism can map control stability. Specifically, when mapping a space using diffeomorphism, a region where control stability is obtained in the state space before mapping is mapped by the diffeomorphism to a region where control stability is obtained in the state space after mapping.

[0050] The latent state data acquisition unit 182 maps the real environment (the state space of the real state) to the latent environment (the state space of the latent state) using a diffeomorphism, thereby replacing the acquisition of the Lyapunov function in the real environment with the acquisition of the Lyapunov function in the latent environment. For example, if a region satisfying the asymptotic stability condition can be detected in the latent environment, control stability can be obtained in the real environment obtained by applying the inverse mapping of the mapping from the real environment to the latent environment to that region.

[0051] Furthermore, it is expected that the real environment can be mapped into an environment in which it is relatively easy to acquire a Lyapunov function by the mapping performed by the latent state data acquisition unit 182. For example, by mapping from real space to latent space, a complex distribution of trajectories (trajectories of change in x indicating the real state) determined by the initial state (initial condition) and the vector field dx / dt in the real space is mapped into a simpler distribution of trajectories (trajectories of change in z indicating the latent state) determined by the initial state and the vector field dz / dt in the latent space, and it is expected that calculation of loss in the Lyapunov function will become relatively easy.

[0052] However, the control device 100 may learn a control rule for the control object 900 in the real state. In this case, the control device 100 does not need to include the latent state data acquisition unit 182. When the control device 100 learns a control rule for the control object 900 in the real state, z representing the latent state in equations (1) to (4) is replaced with x representing the real state.

[0053] The transition model acquisition unit 185 learns a latent state transition model. The latent state transition model is a model that indicates transitions of latent states in response to control of the control target 900, and outputs latent state data that indicates a next state in a latent state in response to input of latent state data and a control command.

[0054] For example, the transition model acquisition unit 185 acquires latent state data z t and a control command u for the control object 900 at time t. t and the latent state z t Next state z t+1 The latent state data z t+1 Triple data (z t , u t , z t+1 ) Dataset D z The latent state transition model is trained using the above as training data.

[0055] Dataset D z The latent state data acquisition unit 182 acquires the data set D x The data set D is obtained by converting the actual state data contained in z corresponds to an example of training data that indicates, by latent state data, state transitions under control of the control object 900. Training data that indicates, by latent state data, state transitions under control of the control object 900 is also referred to as training data based on latent states.

[0056] The transition model acquisition unit 185 corresponds to an example of a transition model acquisition means. The transition model acquisition unit 185 uses the obtained latent state transition model to generate latent state data z t For the input oft Next state z t+1 The latent state data z t+1 Calculate and output.

[0057] In the following, an example will be described in which the latent state transition model is configured to include a vector field indicating the time derivative of the latent state and a numerical integral of the time derivative of the latent state indicated by the vector field.

[0058] The vector field calculation unit 186 learns a vector field that indicates the time derivative of the latent state, and calculates the time derivative of the latent state using the obtained vector field. In particular, the vector field calculation unit 186 calculates the time derivative of the latent state using the latent state data z t and the control command u for the control object 900 t Upon receiving the input, the vector field calculation unit 186 learns a vector field indicating the time derivative of the latent state under control of the control object 900. Then, using the obtained vector field, the vector field calculation unit 186 calculates the time derivative of the latent state under control of the control object 900. The vector field here indicates the value of the time derivative dz / dt of the latent state for each latent state (value of latent variable z) indicated by a vector. This vector field can be considered as a model representing the ordinary differential equation shown in the above formula (1).

[0059] The numerical integration unit 187 performs numerical integration of the time derivative of the latent state output by the vector field calculation unit 186. In particular, the numerical integration unit 187 performs numerical integration of the latent state data z 0 By receiving the input of t+1 The combination of the vector field calculation unit 186 and the numerical integration unit 187 corresponds to an example of a latent state transition model.

[0060] The method by which the transition model acquisition unit 185 learns the latent state transition model is not limited to a specific method. For example, the vector field calculation unit 186 may be configured to include a neural network, and the transition model acquisition unit 185 may learn the latent state transition model using a known method of learning ordinary differential equations with a neural network under conditioning by actions in reinforcement learning.

[0061] The function acquisition unit 191 learns the Lyapunov function. Specifically, in an optimization problem using an objective function indicating the stability condition by the Lyapunov function, the function acquisition unit 191 searches for a function included in the objective function so as to improve the stability evaluation by the objective function as much as possible. The function acquisition unit 191 corresponds to an example of function acquisition means. The function shown in Equation (5) can be used as the objective function indicating the stability condition by the Lyapunov function.

[0062]

[0063] The function V corresponds to an example of a function included in the objective function. "(∂V / ∂z)(dz / dt)" in formula (5) is a formula obtained by transforming "dV(z) / dt" in formula (4) above using the chain rate shown in formula (6).

[0064]

[0065] max is a function that outputs the maximum value of its arguments. The value of "max(0, (∂V / ∂z)(dz / dt))" is 0 when (∂V / ∂z)(dz / dt)≦0, and is greater than 0 when (∂V / ∂z)(dz / dt)>0. Therefore, when the above formula (4) holds, max(0, (∂V / ∂z)(dz / dt)) = 0. On the other hand, when dV(z) / dt>0, max(0, (∂V / ∂z)(dz / dt))>0.

[0066] The value of "max(0, -V(z))" is 0 when the above formula (3) is satisfied, and is greater than 0 when V(z)<0. 2 The value of "(0)" is 0 when the above formula (2) is satisfied, and is greater than 0 when the formula (2) is not satisfied.

[0067] Equation (5) takes on 0 or a negative value, and the larger the value of equation (5) (i.e., the closer it is to 0), the better the stability evaluation. When all of the conditions shown in equations (2) to (4) are met, equation (5) takes on the maximum value of 0. The value of equation (5) is also called the negative Lyapunov reward (Lyapunov penalty) and is represented by r.

[0068] For example, consider a case where both the function acquisition unit 191 and the behavior determination unit 193 are configured using neural networks. When the function acquisition unit 191 uses a neural network to learn the Lyapunov function, various neural networks with differentiable and continuous activation functions can be used, and the Lyapunov function is differentiable. Furthermore, the ordinary differential equation learned by the transition model acquisition unit 185 is also differentiable.

[0069] Since both the Lyapunov function and the ordinary differential equation are differentiable, the objective function shown in equation (5) becomes an immediate reward that is differentiable with respect to the control rule for the controlled object. Specifically, the objective function shown in equation (5) is differentiable with respect to the parameters of the neural network that constitutes the function acquisition unit 191 and the parameters of the neural network that constitutes the action determination unit 193. This allows the neural network that constitutes the function acquisition unit 191 and the neural network that constitutes the action determination unit 193 to be trained using a gradient method such as backpropagation. For example, if the negative Lyapunov reward r at each time t is t N-step discount reward and R t (N) is expressed as in equation (7).

[0070]

[0071] t 0 indicates the start time of the N steps for which the discounted reward sum is to be calculated. N is an integer constant of N≧1. γ is the time t 0 is a constant that indicates the discount rate of future reward values. t At time t 0Time from to time t (number of steps in the time step) t-t 0 γ is raised to the power r t - t0 are multiplied.

[0072] N-step discount reward and R t (N) can be treated as a function of the parameters of the neural network that constitutes the function acquisition unit 191 and the parameters of the neural network that constitutes the behavior decision unit 193, and can be differentiated with respect to these parameters. Here, the parameters of the neural network that constitutes the function acquisition unit 191 and the parameters of the neural network that constitutes the behavior decision unit 193 are represented by θ. The gradient ∂R t (N) / ∂θ, the N-step discounted reward sum R t (N) By searching for the value of θ so as to maximize the sum of N-step discounted rewards R t (N) is an example of maximizing the negative Lyapunov reward r.

[0073] Similar calculations can be used to learn the latent state transition model learned by the transition model acquisition unit 185. Learning the latent state transition model corresponds to correcting the dynamics itself so that the dynamics in the latent space satisfy the stability condition.

[0074] Note that even if the function acquisition unit 191 searches for the function V using the objective function shown in Equation (5), it is not necessarily possible to obtain a function that satisfies the global asymptotic stability condition or a function that satisfies the asymptotic stability condition. For example, there may be a case where the function acquisition unit 191 cannot obtain a function V that makes the negative Lyapunov reward r equal to 0.

[0075] Even if the function acquisition unit 191 cannot obtain a function V such that the negative Lyapunov reward r is 0, by acquiring a function V such that the negative Lyapunov reward r is as large (close to 0) as possible, the control device 100 is expected to be able to learn relatively stable control. For example, consider a case where equations (2) and (3) hold over the entire region of the latent state that can be subject to control for the control object 900, and equation (4) includes a region where dV(z) / dt<0 and a region where dV(z) / dt≧0. In this case, it is considered that the latent state approaches the origin z=0 in the region where dV(z) / dt<0, and moves away from the origin z=0 in the region where dV(z) / dt≧0.

[0076] At this time, by the function acquisition unit 191 acquiring a function V that maximizes the negative Lyapunov reward r, it is conceivable that the value of dV(z) / dt will be relatively small even in the region where dV(z) / dt≧0, and the distance by which the latent state moves away from the origin z=0 will be relatively small. The distance by which the latent state approaches the origin z=0 in the region where dV(z) / dt<0 is greater than the distance by which the latent state moves away from the origin z=0 in the region where dV(z) / dt≧0, and therefore it is expected that the latent state will approach the origin z=0.

[0077] Furthermore, the condition that there exists a function V that satisfies equations (2), (3), and (8) in a neighborhood B of the origin z = 0 corresponds to the condition of Lyapunov stability, and in this case too, the function V is called a Lyapunov function.

[0078]

[0079] When a function V that satisfies the conditions of Lyapunov stability exists, the control device 100 can control the control object 900 so that the time series of the latent state continues to remain within a certain vicinity of the origin z = 0. In this case, the control device 100 can also control the control object 900 using any value, not limited to 0, as the target value. Therefore, when the function acquisition unit 191 acquires a function that satisfies the conditions of Lyapunov stability, it is expected that the control device 100 can control the latent state so that it continues to remain within a certain vicinity of the origin z = 0, even if it is not possible to control the latent state so that it converges to the origin z = 0.

[0080] Furthermore, the control device 100 may narrow the area for learning the Lyapunov function toward the equilibrium point, for example, by limiting the area for learning the Lyapunov function to an area within a predetermined distance from the equilibrium point. Specifically, the control device 100 may limit the data used for learning the Lyapunov function to data within a predetermined area relatively close to the equilibrium point. This reduces the calculation load in learning the Lyapunov function. On the other hand, if the control device 100 targets a wider area for learning the Lyapunov function, control stability may be ensured over a wider area than when the area is limited.

[0081] The objective function used by the function acquisition unit 191 is not limited to that shown in equation (5). For example, the function acquisition unit 191 may use an objective function in which the smaller the value of the objective function, the better the evaluation of stability. The value of the objective function indicating the evaluation of stability is also collectively referred to as a Lyapunov reward.

[0082] The evaluation value calculation unit 192 calculates the value of the objective function. For example, when the function acquisition unit 191 uses the objective function shown in Equation (5), the evaluation value calculation unit 192 calculates a negative Lyapunov reward r. For example, the evaluation value calculation unit 192 acquires the value of the differential dz / dt calculated by the vector field calculation unit 186 and the function V acquired by the function acquisition unit 191, and calculates the value of the objective function.

[0083] The behavior decision unit 193 learns a control rule for the control object 900. In particular, the behavior decision unit 193 uses the same objective function as the objective function used by the function acquisition unit 191 to learn the Lyapunov function to search for a control rule for the control object 900 so as to improve the stability evaluation by the objective function as much as possible. For example, the behavior decision unit 193 searches for a control rule by model-based reinforcement learning using the objective function shown in equation (5) so as to maximize the negative Lyapunov reward r. The behavior decision unit 193 determines a control command for the control object 900 using the obtained control rule. The behavior decision unit 193 corresponds to an example of behavior decision means.

[0084] The control execution unit 194 controls the control target 900 based on the control command determined by the action determination unit 193. For example, the control execution unit 194 controls the communication unit 110 to transmit the control command to the control target 900. The control execution unit 194 corresponds to an example of a control execution means. The combination of the action determination unit 193 and the control execution unit 194 corresponds to an example of an agent in reinforcement learning.

[0085] 5 is a diagram showing an example of the flow of data when the action decision unit 193 learns a control rule for the control object 900 in the latent state. In the example of Fig. 5, the action decision unit 193 learns a control rule for the control object 900. In addition, the function acquisition unit 191 learns a Lyapunov function for evaluating the stability of the control learned by the action decision unit 193.

[0086] Furthermore, prior to the learning of the control rules by the action determining unit 193 and the learning of the Lyapunov function by the function obtaining unit 191, the transition model obtaining unit 185 performs learning of a latent state transition model for calculating a next state in a latent state according to the control of the control object 900 determined by the action determining unit 193. Alternatively, the transition model obtaining unit 185 may perform learning of the latent state transition model in parallel with the learning of the control rules by the action determining unit 193 and the learning of the Lyapunov function by the function obtaining unit 191.

[0087] In the example of FIG. 5, the behavior decision unit 193 determines the latent state z t The control command u for the control object 900 in response to t The transition model acquisition unit 185 determines the control command u determined by the control object 900. t Depending on t The potential state z at time step t+1 is the next state for t+1 Specifically, the vector field calculation unit 186 calculates the latent state z t and control command u t The numerical integration unit 187 calculates the value of the time derivative dz / dt of the latent variable z according to the initial state z 0 Based on this, the numerical integration unit 187 numerically integrates the time series of the time derivative dz / dt of the latent variable z. t+1 (latent state z t+1 The behavior determination unit 193 and the transition model acquisition unit 185 repeat, for each time, the determination of a control command according to the latent state and the calculation of a next state according to the control command.

[0088] Furthermore, the function acquisition unit 191 learns a Lyapunov function based on the latent state calculated by the transition model acquisition unit 185 and the negative Lyapunov reward r calculated by the evaluation value calculation unit 192, and calculates the function V. The purpose of the learning by the function acquisition unit 191 is to acquire a Lyapunov function, but as described above, the function V is not necessarily a Lyapunov function.

[0089] The evaluation value calculation unit 192 receives as input the function V calculated by the function acquisition unit 191 and the value of dz / dt calculated by the vector field calculation unit 186, and calculates a negative Lyapunov reward using the function V. In relation to the learning of the Lyapunov function performed by the function acquisition unit 191, the negative Lyapunov reward r can be considered an evaluation index indicating the degree to which the function V calculated by the function acquisition unit 191 satisfies the conditions for being a Lyapunov function. The function acquisition unit 191 learns the Lyapunov function so that the value of the negative Lyapunov reward r becomes as large as possible. The function acquisition unit 191 and the evaluation value calculation unit 192 repeat, at each time point, learning the Lyapunov function, calculating the function V, and calculating the negative Lyapunov reward r using the function V.

[0090] The evaluation value calculation unit 192 also outputs the calculated negative Lyapunov reward r to the action decision unit 193. In relation to the learning of the control rule for the controlled object 900 performed by the action decision unit 193, the negative Lyapunov reward r can be considered as an evaluation index indicating the degree to which the control of the controlled object 900 by the action decision unit 193 satisfies the stability condition based on the Lyapunov function. The action decision unit 193 learns the control rule for the controlled object 900 so that the value of the negative Lyapunov reward r becomes as large as possible.

[0091] In this way, the action decision unit 193 determines the control command u t Latent state data z according to t+1 is acquired from the transition model acquisition unit 185, and the latent state data z t+1 and a negative Lyapunov reward r according to the value of the differential dz / dt indicating the state transition, are obtained from the evaluation value calculation unit 192. The action decision unit 193 uses these data to search for control rules by model-based reinforcement learning.

[0092] After the learning of the control rule for the control object 900 in the latent state is completed, the control device 100 may perform additional learning for fine tuning of the control for the control object 900 in the real environment.

[0093] 6 is a diagram showing an example of the flow of data in the control device 100 during additional learning and inference (when the control system 1 is in operation). In the example of FIG. 6, the action decision unit 193 issues a control command u t The determined control command u t The control command u is sent to the control target 900 via the control execution unit 194. t The control object 900 operates in accordance with the above, and the actual state transitions according to the operation of the control object 900.

[0094] The latent state data acquisition unit 182 acquires real state data x obtained by observing the real environment. t+1 The latent state data z t+1 The behavior decision unit 193 converts the latent state data z t+1 The control command u for the control object 900 in response to t+1 The determined control command u t+1 The behavior decision unit 193, the control object 900, and the latent state data acquisition unit 182 repeat, at each time, determining a control command for the control object 900, notifying the control command, operating in accordance with the control command, and converting the actual state data into the latent state data. t+1 The latent state data z t+1 By converting the control rule into the control rule, the behavior decision unit 193 can control the control target 900 using the control rule obtained by learning in the example of FIG.

[0095] During additional learning, the behavior decision unit 193 learns control of the control object 900 in addition to controlling the control object 900. In this case, the behavior decision unit 193 may learn the control rule using an objective function different from the objective function used for learning in the example of FIG. 5 . For example, if the control object 900 is an air conditioning facility and the control object 900 is controlled so that the air temperature to be adjusted maintains a set temperature, the behavior decision unit 193 may use an objective function that indicates a better evaluation the closer the measured air temperature is to the set temperature. In this case, an evaluation function using air temperature data in the actual environment may be used, or an evaluation function using air temperature data in the potential environment may be used. Alternatively, the behavior decision unit 193 may learn the control rule using the negative Lyapunov reward r calculated by the evaluation value calculation unit 192, as in the example of FIG. 5 .

[0096] In addition, the action decision unit 193 determines the control command u t to the transition model acquisition unit 185. The data flow and processing in the transition model acquisition unit 185, function acquisition unit 191, and evaluation value calculation unit 192 are the same as those in FIG. 5 . The transition model acquisition unit 185 calculates the next state in the latent state at each time point in accordance with the control command determined by the controlled object 900. The function acquisition unit 191 and the evaluation value calculation unit 192 repeat, at each time point, learning the Lyapunov function, calculating the function V, and calculating the negative Lyapunov reward r using the function V. The user can confirm the evaluation of the control stability by referring to the negative Lyapunov reward r calculated by the evaluation value calculation unit 192.

[0097] 7 is a diagram showing an example of a procedure for the control device 100 to learn control of the control target 900. In the process of FIG. 7, the actual state data acquisition unit 181 acquires training data based on the actual state (step S101).

[0098] Next, the latent state data acquisition unit 182 learns a diffeomorphism that maps the actual state to the latent state (step S102).The latent state data acquisition unit 182 then uses the learned mapping to acquire training data for the latent state (step S103).Next, the transition model acquisition unit 185 uses the training data for the latent state to train an environment model (step S104).

[0099] Next, the function acquisition unit 191 learns the Lyapunov function, and the behavior determination unit 193 learns the control rule for the control object 900 (step S105). In particular, the function acquisition unit 191 calculates the function V based on the latent state data calculated by the transition model acquisition unit 185, and learns the Lyapunov function so that the negative Lyapunov reward r calculated by the evaluation value calculation unit 192 is as large as possible. Furthermore, the behavior determination unit 193 calculates a control command based on the latent state data calculated by the transition model acquisition unit 185, and learns the control rule for the control object 900 so that the negative Lyapunov reward r calculated by the evaluation value calculation unit 192 is as large as possible.

[0100] The function acquisition unit 191 and the behavior determination unit 193 repeat the learning of the Lyapunov function and the learning of the control rule for the control object 900 until a termination condition for learning in the latent environment of control for the control object 900 is met. The termination condition here is not limited to a specific one. For example, the termination condition here may be a condition that learning in the latent environment of control for the control object 900 has been repeated a predetermined number of times in the time step. Alternatively, the termination condition here may be a condition that the value of the negative Lyapunov reward r is greater than a predetermined threshold.

[0101] After the processing unit 180 determines that the learning termination condition for the control of the control object 900 in the latent environment is satisfied, the action determination unit 193 performs fine tuning of the control of the control object 900 (step S106). As described with reference to FIG. 6 , in fine tuning, the action determination unit 193 determines a control command and learns a control rule using latent state data obtained by the latent state data acquisition unit 182 mapping the actual state data. The function acquisition unit 191 also learns the Lyapunov function in step S106. The evaluation value calculation unit 192 also calculates the negative Lyapunov reward r in step S106.

[0102] The function acquisition unit 191 and the behavior determination unit 193 repeat the learning of the Lyapunov function and the fine tuning of the control of the controlled object 900 until a termination condition for fine tuning is met. The termination condition here is not limited to a specific one. For example, the termination condition here may be a condition that the learning of the control rule for the controlled object 900 in fine tuning has been repeated a predetermined number of times in the time steps. Alternatively, the termination condition here may be a condition that the control rule used by the behavior determination unit 193 to calculate a control command has not been changed for more than a predetermined number of times in the time steps. After step S106, the control device 100 terminates the processing of FIG. 7.

[0103] FIG. 8 is a diagram showing another example of the flow of data in the control device 100 during operation of the control system 1. FIG. 8 shows an example in which the Lyapunov reward is not calculated. In the example of FIG. 8, the data flow and processing in the action decision unit 193, the control object 900, and the latent state data acquisition unit 182 are the same as those in FIG. 6. The action decision unit 193, the control object 900, and the latent state data acquisition unit 182 repeat, at each time point, the determination and transmission of a control command for the control object 900, the operation in accordance with the control command, and the conversion of actual state data into latent state data.

[0104] On the other hand, the example of Fig. 8 does not show the transition model acquisition unit 185, function acquisition unit 191, and evaluation value calculation unit 192, among the units shown in Fig. 6. These units acquire a function V according to a control command calculated by the controlled object 900, and perform various processes for calculating a negative Lyapunov reward r using the obtained function V. In contrast, in the example of Fig. 8, since calculation of the negative Lyapunov reward r is not performed, as described above, the transition model acquisition unit 185, function acquisition unit 191, and evaluation value calculation unit 192 are not shown.

[0105] Even when the calculation of the negative Lyapunov reward r is not performed in this way, the control device 100 shown in Fig. 2 may be used during operation of the control system 1. In this case, the control device 100 may perform processing during operation using the units shown in Fig. 8 out of the units shown in Fig. 2. Alternatively, a control device 200 may be provided separately from the control device 100 and dedicated to operation.

[0106] 9 is a diagram showing an example of the configuration of a control device dedicated to operation when the Lyapunov reward is not calculated. In the configuration shown in Fig. 9, the control device 200 includes a communication unit 110, a display unit 120, an operation input unit 130, a storage unit 170, and a processing unit 280. The processing unit 280 includes a latent state data acquisition unit 182, a behavior determination unit 193, and a control execution unit 194.

[0107] 9, parts having the same functions as those in FIG. 1 are denoted by the same reference numerals (110, 120, 130, 170, 182, 193, 194), and detailed description thereof will be omitted here. The control device 200 differs from the control device 100 in that the processing unit 280 includes only some of the parts of the processing unit 180 of the control device 100. In other respects, the control device 200 is similar to the control device 100.

[0108] 9, the processing unit 280 includes the behavior determination unit 193 and latent state data acquisition unit 182 shown in FIG. 8, and a control execution unit 194 that controls the control target 900 based on the control command determined by the behavior determination unit 193. In the example of FIG. 9, it is assumed that the latent state data acquisition unit 182 has already acquired a diffeomorphism. Therefore, in the example of FIG. 9, the latent state data acquisition unit 182 does not need to have a function for learning a diffeomorphism.

[0109] As described above, the transition model acquisition unit 185, the function acquisition unit 191, and the evaluation value calculation unit 192 acquire the function V according to the control command calculated by the controlled object 900, and perform various processes for calculating the negative Lyapunov reward r using the obtained function V. These are not required if the calculation of the negative Lyapunov reward r is not performed, and are therefore not shown in FIG.

[0110] When the control device 200 is used instead of the control device 100 during operation of the control system 1, the settings of each unit of the control device 200 may be made based on the learning results of the control device 100. In particular, the mapping obtained through learning by the latent state data acquisition unit 182 of the control device 100 may be set in the latent state data acquisition unit 182 of the control device 200. Furthermore, the control rule obtained through learning by the action decision unit 193 of the control device 100 may be set in the action decision unit 193 of the control device 200.

[0111] As described above, the function acquisition unit 191 uses an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation using the objective function as much as possible. The behavior determination unit 193 searches for control rules for the controlled object 900 so as to improve the stability evaluation using the objective function as much as possible, and determines a control command for the controlled object based on the obtained control rules. The control execution unit 194 controls the controlled object 900 based on the obtained control commands.

[0112] According to the control device 100, the function acquisition unit 191 searches for a function using an objective function that indicates the condition for stability using a Lyapunov function, and thus it is expected that control stability can be obtained without the need to manually discover a Lyapunov function in advance. Even if a Lyapunov function cannot be obtained by the function search performed by the function acquisition unit 191, the function acquisition unit 191 searches for a function that improves the evaluation of stability using the objective function as much as possible, and thus it is expected that control stability can be obtained as described above.

[0113] Furthermore, the real state data acquiring unit 181 acquires training data indicating state transitions of the control object 900 under control of the control object 900 in the form of real state data indicating real states, which are states of the control object 900 in a real environment. The latent state data acquiring unit 182 converts the training data indicating state transitions of the control object 900 under control of the control object 900 in the form of real state data into training data indicating state transitions of the control object 900 under control of the control object 900 in the form of latent state data indicating latent states, which are states in a virtual environment. The transition model acquiring unit 185 performs training of the latent state transition model using training data indicating state transitions of the control object 900 under control of the control object 900 in the form of latent state data. The latent state transition model is a model that calculates state transitions of the latent state under control of the control object 900. The latent state is a state indicated by the latent state data. The function acquiring unit 191 searches for a function using the latent state data output by the latent state transition model. The behavior decision unit 193 searches for a control rule for the control object 900 using the latent state data output by the latent state transition model.

[0114] According to the control device 100, the function acquisition unit 191 can search for the function V in the latent environment. In this respect, according to the control device 100, it is expected that the function acquisition unit 191 can learn the Lyapunov function in an environment in which it is relatively easy to acquire the Lyapunov function.

[0115] Furthermore, the latent state data acquisition unit 182 performs diffeomorphism learning and uses the obtained diffeomorphism to convert training data indicating, as real state data, state transitions of the controlled object 900 under control of the controlled object 900 into training data indicating, as latent state data, state transitions of the controlled object 900 under control of the controlled object 900. The control device 100 can map the real state to the latent state using the diffeomorphism, and can replace the acquisition of the Lyapunov function in the real environment with the acquisition of the Lyapunov function in the latent environment. For example, if a region satisfying the condition for asymptotic stability based on the Lyapunov function can be detected in the latent environment, control stability can be obtained in a region in the real environment obtained by applying the inverse mapping of the diffeomorphism from the real environment to the latent environment to that region.

[0116] The latent state transition model also includes a vector field indicating the time derivative of the latent state and a numerical integral of the time derivative of the latent state indicated by the vector field. The transition model acquisition unit 185 learns this vector field. According to the control device 100, the time derivative of the latent state indicated by the vector field can be used to calculate the value of the objective function, eliminating the need to separately calculate the time derivative of the latent state to calculate the value of the objective function. In this respect, the control device 100 can relatively reduce the calculation load.

[0117] Furthermore, the behavior decision unit 193 further searches for a control rule for the control object 900, using the latent state data converted by the latent state data acquisition unit 182 from the actual state data obtained under control of the control object 900 in the real environment. It is expected that the control device 100 will be able to search for a control rule for the control object 900 with higher accuracy.

[0118] FIG. 10 is a diagram illustrating another exemplary configuration of a control device according to some embodiments of the present disclosure. In the configuration illustrated in FIG. 10 , the control device 610 includes a function acquisition unit 611, an action determination unit 612, and a control execution unit 613. In this configuration, the function acquisition unit 611 uses an objective function indicating stability conditions based on a Lyapunov function to search for a function included in the objective function so as to improve the stability evaluation based on the objective function. The action determination unit 612 searches for a control rule for the control object so as to improve the stability evaluation based on the objective function. The control execution unit 613 controls the control object based on the obtained control rule. The function acquisition unit 611 is an example of a function acquisition means. The action determination unit 612 is an example of a action determination means. The control execution unit 613 is an example of a control execution means.

[0119] According to the control device 610, the function acquisition unit 611 searches for a function using an objective function that indicates the condition for stability using the Lyapunov function, and it is expected that control stability can be obtained without the need to manually discover the Lyapunov function in advance. Even if the Lyapunov function cannot be obtained by the function search performed by the function acquisition unit 611, it is expected that control stability can be obtained by the function acquisition unit 611 searching for a function that improves the evaluation of stability using the objective function as much as possible.

[0120] The function acquisition unit 611 can be realized, for example, by using the functions of the function acquisition unit 191 in Fig. 2, etc. The behavior decision unit 612 can be realized, for example, by using the functions of the behavior decision unit 193 in Fig. 2, etc. The control execution unit 613 can be realized, for example, by using the functions of the control execution unit 194 in Fig. 2, etc.

[0121] FIG. 11 is a diagram illustrating an example of the configuration of a learning device according to some embodiments of the present disclosure. In the configuration illustrated in FIG. 11 , the learning device 620 includes a function acquisition unit 621 and an action determination unit 622. In this configuration, the function acquisition unit 621 uses an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation using the objective function as much as possible. The action determination unit 622 searches for a control rule for the control object so as to improve the stability evaluation using the objective function as much as possible. The function acquisition unit 621 corresponds to an example of function acquisition means. The action determination unit 622 corresponds to an example of action determination means.

[0122] According to the learning device 620, the function acquisition unit 621 searches for a function using an objective function that indicates the condition for stability using the Lyapunov function, and it is expected that control stability can be obtained without the need to manually discover the Lyapunov function in advance. Even if the Lyapunov function cannot be obtained by the function search performed by the function acquisition unit 621, it is expected that control stability can be obtained by the function acquisition unit 621 searching for a function that improves the evaluation of stability using the objective function as much as possible.

[0123] The function acquisition unit 621 can be realized using, for example, the function of the function acquisition unit 191 in Fig. 2. The behavior decision unit 622 can be realized using, for example, the function of the behavior decision unit 193 in Fig. 2.

[0124] 12 is a diagram illustrating an example of a processing procedure in a control method according to some embodiments of the present disclosure. The control method illustrated in FIG. 12 includes obtaining a function (step S611), determining an action (step S612), and executing control (step S613). In obtaining a function (step S611), the computer uses an objective function indicating stability conditions based on a Lyapunov function to search for a function included in the objective function so as to improve the stability evaluation based on the objective function. In determining an action (step S612), the computer searches for a control rule for the control object so as to improve the stability evaluation based on the objective function. In executing control (step S613), the computer controls the control object based on the obtained control rule.

[0125] According to the control method shown in Fig. 12, by searching for a function using an objective function that indicates the stability condition based on the Lyapunov function, it is expected that control stability can be obtained without the need to manually find the Lyapunov function in advance. Even if the Lyapunov function cannot be obtained by searching for a function, it is expected that control stability can be obtained by searching for a function that improves the stability evaluation based on the objective function as much as possible.

[0126] 13 is a diagram illustrating an example of a processing procedure in a learning method according to some embodiments of the present disclosure. The learning method illustrated in FIG. 13 includes acquiring a function (step S621) and determining an action (step S622). In acquiring a function (step S621), the computer uses an objective function indicating stability conditions based on a Lyapunov function to search for a function included in the objective function so as to improve the stability evaluation based on the objective function. In determining an action (step S622), the computer searches for a control rule for the controlled object so as to improve the stability evaluation based on the objective function.

[0127] According to the learning method shown in Fig. 13, by searching for a function using an objective function that indicates the stability conditions based on the Lyapunov function, it is expected that control stability can be obtained without the need to manually find the Lyapunov function in advance. Even if a Lyapunov function cannot be obtained by searching for a function, it is expected that control stability can be obtained by searching for a function that improves the evaluation of stability based on the objective function as much as possible.

[0128] 14 is a schematic block diagram illustrating the configuration of a computer according to at least one embodiment. In the configuration shown in FIG. 14, a computer 700 includes a CPU 710, a main memory device 720, an auxiliary memory device 730, an interface 740, and a non-volatile recording medium 750.

[0129] One or more of the control device 100, control device 200, control device 610, and learning device 620, or a part thereof, may be implemented in a computer 700. In this case, the operation of each of the above-described processing units is stored in the auxiliary storage device 730 in the form of a program. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program. The CPU 710 also allocates storage areas in the main storage device 720 corresponding to each of the above-described storage units in accordance with the program. Communication between each device and other devices is performed by an interface 740 having a communication function and performing communication under the control of the CPU 710. The interface 740 also has a port for a non-volatile recording medium 750, and reads information from the non-volatile recording medium 750 and writes information to the non-volatile recording medium 750.

[0130] When the control device 100 is implemented in a computer 700, the operations of the processing unit 180 and each of its units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0131] Furthermore, the CPU 710 allocates a storage area for the storage unit 170 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0132] When the control device 200 is implemented in a computer 700, the operations of the processing unit 280 and each of its units are stored in the form of a program in an auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0133] Furthermore, the CPU 710 allocates a storage area for the storage unit 170 in the main storage device 720 in accordance with the program. Communication with other devices by the communication unit 110 is performed by the interface 740 having a communication function and operating under the control of the CPU 710. Display of images by the display unit 120 is performed by the interface 740 having a display device and displaying various images under the control of the CPU 710. Reception of user operations by the operation input unit 130 is performed by the interface 740 having an input device and receiving user operations under the control of the CPU 710.

[0134] When the control device 610 is implemented in the computer 700, the operations of the function acquisition unit 611, the behavior determination unit 612, and the control execution unit 613 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-mentioned processing in accordance with the program.

[0135] Furthermore, the CPU 710 allocates a storage area in the main storage device 720 for the control device 610 to perform processing in accordance with the program. Communication between the control device 610 and other devices is performed by the interface 740, which has a communication function and operates under the control of the CPU 710. Interaction between the control device 610 and a user is performed by the interface 740, which has an input device and an output device, presenting information to the user via the output device under the control of the CPU 710 and accepting user operations via the input device.

[0136] When the learning device 620 is implemented in the computer 700, the operations of the function acquisition unit 621 and the behavior determination unit 622 are stored in the form of a program in the auxiliary storage device 730. The CPU 710 reads the program from the auxiliary storage device 730, loads it into the main storage device 720, and executes the above-described processing in accordance with the program.

[0137] Furthermore, CPU 710 allocates a storage area in main memory 720 for learning device 620 to perform processing in accordance with the program. Communication between learning device 620 and other devices is performed by interface 740, which has a communication function and operates under the control of CPU 710. Interaction between learning device 620 and a user is performed by interface 740, which has an input device and an output device, presenting information to the user via the output device under the control of CPU 710 and accepting user operations via the input device.

[0138] One or more of the above-described programs may be recorded on nonvolatile recording medium 750. In this case, interface 740 may read the programs from nonvolatile recording medium 750. Then, CPU 710 may directly execute the programs read by interface 740, or may temporarily store the programs in main storage device 720 or auxiliary storage device 730 and then execute them.

[0139] In addition, a program for executing all or part of the processing performed by the control device 100, control device 200, control device 610, and learning device 620 may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be loaded into a computer system and executed to perform the processing of each unit. Note that the term "computer system" here includes hardware such as an operating system (OS) and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, read-only memories (ROMs), and compact disc read-only memories (CD-ROMs), as well as storage devices such as hard disks built into computer systems. The program may be designed to implement part of the aforementioned functions, or may be capable of implementing the aforementioned functions in combination with a program already recorded on the computer system.

[0140] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention.

[0141] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0142] (Supplementary Note 1) A control device comprising: a function acquisition means for searching for a function included in an objective function, using an objective function indicating stability conditions using a Lyapunov function, so as to improve the stability evaluation by the objective function as much as possible; an action decision means for searching for a control rule for a controlled object, so as to improve the stability evaluation by the objective function as much as possible, and determining a control command for the controlled object based on the obtained control rule; and a control execution means for controlling the controlled object based on the control command.

[0143] (Supplementary Note 2) A control device according to Supplementary Note 1, comprising: actual state data acquisition means for acquiring training data indicating state transitions of the control object under control of the control object in the form of actual state data indicating actual states, which are states of the control object in a real environment; latent state data acquisition means for converting the training data indicating state transitions of the control object under control of the control object in the form of actual state data, into training data indicating state transitions of the control object under control of the control object in the form of latent state data indicating state transitions in a virtual environment; and transition model acquisition means for learning a latent state transition model, which is a model for calculating the transitions of the latent states under control of the control object, using the training data indicating state transitions of the control object under control of the control object in the form of latent state data, wherein the function acquisition means searches for the function using the latent state data output by the latent state transition model, and the action decision means searches for the control rule using the latent state data output by the latent state transition model.

[0144] (Supplementary Note 3) The control device according to Supplementary Note 2, wherein the latent state data acquisition means learns a diffeomorphism and converts training data indicating, as actual state data, state transitions of the control object under control of the control object into training data indicating, as latent state data, state transitions of the control object under control of the control object using the obtained diffeomorphism.

[0145] (Supplementary Note 4) The control device according to Supplementary Note 2 or Supplementary Note 3, wherein the latent state transition model includes a vector field indicating a time derivative of the latent state and a numerical integral of the time derivative of the latent state indicated by the vector field, and the transition model acquisition means learns the vector field.

[0146] (Supplementary Note 5) The control device according to any one of Supplementary Notes 2 to 4, wherein the behavior decision means further searches for the control rule using latent state data obtained by the latent state data acquisition means converting actual state data obtained under control of the control object in a real environment.

[0147] (Supplementary Note 6) A learning device comprising: a function acquisition means for searching for a function included in an objective function using an objective function indicating stability conditions using a Lyapunov function so as to improve the stability evaluation by the objective function; and a behavior decision means for searching for a control rule for a controlled object so as to improve the stability evaluation by the objective function, and for deciding a control command for the controlled object based on the obtained control rule.

[0148] (Supplementary Note 7) A control method including: a computer uses an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation by the objective function; a computer searches for control rules for a controlled object so as to improve the stability evaluation by the objective function; a computer determines control commands for the controlled object based on the control rules; and a computer controls the controlled object based on the control commands.

[0149] (Supplementary Note 8) A learning method including: a computer uses an objective function indicating stability conditions using a Lyapunov function to search for functions included in the objective function so as to improve the stability evaluation by the objective function as much as possible; a computer searches for control rules for a controlled object so as to improve the stability evaluation by the objective function as much as possible; and a computer determines a control command for the controlled object based on the control rules obtained.

[0150] (Supplementary Note 9) A recording medium storing a program for causing a computer to execute the following: using an objective function indicating stability conditions using a Lyapunov function, searching for functions included in the objective function so as to improve the stability evaluation by the objective function as much as possible; searching for control rules for a controlled object so as to improve the stability evaluation by the objective function as much as possible, and determining control commands for the controlled object based on the obtained control rules; and controlling the controlled object based on the control commands.

[0151] (Supplementary Note 10) A recording medium storing a program for causing a computer to execute the following: using an objective function indicating stability conditions using a Lyapunov function, searching for functions included in the objective function so as to improve the stability evaluation using the objective function; searching for control rules for a controlled object so as to improve the stability evaluation using the objective function, and determining a control command for the controlled object based on the obtained control rules.

[0152] This application claims priority based on Japanese Patent Application No. 2022-175607, filed November 1, 2022, the disclosure of which is incorporated herein by reference in its entirety.

[0153] The present disclosure may be applied to a control device, a learning device, a control method, a learning method, and a recording medium.

[0154] DESCRIPTION OF SYMBOLS 1 Control system 100, 200, 610 Control device 110 Communication unit 120 Display unit 130 Operation input unit 170 Memory unit 180, 280 Processing unit 181 Actual state data acquisition unit 182 Latent state data acquisition unit 185 Transition model acquisition unit 186 Vector field calculation unit 187 Numerical integration unit 191, 611, 621 Function acquisition unit 192 Evaluation value calculation unit 193, 612, 622 Action decision unit 194, 613 Control execution unit 620 Learning device 900 Control target

Claims

1. Function acquisition means for searching for a function included in the objective function so that the evaluation of stability by the objective function becomes as good as possible, using an objective function indicating the stability condition by a Lyapunov function; Action determination means for searching for a control rule for a control target so that the evaluation of stability by the objective function becomes as good as possible, and determining a control command for the control target based on the obtained control rule; Control execution means for performing control on the control target based on the control command; A control device comprising:

2. Actual state data acquisition means for acquiring training data indicating the transition of the state of the control target under the control of the control target, with the actual state data indicating the actual state of the control target in the actual environment; Latent state data acquisition means for converting training data indicating the transition of the state of the control target under the control of the control target, with the actual state data indicating the transition of the state of the control target under the control of the control target, into training data indicating the transition of the state of the control target under the control of the control target, with the latent state data indicating the latent state in a virtual environment; Transition model acquisition means for learning a latent state transition model, which is a model for calculating the transition of the latent state under the control of the control target, using training data indicating the transition of the state of the control target under the control of the control target, with the latent state data; Comprising: The function acquisition means searches for the function using the latent state data output by the latent state transition model; The action determination means searches for the control rule using the latent state data output by the latent state transition model; The control device according to claim 1.

3. The latent state data acquisition means performs learning of a diffeomorphism, and uses the obtained diffeomorphism to convert training data indicating the transition of the state of the control target under the control of the control target, with the actual state data, into training data indicating the transition of the state of the control target under the control of the control target, with the latent state data; The control device according to claim 2.

4. The latent state transition model includes a vector field indicating the time derivative of the latent state and a numerical integration of the time derivative of the latent state indicated by the vector field; The transition model acquisition means performs learning of the vector field; The control device according to claim 2.

5. The action decision means further performs search for the control rule using the potential state data converted by the potential state data acquisition means and the actual state data obtained under the control of the control target in the actual environment. The control device according to claim 2.

6. Function acquisition means for performing search for a function included in the objective function so that the evaluation of stability by the objective function becomes as good as possible using an objective function indicating the stability condition by the Lyapunov function; Action decision means for performing search for a control rule for the control target so that the evaluation of stability by the objective function becomes as good as possible, and determining a control command for the control target based on the obtained control rule; A learning device comprising:

7. A computer, performs search for a function included in the objective function so that the evaluation of stability by the objective function becomes as good as possible using an objective function indicating the stability condition by the Lyapunov function, performs search for a control rule for the control target so that the evaluation of stability by the objective function becomes as good as possible, determines a control command for the control target based on the obtained control rule, and performs control for the control target based on the control command. A control method including this.

8. A computer, performs search for a function included in the objective function so that the evaluation of stability by the objective function becomes as good as possible using an objective function indicating the stability condition by the Lyapunov function, performs search for a control rule for the control target so that the evaluation of stability by the objective function becomes as good as possible, and determines a control command for the control target based on the obtained control rule. A learning method including this.

9. On a computer, performing search for a function included in the objective function so that the evaluation of stability by the objective function becomes as good as possible using an objective function indicating the stability condition by the Lyapunov function, performing search for a control rule for the control target so that the evaluation of stability by the objective function becomes as good as possible, and determining a control command for the control target based on the obtained control rule, and performing control for the control target based on the control command. A program for causing the above to be executed.

10. On a computer, performing search for a function included in the objective function so that the evaluation of stability by the objective function becomes as good as possible using an objective function indicating the stability condition by the Lyapunov function, Search for control rules for the controlled object so that the evaluation of stability by the objective function becomes as good as possible, and determine a control command for the controlled object based on the obtained control rules; A program for causing the above to be executed.