Controller, control method, and computer program
The control device employs the proximal operator of a closed proper convex function to convert generalized equations into equivalent equations, allowing real-time calculation of control inputs with specific properties like sparsity, addressing the limitations of existing algorithms in handling non-smooth evaluation functions.
Patent Information
- Application Number
- JP2023206081
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-18
AI Technical Summary
Existing real-time algorithms for model predictive control, such as the continuous deformation method, are unable to handle non-smooth evaluation functions that result in generalized equations, making it difficult to obtain optimal control inputs with specific properties like sparsity.
A control device that uses the proximal operator of a closed proper convex function to convert generalized equations into equivalent equations, allowing the application of real-time algorithms like the continuous deformation method or Newton's method to efficiently determine control inputs with specific properties.
Enables real-time calculation of control inputs with specific properties, such as sparsity, that cannot be achieved by substituting non-differentiable norms with differentiable functions, improving the stability and efficiency of sparse feedback control.
Smart Images

Figure 2025091084000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a control device, a control method, and a computer program.
Background Art
[0002] An optimal control problem of optimizing an input to a control target using feedback control is known (see, for example, Non-Patent Documents 1 to 5). In Non-Patent Document 1, when realizing feedback control by a model predictive control method, a real-time algorithm for tracking the time change of an optimal control input by a continuous deformation method is disclosed. Non-Patent Document 2 discloses a method for solving an optimal control problem with improved sparsity using differential dynamic programming. In this method, in order to cope with the problem that the L1 norm generally introduced to improve sparsity is non-differentiable, the input to the control target is optimized by substituting the L1 norm with a differentiable function. Non-Patent Document 3 discloses a method for performing model predictive control based on optimal control that minimizes the L1 norm for a linear time-invariant plant system to realize sparse feedback control. In this method, a plurality of previously obtained optimal control input candidates are stored, and when controlling the control target, the stored optimal control is searched and used as an input.
[0003] Non-Patent Document 4 describes a control device that acquires the attitude, position, and speed of a host vehicle as sensor information using LIDAR (Light Detection And Ranging), self-position estimation technology, GPS (Global Positioning System) information, etc., and determines an accelerator amount and a steering amount. In Non-Patent Document 4, sparsity of the steering amount is imparted by adding an L1 norm regularization term to the evaluation function with respect to the steering amount. Non-Patent Document 5 describes a control device for a spacecraft that estimates the position and speed from sensor information and determines the thruster injection timing. In Non-Patent Document 5, in the startup design of the spacecraft, a control problem of minimizing the L1 norm of the thruster input is solved for the purpose of making the thruster injection pattern sparse.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the field of optimization, a method is known that gives specific properties to an optimal solution by using a non-smooth evaluation function. A typical example is a method of imparting sparsity (the property that many elements are zero) to an optimal solution by using an evaluation function that additively includes the L1 norm, and it also has applications in the field of control. Model predictive control is a feedback control method that determines a control input based on an optimal control problem, and since it is necessary to obtain optimal control with a short control period, the real-time performance of the algorithm becomes an issue. As a real-time algorithm for model predictive control, there is the continuous deformation method, but in this method, the first-order optimality condition of the optimal control problem is used as the equation for determining the optimal control. On the other hand, when it is desired to impart specific properties to the optimal control by using a non-smooth evaluation function, the first-order optimality condition of the optimal control problem becomes a generalized equation rather than an equation, so there has been a problem that real-time algorithms such as the continuous deformation method cannot be used. When the evaluation function is non-smooth and the first-order optimality condition becomes a generalized equation, the input cannot be obtained by the model predictive control described in Non-Patent Document 1. Also, in the method described in Non-Patent Document 2, since the L1 norm is replaced with a differentiable function, there is a possibility that the obtained input may not become strictly sparse. In the method described in Non-Patent Document 3, a plurality of candidate optimal control inputs are stored, and it is necessary to search for the optimal control input for each state quantity at a specific time. The number of candidate optimal control inputs to be stored increases exponentially according to the dimension of the state quantity. Therefore, for a plant with a large dimension of state quantity, since the number of candidate optimal control inputs is very large, it is difficult to search for the optimal control input for the state quantity in real time. In Non-Patent Documents 4 and 5, when performing model predictive control, since the non-smoothness of the evaluation function appears, there may be a case where the L1 norm regularization term is non-differentiable, and there is a possibility that the optimal control input may not be performed.
[0006] The present invention has been made to solve at least a part of the above-described problems, and an object thereof is to provide a method for calculating a control input in real time when imparting specific properties to the control input by using a closed proper convex function.
Means for Solving the Problems
[0007] The present invention is made to solve at least a part of the above-described problems and can be realized in the following forms.
[0008] (1) According to one aspect of the present invention, a control device that calculates a control input for a control target is provided. This control device includes an acquisition unit that acquires an equation using the proximal operator of a closed proper convex function, an update unit that updates internal variables using the equation, and an input determination unit that determines the control input using the updated internal variables.
[0009] According to this configuration, when imparting specific properties to the control input using a closed proper convex function, the conditional expression for determining the input may be a generalized equation derived from the subdifferential of the closed proper convex function (for example, when imparting sparsity to the input, a generalized equation using the subdifferential of the L1 norm). Since real-time algorithms such as the continuous deformation method and the Newton method cannot be used for generalized equations that are not equations, there is a possibility that a control input with specific properties cannot be obtained. On the other hand, in this configuration, the generalized equation derived from the subdifferential of the closed proper convex function is converted into an equation expressed as an equation by using the proximal operator of the closed proper convex function. By using a real-time algorithm for the converted equation, a control input with specific properties is efficiently determined. It becomes possible to calculate in real time a control input with specific properties imparted using a closed proper convex function. As a result, for example, when imparting sparsity to the control input using the L1 norm as a closed proper convex function, a strictly sparse control input that cannot be achieved when the control input is determined by substituting the L1 norm with a differentiable function is calculated. That is, in this configuration, when imparting specific properties to the control input using a closed proper convex function, a method for calculating the control input in real time becomes applicable by transforming the generalized equation for determining the control input into an equation using the proximal operator.
[0010] (2) In the control device according to the above aspect, the acquisition unit acquires, as the equation to be acquired, an equation obtained by equivalently transforming a generalized equation included in the first-order optimality condition of the optimal control problem using a proximal operator, and the evaluation function of the optimal control problem may additively include the closed proper convex function. According to this configuration, the solution of the optimal control problem having a non-smooth evaluation function can be searched for using an equation obtained by transforming the first-order optimality condition. Therefore, based on the optimal control problem with a non-smooth evaluation function, the control input can be calculated in real time.
[0011] (3) In the control device according to the above aspect, the update unit may update the internal variable using Newton's method. According to this configuration, by updating the internal variable using Newton's method, which is a well-known real-time algorithm, for the transformed equation, the control input can be efficiently improved based on the optimal control problem.
[0012] (4) In the control device according to the above aspect, further, a state quantity acquisition unit for acquiring the state quantity of the controlled object is provided, and the acquisition unit may acquire the equation including the acquired state quantity as a parameter. According to this configuration, the state quantity of the controlled object detected by the state quantity acquisition unit can be used as an input, and the control input to the controlled object can be output as an output. According to this configuration, feedback control is performed according to the acquired state quantity.
[0013] (5) In the control device according to the above aspect, the acquisition unit acquires a plant model for predicting the dynamics of the controlled object, and the update unit may update the internal variable using the continuous deformation method. According to this configuration, by using the plant model, the change amount of the state quantity after a predetermined time elapses can be predicted based on the state quantity acquired at a specific time. According to this configuration, using the continuous deformation method, the change amount of the control input that satisfies the equation can be calculated from the predicted change amount of the state quantity, and thereby, the control input that satisfies the equation after a predetermined time elapses from a specific time can be calculated.
[0014] Note that the present invention can be implemented in various forms, for example, a control device, an automatic driving device, a control system, and devices and systems including these, a control method, a computer program for executing these systems and methods, a server device for distributing this computer program, a non-transitory storage medium storing the computer program, and the like.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Mode for Carrying Out the Invention
[0016] <Embodiment> 1. Configuration of the control device: FIG. 1 is a block diagram of a control system 100 including a control device 10 according to an embodiment of the present invention. The control system 100 includes a control target 50 that receives and drives a control signal as a control input (hereinafter, also simply referred to as "input"), and a control device 10 that transmits a control signal of the input to the control target 50. The control device 10 of the present embodiment determines an input to the control target 50 using a plurality of state quantities of the detected control target 50. The control device 10 performs sparse control of the control target 50 by adding the L1 norm as an evaluation function of the optimal control problem in order to impart sparsity to the optimal control input, and updating internal variables using an equation obtained by transforming the first-order optimality condition with a proximal operator.
[0017] As shown in FIG. 1, the controlled object 50 includes a driving unit 51 for driving the controlled object 50, a detection unit 52 for detecting a plurality of state quantities of the controlled object 50, and a communication unit 53 for transmitting and receiving information with the control device 10. Examples of the controlled object 50 include an automobile that performs autonomous driving. In this case, the driving unit 51 is a mechanism for changing the position and attitude of the automobile, and examples thereof include a gasoline engine or an electric motor for rotating the tires, an accelerator and a brake for changing the rotational speed of the tires, and a steering wheel for changing the direction of the tires.
[0018] Examples of the state quantities of the controlled object 50 detected by the detection unit 52 are parameters representing the position and attitude of the automobile when the controlled object 50 is an automobile. Examples of the detection unit 52 for detecting these parameters include LIDAR, radar, in-vehicle cameras, GPS, speed sensors, acceleration sensors, and the like. The communication unit 53 transmits the state quantities of the controlled object 50 detected by the detection unit 52 to the control device 10. Further, the communication unit 53 receives a control signal as an input to the driving unit 51 transmitted from the control device 10. Examples of the input to the controlled object 50 include the input power to the electric motor, the accelerator amount, and the steering wheel amount.
[0019] The control device 10 is a computer for controlling the controlled object 50. When the controlled object 50 is an automobile, the control device 10 may be a computer mounted in the automobile, or may be a server or the like that performs wireless communication with the controlled object 50. As shown in FIG. 1, the control device 10 includes a CPU (Central Processing Unit) 20, an input unit 31, an output unit 32, and a storage unit 33.
[0020] The input unit 31 is a user interface such as a keyboard, a microphone, and a mouse that accepts inputs from a user or the like. The output unit 32 outputs images and sounds according to control signals from the CPU 20. The output unit 32 includes a monitor capable of displaying images and a speaker capable of outputting sounds. The storage unit 33 is composed of a hard disk drive (HDD) or the like. The storage unit 33 stores a plant model that predicts the time evolution of the control target 50 controlled by the control device 10 and an equation used to update the input to the control target 50. The stored equation is an equation using the proximal operator of the L1 norm. The plant model of the present embodiment includes a prediction formula that predicts the state quantity after a predetermined time has elapsed from the current state quantity. Details of the plant model and the equation used for input update will be described later.
[0021] The CPU 20 is connected to each part of the control device 10 and a ROM (Read Only Memory) and a RAM (Random Access Memory) not shown in the figure, and controls each part of the control device 10 by expanding and executing the computer program stored in the ROM in the RAM. In addition, as shown in FIG. 1, the CPU 20 also functions as an acquisition unit 21, an update unit 22, an input determination unit 23, and a communication unit (state quantity acquisition unit) 24.
[0022] The acquisition unit 21 acquires the plant model stored in the storage unit 33 and the equation used for input update. The plant model and the equation are also used when improving the input by the continuous deformation method. The plant model and the equation will be described in the first embodiment below. The communication unit (state quantity acquisition unit) 53 acquires the state quantity detected by the detection unit 52 of the control target 50.
[0023] The acquisition unit 21 acquires, from the storage unit 33, a generalized equation representing the first-order optimality condition for an optimal control problem having an evaluation function including the L1 norm as a regularization term for performing sparse control on the control target 50, and an equation obtained by transforming the generalized equation using the proximal operator of the L1 norm.
[0024] Here, take a positive constant γ and a positive definite matrix P, and take a subset of the set of all real d-dimensional vectors as the admissible set C. Note that hereinafter, the fact that the admissible set C is a subset of the set of all real d-dimensional vectors is expressed as in the following formula (1). Also, the set of all real d-dimensional vectors is denoted as (RR d ), and the set of all real numbers is denoted as (RR). Furthermore, let the function r: (RR d ) → (RR) ∪ {+∞} be a closed proper convex function with the variable w (∈ (RR d )) as an argument. Then, the proximal operator of the closed proper convex function is expressed as in the following formula (2).
[0025]
Number
Number
[0026] From the above formula (1), the equation using the proximal operator of the closed proper convex function is expressed as in the following formula (3).
Number
[0027] The update unit 22 updates the internal variable z using the continuous deformation method based on the equation and the prediction formula acquired by the acquisition unit 21. The update unit 22 further updates the internal variable in the Newton direction using the Newton method.
[0028] The input determination unit 23 determines the control input from the internal variable updated by the update unit 22. The internal variable z includes candidates for the input to the control target 50, and the input is determined from the updated candidates for the input.
[0029] Here, the internal variables are z∈(RR d ), and the equation used to update the internal variable z is defined as the following formula (4) using the function W. The value of the function W that changes over time during the control period is ΔW:(RR d ) → (RR d ) where, there are two functions W, ΔW:(RR d ) → (RR d ) are d-dimensional vector-valued functions that take the internal variable z as an argument. C Taking the above, the update step of the internal variable z by the continuation method is expressed as the following equation (5).
[0030]
number
number
[0031] In the Newton method, as in the continuation method, the internal variables are z∈(RR d ), and the function W(z) used to update the internal variable z is defined as in the above formula (4). d ) → (RR d ) is a d-dimensional vector-valued function of the internal variable z. N Taking the above, the update step of the internal variable z by the Newton method is expressed as the following equation (6).
number
[0032] 2. First Example: The optimization problem to be solved in this embodiment is expressed as the following formula (7) or (8).
number
number
[0033] However, K a ⊂ (RR d ) is an admissible set, and the function f a : (RR d ) → (RR) is a scalar-valued function that takes the variable w as an argument. Also, c ∈ (RR s ). K b ⊂ (RR d ) is an admissible set, and the function f b : (RR d ) × (RR s ) → (RR) is a scalar-valued function that takes the variables w and s as arguments. Also, the functions f a , f b , the admissible sets K a , K b , and the constant γ appropriately depend on another p variables θ ∈ (RR p ).
[0034] The acquisition unit 21 of the control device 10 according to the present embodiment acquires a plant model represented by a generalized equation as in the following formula (9) from the storage unit 33. In the following formula (9), ∂r is a function that takes the variable w as an argument and returns a vector-valued set, and is the subdifferential of the closed proper convex function r.
Equation
[0035] Also, the acquisition unit 21 acquires from the storage unit 33 an equation obtained by transforming a generalized equation included in the first-order optimality condition of the optimal control problem represented by the following formula (10) using the proximal operator. Note that, among the following formula (10), the following formula (11) represented by (P) corresponds to the plant model. In the following formula (10), the time step width is Δt, and the number of steps representing time is defined as k. The time t k at each step is represented as in the following formula (12).
[0036]
Equation
Equation
Mathematics
[0037] The function Φ in the above formula (10): (RR d ) → (RR) is a differentiable scalar - valued function with the terminal state x T as an argument. The function L k : (RR n ) × (RR m ) → (RR) is a differentiable scalar - valued function with the state quantity x and the input u as arguments. The function f k of the plant model in the above formula (11): (RR n ) × (RR m ) → (RR n ) is a function that returns the state of the next step with the state quantity x and the input u as arguments. The plant model of this embodiment is a non - linear state - space model having a 5 - dimensional state variable (n = 5) and a 2 - dimensional input (m = 2). As a constraint condition, an input upper - lower limit constraint of - 1 or more and 1 or less is provided. The Lagrange multiplier μ of the constraint condition is expressed as in the following formula (13). Therefore, the primary optimality condition obtained by the acquisition unit 21 is expressed as in the following formulas (14) to (16) including the generalized equation.
[0038]
Mathematics
Mathematics
Mathematics
Mathematics
[0039] The function J(u 1:T-1 , x obs) is a function obtained by substituting the plant model (the above formula (11)) into a cost function other than the regularization term of the L1 norm, and differentiating it with respect to the input sequence u 1:T-1 obtained by differentiating with respect to, and is an m(T - 1)-dimensional vector-valued function. The function φ in the above formula (15) is a function that applies a complementary function to each element of the vector, and is used to transform the complementary condition derived from the inequality constraint into an equation.
[0040] The equation of the above formula (14) is in the form of a generalized equation. When P is the identity matrix, the proximity operator of the L1 norm becomes the soft threshold function Sγ of the threshold γ (positive constant γ). Therefore, according to formulas (7) and (8), the formula (14) represented by the generalized equation is equivalently transformed into the equation represented by the following formula (17).
Equation
[0041] Regarding the above equations (15) and (17), let the system of equations W(z, x 1:T-1 , μ) = 0 for determining the internal variables z := (u obs ). The update unit 22 uses the plant model described in the above formula (11) to obtain the current state quantity x obs detected by the detection unit 52, and the control input u in applied to the control target simultaneously with the detection, and predicts the state quantity x pred after a predetermined time as shown in the following formula (18).
Equation
[0042] Furthermore, the solution of the equation is tracked by the continuous deformation method. Specifically, the update unit 22 determines a function ΔW representing the time change of the function W as shown in the following formula (19).
Equation
[0043] The update unit 22 is a positive constant ζ CTake it and update the internal variable z according to the continuous deformation method. Specifically, update the internal variable z according to the update formula of the following formula (20).
Number
[0044] The update unit 22 further takes a positive constant ζ N and updates the internal variable z according to the Newton method. Specifically, update once in the Newton direction as shown in the following formula (21).
Number
[0045] Figure 2 is an explanatory diagram of the time-series changes of two inputs u1 and u2 in this embodiment. As shown in Figure 2, for the input u1 (solid line) of this embodiment, sparse control is performed such that the input value becomes zero after t = 9. Also, for the input u2 (dashed-dotted line), sparse control is performed such that the input value becomes zero after t = 13.
[0046] Figure 3 is an explanatory diagram of the time-series changes of two inputs u 1x , u 2x in the comparative example. In the comparative example, the L1 norm regularization term described in Non-Patent Document 1 and Non-Patent Document 2 is replaced with a differentiable function, and feedback control by model predictive control is performed by the continuous deformation method. That is, by approximating the non-differentiable L1 norm with a differentiable function, the introduction of the subdifferential is avoided, and the first-order optimality condition is made into an equation to enable the application of the continuous deformation method. For the input u 1x (solid line) in the comparative example, as shown in Figure 3, although it approaches zero after t = 10, it does not completely become zero until t = 20. For the input u 2x (dashed-dotted line), although it approaches zero after t = 13, similar to the input u 1x , it does not completely become zero until t = 20. Furthermore, the input u 1x is between t = 3 and t = 6, and the input u 2xIt fluctuates violently when t is between 2 and 4. That is, it can be seen that the calculation of the input is numerically unstable. This destabilization of the input calculation, unlike in this embodiment, is due to the fact that the non-differentiable L1 norm is approximated by a differentiable function, resulting in an increase in the condition number of the coefficient matrix of the system of linear equations obtained by the continuous deformation method.
[0047] Figure 4 is a flowchart of the control method for the control object 50 of this embodiment. In the control flow shown in Figure 4, first, an acquisition step is performed in which the acquisition unit 21 acquires an equation using the proximal operator of the closed proper convex function from the storage unit 33 (step S1). The acquisition unit 21 of this embodiment acquires a plant model in addition to the equation. The acquisition unit 21 further acquires a positive constant γ (threshold γ (the above formula (17))), ζ C (the above formula (20)), and ζ N (the above formula (21)). The equation acquired by the acquisition unit 21 from the storage unit 33 is a transformed equation using the proximal operator of the L1 norm for a generalized equation representing the first-order optimality condition for an optimal control problem having an evaluation function including the L1 norm as a regularization term in order to perform sparse control on the control object 50.
[0048] The input sequence u 1:T-1 and the Lagrange multiplier μ in the transformed equivalent equation (the above formula (17)) are initialized (step S2). In this embodiment, the initial values of the input sequence u 1:T-1 and the Lagrange multiplier μ are stored in the storage unit 33. Therefore, by the acquisition unit 21 acquiring the initial values, the initialization of the input sequence u 1:T-1 and the Lagrange multiplier μ is executed. In this embodiment, as the initial values set for initialization, the optimal control obtained by methods such as the semi-smooth Newton method and the ADMM algorithm is used.
[0049] The input determination unit 23 determines the first input u1 of the input sequence u 1:T-1 as the input to be applied to the control object 50 (step S3). The communication unit 24 transmits the input determined by the input determination unit 23 to the drive unit 51, and the detection unit 52 detects the current state quantity x obsMeasure, and the communication unit 24 receives the state quantity measured by the detection unit 52 (step S4). The update unit 22 determines the function W using the measured state quantity x obs (step S5). In this embodiment, the update unit 22 uses the state quantity x obs and the state equation to determine the function ΔW (the above formula (19)) representing the time change of the function W. In addition to the method of substituting the plant model into the evaluation function, in other embodiments, adjoint variables may be introduced to obtain an equation. Note that the process of step S3 corresponds to the input determination step.
[0050] The update unit 22 updates the input sequence u 1:T-1 and the Lagrange multiplier μ by the continuous deformation method (step S6). The update unit 22 uses the above formula (20) and the positive constant ζ C obtained in the acquisition step to update the input sequence u 1:T-1 and the Lagrange multiplier μ. In this process, since it is necessary to calculate the product of the inverse matrix and the vector, it is necessary to solve a system of linear equations. As an efficient solution method, the FDGMRES method or the like can be used.
[0051] Next, the update unit 22 updates the input sequence u 1:T-1 and the Lagrange multiplier μ in the Newton direction (step S7). In this embodiment, the update unit 22 descends in the Newton direction the number of times that can meet the control period. The update unit 22 determines whether to end the update process (step S8). If the current time is before the control end time stored in the storage unit 33 (step S8: NO), the processes after step S3 are repeated. If the current time is after the control end time stored in the storage unit 33 (step S8: YES), the control flow of the control target 50 ends. Note that the processes of step S6 and step S7 correspond to the update step.
[0052] As described above, the acquisition unit 21 of the present embodiment acquires an equation using the proximal operator of a closed proper convex function. The update unit 22 updates the internal variable z included in the acquired equation. The input determination unit 23 determines the input using the updated internal variable. In the present embodiment, when imparting specific properties to the input using a closed proper convex function, the conditional expression for determining the input may be a generalized equation derived from the subdifferential of the closed proper convex function (for example, when it is desired to impart sparsity to the input, a generalized equation using the subdifferential of the L1 norm). Since real-time algorithms such as the continuous deformation method and the Newton method cannot be used for the generalized equation, there is a risk that an appropriate input cannot be obtained. On the other hand, in the present embodiment, the generalized equation derived from the subdifferential of the closed proper convex function is converted into an equation of an equality such as the above formula (17) by using the proximal operator of the closed proper convex function. By applying a real-time algorithm such as the continuous deformation method or the Newton method to the converted equation, an appropriate input for the control target 50 is efficiently determined. As a result, for example, strict sparse feedback control that cannot be achieved when substituting a differentiable function for the L1 norm to avoid the introduction of the subdifferential as in the comparative example is performed on the control target 50. That is, by using the control device 10 of the present embodiment, the input can be efficiently improved based on an optimal control problem with a non-smooth evaluation function, and an input with specific properties can be calculated in real time.
[0053] Further, the equation acquired by the acquisition unit 21 from the storage unit 33 is an equation obtained by transforming, using the proximal operator, a generalized equation included in the first-order optimality condition of an optimal control problem having an evaluation function that additively includes a closed proper convex function into an equality. In the present embodiment, an equation obtained by converting a generalized equation that was conventionally included in the first-order optimality condition of an optimal control problem having a non-smooth evaluation function into an equality is used. Therefore, the input can be efficiently improved based on an optimal control problem with a non-smooth evaluation function, and an input with specific properties can be calculated in real time.
[0054] Also, the update unit 22 of the present embodiment updates the internal variable z in the Newton direction using the Newton method. In the present embodiment, by updating the internal variable using the Newton method, which is a well-known real-time algorithm, for the equivalent equation after transformation, the input can be efficiently improved based on the optimal control problem.
[0055] Also, the communication unit 53 of the present embodiment acquires the state quantity x detected by the detection unit 52 of the control target 50. obs In the present embodiment, since the optimal control problem shown in the above formula (10) is determined according to the acquired state quantity x, obs the control device 10 can output the input to the control target 50 as an output with the state quantity x of the control target 50 detected by the detection unit 52 as an input. Therefore, feedback control is performed according to the detected state quantity x. obs obs
[0056] Also, the acquisition unit 21 of the present embodiment acquires the plant model represented by the above formula (11). The update unit 22 updates the internal variable z using the continuous deformation method. In the present embodiment, by using the plant model, the state quantity after a lapse of a predetermined time from the current state is predicted based on the current state quantity x. obs obs Therefore, since feedback control to the control target 50 is performed using the predicted state quantity x, more appropriate control is performed.
[0057] <Second Embodiment> In the second embodiment, instead of the above formula (3), the input to the control target 50 may be optimized using the equation of the following formula (22).
Equation
[0058] Here, z ∈ (RR d ) is an internal variable. The function f: (RR d ) → (RR) is a scalar value function that takes the variable w as an argument. The function D: (RR d ) → (RRd × d ) is a matrix-valued function that takes the variable w as an argument. The equation represented by the above formula (22) is equivalent to the generalized equation represented by the following formula (23).
Number
[0059] On the other hand, the equation represented by the above formula (3) is equivalent to the following formula (24).
Number
[0060] When the function D has singular values smaller than 1, the solution of the generalized equation in the form of the above formula (23) is close to the solution of the following formula (25) from which the subdifferential of the closed proper convex function is removed. Also, it can be expected that this solution still endows the function D1 with specific properties (for example, sparsity when the function r is the L1 norm) due to the non-smoothness of the function r. However, in the above formula (23), since the function D depends on the variable w, the first term D(w)∇f(w) on the left side is not represented by the gradient of some function f a as in the above formula (24). Therefore, the above formula (23) is not a generalized equation obtained as the first-order optimality condition of some optimal control problem.
Number
[0061] <Third Embodiment> When the control device 10 of the above embodiment takes the autonomous vehicle described in Non-Patent Document 4 as the control target 50, it can control the input to the autonomous vehicle. In this case, the state quantity x of the autonomous vehicle obsAs such, three pieces of information are acquired: the position of the autonomous vehicle obtained from LIDAR and GPS information, the rotation speed of the tires, and the amount of steering. Also, as controlled inputs to the autonomous vehicle, two pieces of information are controlled: the amount of accelerator and the amount of steering. Therefore, the plant model for predictive control in the third embodiment is a non-linear state space model having three-dimensional state variables (n = 3) and two-dimensional inputs (m = 2). Constraints such as the above formula (16) are provided for the constraint conditions of the optimal control problem in the third embodiment. The acquisition unit 21 acquires, using the proximity operator, an equation obtained by transforming a generalized equation, which is a first-order optimality condition of an optimal control problem having an evaluation function that additively includes the L1 norm, in order to perform sparse control on the autonomous vehicle. Also, the acquisition unit 21 acquires a plant model representing the change in three-dimensional state variables with respect to two-dimensional inputs.
[0062] The update unit 22 optimizes the two inputs by updating internal variables using the equation and the plant model acquired by the acquisition unit 21. The input determination unit 23 determines the control input from the internal variables updated by the update unit 22. By applying the determined input to the autonomous vehicle, which is the control target 50, sparse control is performed.
[0063] <Fourth Embodiment> When the control device 10 of the above embodiment uses the thruster of the spacecraft described in Non-Patent Document 5 as the control target 50, it can control the input to the thruster jet. In this case, as the state quantity x of the spacecraft obs four pieces of information are acquired: the position and velocity of the spacecraft, and the temperature and pressure of the gas emitted from the thruster. Also, as controlled inputs to the thruster, two pieces of information are controlled: the flow rate of the gas ejected by the thruster and the rotation amount (direction) of the thruster. Therefore, the plant model for predictive control in the fourth embodiment is a non-linear state space model having four-dimensional state variables (n = 4) and two-dimensional inputs (m = 2). Constraints such as the above formula (16) are provided for the constraint conditions of the optimal control problem in the fourth embodiment. Also, the acquisition unit 21 acquires a prediction formula representing the change in four-dimensional state variables with respect to two-dimensional inputs.
[0064] The acquisition unit 21 obtains an equation obtained by transforming, using the proximity operator, a generalized equation obtained as a first-order optimality condition of an optimal control problem having an evaluation function that additively includes the L1 norm, in order to perform sparse control on the thruster. The update unit 22 updates the internal variables using the transformed equivalent equation and the plant model, thereby optimizing the two inputs. The input determination unit 23 determines the control input from the internal variables updated by the update unit 22. By applying the optimized input to the thruster of the spacecraft that is the control target 50, sparse control is performed.
[0065] <Modification Example of Embodiment> The present invention is not limited to the above-described embodiment, and can be implemented in various modes without departing from the gist thereof. For example, the following modifications are possible. Also, in the above embodiment, a part of the configuration assumed to be realized by hardware may be replaced with software, or conversely, a part of the configuration assumed to be realized by software may be replaced with hardware.
[0066] <Modification Example 1> The control device 10 of the above embodiment is an example, and a generalized equation using a closed proper convex function can be transformed using the proximity operator of the closed proper convex function, and the internal variables can be updated and the input can be determined based on the transformed equation. It is deformable within the range. For example, in order to optimize the input, the current state quantity x obs from, a prediction formula for obtaining the state quantity x obs after a predetermined time may not be used. As the closed proper convex function, other functions such as the L2 norm, the absolute value sum function, and the indicator function of a non-empty closed convex set may be adopted according to the properties to be imparted to the control input, in addition to the L1 norm. As the real-time algorithm, only one of the continuous deformation method and the Newton method may be used instead of both. Also, instead of the continuous deformation method and the Newton method, the gradient method represented by the following formula (26) may be used to update the internal variable z. Note that ζ g in the following formula (26) is a positive constant.
Equation
[0067] In the above embodiment, the control device 10 has acquired the state quantity x detected by the detection unit 52 of the control target 50 via the communication unit 24 obs ; however, the detection unit 52 may not be provided. Further, the control device 10 may acquire the state quantity x obs from the storage unit 33 that stores it as data. The storage unit 33 is not a configuration included in the control device 10 but an independent device such as a server, and the control device 10 may acquire information such as equations and plant models from the server.
[0068] The input sequence u at step S2 in the control flow shown in FIG. 4 1:T-1 and the initialization process of the Lagrange multiplier μ may be replaced with what is called a hot start. In this modification, a short prediction horizon is set in the process of step S2, and the prediction horizon may be gradually increased after the start of control of the control target 50. As a result, optimal control is quickly required after measuring the initial value of the state.
[0069] The sparsity of the Jacobian matrix of the equation may be utilized. In the continuous deformation method and the Newton method, it is necessary to solve a system of linear equations with the Jacobian matrix of the function W that defines the equation W(z, x obs ) = 0 as the coefficient matrix. When the non-smooth regularization function r(u 1:T-1 ) with respect to the input sequence u 1:T-1 is represented by the sum of the regularization functions for the input u k at each time step (the following formula (27)), similar to the non-smooth optimal control problem, by introducing the adjoint equation, a structure as a sparse matrix appears in the Jacobian matrix. By actively using this structure, the solution of the system of linear equations is efficiently performed. Note that the adjoint equation is an equation for determining the adjoint variable.
[0070]
Equation
[0071] Although the present aspect has been described based on the embodiments and modification examples above, the embodiments of the above-described aspects are for facilitating the understanding of the present aspect and do not limit the present aspect. The present aspect can be changed and improved without departing from its spirit and the scope of the claims, and equivalents thereof are included in the present aspect. Further, if its technical features are not described as essential in this specification, they can be deleted as appropriate.
[0072] The present invention can also be realized in the following forms. [Application Example 1] A control device that calculates a control input for a control target, an acquisition unit that acquires an equation using the proximal operator of a closed proper convex function, an update unit that updates internal variables using the equation, an input determination unit that determines the control input using the updated internal variables, A control device comprising: [Application Example 2] The control device according to Application Example 1, wherein the acquisition unit acquires, as the equation to be acquired, an equation obtained by equivalently transforming a generalized equation included in the first-order optimality condition of an optimal control problem using a proximal operator, and the evaluation function of the optimal control problem additively includes the closed proper convex function. [Application Example 3] The control device according to Application Example 1 or Application Example 2, wherein the update unit updates the internal variables using Newton's method. [Application Example 4] The control device according to any one of Application Examples 1 to 3, further comprising a state quantity acquisition unit that acquires the state quantity of the control target, wherein the acquisition unit acquires the equation including the acquired state quantity as a parameter. [Application Example 5] The control device according to any one of Application Examples 1 to 4, The acquisition unit acquires a plant model that predicts the dynamics of the control target, The update unit is a control device that updates the internal variable using the continuous deformation method. [Application Example 6] A control method for calculating a control input to a control target, wherein a computer an acquisition step of obtaining an equation using the proximity operator of a closed proper convex function; an update step of updating an internal variable using the equation; an input determination step of determining the control input using the updated internal variable; A control method that realizes the above. [Application Example 7] A computer program for calculating a control input to a control target, wherein an acquisition function of obtaining an equation using the proximity operator of a closed proper convex function; an update function of updating the internal variable using the equation; an input determination function of determining the control input using the updated internal variable; A computer program that causes a computer to realize the above.
Explanation of Signs
[0073] 10…Control device 20…CPU 21…Acquisition unit 22…Update unit 23…Input determination unit 24…Communication unit (state quantity acquisition unit) 31…Input unit 32…Output unit 33…Storage unit 50…Control target 51…Drive unit 52…Detection unit 53…Communication unit 100…Control system
Claims
1. A control device that calculates a control input for a control target, an acquisition unit that acquires an equation using the proximity operator of a closed proper convex function, an update unit that updates internal variables using the equation, an input determination unit that determines the control input using the updated internal variables, and includes a control device.
2. The control device according to claim 1, wherein the acquisition unit acquires, as the equation to be acquired, an equation obtained by equivalently transforming a generalized equation included in the first-order optimality condition of an optimal control problem using a proximity operator, and the evaluation function of the optimal control problem includes the closed proper convex function additively, a control device.
3. The control device according to claim 1, wherein the update unit updates the internal variables using the Newton method, a control device.
4. The control device according to any one of claims 1 to 3, further comprising a state quantity acquisition unit that acquires a state quantity of the control target, and the acquisition unit acquires the equation including the acquired state quantity as a parameter, a control device.
5. The control device according to claim 4, wherein the acquisition unit acquires a plant model that predicts the dynamics of the control target, and the update unit updates the internal variables using the continuous deformation method, a control device.
6. A control method for calculating a control input for a control target, wherein a computer performs an acquisition step of acquiring an equation using the proximity operator of a closed proper convex function, an update step of updating internal variables using the equation, and an input determination step of determining the control input using the updated internal variables, A control method for realizing
7. A computer program for calculating a control input for a control target, An acquisition function for acquiring an equation using the proximity operator of a closed proper convex function, An update function for updating internal variables using the equation, An input determination function for determining the control input using the updated internal variables, A computer program for causing a computer to realize
Citation Information
Patent Citations
Smoothed and regularized fischer-burmeister solver for embedded real-time constrained optimal control problems in autonomous systems
JP2020040651A