Predictive control system, predictive control method and program

The predictive control system addresses high computational loads and risk management in data-driven control by using neural networks to solve optimal control problems with risk tolerance parameters, achieving efficient and safe path transitions.

JP7799895B1Active Publication Date: 2026-01-15TOKYO GAS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025153170
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-01-15
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing data-driven predictive control methods lack the ability to manage risk tolerance of transition paths and have high computational loads, limiting their online implementation.

Method used

A predictive control system that incorporates a risk tolerance parameter into a machine learning control rule, using neural networks to solve optimal control problems, allowing for safe and efficient path management with reduced calculation load.

Benefits of technology

Enables predictive control with managed risk tolerance and reduced computational requirements, ensuring safe and efficient path transitions with minimal calculation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799895000001_ABST
    Figure 0007799895000001_ABST
Patent Text Reader

Abstract

This realizes predictive control that has a low computational load and can manage the risk tolerance of transition paths. [Solution] A predictive control system having a processor, which provides an input pair consisting of a state obtained from a target system and a risk tolerance parameter to a machine learning control law incorporating the risk tolerance parameter, and provides the output of the machine learning control law to the target system as a control input.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a predictive control system, a predictive control method, and a program. [Background technology]

[0002] Model predictive control is a model-driven control that is applied to processes that require high control accuracy. Model predictive control can achieve high-precision control. However, the computational load required to solve complex optimal control problems is a hindrance to model predictive control, limiting its online implementation. Therefore, data-driven control, which uses machine learning to approximate model predictive control laws, has been proposed as a way to implement model predictive control online. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-195307 Summary of the Invention [Problem to be solved by the invention]

[0004] However, further improvement is required for the data-driven control described in Patent Document 1. For example, in the case of Patent Document 1, the user cannot manage the risk tolerance of the path leading to the goal state.

[0005] An object of the present invention is to realize predictive control that has a small calculation load and is capable of managing the risk tolerance of transition paths. [Means for solving the problem]

[0006] The invention described in claim 1 is a predictive control system having a processor that provides an input pair, which is a state obtained from a target system and a risk tolerance parameter, to a machine learning control rule that incorporates the risk tolerance parameter, and provides the output of the machine learning control rule to the target system as a control input. The invention described in claim 2 is the predictive control system described in claim 1, wherein the machine learning control law includes a trained model obtained by machine learning a model predictive control law obtained for an optimal control problem incorporating a risk tolerance parameter. The invention described in claim 3 is a predictive control system described in claim 2, in which the trained model is obtained by machine learning the results of solving an optimal control problem that incorporates a control barrier function that includes a risk tolerance parameter. The invention described in claim 4 is the predictive control system described in claim 3, in which a neural network is used as the machine learning algorithm of the machine learning control law. The invention described in claim 5 is the predictive control system described in claim 3, in which the machine learning algorithm of the machine learning control law uses neural network-related technology including a recurrent neural network. The invention described in claim 6 is the predictive control system described in claim 1, further comprising a first model predictive control law that solves a first optimal control problem incorporating the machine learning control law into a stage cost function, wherein the first optimal control problem achieves safe asymptotic approach to a target state. The invention described in claim 7 is the predictive control system described in claim 6, further comprising a second model predictive control law that solves a second optimal control problem including the machine learning control law, and the processor switches the input destination of the input pair to the first model predictive control law or the second model predictive control law depending on the value of the input pair. An invention described in claim 8 is the predictive control system described in claim 7, wherein, when the input pair is provided to the first model predictive control law or the second model predictive control law, the processor solves the first optimal control problem or the second optimal control problem by dividing the prediction period into prediction periods that can be implemented online. The invention described in claim 9 is the predictive control system described in claim 7, in which the processor controls the switching of the input destination using a trained model learned by reinforcement learning. The invention described in claim 10 is a predictive control system described in claim 1, in which the processor changes the value of a risk tolerance parameter to be assigned to the machine learning control law in accordance with constraints on the path from the initial state to the transition of the state to the target state. The invention described in claim 11 is a predictive control system described in claim 10, wherein the processor changes the value of the risk tolerance parameter to a first value when prioritizing the safety of the path to the target state, and changes the value of the risk tolerance parameter to a second value greater than the first value when prioritizing the distance or time of the path to the target state. The invention of claim 12 is the predictive control system of claim 10 or 11, wherein the processor changes the value of the risk tolerance parameter while controlling the target system. The invention described in claim 13 is a predictive control method in which a processor executes a process of providing an input pair, which is given by a state obtained from a target system and a risk tolerance parameter, to a machine learning control rule incorporating the risk tolerance parameter, and a process of providing the output of the machine learning control rule to the target system as a control input. The invention described in claim 14 is a program for causing a computer to realize the following functions: providing an input pair, given as a state obtained from a target system and a risk tolerance parameter, to a machine learning control rule incorporating the risk tolerance parameter; and providing the output of the machine learning control rule to the target system as a control input. [Effects of the Invention]

[0007] According to the present invention, predictive control can be realized that has a small calculation load and is capable of managing the risk tolerance of transition paths. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating a schematic configuration of an information processing system assumed in a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of a prediction control server. [Figure 3]FIG. 2 is a diagram illustrating a conceptual configuration of predictive control by a predictive control server used in the first embodiment. [Figure 4] FIG. 10 is a diagram illustrating a simulation result when an optimal control problem that does not include a risk tolerance parameter is used. [Figure 5] This is a diagram explaining the transition path and the location of the constraint violation when one initial state [-5,-5,0,0] is given to a machine learning control rule that does not include a risk tolerance parameter. [Figure 6A] FIG. 10 is a diagram illustrating differences in transition paths depending on the value of a risk tolerance parameter. [Figure 6B] FIG. 10 is a diagram illustrating state transitions and locations of constraint violations according to a machine learning control rule including a risk tolerance parameter γ. [Figure 7] FIG. 10 is a diagram illustrating a conceptual configuration of predictive control by a predictive control server used in the second embodiment. [Figure 8] FIG. 10 illustrates simulation results for achieving safe asymptotic approach to a target state including risk tolerance parameters. [Figure 9A] This is a diagram illustrating the transition path and the location of the constraint violation when one initial state [-5,-5,0,0] is given to an optimal control problem that achieves safe asymptotic approach to the target state. [Figure 9B] This is a diagram illustrating the transition path and the location of the constraint violation when one initial state [-5,-5,0,0] is given to an optimal control problem that achieves safe asymptotic approach to the target state. [Figure 10] FIG. 11 is a diagram illustrating a conceptual configuration of predictive control by a predictive control server used in the third embodiment. [Figure 11] FIG. 10 illustrates simulation results that achieve safe asymptosis and convergence to the target state. [Figure 12] 10A and 10B are diagrams illustrating state transitions and switching control when achieving convergence to a target state through a highly secure transition path. [Figure 13] 10A and 10B are diagrams illustrating state transitions and switching control when convergence to a target state is achieved through a transition path that achieves safe asymptotic approach. [Figure 14]10 is a diagram illustrating the number of simulation steps required to converge to a target state. [Figure 15] 1 is a diagram illustrating the maximum calculation time required to solve a single optimal control problem. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. <Embodiment 1: Risk Tolerance> <System configuration> FIG. 1 is a diagram illustrating a schematic configuration of an information processing system 1 assumed in the first embodiment. The information processing system 1 shown in FIG. 1 includes a predictive control server 10 and a plurality of target systems 20 connected via a network N.

[0010] 1, the target system 20 of client A is denoted as target system 20A, and the target system 20 of client B is denoted as target system 20B. The target system 20 includes, for example, energy-related facilities and logistics systems.

[0011] Energy-related facilities are devices and systems for generating, converting, storing, or utilizing energy, including, for example, power generation facilities, power substations, power transmission lines, gas pipelines, renewable energy facilities, and energy-saving equipment.

[0012] A logistics system is a system that manages the movement of goods. The objects of management of a logistics system include logistics facilities (such as ships, trucks, and aircraft), storage facilities, packaging facilities, and distribution processing facilities.

[0013] The predictive control server 10 is a server used by a business providing a predictive control service (hereinafter referred to as a "service business") to provide the service. The predictive control server 10 may be an on-premise server or a cloud server. 1, there is one predictive control server 10, but there may be multiple predictive control servers 10. The multiple predictive control servers 10 may be operated by different businesses or the same business. The predictive control server 10 is an example of a reservation control system.

[0014] The target system 20 is a system operated by a client who uses the predictive control service. While Fig. 1 shows an example in which one client operates one target system 20, one client may operate multiple target systems 20. Furthermore, the target systems 20 that use the predictive control service may all be the same type of system or different types of systems.

[0015] <Configuration of predictive control server> FIG. 2 is a diagram illustrating an example of the configuration of the predictive control server 10. As shown in FIG. The predictive control server 10 has a processor 11 that controls the operation of the entire server, a semiconductor memory 12, an auxiliary storage device 13 that records various data, and a communication interface 14 used for communication with the target system 20, etc. The processor 11 and other processing units are interconnected via a bus or other signal lines 15.

[0016] The processor 11 is a device that realizes various functions by executing programs. Here, the program is a general term for the OS (=Operating System) and application programs. For example, the processor 11 controls the operation of the target system 20 online (i.e., in real time) through the execution of the programs.

[0017] The processor 11 in this embodiment is assumed to be one of the programs for realizing the function of providing an input pair s given by a state x obtained from the target system 20 and a risk tolerance parameter γ to a control law that has been machine-learned by incorporating the risk tolerance parameter γ, and the function of providing the output of this control law to the target system as a control input.

[0018] The semiconductor memory 12 includes a ROM (=Read Only Memory) in which, for example, a UEFI (=Unified Extensible Firmware Interface) or the like is stored, and a RAM (=Random Access Memory) used as a work area for the processor 11. Here, the processor 11 and the semiconductor memory 12 operate as a computer. The auxiliary storage device 13 is configured by, for example, a hard disk device or semiconductor storage, and stores programs and other data.

[0019] <Conceptual structure of predictive control> Fig. 3 is a diagram illustrating a conceptual configuration of predictive control by the predictive control server 10 used in the first embodiment. The conceptual configuration shown in Fig. 3 is realized through the execution of a program. The execution of the conceptual configuration shown in Fig. 3 is an example of a predictive control method. The predictive control server 10 shown in FIG. 3 includes an initial state update unit 101, a generation unit 102 for generating input pairs of the initial state and the risk tolerance parameter, and a machine learning control rule 103.

[0020] The initial state update unit 101 has a function of giving the target system 20 an initial state x0 and a function of updating the latest state x of the target system 20. + The updated initial state x0 is provided to the input pair generation unit 102.

[0021] The input pair generation unit 102 receives the initial state x0 and a risk tolerance parameter γ to generate an input pair s=(x0, γ). The risk tolerance parameter γ is an example of information that defines the safety level of the path from the initial state x0 to the transition to the goal state. The value of the risk tolerance parameter γ can be changed according to a request. The request for change may be issued, for example, from the target system 20. Alternatively, the request may be issued by the service provider operating the predictive control server 10 based on an agreement with the client.

[0022] In this embodiment, the risk tolerance parameter γ is defined as a value between 0.0 and 1.0. The smaller the value, the lower the risk tolerance of the route, and a safe route to the goal state is prioritized. The larger the value, the higher the risk tolerance, and a shorter route distance to the goal state and a shorter time to reach the goal state are prioritized. The value of the risk tolerance parameter γ (for example, 0.2) used when prioritizing the safety of the path leading to the goal state is an example of the first value.

[0023] The value of the risk tolerance parameter γ (e.g., 0.6) used when prioritizing shortening the distance of the route to the target state or the time to reach the target state (i.e., when prioritizing proactive control) is an example of the second value. In this embodiment, the second value is greater than the first value. The input pair s is provided to the machine learning control law 103 .

[0024] The machine learning control law 103 in this embodiment solves an optimal control problem P cbf The solution of (s) is given by a control law (i.e., a learned model) that is machine-learned. The machine-learned control law 103 is a control input u corresponding to the input pair s = (x0, γ). in is output to the target system 20.

[0025] For example, a neural network or a neural network-related technology is used as the machine learning algorithm of the machine learning control law 103. The neural network-related technology includes, for example, a recurrent neural network, a Long Short-Term Memory (LSTM), and a Gated Recurrent Unit (GRU).

[0026] <Examples of predictive control> For example, in the case of a machine learning control law (different from the machine learning control law 103 in this embodiment) that approximates a model predictive control law for an optimal control problem P(x0) that does not include a risk tolerance parameter γ, the control input u inIn other words, this type of machine learning control rule cannot adjust the risk intensity.

[0027] On the other hand, in the case of this embodiment, a control barrier function including a risk tolerance parameter γ is added as a constraint to the optimal control problem P(x0). The control barrier function for state x(k) at time k is given by the following equation: h(x(k+1))-h(x(k))≧-γ(h(x(k))) where state x(k)∈R m and the control input u(k)∈R n and the risk tolerance parameter γ∈K ∞ is.

[0028] In this embodiment, as described above, the input pair of the initial state x0 and the risk tolerance parameter γ is set to s=(x0,γ). In this case, a new optimal control problem P cbf (s) is formulated as follows: TIFF0007799895000002.tif15147where x(k+1)=f(x(k),u(k)), x(0)=x0, x(k)∈X, x(N)∈X f , u(k)∈U, h(x(k+1))-h(x(k))≧-γ(h(x(k))).

[0029] Note that l(x(k),u(k)) is the stage cost function. f (x(N)) is the terminal penalty function. X is the set of constraints on the states. X f is a set of constraints on the final state. U is the control input u in is the set of constraints on N, and N is the forecast period.

[0030] The above optimal control problem P cbf The model predictive control law obtained by solving (s) is given by the following equation: κ cbf (s)=u* (0;s) In addition, u * (i;x0) is the control input at the i-th future step relative to the initial state x0. The machine learning control law 103 that approximates this model predictive control law solves the optimal control problem P for various input pairs s. cbf (s) is solved by the pair (s,u * This is a model trained using (0;s)). The machine learning control rule 103 in this embodiment is expressed by the following equation. κ cbf ML (s)=u

[0031] That is, the machine learning control rule 103 includes the input pair s including the risk tolerance parameter γ as one of its parameters. Therefore, by adjusting the risk tolerance parameter γ of the input pair s to an appropriate value, the machine learning control rule 103 can obtain a safe control input u in can be output. In addition, the safe control input u in This means that the path to the goal state is safe, or in other words, that the state transition path does not pass near the constraint.

[0032] For example, by reducing the value of the risk tolerance parameter γ, the safety of the path leading to the goal state is fundamentally increased. As a result, the possibility of a constraint violation occurring is reduced (i.e., the risk of a constraint violation is reduced), making the machine learning control rule 103 more practically useful than a machine learning control rule that does not include the risk tolerance parameter γ.

[0033] On the other hand, increasing the value of the risk tolerance parameter γ enables aggressive control of constraints. Aggressive control prioritizes the distance or time of the path to the goal state. Therefore, it is also acceptable for the state transition path to pass near the constraints.

[0034] <Verification by simulation> Below, the difference between the case where a control law that does not include the risk tolerance parameter γ is used and this embodiment will be verified by simulation. In order to verify the effectiveness of this embodiment, a simulation is performed in which a system that involves avoidance of obstacles (i.e., constraints) approaches a target state.

[0035] In this case, the present embodiment is given below. TIFF0007799895000003.tif7145where matrices A and B are given by: TIFF0007799895000004.tif30142

[0036] Also, the stage cost functions l(x(k),u(k)) and l ML (x(k),u(k)) and the terminal penalty function V f (x(N)) are given by the following equations: TIFF0007799895000005.tif13137

[0037] Also, the optimal control problems P(x0) and P cbf The weighting matrices Q, R, and P of (s) are given by the following equations, respectively. TIFF0007799895000006.tif58169

[0038] In these cases, the state x and the control input u in The constraints are given by the following equations: X={x(k)⊂R n :x min ≦x≦x max} U={u(k)⊂R m :u min ≦u≦u max} x min =[-6,‐6,‐1,‐1], x max =[1,1,1,1], u min =[‐10,‐10], u max =[10,10]

[0039] The constraint here (hereinafter referred to as *1) can be replaced by the following two control barrier functions *2 and *3, which have a risk tolerance parameter γ. h i (x(k+1))-h i (x(k))≧-h i (x(k)), γ∈K ∞ *2 h i (x(k))=(x1(k)-x1 obs,i ) 2 +(x2(k)-x2 obs,i ) 2 -(r obs ) *3 The problem with *1 as a constraint is the optimal control problem P(x0) that does not include the risk tolerance parameter γ. *2 and *3 are the optimal control problem P cbf (s) and the optimal control problem P cbf ML (s) is granted.

[0040] The constraints on the obstacles (i.e., constraints) that the agent must avoid are given by: ((x1(k)-x1 obs,i ) 2 +(x2(k)-x2 obs,i ) 2 >(r obs ) 2 *1 In this embodiment, for convenience of explanation, all obstacles are represented by seven circles with a radius of 0.5. The center coordinates of the seven circles that provide obstacles are given below. [x1 obs,i ,x2 obs,i ]=[-4.2,-3.9],[-3.4,-2.7],[-2.8,-3.8],[-2.1,-2.2],[-2.5,-1.0],[-0.9,-2.0],[-1.2,-0.8] By the way, i=1.2.…,7.

[0041] <Simulation results for the optimal control problem P(x0) that does not include the risk tolerance parameter γ> First, we will explain the simulation results of the optimal control problem P(x0) that does not include the risk tolerance parameter γ, for a system that involves avoidance of obstacles (i.e., constraints). Here, the goal state is x r = [0,0,0,0], the target input is u r =[0,0,0,0].

[0042] First, we solve the optimal control problem P(x0) offline for various initial states and obtain the pair (x0,u * (0,x0)). Next, the accumulated pair (x0,u * (0,x0)) and machine learning control law κ ML Learn (·). After this, the initial state is x0=[-5,-5,0,0], and the model predictive control law and machine learning control law κ ML Simulations were performed for each of the (·).

[0043] Figure 4 illustrates the simulation results when using an optimal control problem that does not include a risk tolerance parameter. (A) shows the state transitions using a model predictive control rule that does not include a risk tolerance parameter, and (B) shows the state transitions using a machine learning control rule that does not include a risk tolerance parameter. The user cannot manage the transition paths for either the model predictive control rule or the machine learning control rule. Figure 5 is a diagram explaining the transition path and the location of the constraint violation when a single initial state [-5, -5, 0, 0] is given to a machine learning control rule that does not include a risk tolerance parameter. (A) shows the state transition, and (B) shows the constraint violation. In Figure 5(B), there is a location where the constraint violation occurs.

[0044] <Simulation results of machine learning control law including risk tolerance parameter γ> Next, a simulation result of this embodiment (that is, using the machine learning control law 103 including the risk tolerance parameter γ) will be described. Again, the goal state is x r = [0,0,0,0], the target input is u r =[0,0,0,0].

[0045] Figure 6A is a diagram explaining the difference in transition paths depending on the value of the risk tolerance parameter γ. Figure 6B is a diagram explaining the state transitions and the locations of constraint violations according to the machine learning control rules including the risk tolerance parameter γ. (A) shows the state transitions, and (B) shows the constraint violations. The risk tolerance parameter γ in FIGS. 6A and 6B is 0.2. As shown in FIGS. 6A and 6B, in this embodiment, by specifying the value of the risk tolerance parameter γ, it is possible to manage the safety of the transition path in relation to the constraint conditions.

[0046] For example, by specifying a small value as the risk tolerance parameter γ, it is possible to prioritize the safety of the transition path, or by specifying a large value, it is possible to prioritize the distance or time of the path to the goal state.

[0047] <Summary> In the information processing system 1 according to the first embodiment, a risk tolerance parameter γ is adopted, and a pair s of the initial state x0 and the risk tolerance parameter γ is given to the machine learning control law 103 to generate a control input for the target system 20. This allows for predictive control that requires a small calculation load and is capable of managing the risk tolerance of the transition path.

[0048] <Embodiment 2> This embodiment also uses the information processing system 1 (see FIG. 1) assumed in the first embodiment. The difference between this embodiment and the first embodiment is the content of predictive control by the predictive control server 10 (see FIG. 1).

[0049] In the case of the first embodiment described above, specifying a large value as the risk tolerance parameter γ enables active control over the constraint conditions. However, actively moving close to the constraint conditions potentially involves the possibility of violating the constraint conditions. This is because the machine learning control rule 103 does not have a mechanism for actively avoiding the constraint conditions. Therefore, when the value of the risk tolerance parameter γ is large, there is a possibility that a constraint violation such as that shown in FIG. 4(B) may occur.

[0050] Therefore, in this embodiment, a mechanism is added to enhance safety even when active control is selected for the constraint conditions. Specifically, for the optimal control problem P cbf A new optimal control problem P is created by incorporating the machine learning control law 103 (see Figure 3) into the stage cost function of (s). cbf ML (s) is given by the following equation: TIFF0007799895000007.tif15133where x(k+1)=f((x(k),u(k)), x(0)=x0, x(k)∈X, x(N)∈X f , u(k)∈U, let h(x(k+1))-h(x(k))≧-γ(h(x)).

[0051] The above optimal control problem P cbf ML By solving (s), the control input u of the machine learning control law 103 is obtained. in The control input u is as close as possible to and satisfies the constraints. in can be calculated. In addition, the stage cost function l ML (x(k),u(k)) is given by the following equation: l ML (x(k),u(k))=(u(k)-κ cbf ML (s))R(u(k)-κ cbf ML (s))'

[0052] where R is the control input u in is an n × n weight matrix that represents the priority of the estimated control input κ cbfML To prioritize the asymptotic approach to (s), we use the terminal penalty function V f (x(N)) is set to 0. In this embodiment, the optimal control problem P cbf ML (s) is also called an optimal control problem involving machine learning control laws 103.

[0053] <Conceptual structure of predictive control> Fig. 7 is a diagram illustrating the conceptual configuration of predictive control by the predictive control server 10 (see Fig. 1) used in embodiment 2. In Fig. 7, parts corresponding to those in Fig. 3 are assigned the same reference numerals. The predictive control server 10 shown in FIG. 7 differs from the predictive control server 10 shown in FIG. 3 in that an optimal control problem 110 is used instead of the machine learning control law 103.

[0054] The optimal control problem 110 here is, as mentioned above, the optimal control problem P cbf A new optimal control problem P is created by incorporating the machine learning control law 103 (see Figure 3) into the stage cost function of (s). cbf ML (s). In Fig. 7, this is expressed as an optimal control problem 110 that achieves safe asymptotic approach to the target state. cbf ML (s) is an example of a first optimal control problem.

[0055] The optimal control problem 110 used in this embodiment is an optimal control problem P cbf ML (s) is divided into prediction periods N and solved. The prediction period N is set to a small value that allows for online implementation. By making the prediction period N small (or short), the optimal control problem 110 can be solved by dividing the control input u corresponding to the input pair s = (x0, γ) into in is calculated in real time, and the calculated control input u in can be provided to the target system 20. This control law is an example of a first model predictive control law. Implementation of the conceptual configuration shown in FIG. 7 is an example of a predictive control method.

[0056] <Verification by simulation> Below, the difference between the case where a control law that does not include a risk tolerance parameter is used and this embodiment will be verified by simulation. In addition, the optimal control problem P cbf ML The weighting matrix R of (s) is given by the following equation: TIFF0007799895000008.tif9133

[0057] FIG. 8 is a diagram illustrating the simulation results for achieving safe asymptotic approach to the target state including the risk tolerance parameter. (A) shows the optimal control problem P cbf ML (B) shows the optimal control problem P that achieves safe asymptotic approach to the target state. cbf ML (s) shows the state transition (γ=0.6). In the case of Figure 8, no violation of the constraint occurs even when actively moving near the constraint.

[0058] Figure 9A shows the optimal control problem P cbf ML This figure explains the transition paths and locations of constraint violations when one initial state [-5,-5,0,0] is given to (s). Note that the risk tolerance parameter γ=0.2. Figure 9B shows the optimal control problem P cbf ML This figure explains the transition paths and the locations of constraint violations when one initial state [-5,-5,0,0] is given to (s). Note that the risk tolerance parameter γ=0.6. In FIGS. 9A and 9B, (A) indicates a state transition, and (B) indicates a constraint violation. In both the cases of FIG. 9A and FIG. 9B, no constraint violation occurs.

[0059] <Summary> In the information processing system 1 according to the second embodiment, the optimal control problem P cbf A new optimal control problem P is created by incorporating the machine learning control law 103 (see Figure 3) into the stage cost function of (s). cbf ML (s) (i.e., the optimal control problem 110) is used to generate a control input for the target system 20. This makes it possible to achieve a safe asymptotic approach to the target state with a small computational load (within a time that can be implemented) not only when the value of the risk tolerance parameter γ is small but also when the value is large.

[0060] <Third embodiment: Ensuring convergence> This embodiment also uses the information processing system 1 (see FIG. 1) assumed in the first embodiment. The difference between this embodiment and the first embodiment is the content of predictive control by the predictive control server 10 (see FIG. 1).

[0061] In the case of the above-described second embodiment, safe asymptotic approach to the target state can be achieved regardless of the magnitude of the risk tolerance parameter γ. On the other hand, the predictive control described in the second embodiment does not have a mechanism to guarantee convergence to the target state. Therefore, in the case of the predictive control described in the second embodiment, the state may converge to the target state in some cases, but may not converge in other cases. In the cases of FIGS. 9A and 9B described above, the state that is finally reached varies from the target state. Therefore, a mechanism for ensuring convergence is added to the predictive control in this embodiment.

[0062] <Conceptual structure of predictive control> Fig. 10 is a diagram illustrating the conceptual configuration of predictive control by the predictive control server 10 (see Fig. 1) used in the third embodiment. In Fig. 10, parts corresponding to those in Fig. 7 are assigned the same reference numerals. Execution of the conceptual configuration shown in Fig. 10 is an example of a predictive control method. The predictive control server 10 shown in FIG. 10 differs from the predictive control server 10 shown in FIG. 7 in that it adds an optimal control problem 120 for convergence and a switch 131 for switching between two optimal control problems and a switching control unit 132 for controlling the switching.

[0063] Therefore, the input pair s is provided to both the switch 131 and the switching control section 132 . The switch 131 is a switch that switches between optimal control problems that provide the input pair s. In other words, the switch 131 is a switch that switches the input destination of the input pair s. In this embodiment, the input pair s is given to the optimal control problem 110 or the optimal control problem 120 in accordance with an instruction from the switching control unit 132 .

[0064] The switching control unit 132 is a functional unit that controls switching of the control law based on the switching strategy π(s). For example, the switching strategy π(s) is a trained model that uses reinforcement learning to estimate the choice of control law (i.e., action a) from the state s = (x0, γ) so as to maximize the reward r(s, a).

[0065] The reward r(s, a) here is designed to increase as the target state is approached. In this embodiment, the action a is either "convergence" or "safe asymptotic." Convergence (ie, a=0) is associated with the choice of optimal control problem 120, and safe asymptotics (ie, a=1) are associated with the choice of optimal control problem 110. The optimal control problem 120 (shown in FIG. 10 as an "optimal control problem that realizes convergence to the target state") is also an optimal control problem P cbf It is given as (s).

[0066] Here, the optimal control problem 120 (i.e., the optimal control problem P cbf (s)) is an example of a second optimal control problem. Moreover, the control law that solves the optimal control problem 120 is an example of a second model predictive control law. In this embodiment, the optimal control problem 120 is solved by a terminal penalty function V f Optimal control problem P involving (x(N)) cbf Model predictive control law κ resulting from solving (s) cbf Use (·). In addition, the model predictive control law κ cbf Optimal control problem P for solving (·) cbf The prediction period N (s) is also set to a small value that allows for online implementation.

[0067] Incidentally, the state x is high-dimensional and continuous, and there are a huge number of patterns. Therefore, the optimal control problem 110 (i.e., P cbf ML (s)) to the optimal control problem 120 (i.e., P cbf It is generally difficult to determine when to switch to (s). Therefore, in this embodiment, a switching strategy π(s) that estimates an action a representing the selection of a control law from the input pair s = (x0, γ) so as to maximize the reward r(s, a) is learned in advance using reinforcement learning.

[0068] When the state x is high-dimensional and continuous, and the action a to be selected is a finite integer, reinforcement learning of the switching policy π(s) uses, for example, a Deep Q-Network, in which the switching policy π(s) is expressed as a neural network. The Deep Q-Network is an example of a neural network-related technology. In this embodiment, one epoch is 300 steps, and reinforcement learning is performed offline for 500 epochs to learn a switching policy π(s) for switching between two control laws.

[0069] To achieve rapid convergence to the goal state, the reward r(s, a) is designed to increase as the goal state is approached. For example, the reward r(s, a) is designed as follows: r(s,a)=((x d,max -||xx r ||) / x d,max ) p

[0070] By the way, x r is the goal state. d,max is ||max(|x min -x r |,|x max -x r |)-x r |)-x r ||. Note that || || are norms. Also, the value in parentheses of the reward r(s, a) is normalized to the range of 0 to 1. Therefore, the larger the value of the exponent p, the higher the reward that can be obtained when the agent approaches the goal state, leading to earlier convergence.

[0071] For example, the exponent p that gives the reward r(s, a) is set to 10. Then, online simulations were performed using two control laws and a switching strategy π(s) for two cases: risk tolerance parameter γ = 0.2 (for safe control) and γ = 0.6 (for proactive control).

[0072] <Verification by simulation> Below, verification will be performed by simulation in this embodiment. As described above, in this embodiment, of the two control laws, the control law to which an input pair s is given is switched by a switching policy π(s) that has been reinforced learned so as to converge to a target state.

[0073] 11 shows the simulation results of achieving safe asymptosis and convergence to the target state. (A) shows a state transition (γ=0.2) that achieves convergence to the target state via a highly safe transition path, and (B) shows a state transition (γ=0.6) that achieves convergence to the target state in addition to safe asymptosis.

[0074] 12A and 12B are diagrams illustrating state transitions and switching control when achieving convergence to a target state via a highly secure transition path. (A) shows state transitions when γ=0.2, and (B) shows constraint violations and switching control when γ=0.2. As shown in Figure 12(B), in this embodiment, optimal control problem 110 (i.e., safe asymptotic) is used up to around 50 steps, and optimal control problem 120 (i.e., convergence) is used thereafter. As a result, no constraint violations occur during the transition. Also, as shown in Figure 12(A), convergence to the target state occurs for all initial states.

[0075] 13A and 13B are diagrams illustrating state transitions and switching control when achieving convergence to the target state along a transition path that achieves safe asymptotic approach. (A) shows state transitions when γ=0.6, and (B) shows constraint violations and switching control when γ=0.6. In the case of Figure 13(B), optimal control problem 110 (i.e., safe asymptotic) is used up to around 50 steps, and then optimal control problem 120 (i.e., convergence) is used thereafter. As a result, no constraint violations occur during the transition. Also, as shown in Figure 13(A), convergence occurs to the target state for all initial states.

[0076] <Convergence and computation time> FIG. 14 is a diagram illustrating the number of simulation steps required to converge to the target state. Figure 14 shows the relationship between the convergence of each control law and the number of steps for each value of the risk tolerance parameter γ. Note that "convergence" was determined when the absolute value of the difference between the target state and all elements of the state was 0.005 or less.

[0077] For example, in the case of a model predictive control law including a risk tolerance parameter, convergence to the target state occurs in 59 steps when the risk tolerance parameter γ=0.2, and in 52 steps when the risk tolerance parameter γ=0.6. On the other hand, in the case of the second embodiment (FIG. 7), the target state was not converged to when the risk tolerance parameter γ=0.2 or γ=0.6. In the case of the third embodiment (FIG. 10), when the risk tolerance parameter γ=0.2, the system converges to the target state in 237 steps, and when the risk tolerance parameter γ=0.6, the system converges to the target state in 220 steps.

[0078] FIG. 15 is a diagram illustrating the maximum calculation time required to solve one optimal control problem. The model predictive control law including the risk tolerance parameter was able to solve the optimal control problem in 0.2 seconds when the risk tolerance parameter γ = 0.2, but it took 1.531 seconds to solve the optimal control problem when the risk tolerance parameter γ = 0.6. This is presumably because it was necessary to carefully consider the constraints in order to select an aggressive route.

[0079] On the other hand, in this embodiment, when the risk tolerance parameter γ=0.2, the time required to solve the optimal control problem 120 was 0.034 seconds, and the time required to solve the optimal control problem 110 was 0.044 seconds. Both control laws are shorter than the model predictive control law. One of the reasons for this is thought to be that the prediction period N is set small (or short) to enable online implementation.

[0080] Furthermore, when the risk tolerance parameter γ=0.6, the time required to solve the optimal control problem 120 was 0.026 seconds, and the time required to solve the optimal control problem 110 was 0.046 seconds. These times are equivalent to those when the risk tolerance parameter γ=0.2. Incidentally, with the model predictive control law, when the risk tolerance parameter γ=0.6, it took 1.531 seconds to solve the optimal control problem. Therefore, it can be seen that in this embodiment, the calculation time when the risk tolerance parameter γ=0.6 is significantly reduced compared to the model predictive control law (offline). This indicates that this embodiment has high online applicability.

[0081] <Summary> As described above, in this embodiment, an input pair s including a risk tolerance parameter γ is given to either an optimal control problem 110 that achieves safe asymptotic approach or an optimal control problem 120 that achieves convergence to a target state, and a control input for the target system 20 is generated. This allows for safe predictive control with a small calculation load.

[0082] Furthermore, in this embodiment, two control laws are switched and controlled using a switching strategy π(s) that has been reinforced learned so that the closer the state approaches the target state, the greater the reward. Therefore, in this embodiment, it is possible to achieve both safety and convergence of the transition path.

[0083] <Other embodiments> (1) Although the embodiments of the present invention have been described above, the technical scope of the present invention is not limited to the scope of the above-described embodiments. For example, it is clear from the claims that various modifications or improvements to the above-described embodiments are also included in the technical scope of the present invention.

[0084] (2) In the above embodiment, the case where the predictive control server 10 (see FIG. 1) controls the state of the target system 20 (see FIG. 1) through the network N has been described. However, the functions of the predictive control server 10 described above may be placed in the target system 20 or in the facility of the business operator that operates the target system 20. In this case, the predictive control server 10 or a terminal placed in the target system 20 or in the facility of the business operator that operates the target system 20 is an example of a predictive control system.

[0085] (3) In the above-described embodiment, the case where a program having the conceptual configuration shown in any one of FIG. 3, FIG. 7, and FIG. 10 is installed in the predictive control server 10 has been described. Incidentally, this program may be distributed to businesses that operate the target systems 20 via a network or recording medium.

[0086] (4) The predictive control process described in the above embodiments can be executed by any computer. Furthermore, any computer may execute the reservation control process using a processor as hardware, a program as software, or a combination of these. In this case, the processor may be configured to execute the predictive control process described in the embodiments in cooperation with the program, and may function as each unit or means described in the embodiments.

[0087] Furthermore, the order in which the processor executes the predictive control processes is not limited to the order described in the embodiment, and may be changed as appropriate. Furthermore, a plurality of processors may share the execution of the predictive control process described above. Any computer may be a general purpose computer, a special purpose computer, a workstation, or any other system capable of performing the processes.

[0088] The processor may be configured with one or more pieces of hardware. Note that the type of the one or more pieces of hardware is not limited. For example, the processor may be a programmable logic device such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or an FPGA (Field Programmable Gate Array), a dedicated circuit for executing specific processing such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), an NPU (Neural Processing Unit), or a TPU (Tensor Processing Unit).

[0089] The type of hardware may be a combination of different types of hardware. When multiple pieces of hardware are configured to execute one or more processes of a certain processor, the multiple pieces of hardware may exist in devices physically separated from each other, or may exist in the same device. Furthermore, the order of each process performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is configured by an electric circuit or the like that combines circuit elements such as semiconductor elements.

[0090] Further, the program may be software such as firmware or microcode. Alternatively, the program may be, for example, a group of program modules, the functions of which may be implemented by a processor configured to perform the respective functions. The program may be a program code or a number of code segments stored in one or more non-transitory computer-readable media (eg, storage media or other storages).

[0091] The program may be stored in multiple non-transitory computer-readable media that are physically separate from each other. A program code or code segment may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, instructions, data structures, or program statements. A program code or code segment may be connected to another code segment or a hardware circuit by sending or receiving information, data, arguments, parameters, or memory contents. [Explanation of symbols]

[0092] 1...information processing system, 10...predictive control server, 20, 20A, 20B...target system, 11...processor, 12...semiconductor memory, 13...auxiliary storage device, 14...communication interface, 15...signal line, 101...initial state update unit, 102...unit for generating input pairs of initial state and risk tolerance parameter, 103...machine learning rule including risk tolerance parameter, 110...optimal control problem realizing safe asymptotic approach to target state, 120...optimal control problem realizing convergence to target state, 131...switch, 132...switching control unit

Claims

1. a processor; The processor: An input pair, which is a state obtained from the target system and a risk tolerance parameter, is provided to a machine learning control rule incorporating the risk tolerance parameter; providing the output of the machine learning control law as a control input to the target system; Predictive control systems.

2. The machine learning control law includes a trained model that has been machine-learned to obtain a model predictive control law obtained for an optimal control problem incorporating a risk tolerance parameter. The predictive control system of claim 1 .

3. The trained model is obtained by machine learning a result of solving an optimal control problem incorporating a control barrier function including a risk tolerance parameter. The predictive control system of claim 2 .

4. A neural network is used as the machine learning algorithm of the machine learning control law. The predictive control system of claim 3 .

5. The machine learning algorithm of the machine learning control law uses neural network-related technology including a recurrent neural network. The predictive control system of claim 3 .

6. The system further includes a first model predictive control law that solves a first optimal control problem incorporating the machine learning control law into a stage cost function, The first optimal control problem achieves safe asymptotic approach to the target state. The predictive control system of claim 1 .

7. a second model predictive control law that solves a second optimal control problem including the machine learning control law; The processor: switching an input destination of the input pair between the first model predictive control law and the second model predictive control law according to the value of the input pair; The predictive control system of claim 6 .

8. The processor: When the input pair is given to the first model predictive control law or the second model predictive control law, the first optimal control problem or the second optimal control problem is solved within a prediction period that can be implemented online. The predictive control system of claim 7.

9. The processor: The switching of the input destination is controlled using a trained model trained by reinforcement learning. The predictive control system of claim 7.

10. The processor: changing a value of a risk tolerance parameter to be assigned to the machine learning control rule in accordance with a constraint on a path from the initial state to the target state; The predictive control system of claim 1 .

11. The processor: When priority is given to the safety of the path to the goal state, the value of the risk tolerance parameter is changed to a first value; If the distance or time of the route to the goal state is prioritized, the value of the risk tolerance parameter is changed to a second value greater than the first value. The predictive control system of claim 10.

12. The processor: changing the value of the risk tolerance parameter while controlling the target system; The predictive control system according to claim 10 or 11.

13. The processor: A process of providing an input pair, which is a state obtained from the target system and a risk tolerance parameter, to a machine learning control rule incorporating the risk tolerance parameter; A process of providing the output of the machine learning control rule as a control input to the target system; A predictive control method that performs

14. On the computer, A function of providing an input pair, which is a state obtained from the target system and a risk tolerance parameter, to a machine learning control rule incorporating the risk tolerance parameter; a function of providing the output of the machine learning control law as a control input to the target system; A program to achieve this.

Citation Information

Patent Citations

  • Method and system for providing optimized control of complex dynamical system

    JP2018195307A

  • Controlling a Moveable Device Utilizing Risk Control Barrier Functions

    US20230384790A1