Management device, management method, and control system

The management device utilizes state transition probabilities and evaluation values to determine optimal operation amounts, addressing the limitations of existing model predictive control systems by enhancing control precision and effectiveness.

JP2025126758APending Publication Date: 2025-08-29HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024023160
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Existing model predictive control systems for control targets lack the ability to appropriately control the target object, as they do not effectively utilize state transition probabilities and evaluation values to determine optimal operation amounts.

Method used

A management device that includes a control characteristic estimation unit to calculate evaluation values based on state transition probabilities and a control law determination unit to determine optimal operation amounts for each state, enhancing the control system's ability to manage and control the target object.

Benefits of technology

The management device provides a system that can appropriately generate information for controlling the control target, improving the precision and effectiveness of control operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126758000001_ABST
    Figure 2025126758000001_ABST
Patent Text Reader

Abstract

To provide a management device, a management method, and a control system that appropriately generate information used for controlling a control target.SOLUTION: A management device 10 comprises: a control characteristic estimation unit 14d which calculates an evaluation value of each operation amount when reaching a target candidate state which is a candidate of a target state from a plurality of states which the control target can take based on a state transition probability which is probability of a state transition when the control target is operated with a predetermined operation amount; and a control law determination unit 14e which determines the optimum operation amount for each state of the control target as a control law based on the evaluation value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a management device, a management method, and a control system. [Background technology]

[0002] Regarding model predictive control of a control object, for example, a technique described in Patent Document 1 is known. That is, Patent Document 1 describes "a method for improving the performance of a method for predicting the future state of a target object by calculation, which evaluates the quality of the prediction by comparing the observed actual state with the prediction result (20)." [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-212872 Summary of the Invention [Problem to be solved by the invention]

[0004] The technology described in Patent Document 1 predicts the future state of the control target (target object) based on a model that simulates the behavior of the control target, and calculates the optimal control method from the predicted future state, but there is room for improvement in terms of appropriately controlling the control target.

[0005] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide a management device, a management method, and a control system that appropriately generate information used to control a control target. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems, the management device according to the present disclosure includes a control characteristic estimation unit that calculates an evaluation value of each operation amount when the controlled object reaches a target candidate state, which is a candidate for the target state, from multiple states that the controlled object can take, based on a state transition probability, which is the probability of a state transition when the controlled object is operated with a predetermined operation amount, and a control law determination unit that determines, based on the evaluation value, an optimal operation amount for each state of the controlled object as a control law. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide a management device, a management method, and a control system that appropriately generate information used for controlling a control target. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a configuration diagram of a control system including a management device according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of a management apparatus according to an embodiment. [Figure 3] FIG. 2 is an explanatory diagram illustrating a specific example of a plant that is a control target of the management device according to the embodiment. [Figure 4] 4 is an explanatory diagram showing a change in steam temperature of the plant in the example of FIG. 3 in the management device according to the embodiment. FIG. [Figure 5] 10 is a flowchart illustrating a process flow of a management device according to an embodiment. [Figure 6] FIG. 10 is an explanatory diagram of a state transition probability matrix in the management device according to the embodiment. [Figure 7] FIG. 3 is an explanatory diagram of a control characteristic matrix in the management device according to the embodiment. [Figure 8] FIG. 2 is an explanatory diagram of a control rule in a management device according to an embodiment. [Figure 9] FIG. 2 is an explanatory diagram relating to a definition of a state of a plant in the management device according to the embodiment. [Figure 10] 4 is an explanatory diagram showing the correspondence relationship between an operation amount and a valve opening adjustment amount in the management device according to the embodiment. FIG. [Figure 11A] 10 is a flowchart relating to estimation of a control characteristic in the management device according to the embodiment. [Figure 11B] 10 is a flowchart relating to estimation of a control characteristic in the management device according to the embodiment. [Figure 12] 10 is a flowchart relating to calculation of a control law in a management device according to an embodiment. [Figure 13] FIG. 10 is an explanatory diagram illustrating a reward function in the management device according to the embodiment. [Figure 14] FIG. 10 is an explanatory diagram of a matrix that is a calculation result of the product of a control characteristic matrix and a reward function vector in the management device according to the embodiment. [Figure 15] 10 is an explanatory diagram showing the correspondence relationship between the maximum value of the evaluation value for each state and the optimal operation amount in the management device according to the embodiment. FIG. [Figure 16] 10 is a diagram illustrating an example of a display screen related to processing of a management device according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] <Embodiment> In the following, as an example, a case will be described in which the management object (control object) of the management device 10 (see FIG. 1) is a predetermined plant 20 (see FIG. 1). Examples of such plants 20 include power plants, chemical plants, oil refineries, steel plants, food plants, pharmaceutical plants, and water treatment plants. Note that the management object (control object) of the management device 10 (see FIG. 1) is not limited to plants, and may also be a robot, a vehicle, a ship, an aircraft, or the like.

[0010] <Control system configuration> FIG. 1 is a configuration diagram of a control system 100 including a management device 10 according to an embodiment. A control system 100 shown in Fig. 1 is a system that controls equipment in a plant 20 using a management device 10. As shown in Fig. 1, the control system 100 includes the management device 10, a sensor 30, an input device 40, a display device 50, and a cloud 60, which are connected in a predetermined manner by wire or wirelessly. Below, a brief description will be given of the plant 20 that is the control target of the management device 10, as well as the sensor 30, the input device 40, the display device 50, etc., and then a detailed description will be given of the management device 10.

[0011] The plant 20 is a facility provided with control elements such as valves and heaters, such as a power plant. The sensor 30 acquires predetermined detected values ​​and environmental information from the equipment of the plant 20. As such a sensor 30, for example, a camera, a distance sensor, a radar, a force sensor, a temperature sensor, an angle sensor, a flow rate sensor, a pressure sensor, a voltage sensor, a current sensor, or the like may be used as appropriate.

[0012] The sensor 30 may be installed inside the plant 20 or outside the plant 20. An example of a sensor 30 installed outside the plant 20 is an outside air temperature sensor that detects the temperature of the environment surrounding the plant 20. In addition, the sensor 30 also includes devices for acquiring data such as power demand related to the control of the plant 20 from outside the plant 20.

[0013] The input device 40 is used when the operator M1 inputs predetermined management information and commands to the management device 10. As such an input device 40, for example, a keyboard, a mouse, or a joystick is used. The display device 50 is used to present the calculation results of the management device 10 to the operator M1. For example, a liquid crystal display is used as such a display device 50. Note that a touch panel type mobile terminal that combines the functions of the input device 40 and the display device 50, such as a smartphone or tablet, may also be used. The cloud 60 has a server (not shown) that communicates in a predetermined manner with the input device 40, the display device 50, and the management device 10 via a network. The processing results of the server are provided to the management device 10 via the network and are also displayed appropriately on the display device 50.

[0014] <Configuration of management device> The management device 10 is a device that controls and manages the plant 20. A computer such as a personal computer, a tablet, or a smartphone may be used as such a management device 10. The management device 10 may also be configured by connecting multiple computers in a predetermined manner via a communication line or a network. For example, the functions of the management device 10 may be distributed among multiple computers such as a cloud server or an edge server.

[0015] 1, the management device 10 includes a data acquisition unit 11, a storage unit 12, a communication unit 13, and a calculation unit 14. The data acquisition unit 11, the storage unit 12, the communication unit 13, and the calculation unit 14 are connected in a predetermined manner via an internal bus (not shown).

[0016] FIG. 2 is a diagram illustrating an example of the hardware configuration of the management device 10. As shown in FIG. As shown in FIG. 2, the management device 10 has a hardware configuration including a processor 10a, a RAM 10b (Random Access Memory), a ROM 10c (Read Only Memory), a HDD 10d (Hard Disk Drive), a communication interface 10e, an input / output interface 10g, and a media interface 10h, which are connected in a predetermined manner via an internal bus 10i.

[0017] 2 is hardware of the calculation unit 14 (see FIG. 1) of the management device 10. RAM 10b, ROM 10c, and HDD 10d are hardware of the storage unit 12 (see FIG. 1) of the management device 10. The processor 10a reads out a predetermined program stored in the ROM 10c or the HDD 10d and loads it into the RAM 10b, thereby executing a predetermined process.

[0018] 2 is used for performing predetermined communications with devices and sensors 30 of the plant 20 and the cloud 60. The input / output interface 10g receives data from the input device 40 and outputs data to the display device 50. The communication interface 10e and the input / output interface 10g function as the communication unit 13 of the management device 10 (see FIG. 1).

[0019] The media interface 10h is connected to an external recording medium 70 such as an optical disc 70a or a USB memory 70b as appropriate, and functions as the data acquisition unit 11 (see FIG. 1) of the management device 10. As such a media interface 10h, an optical disc drive that plays the optical disc 70a or a USB interface that reads the USB memory 70b may be used. Note that "USB" is a registered trademark.

[0020] Returning to Figure 1 again, the explanation will continue. The data acquisition unit 11 shown in Fig. 1 acquires predetermined data from an external recording medium 70 such as an optical disc 70a or a USB memory 70b. Various data are stored in the storage unit 12. As shown in Fig. 1, the storage unit 12 includes an action storage unit 12a, a control characteristic storage unit 12b, a reward storage unit 12c, and a control law storage unit 12d. These will be described in detail later, but in summary, they have the following functions.

[0021] That is, the operation storage unit 12a stores the calculation results of the model estimation unit 14c as data indicating the operation of the plant 20. Note that data (data specifying a model of the plant 20) acquired from the cloud 60 or the external recording medium 70 may be stored in the operation storage unit 12a. The format of the data stored in the operation storage unit 12a may be a model or simulator that simulates the operation of the plant 20, or may be a flow chart, table, graph, or other predetermined calculation formula that indicates the behavior of the plant 20.

[0022] The control characteristic storage unit 12b stores the calculation results of the control characteristic estimation unit 14d. That is, data such as evaluation values ​​of manipulated variables in each state with respect to the control target of the plant 20 is stored in the control characteristic storage unit 12b in the form of a table, graph, or calculation formula.

[0023] The reward storage unit 12c stores a predetermined reward function that represents a control goal in the form of a mathematical formula, table, vector, matrix, etc. Details of the reward function will be described later. The control law storage unit 12d stores the control law, which is the calculation result of the control law determination unit 14e. Here, the "control law" refers to a rule (law) for making the manipulated variable of the controlled object follow a predetermined target value in order to achieve the control objective. For example, if the control objective is to stabilize the steam temperature in a thermal power plant, the rule (law) for controlling the temperature control valve to bring the steam temperature closer to the predetermined target value is set as the control law.

[0024] The communication unit 13 communicates in a predetermined manner with the plant 20, the sensor 30, the input device 40, the display device 50, and the cloud 60. The communication unit 13 is connected to the storage unit 12 and the calculation unit 14 via an internal bus. Data acquired from the plant 2, etc. via the communication unit 13 is stored in the storage unit 12, and calculation results of the calculation unit 14 are transmitted to the plant 20, etc. via the communication unit 13.

[0025] The calculation unit 14 generates predetermined data related to the autonomous control of the plant 20. At least a part of the processing of the calculation unit 14 may be performed by AI (Artificial Intelligence). As shown in FIG. 1, the calculation unit 14 includes an input control unit 14a, an output control unit 14b, a model estimating unit 14c, a control characteristic estimating unit 14d, a control law determining unit 14e, and an equipment control unit 14f. Of the functional units of the calculation unit 14, the functions of the model estimating unit 14c, the control characteristic estimating unit 14d, and the control law determining unit 14e may be performed by a cloud 60.

[0026] The details of each functional unit of the calculation unit 14 will be described later, but in summary, it has the following functions: The input control unit 14a transfers data input via the data acquisition unit 11 or the communication unit 13 to the storage unit 12 or an appropriate functional unit of the calculation unit 14.

[0027] The output control unit 14b outputs data from the storage unit 12 and the calculation unit 14 to the plant 20, the display device 50, the cloud 60, or the like via the communication unit 13. The output control unit 14b also generates image information for displaying data related to the processing of the management device 10 in a predetermined manner on the display device 50. Note that a predetermined computer (not shown) connected to the display device 50 may generate the image information.

[0028] The model estimation unit 14c estimates a model for estimating the operation of the plant 20. The operation of the plant 20 may be represented by a predetermined model or simulator, or may be represented in the form of a flow chart, table, graph, or calculation formula. The estimation result of the model estimation unit 14c is stored in the operation storage unit 12a.

[0029] The control characteristic estimation unit 14d estimates an evaluation value of the manipulated variable in each state with respect to the control target of the plant 20. In addition to the evaluation value of the manipulated variable, the control characteristic estimation unit 14d may also calculate the state transition of the plant 20 associated with the manipulated variable and the state transition probability. The evaluation value, etc., which is the calculation result of the control characteristic estimation unit 14d, is expressed in the form of a table, graph, or calculation formula. The estimation result of the control characteristic estimation unit 14d is stored in the control characteristic storage unit 12b.

[0030] Based on the data in the control characteristic storage unit 12b and the reward storage unit 12c, the control law determination unit 14e determines an optimal control law for controlling the plant 20. The data of the control law determined by the control law determination unit 14e is stored in the control law storage unit 12d.

[0031] The equipment control unit 14f calculates the manipulated variable of the equipment of the plant 20 based on the optimal control law stored in the control law storage unit 12e. The manipulated variable calculated by the equipment control unit 14f is transmitted to the plant 20 via the output control unit 14b and the communication unit 13 in this order. As a result, the equipment of the plant 20 is controlled with the predetermined manipulated variable.

[0032] FIG. 3 is an explanatory diagram showing a specific example of a plant 20 that is the control target of the management device. 3 shows an example in which the plant 20 is a thermal power plant. In the plant 20 having the configuration shown in FIG. 3, a boiler drum 21, a first superheater 22, a second superheater 23, and a steam turbine 24 are connected in this order via steam piping.

[0033] The steam in the boiler drum 21 is superheated in the first superheater 22, and then further superheated in the downstream second superheater 23. The steam thus superheated and heated in the second superheater 23 is guided to the steam turbine 24. The steam turbine 24 is rotated by the wind pressure of the steam, and the rotor of a generator (not shown) rotates integrally with the steam turbine 24, thereby generating electricity.

[0034] As shown in Fig. 3, the downstream end of a spray pipe 25 is connected to the steam pipe between the first superheater 22 and the second superheater 23. A nozzle-shaped spray 26 is provided at the downstream end of the spray pipe 25. A regulating valve 27 is provided on the spray pipe 25 to adjust the amount of compressed water injected. The compressed water injected via the spray 26 mixes with the steam, thereby lowering the temperature of the steam. Note that the greater the amount of compressed water injected via the spray 26 (i.e., the greater the opening of the regulating valve 27), the greater the amount of decrease in the temperature of the steam.

[0035] A temperature sensor 28 is provided in the steam pipe between the second superheater 23 and the steam turbine 24. The opening of the adjustment valve 27 is adjusted so that the detected value (steam temperature) of the temperature sensor 28 approaches a predetermined target temperature. The temperature sensor 28 corresponds to the above-mentioned sensor 30 (see FIG. 1).

[0036] FIG. 4 is an explanatory diagram showing changes in steam temperature in the plant in the example of FIG. In Fig. 4, the horizontal axis represents the elapsed time from a predetermined time, and the vertical axis represents the steam temperature (the temperature at the location where the temperature sensor 28 is installed). States s1 to s6 shown in Fig. 4 indicate states associated with the level of the steam temperature. Furthermore, the solid and dashed arrows shown in Fig. 4 each indicate a state transition based on a predetermined mathematical model.

[0037] In the example of FIG. 4, the target steam temperature is set to 500°C, and the target temperature change rate for the steam temperature is set to approximately 0°C / control step. The initial steam temperature is 495°C, and the initial temperature change rate is set to approximately 2°C / control step. Here, the "control step" refers to the cycle at which the manipulated variable is updated based on a predetermined control law. Note that the "control step" may be set separately from the control cycle of the adjustment valve 27 (see FIG. 3) based on PID control (Proportional-Integral-Differential Controller).

[0038] In the control pattern indicated by the solid arrows in Figure 4, as the control steps progress, the transitions are as follows: State s1 → State s2 → State s3 → State s4 → State s4 → State s4. Note that State s4 corresponds to the target steam temperature of 500°C.

[0039] <Management device processing> FIG. 5 is a flowchart showing the flow of processing by the management device (see also FIG. 1 as appropriate). 5 is started by a user operation via the input device 40. For example, when operation of the plant 20 is started or when control is switched from predetermined PID control to control by the management device 10, the series of processes shown in FIG. 5 is started by a user operation via the input device 40.

[0040] In step St1 of Fig. 5, the management device 10 estimates a model of the plant 20 using the model estimation unit 14c. Here, the "model" is a mathematical model used to simulate the behavior of the plant 20. To specifically explain the processing of step St1, the management device 10 first acquires operation data and simulation data of the plant 20 via the data acquisition unit 11 and the communication unit 13. The operation data is mainly data accumulated during past operations of the plant 20.

[0041] For example, the operation data may include information that when a predetermined plant 20 is in state s1 (see FIG. 4), a transition to state s2 occurs when the regulating valve 27 (see FIG. 3) is set to a predetermined opening. Also, predetermined simulation data may be used together with (or instead of) the operation data. For example, simulation data may be used when the amount of data obtained from past operation data alone is insufficient. In this embodiment, in the process of step St1, the model estimation unit 14c estimates a model of the plant 20 in the form of a state transition probability matrix Mx1 (see FIG. 6).

[0042] FIG. 6 is an explanatory diagram of the state transition probability matrix Mx1 (also see FIG. 1 as appropriate). 6 shows a three-dimensional representation of a state transition probability matrix Mx1 including a "current state," a "transition destination state," and an "operation amount." The vertical axis of the state transition probability matrix Mx1 shown in FIG. 6 represents a "current state" before a predetermined operation is performed on the plant 20. In FIG. 6, the vertical axis represents the direction in which rows are arranged, and is the first row, second row, third row, etc. from top to bottom. The horizontal axis represents the direction in which columns are arranged, and is the first column, second column, third column, etc. from left to right. Incidentally, the term "current state" is not limited to the meaning of the current state in real time in the plant 20, but is used to mean the state one step before in the control cycle with respect to the state to which the transition is made.

[0043] The horizontal axis of the state transition probability matrix Mx1 represents the "destination state" after a predetermined operation is performed on the plant 20. The depth axis of the state transition probability matrix Mx1 represents the manipulated variable at a predetermined operating terminal of the plant 20. For example, the valve opening adjustment variable of the adjustment valve 27 shown in FIG. 3 corresponds to the "manipulated variable" in FIG. 6.

[0044] In the state transition probability matrix Mx1, a state transition probability value is associated with each of the elements identified by the current state, the destination state, and the manipulated variable. Here, the "state transition probability" refers to the probability that a predetermined destination state will be reached when control is performed from the current state with a predetermined manipulated variable. For example, the value of 0.9 is associated with the element in the first row and second column of the state transition probability matrix Mx1, whose depth corresponds to the manipulated variable a1. This means that when the current state is state s1, if the manipulated variable of the plant 20 is controlled with the manipulated variable a1, there is a 90% probability of transitioning to state s2.

[0045] Such a state transition probability matrix Mx1 is generated by the management device 10 (or the cloud 60) performing statistical processing using a huge amount of past operation data of the plant 20. The data of the state transition probability matrix Mx1 estimated as a "model" in step St1 of Fig. 5 is stored in the operation storage unit 12a.

[0046] Next, in step St2 of Fig. 5, the management device 10 estimates the control characteristics of the plant 20 using the control characteristics estimation unit 14d (control characteristics estimation step). In this embodiment, the control characteristics estimation unit 14d estimates the control characteristics of the plant 20 based on the state transition probability matrix Mx1 (see Fig. 6) stored in the operation storage unit 12a. In this embodiment, as the processing of step St2 of Fig. 5, the control characteristics estimation unit 14d represents the evaluation value of each manipulated variable (i.e., the control characteristics of the plant 20) in the form of a control characteristics matrix Mx2 (see Fig. 7).

[0047] FIG. 7 is an explanatory diagram of the control characteristic matrix Mx2 (also see FIG. 1 as appropriate). The vertical axis of the control characteristic matrix Mx2 shown in FIG. 7 is the "current state" before a predetermined operation is performed on the plant 20. The horizontal axis of the control characteristic matrix Mx2 is the "target state" indicating a target state of the plant 20. The values ​​of the elements in the first depth layer of the control characteristic matrix Mx2 are "evaluation values" associated with combinations of the current state and the target state. Here, the "evaluation value" is an expected value when the plant 20, which is the control target, reaches a target candidate state (a candidate for the target state) from a predetermined state.

[0048] The values ​​of the elements in the second depth layer of the control characteristic matrix Mx2 are "optimal operation amounts" associated with combinations of the current state and the target state. Here, the "optimal operation amount" refers to the operation amount having the largest evaluation value. For example, in the first row and second column of the control characteristic matrix Mx2, a value of 1.0 is associated with the first depth layer as the evaluation value, and a2 is associated with the second layer as the optimal operation amount. This indicates that when state s2 is the target state, the optimal operation amount in state s1 is a2, and the evaluation value of this operation amount a2 is 1.0. Such a control characteristic matrix Mx2 is created based on the state transition probability matrix Mx1 (see FIG. 6), and details of this process will be described later.

[0049] Returning to FIG. 5 again, the explanation will be continued. In step St3, the management device 10 calculates a control law by the control law determination unit 14e (control law determination step). As described above, a "control law" is a rule for making an manipulated variable follow a predetermined target value. To explain the processing of step St3 in more detail, the control law determination unit 14e derives an optimal control law based on the control characteristic matrix Mx2 (see FIG. 7) stored in the control characteristic storage unit 12b and the reward function stored in the reward storage unit 12c. The control law derived in this manner is stored in the control law storage unit 12d. The derivation of the control law based on the control characteristic matrix Mx2 and the reward function will be described in detail later.

[0050] FIG. 8 is an explanatory diagram of the control law. 8, the control law is shown in a table format in which the state of the plant 20 is associated with the optimal manipulated variable. For example, when the plant 20 is in the state s1, the control law indicates that it is optimal (highest evaluation value) to control a predetermined manipulated variable with the manipulated variable a2.

[0051] Next, in step St4 of Fig. 5, the management device 10 measures the state of the plant 20. Specifically, the management device 10 acquires the detection values ​​of the sensor 30 and the like via the communication unit 13. Here, an example of a definition regarding the state of the plant 20 (hereinafter referred to as a state definition) will be described with reference to Fig. 9.

[0052] FIG. 9 is an explanatory diagram regarding the state definition of the plant. In the example of FIG. 9, states s1 to s2 are determined by a combination of the "temperature" of the steam detected by the temperature sensor 28 (see FIG. 3) and the "rate of change" (speed of change) of the steam temperature. nis defined. Specifically, the steam temperature is divided into a range of 494.5°C to 550.5°C in 1°C increments. A range of the steam temperature change rate is set corresponding to each range of the steam temperature. The "step" included in the unit of the steam temperature change rate is the control period (for example, 1 minute) of the control valve 27 (see FIG. 3). In the example of FIG. 9, the state when the steam temperature is 494.5°C to 495.5°C and the steam temperature change rate is in the range of 1.5°C / step to 2.5°C / step is defined as state s1. Similarly, other states s2, s3, . . . , s n Such state definitions are set in advance by a user through an operation via the input device 40 (see FIG. 1).

[0053] Returning to FIG. 5 again, the explanation will be continued. In step St5, the management device 10 uses the equipment control unit 14f to calculate the manipulated variable of the manipulated variable of the plant 20. That is, the equipment control unit 14f calculates the manipulated variable of the manipulated variable of the plant 20 based on the optimal control law stored in the control law storage unit 12e. For example, when the plant 20 is in state s1, the management device 10 identifies the optimal manipulated variable a2 based on a predetermined control law (see FIG. 8).

[0054] FIG. 10 is an explanatory diagram showing the correspondence relationship between the operation amount and the valve opening adjustment amount. In the example of FIG. 10, the manipulated variable a2 corresponds to an operation of setting the valve opening adjustment amount of the regulating valve 207 to 0% (i.e., maintaining the valve opening). The remaining manipulated variables a1 and a3 correspond to valve opening adjustment amounts of -1% and +1%, in that order. Such a correspondence relationship is set in advance by a user operating the input device 40 (see FIG. 1). Note that the correspondence relationship shown in FIG. 10 is an example and is not limited to this.

[0055] 5, the management device 10 determines whether or not to end the control of the plant 20. If the control of the plant 20 is to be ended in step St6 (St6: Yes), the management device 10 ends the series of processes (END). For example, the control by the management device 10 is ended when the operation of the plant 20 is to be ended, when the operation is to be suspended for maintenance, or when the control by the management device 10 is to be switched to PID control or the like. The trigger for ending control by the management device 10 may be an operation of the input device 40 by an operator, or may be based on a judgment by AI.

[0056] Furthermore, in step St6, if the control of the plant 20 is to be continued without being terminated (St6: No), the processing of the management device 10 proceeds to step St7. In step St7, the management device 10 determines whether or not a predetermined reward function has been updated. Whether or not the reward function needs to be updated is determined by an operator of the plant 20 or an AI. Details of the reward function will be described later.

[0057] If the reward function has been updated in step St7 (St7: Yes), the management device 10 returns to step St3. In this case, the control law is calculated again based on the updated reward function (St3). If the reward function has not been updated in step St7 (St7: No), the management device 10 proceeds to step St8.

[0058] In step St8, the management device 10 determines whether or not the model of the plant 20 has been updated. Since operation data of the plant 20 is accumulated daily, it is desirable to periodically update the state transition probability matrix Mx1 (model) shown in Fig. 6. Note that whether or not the model needs to be updated is determined by the operators of the plant 20 or by AI.

[0059] If the model has been updated in step St8 (St8: Yes), the processing of the management device 10 returns to step St1. In this case, the estimation of the control characteristics (St2) and the calculation of the control law (St3) are performed again based on the updated model. On the other hand, if the model has not been updated in step St8 (St8: No), the processing of the management device 10 returns to step St4. In this case, the manipulated variable of the plant 20 is calculated based on the current control characteristics and control law. Note that the order of steps St7 and St8 in FIG. 5 may be reversed.

[0060] <Estimation of control characteristics> 11A and 11B are flowcharts relating to the estimation of control characteristics (see also FIG. 1 as appropriate). 11A and 11B show details of the process (estimation of control characteristics) of step St2 in Fig. 5. That is, the flowcharts shown in Fig. 11A and 11B show a series of processes when the management device 10 calculates the control characteristic matrix Mx2 (see Fig. 7) from the state transition probability matrix Mx1 (see Fig. 6) using the control characteristic estimation unit 14d. This series of processes is performed as the calculation of an evaluation function (action value function) in reinforcement learning when the i-th state is assumed to be the target state.

[0061] In this embodiment, when calculating the control characteristic matrix Mx2, the management device 10 calculates in advance the evaluation values ​​of the manipulated variables for each of a plurality of assumed target states. This simplifies the process of deriving the control law (the series of processes in FIG. 12), which will be described later, and increases the calculation speed.

[0062] First, in step St201, the management device 10 generates a counter i and assigns 1 to the value of this counter i. Note that the counter i is used to calculate the evaluation value of the control characteristic matrix Mx2 (see FIG. 7) and the optimal manipulated variable to the target state s i When calculating each time (i.e., each column in Figure 7), the target state s i Used to specify

[0063] In step St202, the management device 10 generates a predetermined vector Q and assigns 1 to the value of the i-th element of this vector Q. Here, the vector Q is a state s x The element values ​​of vector Q (i.e., the value) are the maximum value V max It is used to estimate (St208).

[0064] The number of elements of vector Q is equal to the number of states that the plant 20 can take. For example, if there are nine states s1 to s9 that the plant 20 can take, the number of elements of vector Q will be nine. Incidentally, the processing of step St202 is repeated as the value of i is incremented (St220), but in the first step St202, all elements except for the ith element of vector Q (the first element since i=1) are empty.

[0065] In step St203, the management device 10 generates a state list Zlist1 and sets the i-th (value of counter i) state s as a goal state that is an element of this state list Zlist1. i As will be described in detail later, the state list Zlist1 is a list for saving the target states whose values ​​are to be recalculated in this loop (the loop accompanying the increment of i in St220). In the first step St203, the only element contained in the state list Zlist1 is state s1.

[0066] In step St204, the management device 10 generates a counter j1 and assigns the value of this counter j1 to 1. The counter j1 is used to specify an element of the above-mentioned state list Zlist1. In step St205, the management device 10 generates a counter j2 and assigns the value of this counter j2 to 1. The counter j2 is used to specify elements of a status list Zlist2, which will be described next.

[0067] In step St206, the management device 10 generates a state list Zlist2 and also generates another state list Znxt. The state list Zlist2 is a list of the state s stored in the j1th position (value of the counter j1) of the state list Zlist1. x In other words, the list of all possible states to be transitioned to is the "transition destination state" shown in Figure 6. x The management device 10 specifies this state s x A list of "current states" that can transition to (those with a transition probability greater than 0) is saved in the state list Zlist2.

[0068] The state list Znxt in step St206 is a list for saving the target state whose value will be recalculated in the next loop when the value of i is incremented (St220). In the first processing of St206, the state list Znxt has no elements (empty state).

[0069] In step St207, the management device 10 sets the j2-th element (the value of the counter j2) of the state list Zlist2 to the state s j2 For example, if the transition destination state targeted by the state list Zlist2 is state s2 in FIG. 6, the states that may transition to state s2 with a predetermined operation amount a1 are states s1 and s5 as the current states. In this case, when the value of counter j2 is 1, state s1, which is the first of states s1 and s5, is designated as the current state.

[0070] Next, in step St208, the management device 10 calculates the maximum value V max where the maximum value V max is the state s specified in step St207. j2 It is the weighted average of the maximum value of the manipulated variable for all states that can be transitioned from.

[0071]

number

[0072] In addition, N included in formula (1) k is the total number of states that the plant 20 can take. For example, if there are nine states s1 to s9 that the plant 20 can take, N k The value of is 9. Also, γ included in formula (1) is a distance attenuation rate. Here, the "distance attenuation rate" is a preset constant (e.g., 0.9) that decreases the evaluation value as the number of times the controlled element is operated (i.e., the longer the time required) to reach a predetermined target state increases.

[0073] p in equation (1) is the state transition probability. j2 In this case, when the control element is operated with a predetermined control amount a, the state s k The probability of transitioning to is the transition probability p. The value of this transition probability p is acquired by referring to the state transition probability matrix Mx1 (see FIG. 6). Specifically, the management device 10 acquires the value of the transition probability when the operation amount a is set to maximize the transition probability in the j2th row and the kth column by referring to the state transition probability matrix Mx1 (see FIG. 6).

[0074] Q included in the formula (1) is the state s as an element of the vector Q set in the processing of step St202. k As mentioned above, in the first step St202, the value of the first element in the vector Q (i.e., Q(s1)) is set to 1. Incidentally, the maximum value V max The magnitude of the state s j2 The value will vary depending on the

[0075] Next, in step St209, the management device 10 calculates the state s stored in the j2-th vector Q (the value of the counter j2). j2 The value (i.e., value) of max The process of step St209 determines whether the state s newly calculated in step St208 is smaller than j2 Maximum value at V maxIn step St209 of FIG. 11A, the state s j2 The value of is simply written as "Q".

[0076] In step St209, the state s j2 The value of V is the maximum value max If it is smaller than (St209: Yes), the processing of the management device 10 proceeds to step St210. In step St210, the management device 10 calculates the state s stored in the j2-th vector Q (the value of the counter j2). j2 (shown as "Q" in Figure 11A) to the maximum value V max In other words, the management device 10 substitutes the value of the state s j2 Update the value of to the maximum value V max In this way, the state s j2 Furthermore, the management device 10 updates the value of the updated state s j2 The operation amount that gives the value of this state s j2 The value is stored in association with the value.

[0077] In step St211, the management device 10 adds the state s to the state list Znxt. j2 As mentioned above, the state list Znxt is a list for saving the target state whose value is to be updated in the next loop when the value of i is incremented (St220). j2 With the update of the value (St210), the maximum value V max Since the value of may change, the process of step St211 is performed. After the process of step St211 is performed, the process of the management device 10 proceeds to step St212.

[0078] In step St209, the state s of the vector Q j2 The value of V is the maximum value maxIf it is equal to or greater than this (St209: No), the management device 10 proceeds to step St212. In this case, the maximum value of the value calculated up to that point is the maximum value calculated as the calculation result of step St208. Vmax Since this is the case, state s j2 There is no particular need to update the values ​​of other states that may transition from.

[0079] In step St212, the management device 10 determines whether the value of the counter j2 is equal to or less than the size Nz2 (number of elements) of the status list Zlist2. If the value of the counter j2 is equal to or less than the size Nz2 of the status list Zlist2 (St212: Yes), the processing of the management device 10 proceeds to step St213.

[0080] In step St213, the management device 10 increments the value of the counter j2 and returns to the processing of step St207. max Among the list of states to be estimated, the maximum value V max If there is any state for which estimation has not yet been performed (Step 212: Yes), the management device 10 calculates the other states included in the state list Zlist2 by the maximum value V max is set as the estimation target (St213).

[0081] To give a specific example, when the transition destination state of the state list Zlist2 is state s2 in FIG. 6, the states that can transition to state s2 with a predetermined operation amount a1 are states s1 and s5 as the current state. When the value of counter j2 is 1, the maximum value V of the operation amount when state s1 is the current state is max When the value of the counter j2 is 2, the maximum value V of the manipulated variable when the state s5 is the current state is calculated. max In this way, the maximum value V of each state included in the state list Zlist2 is calculated. max are calculated sequentially (Step 208).

[0082] Furthermore, if the value of the counter j2 is greater than the size Nz2 (number of elements) of the status list Zlist2 in step St212 (St212: No), the processing of the management device 10 proceeds to step St214. In step St214, the management device 10 determines whether the value of the counter j1 is less than or equal to the size Nz1 (number of elements) of the state list Zlist1. If the value of the counter j1 is less than or equal to the size Nz1 of the state list Zlist1 (St214: Yes), the processing of the management device 10 proceeds to step St215. In other words, if there is any state in the state list Zlist1 (a list of states whose values ​​are to be recalculated) for which the value has not yet been recalculated, the processing of the management device 10 proceeds to step St215.

[0083] In step St215, the management device 10 increments the value of the counter j1 and returns to the processing of step St205. In short, if there is any one of the multiple goal states that are elements of the state list Zlist1 whose value has not yet been updated (St214: Yes), the management device 10 sequentially updates the values ​​of the other states included in the state list Zlist1. Also, if in step St214 the value of the counter j1 is greater than the size Nz1 (number of elements) of the state list Zlist1 (St214: No), the processing of the management device 10 proceeds to step St216.

[0084] Next, in step St216 of FIG. 11B, the management device 10 substitutes the state list Znxt into the state list Zlist1. In this way, by rewriting the elements of the state list Zlist1 to be updated, the elements included in the new state list Zlist1 are updated to the maximum value V max The calculation is set to be subject to recalculation.

[0085] In step St217, the management device 10 determines whether the size Nz1 (number of elements) of the state list Zlist1 is greater than 0. That is, the management device 10 determines whether any goal states that require recalculation of value remain as elements of the state list Zlist1. If the size Nz1 of the state list Zlist1 is greater than 0 (St217: Yes), the processing of the management device 10 returns to step St204 in FIG. 11A.

[0086] Furthermore, if the size Nz1 of the state list Zlist1 is equal to or less than 0 in step St217 (St217: No), the processing of the management device 10 proceeds to step St218. In this case, all updates related to the elements included in the state list Zlist1 have been completed, and the value of each element of the vector Q (i.e., the evaluation value of the i-th column of the control characteristic matrix Mx2) has been determined.

[0087] In step St218, the management device 10 assigns the elements included in the vector Q to the i-th column (value of counter i) in the evaluation values ​​of the control characteristic matrix Mx2 (see FIG. 7). Also, although omitted in FIG. 11B, the management device 10 assigns the manipulated variable a corresponding to the element (target state) included in the vector Q to the i-th column in the manipulated variables of the control characteristic matrix Mx2 (see FIG. 7).

[0088] In step St219, the management device 10 checks whether the value of the counter i is equal to the number of states N k Here, it is determined whether the number of states N k is the total number of states that the plant 20 can take. In step St219, the value of the counter i is set to the number of states N k If the result is equal to or less than the target state s (St219: No), the processing of the management device 10 proceeds to step St220 in Fig. 11A. In this case, in the control characteristic matrix Mx2 (see Fig. 7), the target state s i (i.e., the column in Figure 7) exists.

[0089] In step St220 of FIG. 11A, the management device 10 increments the value of the counter i and returns to the processing of step St202. In short, the management device 10 increments the value of the counter i in the control characteristic matrix Mx2 (see FIG. 7) in the column (target state s i ) to the next column.

[0090] In step St219, the value of the counter i is set to the number of states N k If it is greater than (St219: Yes), the management device 10 ends the series of processes related to the estimation of the control characteristics (St2 in FIG. 5) (END). In this case, the evaluation values ​​and optimal manipulated variables have been calculated for all columns (target states) of the control characteristics matrix Mx2 (see FIG. 7).

[0091] In this way, the control characteristic estimation unit 14d calculates an evaluation value of each manipulated variable when the plant 20 reaches a target candidate state, which is a candidate for the target state, from a plurality of states that the plant 20 can take, based on the state transition probability, which is the probability of a state transition when the controlled plant 20 is operated with a predetermined manipulated variable. The "target candidate state" mentioned above is a plurality of assumed candidate target states (see FIG. 7).

[0092] <Control law calculation> FIG. 12 is a flowchart relating to the calculation of the control law (see also FIG. 1 as appropriate). Fig. 12 shows details of the process (calculation of the control law) of step St3 in Fig. 5. That is, the flowchart in Fig. 12 shows a series of processes when the management device 10 calculates a predetermined control law from the reward function and the control characteristic matrix Mx2 (see Fig. 7) by the control law determiner 14e.

[0093] In step St301, the management device 10 calculates the product of each row of the matrix in the first depth layer of the control characteristic matrix Mx2 (see FIG. 7) and the vector of the reward function. Explaining this in more detail, the management device 10 first reads the control characteristic matrix Mx2 (see FIG. 7) stored in the control characteristic storage unit 12b, and also reads the reward function stored in the reward storage unit 12c. Here, the "reward function" is a function for providing a predetermined reward that serves as an incentive when a predetermined operation amount is selected.

[0094] FIG. 13 is an explanatory diagram regarding the reward function. In the example of FIG. 13, states s1 to s n The reward function is set in advance as a matrix in which the reward value is associated with the state s5 and state s7. For example, for states s5 and s7, which are the actual goal states, the element values ​​of the reward function are set to positive real numbers such as 1 and 1.1. n For non-goal states, such as the above, the value of the reward function element is set to 0.

[0095] In step St301 in Fig. 12, the product of each row of the matrix in the first depth layer of the control characteristic matrix Mx2 (see Fig. 7) and the reward function vector (the column vector of the "reward" part in Fig. 13) is calculated. This results in a matrix as shown in Fig. 14.

[0096] FIG. 14 is an explanatory diagram of the matrix X, which is the calculation result of the product of the control characteristic matrix and the reward function vector. The vertical axis of the matrix X shown in FIG. 14 represents the "current state" before a predetermined operation is performed on the plant 20. The horizontal axis of the matrix X represents a predetermined "target state" of the plant 20. By performing the processing of step St301 (see FIG. 12) described above, the element values ​​of each column that does not correspond to the actual target state among the multiple target states become 0, while the element values ​​of each column that corresponds to the target state become values ​​that are multiplied by the reward (for example, 1 or 1.1 times) from the element values ​​of the control characteristic matrix Mx2 (see FIG. 7). In the example of FIG. 14, of the evaluation values ​​of the first layer of the control characteristic matrix Mx2 (see FIG. 7), the evaluation values ​​of states s5 and s7 whose reward values ​​(see FIG. 13) are greater than 0 remain, and the remaining states s1 to s4, s6, s8 to s9 remain. n The values ​​of all elements are 0.

[0097] 12, the management device 10 acquires the maximum value of the element value (evaluation value) for each row of the matrix X. In the example of Fig. 14, the maximum value of the element value (evaluation value) in the matrix X when the current state is state s1 is 0.9.

[0098] Next, in step St303 of Fig. 12, the management device 10 acquires the operation amount corresponding to the maximum value of the element value (evaluation value). A vector having the operation amount acquired in this way as an element is set as the optimal operation amount of the second depth layer in the control characteristic matrix Mx2 (see Fig. 7).

[0099] FIG. 15 is an explanatory diagram showing the correspondence relationship between the maximum evaluation value of each state and the optimal operation amount. The first column in FIG. 15 shows the maximum evaluation value for each state, and the second column shows the optimal manipulated variable for each state. For example, the element value for state s1 is 0.9, and the optimal manipulated variable is a1. In this way, the control law determiner 14e determines the optimal manipulated variable for each state of the plant 20, which is the control target, as the control law, based on the evaluation values ​​of the manipulated variables. In other words, the control law determiner 14e determines the control law based on the reward function, which includes the reward when a predetermined manipulated variable is selected, and the evaluation values ​​of the manipulated variables.

[0100] During operation of the plant 20 (control target), the equipment control unit 14f may identify the actual state of the plant 20 based on information including the detection values ​​of the sensors 30 of the plant 20, identify the optimal operation amount for moving from this state to the target state from the control law, and output the optimal operation amount to the plant 20.

[0101] Furthermore, when the reward function is changed, the control law determination unit 14e may perform processing to update the control law based on the changed reward function while the plant 20 (control target) is in operation. Since the amount of calculation required to update the control law (processing similar to that in FIG. 12) is relatively small, the control law can be updated quickly even while the plant 20 is in operation.

[0102] <Display screen example> FIG. 16 shows an example of a display screen relating to the processing of the management device (also see FIG. 1 as appropriate). 16 shows an example in which the control characteristic matrix Mx2 is displayed on the display device 50. As shown in FIG. 16, the output control unit 14b causes the display device 50 to display the evaluation values ​​of the manipulated variables in association with combinations of a plurality of states (current states) that the plant 20 (control object) can take and a plurality of target states. The output control unit 14b also causes the display device 50 to display the optimal manipulated variables in association with combinations of a plurality of states (current states) that the plant 20 (control object) can take and a plurality of target states. This allows the operator to grasp the evaluation values ​​and the optimal manipulated variables at a glance.

[0103] In the example of FIG. 16, a third layer is added to the depth of the control characteristic matrix Mx2. This third layer is configured to show the state with the highest transition probability when the control terminal of the plant 20 is operated with the optimal control amount in the second layer. In this way, the output control unit 14b associates the state with the highest state transition probability when the plant 20 (control target) is operated with the optimal control amount with the optimal control amount and displays it on the display device 50. This makes it possible to present the relationship between the control amount and the state transition to the operator in an easy-to-understand manner.

[0104] In FIG. 16, the evaluation value of the first layer in the control characteristic matrix Mx2 is displayed, but it is also possible for the user to operate the input device 40 to display the value of the optimal operation amount of the second layer or the state of the third layer (the state with the highest transition probability).

[0105] Furthermore, display examples related to the processing of the management device 10 are not limited to the example in FIG. 16. For example, the transition of the state of the plant 20 as shown in FIG. 4 may be displayed on the display device 50. Furthermore, a state transition probability matrix Mx1 (see FIG. 6) may be displayed on the display device 50. That is, the output control unit 14b may cause the display device 50 to display the state transition probability matrix Mx1 (see FIG. 6), which is a matrix in which state transition probabilities are associated with combinations of a plurality of states (current states) that the plant 20 (control target) can take, destination states from each state, and manipulated variables of the plant 20. This allows the user to grasp at a glance the state transition probabilities corresponding to combinations of the current state, destination states, and manipulated variables.

[0106] Furthermore, the control law (see FIG. 8) for controlling the plant 20 may be displayed on the display device 50. That is, the output control unit 14b may display, as a control law, a correspondence between a plurality of states that the plant 20 (control target) can take and the optimal manipulated variable in each state on the display device 50. This allows the operator to understand what control law is currently set.

[0107] In addition, the display device 50 may display the state definition of the plant 20 (see FIG. 9) and the correspondence between the manipulated variable and the valve opening adjustment variable (see FIG. 10). Furthermore, the display device 50 may display the matrix X (see FIG. 14), which is the calculation result of the product of the control characteristic matrix Mx2 (see FIG. 7) and the reward function (see FIG. 13), as well as an explanatory diagram of the reward function (see FIG. 13). By viewing such data, the operator can confirm the entire process from information about the behavior of the controlled object (operation data and simulation data) to the calculation of the control law. Therefore, the operator can confirm the data that serves as the basis for setting the manipulated variable at each moment in the plant 20. Safety standards are particularly high in the field of plant control. As described above, the management device 10 uniquely derives the control law from information about the behavior of the controlled object, and by presenting the calculation process, the basis for the manipulated variable can be presented to a third party.

[0108] <Effects> According to this embodiment, the dimensions of the state transition probability matrix Mx1 (see FIG. 6) include not only the current state and the transition destination state but also the manipulated variable of the plant 20. This allows accurate simulation of the relationship between the manipulated variable of the plant 20 and the state transition, thereby improving the prediction accuracy of the model.

[0109] The management device 10 also generates a control characteristic matrix Mx2 (see FIG. 7) that includes evaluation values ​​of the manipulated variables in each state relative to the control target and optimal manipulated variables that maximize the evaluation values. Furthermore, a control law for bringing the plant 20 to the target state is derived based on the control characteristic matrix Mx2 and the reward function. Because the amount of calculation required to calculate such a control law (flowchart in FIG. 12) is relatively small, the management device 10 can quickly calculate the control law. Therefore, even while the plant 20 is operating, the management device 10 can update the control law as needed and quickly calculate the optimal manipulated variables.

[0110] Furthermore, even if the target state (control target) or predetermined parameters are changed during operation of the plant 20, there is no particular need to recreate the control characteristic matrix Mx2 (see FIG. 7), so the control law can be updated quickly in real time. As described above, the amount of calculation required to update the control law is relatively small, so the control law can be updated quickly even when the plant 20 is large in scale or the amount of data is large.

[0111] Furthermore, conventional model predictive control techniques tend to optimize control that avoids risks, making it difficult to maximize the performance of the controlled object. Conventional model predictive control techniques refer to, for example, techniques that have the function of quickly updating conventional control laws in real time. In contrast, in this embodiment, the manipulated variable that maximizes the evaluation value is identified in the state transition of the plant 20, assuming all possible combinations. This makes it possible to perform control that maximizes the performance of the plant 20.

[0112] Furthermore, even if the control characteristics of the plant 20 change due to aging of the plant 20 or changes in the surrounding environment, the management device 10 can respond appropriately by appropriately updating the state transition probability matrix Mx1 (see Figure 6) and the control characteristic matrix Mx2 (see Figure 7).

[0113] <<Variations>> The management device 10 and the like according to the present disclosure have been described above in the embodiments, but the present disclosure is not limited to these descriptions and can be modified in various ways. For example, in the embodiment, the case where one type of sensor 30 is used to acquire the state of the plant 20 has been described, but this is not limiting. That is, multiple types of sensors (or multiple sensors of the same type) may be provided. In this case, the state of the controlled object is identified in association with a combination of ranges of detected values ​​of the multiple sensors. Furthermore, the number of controlled objects is not limited to one, and the embodiment can also be applied to cases where there are multiple controlled objects.

[0114] In the embodiment, the management device 10 (see FIG. 1) includes the communication unit 13 (see FIG. 1), but the present invention is not limited to this. That is, the communication unit 13 may be omitted as appropriate from the configuration in FIG. 1, and the manipulated variable, which is the calculation result of the equipment control unit 14f, may be transmitted to the plant 20 via the output control unit 14b.

[0115] In the embodiment, the device control unit 14f (see FIG. 1) determines the manipulated variable of the plant 20, but the present invention is not limited to this. For example, the calculation unit 14 may select an optimal control law from among the control laws stored in the control law storage unit 12d, and then transmit the optimal control law to the plant 20.

[0116] In the embodiment, the state transition probability matrix Mx1 (see FIG. 6) is used to estimate the behavior of the controlled object, but the present invention is not limited to this. For example, a neural network may be used, or information corresponding to the state transition probability matrix Mx1 may be expressed by a predetermined mathematical formula.

[0117] Furthermore, in the embodiment, a case where the action value function (Equation (1)) is calculated using a processing flow based on dynamic programming has been described, but this is not limiting, and a predetermined action value function may be calculated using another calculation formula or method. In addition, in the embodiment, the "evaluation value" is defined as the expected value when the controlled object moves from a predetermined state to the target candidate state, but this is not limited to this. In other words, the "evaluation value" may be defined based on a predetermined index different from the expected value.

[0118] Furthermore, in the embodiment, a case has been described in which a signal (control signal) indicating the manipulated variable of the control terminal of the plant 20 is transmitted from the management device 10 to the plant 20, but this is not limiting. For example, a signal indicating the manipulated variable of the control terminal of the plant 20 may be transmitted to a predetermined server or a user's terminal device (i.e., provided as data indicating the manipulated variable). The process (management method) executed by the management device 10 may be executed as a predetermined program on a computer. The program may be provided via a communication line, or may be written to a recording medium such as a CD-ROM and distributed.

[0119] Furthermore, the present disclosure is not limited to the embodiments and includes various modifications. For example, the embodiments have been described in detail to clearly explain the present disclosure, and the present disclosure is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the embodiments with other configurations.

[0120] Furthermore, the above-mentioned configurations, functions, processing units, processing means, etc. may be partly or entirely implemented in hardware, for example, by designing them as integrated circuits. Furthermore, the above-mentioned configurations, functions, etc. may be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.

[0121] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]

[0122] 10 Management device 11 Data Acquisition Section 12 Storage section 12a Operation memory section 12b Control characteristics memory section 12c Reward storage 12d Control law memory section 13 Communications Department 14 Arithmetic section 14a Input control section 14b Output control section 14c Model Estimation Section 14d Control characteristic estimation section 14e Control law determination section 14f Equipment control unit 20 Plant (control object) 30 sensors 40 Input Devices 50 Display device 60 Cloud 70 External storage media Mk1 state transition probability matrix Mk2 control characteristic matrix St2 step (control characteristic estimation step) St3 step (control law determination step)

Claims

1. a control characteristic estimation unit that calculates an evaluation value of each manipulated variable when a target state candidate is reached from a plurality of states that the control object can take, based on a state transition probability that is a probability of a state transition when the control object is manipulated by a predetermined manipulated variable; a control law determination unit that determines, based on the evaluation value, an optimal manipulated variable in each state of the controlled object as a control law.

2. The evaluation value is an expected value when the controlled object reaches the target candidate state from a predetermined state. The management device according to claim 1 .

3. The control law determination unit determines the control law based on a reward function including a reward when a predetermined manipulated variable is selected and the evaluation value. The management device according to claim 1 .

4. When the reward function is changed, the control law determination unit performs a process of updating the control law based on the changed reward function while the controlled object is operating.

4. The management device according to claim 3, wherein:

5. The system is provided with an equipment control unit that, while the controlled object is in operation, identifies the actual state of the controlled object based on information including the detected values ​​of the sensors of the controlled object, identifies the optimal manipulated variable for moving from the actual state to the target state from the control law, and outputs the optimal manipulated variable to the controlled object. The management device according to claim 1 .

6. An output control unit is provided that displays, on a display device, a state transition probability matrix, which is a matrix that associates the state transition probabilities with combinations of a plurality of states that the control object can take, transition destination states from each of the states, and manipulated variables of the control object. The management device according to claim 1 .

7. An output control unit is provided that displays the evaluation values ​​on a display device in association with combinations of a plurality of states that the control object can take and a plurality of target states. The management device according to claim 1 .

8. An output control unit is provided that displays the optimal manipulated variable on a display device in association with a combination of a plurality of states that the controlled object can take and a plurality of target states. The management device according to claim 1 .

9. The output control unit causes the display device to display a state having the highest state transition probability when the controlled object is operated with the optimal operation amount in association with the optimal operation amount. The management device according to claim 8 .

10. An output control unit is provided that displays, on a display device, a correspondence between a plurality of states that the controlled object can take and the optimal manipulated variable in each of the states as the control law. The management device according to claim 1 .

11. a control characteristic estimation step of calculating an evaluation value of each manipulated variable when the controlled object reaches a target candidate state, which is a candidate for the target state, from a plurality of states that the controlled object can take, based on a state transition probability, which is a probability of a state transition when the controlled object is manipulated by a predetermined manipulated variable; and a control law determination step of determining, as a control law, an optimal manipulated variable in each state of the controlled object based on the evaluation value.

12. a control characteristic estimator that calculates an evaluation value of each manipulated variable when a plant, which is a control target, reaches a target candidate state, which is a candidate for a target state, from a plurality of states that the plant can take, based on a state transition probability that is a probability of a state transition when the plant is operated with a predetermined manipulated variable; a control law determination unit that determines an optimal operation amount in each state of the plant as a control law based on the evaluation value; a control system including an equipment control unit that, while the plant is in operation, identifies an actual state of the plant based on information including detection values ​​of sensors in the plant, identifies the optimal manipulated variable for moving from that state to the target state from the control law, and outputs the optimal manipulated variable to the plant.

Citation Information

Patent Citations

  • Method for improving performance of method for computationally predicting future state of target object, driver assistance system, vehicle including such driver assistance system and corresponding program storage medium and program

    JP2016212872A