Management device, management method, and control system

The management device and system address the challenge of generating effective control information by using a control characteristic estimation unit and control law determination unit to optimize control strategies for complex industrial systems, enhancing their management and operation.

WO2025177607A1PCT designated stage Publication Date: 2025-08-28HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/033115
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2024-09-17
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing model predictive control technologies for controlling control targets lack the ability to appropriately generate information for effective control, particularly in managing complex systems like power plants, chemical plants, and other industrial facilities.

Method used

A management device and system that includes a control characteristic estimation unit to calculate evaluation values based on state transition probabilities and a control law determination unit to determine optimal operation amounts, utilizing a state transition probability matrix and control characteristic matrix to derive an optimal control law.

Benefits of technology

Enables appropriate generation of information for controlling control targets, improving the management and operation of complex industrial systems by optimizing control strategies based on predicted future states and state transitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024033115_28082025_PF_FP_ABST
    Figure JP2024033115_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a management device, a management method, and a control system that appropriately generate information used for control of a control target. A management device (10) includes: a control characteristic estimation unit (14d) that, on the basis of a state transition probability that is the probability of a state transition in a case where a control target is operated at a predetermined operation amount, calculates an evaluation value for operation amounts when reaching a target candidate state which is a candidate for the target state, from a plurality of states that can be taken by the control target; and a control rule determination unit (14e) that, on the basis of the evaluation value, determines, as a control rule, an optimal operation amount in each state of the control target.
Need to check novelty before this filing date? Find Prior Art

Description

Management device, management method, and control system

[0001] The present disclosure relates to a management device, a management method, and a control system.

[0002] Regarding model predictive control of a controlled object, for example, a technique described in Patent Document 1 is known. That is, Patent Document 1 describes "a method for improving the performance of a method for predicting the future state of a target object by calculation, in which the quality of the prediction is evaluated by comparing the observed actual state with the prediction result (20)."

[0003] Japanese Patent Application Laid-Open No. 2016-212872

[0004] The technology described in Patent Document 1 predicts the future state of the control target (target object) based on a model that simulates the behavior of the control target, and calculates the optimal control method from the predicted future state, but there is room for improvement in terms of appropriately controlling the control target.

[0005] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide a management device, a management method, and a control system that appropriately generate information used to control a control target.

[0006] In order to solve the above-mentioned problems, the management device according to the present disclosure includes a control characteristic estimation unit that calculates an evaluation value of each operation amount when the controlled object reaches a target candidate state, which is a candidate for the target state, from multiple states that the controlled object can take, based on a state transition probability, which is the probability of a state transition when the controlled object is operated with a predetermined operation amount, and a control law determination unit that determines, based on the evaluation value, an optimal operation amount for each state of the controlled object as a control law.

[0007] According to the present invention, it is possible to provide a management device, a management method, and a control system that appropriately generate information used for controlling a control target.

[0008] 1 is a configuration diagram of a control system including a management device according to an embodiment. FIG. 1 is a diagram illustrating an example of a hardware configuration of a management device according to an embodiment. FIG. 2 is an explanatory diagram illustrating a specific example of a plant that is a control target of the management device according to an embodiment. FIG. 3 is an explanatory diagram illustrating a change in steam temperature of the plant in the example of FIG. 3 in the management device according to an embodiment. FIG. 4 is a flowchart illustrating a processing flow of the management device according to an embodiment. FIG. 5 is an explanatory diagram of a state transition probability matrix in the management device according to an embodiment. FIG. 6 is an explanatory diagram of a control characteristic matrix in the management device according to an embodiment. FIG. 7 is an explanatory diagram of a control law in the management device according to an embodiment. FIG. 8 is an explanatory diagram of a plant state definition in the management device according to an embodiment. FIG. 9 is an explanatory diagram illustrating a correspondence relationship between an operation amount and a valve opening adjustment amount in the management device according to an embodiment. FIG. 10 is a flowchart illustrating estimation of control characteristics in the management device according to an embodiment. FIG. 11 is a flowchart illustrating estimation of control characteristics in the management device according to an embodiment. FIG. 12 is a flowchart illustrating calculation of a control law in the management device according to an embodiment. FIG. 13 is an explanatory diagram of a reward function in the management device according to an embodiment. FIG. 14 is an explanatory diagram of a matrix that is a calculation result of the product of a control characteristic matrix and a reward function vector in the management device according to an embodiment. FIG. 15 is an explanatory diagram illustrating a correspondence relationship between an optimal operation amount and a maximum evaluation value of each state in the management device according to an embodiment. FIG. 16 is an example of a display screen relating to processing of the management device according to an embodiment.

[0009] <<Embodiment>> In the following, as an example, a case will be described in which the management object (control object) of the management device 10 (see FIG. 1) is a predetermined plant 20 (see FIG. 1). Examples of such plants 20 include power plants, chemical plants, oil refineries, steel plants, food plants, pharmaceutical plants, and water treatment plants. Note that the management object (control object) of the management device 10 (see FIG. 1) is not limited to plants, and may also be a robot, a vehicle, a ship, an aircraft, or the like.

[0010] <Configuration of Control System> Fig. 1 is a configuration diagram of a control system 100 including a management device 10 according to an embodiment. The control system 100 shown in Fig. 1 is a system in which the management device 10 controls equipment in a plant 20. As shown in Fig. 1, the control system 100 includes the management device 10, a sensor 30, an input device 40, a display device 50, and a cloud 60, which are connected in a predetermined manner via wire or wirelessly. Below, a brief description will be given of the plant 20, which is the control target of the management device 10, as well as the sensor 30, the input device 40, the display device 50, etc., and then a detailed description will be given of the management device 10.

[0011] The plant 20 is a facility provided with control elements such as valves and heaters, such as a power plant. The sensor 30 acquires predetermined detected values ​​and environmental information from the equipment of the plant 20. As such a sensor 30, for example, a camera, a distance sensor, a radar, a force sensor, a temperature sensor, an angle sensor, a flow rate sensor, a pressure sensor, a voltage sensor, a current sensor, or the like may be used as appropriate.

[0012] The sensor 30 may be installed inside the plant 20 or outside the plant 20. An example of a sensor 30 installed outside the plant 20 is an outside air temperature sensor that detects the temperature of the environment surrounding the plant 20. In addition, the sensor 30 also includes equipment for acquiring data such as power demand related to the control of the plant 20 from outside the plant 20.

[0013] The input device 40 is used when the operator M1 inputs predetermined management information and commands to the management device 10. Examples of such input device 40 include a keyboard, a mouse, and a joystick. The display device 50 is used when presenting the calculation results of the management device 10 to the operator M1. Examples of such display device 50 include a liquid crystal display. Note that a touch panel type mobile terminal that combines the functions of the input device 40 and the display device 50, such as a smartphone or tablet, may also be used. The cloud 60 has a server (not shown) that communicates in a predetermined manner with the input device 40, the display device 50, and the management device 10 via a network. The processing results of the server are provided to the management device 10 via the network and are also displayed appropriately on the display device 50.

[0014] <Configuration of Management Device> The management device 10 is a device that controls and manages the plant 20. A computer such as a personal computer, a tablet, or a smartphone may be used as such a management device 10. The management device 10 may also be configured by connecting multiple computers in a predetermined manner via a communication line or a network. For example, the functions of the management device 10 may be distributed across multiple computers such as a cloud server or an edge server.

[0015] 1, the management device 10 includes a data acquisition unit 11, a storage unit 12, a communication unit 13, and a calculation unit 14. The data acquisition unit 11, the storage unit 12, the communication unit 13, and the calculation unit 14 are connected in a predetermined manner via an internal bus (not shown).

[0016] Fig. 2 is a diagram illustrating an example of the hardware configuration of the management device 10. As shown in Fig. 2, the management device 10 includes, as its hardware configuration, a processor 10a, a RAM 10b (Random Access Memory), a ROM 10c (Read Only Memory), a HDD 10d (Hard Disk Drive), a communication interface 10e, an input / output interface 10g, and a media interface 10h, which are connected in a predetermined manner via an internal bus 10i.

[0017] 2 is hardware of the calculation unit 14 (see FIG. 1) of the management device 10. RAM 10b, ROM 10c, and HDD 10d are hardware of the storage unit 12 (see FIG. 1) of the management device 10. The processor 10a reads out a predetermined program stored in the ROM 10c or the HDD 10d and loads it into the RAM 10b, thereby executing a predetermined process.

[0018] 2 performs predetermined communication with the devices and sensors 30 of the plant 20 and the cloud 60. The input / output interface 10g inputs data from the input device 40 and outputs data to the display device 50. The communication interface 10e and the input / output interface 10g function as the communication unit 13 of the management device 10 (see FIG. 1).

[0019] The media interface 10h is connected to an external recording medium 70 such as an optical disk 70a or a USB memory 70b as appropriate, and functions as the data acquisition unit 11 (see FIG. 1) of the management device 10. As such a media interface 10h, an optical disk drive that plays the optical disk 70a or a USB interface that reads the USB memory 70b may be used. Note that "USB" is a registered trademark.

[0020] Returning to Fig. 1, the explanation will be continued. The data acquisition unit 11 shown in Fig. 1 acquires predetermined data from an external recording medium 70 such as an optical disk 70a or a USB memory 70b. Various data are stored in the storage unit 12. As shown in Fig. 1, the storage unit 12 includes an action storage unit 12a, a control characteristic storage unit 12b, a reward storage unit 12c, and a control law storage unit 12d. Details of these will be described later, but in summary, they have the following functions.

[0021] That is, the operation storage unit 12a stores the calculation results of the model estimation unit 14c as data indicating the operation of the plant 20. Note that data (data specifying a model of the plant 20) acquired from the cloud 60 or the external recording medium 70 may be stored in the operation storage unit 12a. The format of the data stored in the operation storage unit 12a may be a model or simulator that simulates the operation of the plant 20, or may be a flow chart, table, graph, or other predetermined calculation formula that indicates the behavior of the plant 20.

[0022] The control characteristic storage unit 12 b stores the calculation results of the control characteristic estimation unit 14 d. That is, data such as evaluation values ​​of manipulated variables in each state with respect to the control target of the plant 20 is stored in the control characteristic storage unit 12 b in the form of a table, graph, or calculation formula.

[0023] The reward memory unit 12c stores a predetermined reward function representing the control target in the form of a formula, table, vector, or matrix. Details of the reward function will be described later. The control law memory unit 12d stores a control law that is the calculation result of the control law determination unit 14e. Here, the "control measure" refers to a rule (law) for making the manipulated variable of the controlled object follow a predetermined target value in order to achieve the control objective. For example, if the control objective is to stabilize the steam temperature in a thermal power plant, the rule (law) for controlling the temperature control valve to bring the steam temperature closer to the predetermined target value is set as the control law.

[0024] The communication unit 13 performs predetermined communication with the plant 20, the sensor 30, the input device 40, the display device 50, and the cloud 60. The communication unit 13 is connected to the storage unit 12 and the calculation unit 14 via an internal bus. Data acquired from the plant 2, etc. via the communication unit 13 is stored in the storage unit 12, and calculation results of the calculation unit 14 are transmitted to the plant 20, etc. via the communication unit 13.

[0025] The calculation unit 14 generates predetermined data related to autonomous control of the plant 20. At least a part of the processing of the calculation unit 14 may be performed by AI (artificial intelligence). As shown in FIG. 1 , the calculation unit 14 includes an input control unit 14a, an output control unit 14b, a model estimating unit 14c, a control characteristic estimating unit 14d, a control law determining unit 14e, and an equipment control unit 14f. Of the functional units of the calculation unit 14, the functions of the model estimating unit 14c, the control characteristic estimating unit 14d, and the control law determining unit 14e may be performed by the cloud 60.

[0026] The details of each functional unit of the calculation unit 14 will be described later, but in summary, it has the following functions: The input control unit 14a transfers data input via the data acquisition unit 11 or the communication unit 13 to the storage unit 12 or an appropriate functional unit of the calculation unit 14.

[0027] The output control unit 14b outputs data from the storage unit 12 and the calculation unit 14 to the plant 20, the display device 50, the cloud 60, or the like via the communication unit 13. The output control unit 14b also generates image information for displaying data related to the processing of the management device 10 in a predetermined manner on the display device 50. Note that a predetermined computer (not shown) connected to the display device 50 may generate the image information.

[0028] The model estimation unit 14c estimates a model for estimating the operation of the plant 20. The operation of the plant 20 may be represented by a predetermined model or simulator, or may be represented in the form of a flow chart, a table, a graph, or a calculation formula. The estimation result of the model estimation unit 14c is stored in the operation storage unit 12a.

[0029] The control characteristic estimator 14d estimates an evaluation value of the manipulated variable in each state with respect to the control target of the plant 20. In addition to the evaluation value of the manipulated variable, the control characteristic estimator 14d may also calculate a state transition of the plant 20 associated with the manipulated variable and the probability of that state transition. The evaluation value, etc., which is the calculation result of the control characteristic estimator 14d, is expressed in the form of a table, graph, or calculation formula. The estimation result of the control characteristic estimator 14d is stored in the control characteristic storage unit 12b.

[0030] The control law determination unit 14e determines an optimal control law for controlling the plant 20, based on the data in the control characteristic storage unit 12b and the reward storage unit 12c. The data of the control law determined by the control law determination unit 14e is stored in the control law storage unit 12d.

[0031] The equipment control unit 14f calculates the manipulated variable of the equipment of the plant 20 based on the optimal control law stored in the control law storage unit 12e. The manipulated variable calculated by the equipment control unit 14f is transmitted to the plant 20 via the output control unit 14b and the communication unit 13 in this order. As a result, the equipment of the plant 20 is controlled with the predetermined manipulated variable.

[0032] Fig. 3 is an explanatory diagram showing a specific example of a plant 20 that is a control target of the management device. In the example of Fig. 3, the plant 20 is a thermal power plant. In the plant 20 configured as shown in Fig. 3, a boiler drum 21, a first superheater 22, a second superheater 23, and a steam turbine 24 are connected in this order via steam piping.

[0033] The steam in the boiler drum 21 is superheated in the first superheater 22 and then further superheated in the downstream second superheater 23. The steam thus superheated and heated in the second superheater 23 is guided to the steam turbine 24. The steam turbine 24 is rotated by the wind pressure of the steam, and the rotor of a generator (not shown) rotates integrally with the steam turbine 24, thereby generating electricity.

[0034] As shown in Fig. 3, the downstream end of a spray pipe 25 is connected to the steam pipe between the first superheater 22 and the second superheater 23. A nozzle-shaped spray 26 is provided at the downstream end of the spray pipe 25. A control valve 27 is provided in the spray pipe 25 to adjust the amount of compressed water injected. The compressed water injected via the spray 26 mixes with the steam, thereby lowering the temperature of the steam. Note that the greater the amount of compressed water injected via the spray 26 (i.e., the greater the opening of the control valve 27), the greater the amount of decrease in the temperature of the steam.

[0035] A temperature sensor 28 is provided in the steam pipe between the second superheater 23 and the steam turbine 24. The opening of the control valve 27 is adjusted so that the detected value (steam temperature) of the temperature sensor 28 approaches a predetermined target temperature. The temperature sensor 28 corresponds to the sensor 30 (see FIG. 1).

[0036] 4 is an explanatory diagram showing the change in steam temperature of the plant in the example of FIG. 3. The horizontal axis of FIG. 4 represents the elapsed time from a predetermined time, and the vertical axis represents the steam temperature (the temperature at the installation location of the temperature sensor 28). 1 ~s 6 indicates a state associated with the level of the steam temperature. Also, the solid arrows and dashed arrows shown in Fig. 4 each indicate a state transition based on a predetermined mathematical model.

[0037] In the example of FIG. 4, the target steam temperature is set to 500°C, and the target temperature change rate for the steam temperature is set to approximately 0°C / control step. The initial temperature of the steam is 495°C, and the initial temperature change rate is set to approximately 2°C / control step. Here, the "control step" refers to the cycle at which the manipulated variable is updated based on a predetermined control law. Note that the "control step" may be set separately from the control cycle of the adjustment valve 27 (see FIG. 3) based on PID control (Proportional-Integral-Differential Controller).

[0038] In the control pattern shown by the solid arrow in FIG. 4, as the control steps progress, the state s 1 →State S 2 →State S 3 →State S 4 →State S 4 →State S 4 The transition is as follows. 4 corresponds to a target steam temperature of 500°C.

[0039] <Processing of Management Device> Fig. 5 is a flowchart showing the flow of processing by the management device (see also Fig. 1 as appropriate). The series of processing shown in Fig. 5 is started by a user operation via the input device 40. For example, when operation of the plant 20 is started or when control is switched from predetermined PID control to control by the management device 10, the series of processing shown in Fig. 5 is started by a user operation via the input device 40.

[0040] In step St1 of Fig. 5, the management device 10 estimates a model of the plant 20 using the model estimation unit 14c. Here, the "model" is a mathematical model used to simulate the behavior of the plant 20. To specifically explain the processing of step St1, the management device 10 first acquires operation data and simulation data of the plant 20 via the data acquisition unit 11 and the communication unit 13. The operation data is mainly data accumulated during past operations of the plant 20.

[0041] For example, the operating data may include information that when a predetermined plant 20 is in state s1 (see FIG. 4), a transition to state s2 occurs when the regulating valve 27 (see FIG. 3) is set to a predetermined opening. Also, predetermined simulation data may be used together with (or instead of) the operating data. For example, simulation data may be used when the amount of data obtained from past operating data alone is insufficient. In this embodiment, in the process of step St1, the model estimation unit 14c estimates a model of the plant 20 in the form of a state transition probability matrix Mx1 (see FIG. 6).

[0042] FIG. 6 is an explanatory diagram of the state transition probability matrix Mx1 (see also FIG. 1 as appropriate). Note that FIG. 6 three-dimensionally illustrates the state transition probability matrix Mx1, which includes a "current state," a "transition destination state," and an "operation amount." The vertical axis of the state transition probability matrix Mx1 shown in FIG. 6 represents the "current state" before a predetermined operation is performed on the plant 20. Note that in FIG. 6, the vertical axis represents the direction in which rows are arranged, from top to bottom being the first row, the second row, the third row, and so on. The horizontal axis represents the direction in which columns are arranged, from left to right being the first column, the second column, the third column, and so on. Note that the term "current state" is not limited to meaning the current state in real time in the plant 20, but is used to mean one step before the transition destination state in the control period.

[0043] The horizontal axis of the state transition probability matrix Mx1 represents the "destination state" after a predetermined operation is performed on the plant 20. The depth axis of the state transition probability matrix Mx1 represents the manipulated variable at a predetermined manipulated variable of the plant 20. For example, the valve opening adjustment variable of the adjustment valve 27 shown in FIG. 3 corresponds to the "manipulated variable" in FIG. 6.

[0044] In the state transition probability matrix Mx1, a value of a state transition probability is associated with each of the elements specified by the current state, the transition destination state, and the operation amount. Here, the "state transition probability" is the probability that a predetermined transition destination state will be reached when control is performed from the current state with a predetermined operation amount. For example, in the first row and second column of the state transition probability matrix Mx1, the depth is the operation amount a 1 The value of 0.9 is associated with the element of 1 In this case, the operation amount a 1 When the control terminal of the plant 20 is controlled with the 2 This means that the transition to

[0045] Such a state transition probability matrix Mx1 is generated by the management device 10 (or the cloud 60) performing statistical processing using a huge amount of past operation data of the plant 20. The data of the state transition probability matrix Mx1 estimated as a “model” in step St1 of FIG. 5 is stored in the operation storage unit 12a.

[0046] Next, in step St2 of Fig. 5, the management device 10 causes the control characteristic estimator 14d to estimate the control characteristics of the plant 20 (control characteristic estimation step). In this embodiment, the control characteristic estimator 14d estimates the control characteristics of the plant 20 based on the state transition probability matrix Mx1 (see Fig. 6) stored in the operation storage unit 12a. In this embodiment, in the process of step St2 of Fig. 5, the control characteristic estimator 14d represents the evaluation values ​​of each manipulated variable (i.e., the control characteristics of the plant 20) in the form of a control characteristic matrix Mx2 (see Fig. 7).

[0047] Fig. 7 is an explanatory diagram of the control characteristic matrix Mx2 (see also Fig. 1 as appropriate). The vertical axis of the control characteristic matrix Mx2 shown in Fig. 7 represents the "current state" before a predetermined operation is performed on the plant 20. The horizontal axis of the control characteristic matrix Mx2 represents the "target state" indicating a target state of the plant 20. The values ​​of the elements in the first layer of the depth of the control characteristic matrix Mx2 are "evaluation values" associated with combinations of the current state and the target state. Here, the "evaluation value" is an expected value when the plant 20, which is the control target, reaches a target candidate state (a candidate for the target state) from a predetermined state.

[0048] The element values ​​of the second layer in the depth direction of the control characteristic matrix Mx2 are "optimal operation amounts" associated with the combination of the current state and the target state. Here, the "optimal operation amount" means the operation amount having the maximum evaluation value. For example, the first layer in the depth direction, in the first row and second column of the control characteristic matrix Mx2, is associated with a value of 1.0 as the evaluation value, and the second layer is associated with a value of a 2 This corresponds to the state s 2 When the target state is set as 1 The optimal operation amount at a 2 and this operation amount a 2 The evaluation value of this control characteristic matrix Mx2 is 1.0. This control characteristic matrix Mx2 is created based on the state transition probability matrix Mx1 (see FIG. 6), and details of this process will be described later.

[0049] Returning to FIG. 5 , the explanation will be continued. In step St3, the management device 10 calculates a control law by the control law determination unit 14e (control law determination step). As described above, a "control measure" is a rule for making an manipulated variable follow a predetermined target value. To explain the processing of step St3 in more detail, the control law determination unit 14e derives an optimal control law based on the control characteristic matrix Mx2 (see FIG. 7 ) stored in the control characteristic storage unit 12b and the reward function stored in the reward storage unit 12c. The control law derived in this manner is stored in the control law storage unit 12d. Details of the derivation of the control law based on the control characteristic matrix Mx2 and the reward function will be described later.

[0050] 8 is an explanatory diagram of the control law. In the example of FIG. 8, the control law is expressed in a table format in which the state of the plant 20 and the optimal manipulated variable are associated with each other. For example, when the plant 20 is in a state s 1 If so, the manipulated variable a 2 It is shown as a control law that controlling a specified control element with the above is optimal (highest evaluation value).

[0051] Next, in step St4 of Fig. 5, the management device 10 measures the state of the plant 20. Specifically, the management device 10 acquires the detected values ​​of the sensor 30 and the like via the communication unit 13. Here, an example of a definition of the state of the plant 20 (hereinafter referred to as a state definition) will be described with reference to Fig. 9.

[0052] 9 is an explanatory diagram regarding the definition of the plant state. In the example of FIG. 9, the state s is defined by a combination of the "temperature" of the steam detected by the temperature sensor 28 (see FIG. 3) and the "rate of change" (speed of change) of the steam temperature. 1 ~s nSpecifically, the steam temperature is divided into a range of 494.5°C to 550.5°C in 1°C increments. A range of the steam temperature change rate is set in association with each range of the steam temperature. The "step" included in the unit of the steam temperature change rate is the control period (for example, 1 minute) of the control valve 27 (see FIG. 3). In the example of FIG. 9, the state when the steam temperature is 494.5°C to 495.5°C and the steam temperature change rate is in the range of 1.5°C / step to 2.5°C / step is called state s. 1 Similarly, other states s 2 , s 3 , ..., s n These state definitions are set in advance by a user through an operation via the input device 40 (see FIG. 1).

[0053] 5 again. In step St5, the management device 10 calculates the manipulated variable of the manipulated variable of the plant 20 by the equipment control unit 14f. That is, the equipment control unit 14f calculates the manipulated variable of the manipulated variable of the plant 20 based on the optimal control law stored in the control law storage unit 12e. For example, when the plant 20 is in state s 1 If so, the management device 10 calculates the optimum manipulated variable a based on a predetermined control law (see FIG. 8). 2 Identify.

[0054] 10 is an explanatory diagram showing the correspondence relationship between the operation amount and the valve opening adjustment amount. In the example of FIG. 10, 2 corresponds to the operation of setting the valve opening adjustment amount of the adjustment valve 207 to 0% (i.e., maintaining the valve opening). 1 , a 3 are associated with the valve opening adjustment amounts of -1% and +1%, in that order. Such correspondence relationships are set in advance by the user operating the input device 40 (see FIG. 1). Note that the correspondence relationships shown in FIG. 10 are merely an example and are not limiting.

[0055] Next, in step St6 of Fig. 5, the management device 10 determines whether or not to end control of the plant 20. If control of the plant 20 is to be ended in step St6 (St6: Yes), the management device 10 ends the series of processes (END). For example, control by the management device 10 is ended when operation of the plant 20 is to be ended, when operation is to be suspended for maintenance, or when control by the management device 10 is to be switched to PID control or the like. The trigger for ending control by the management device 10 may be an operation of the input device 40 by an operator, or may be based on a judgment by AI.

[0056] Furthermore, in step St6, if the control of the plant 20 is to be continued without being terminated (St6: No), the processing of the management device 10 proceeds to step St7. In step St7, the management device 10 determines whether or not a predetermined reward function has been updated. Note that whether or not the reward function needs to be updated is determined by an operator of the plant 20 or an AI. The details of the reward function will be described later.

[0057] If the reward function has been updated in step St7 (St7: Yes), the management device 10 returns to step St3. In this case, the control law is calculated again based on the updated reward function (St3). If the reward function has not been updated in step St7 (St7: No), the management device 10 proceeds to step St8.

[0058] In step St8, the management device 10 determines whether or not the model of the plant 20 has been updated. Because operational data of the plant 20 is accumulated daily, it is desirable to periodically update the state transition probability matrix Mx1 (model) shown in Fig. 6. Note that whether or not the model needs to be updated is determined by an operator of the plant 20 or an AI.

[0059] If the model has been updated in step St8 (St8: Yes), the processing of the management device 10 returns to step St1. In this case, the estimation of the control characteristics (St2) and the calculation of the control law (St3) are performed again based on the updated model. On the other hand, if the model has not been updated in step St8 (St8: No), the processing of the management device 10 returns to step St4. In this case, the manipulated variable of the plant 20 is calculated based on the current control characteristics and control law. Note that the order of steps St7 and St8 in FIG. 5 may be reversed.

[0060] <Estimation of Control Characteristics> Figures 11A and 11B are flowcharts related to the estimation of control characteristics (see also Figure 1 as appropriate). The flowcharts in Figures 11A and 11B show details of the process of step St2 in Figure 5 (estimation of control characteristics). That is, the flowcharts in Figures 11A and 11B show a series of processes when the management device 10 calculates the control characteristic matrix Mx2 (see Figure 7) from the state transition probability matrix Mx1 (see Figure 6) using the control characteristic estimation unit 14d. This series of processes is performed as the calculation of an evaluation function (action value function) in reinforcement learning when the i-th state is assumed to be the target state.

[0061] In this embodiment, when calculating the control characteristic matrix Mx2, the management device 10 calculates in advance the evaluation values ​​of the manipulated variables for each of a plurality of assumed target states, thereby simplifying the process of deriving the control law (the series of processes in FIG. 12 ), which will be described later, and increasing the calculation speed.

[0062] First, in step St201, the management device 10 generates a counter i and assigns 1 to the value of the counter i. The counter i is used to calculate the evaluation value of the control characteristic matrix Mx2 (see FIG. 7) and the optimal manipulated variable to the target state s i When calculating the target state s i Used to specify

[0063] In step St202, the management device 10 generates a predetermined vector Q and assigns 1 to the value of the i-th element of the vector Q. Here, the vector Q is a state sx The element values ​​of the vector Q (i.e., the value) correspond to the maximum value V max is used to estimate (St208).

[0064] The number of elements of the vector Q is equal to the number of states that the plant 20 can take. For example, the plant 20 can take states s 1 ~s 9 If there are nine states, the number of elements of vector Q will be nine. Incidentally, the processing of step St202 is repeated as the value of i is incremented (St220), but in the first step St202, all elements except the i-th element (the first element since i=1) of vector Q are empty.

[0065] In step St203, the management device 10 generates a state list Zlist1 and sets the i-th (value of counter i) state s as a target state, which is an element of the state list Zlist1. i The state list Zlist1 is a list for saving the target state whose value is to be recalculated in the current loop (the loop associated with the increment of i in St220), as will be described in detail later. In the first step St203, the elements included in the state list Zlist1 are the state s 1 It is only.

[0066] In step St204, the management device 10 generates a counter j1 and assigns 1 as the value of this counter j1. Note that the counter j1 is used when specifying elements of the status list Zlist1 described above. In step St205, the management device 10 generates a counter j2 and assigns 1 as the value of this counter j2. Note that the counter j2 is used when specifying elements of the status list Zlist2 described next.

[0067] In step St206, the management device 10 generates a state list Zlist2 and also generates another state list Znxt. The state list Zlist2 is a list of the state s stored in the j1th position (value of the counter j1) of the state list Zlist1.x In other words, the list stores all the states that can be transitioned to. x The management device 10 specifies this state s x A list of "current states" that may transition to (those with a transition probability greater than 0) is stored in a state list Zlist2.

[0068] The state list Znxt in step St206 is a list for saving the target state whose value will be recalculated in the next loop when the value of i is incremented (St220). In the first processing of St206, the state list Znxt has no elements (empty state).

[0069] In step St207, the management device 10 sets the j2-th element (the value of the counter j2) of the state list Zlist2 to the state s j2 For example, if the transition destination state of the state list Zlist2 is state s 2 If so, the state s 2 The state s as the current state can transition to 1 , s 5 In this case, when the value of the counter j2 is 1, the state s 1 , s 5 The first state s 1 is designated as the current state.

[0070] Next, in step St208, the management device 10 calculates the maximum value V max where the maximum value V max The state s specified in step St207 j2 It is the weighted average of the maximum value of the manipulated variable for all states that can be transitioned from.

[0071]

[0072] In addition, N included in formula (1) k is the total number of states that the plant 20 can take. For example, the plant 20 can take a state s 1 ~s9 If there are nine states, N k The value of is 9. Also, γ included in formula (1) is a distance attenuation rate. Here, the "distance attenuation rate" is a preset constant (e.g., 0.9) for decreasing the evaluation value as the number of times the controlled element is operated (i.e., as the time required) to reach a predetermined target state increases.

[0073] p in Equation (1) is the state transition probability. j2 In this case, when the controlled element is operated by a predetermined operation amount a, the state s k The probability that the state transitions to is the transition probability p. The value of this transition probability p is obtained by referring to the state transition probability matrix Mx1 (see FIG. 6). Specifically, the management device 10 refers to the state transition probability matrix Mx1 (see FIG. 6) and obtains the value of the transition probability when the operation amount a is set to the value that maximizes the transition probability in the j2th row and the kth column.

[0074] Q included in the formula (1) is the state s as an element of the vector Q set in the processing of step St202. k As described above, in the first step St202, the value of the first element in the vector Q (i.e., Q(s 1 )) is set to 1. Incidentally, the maximum value V max The magnitude of the state s j2 The value will vary depending on the

[0075] Next, in step St209, the management device 10 calculates the state s stored in the j2-th vector Q (the value of the counter j2). j2 The value (i.e., value) of the maximum value V max The process in step St209 determines whether the state s newly calculated in step St208 is smaller than j2 Maximum value at V max In step St209 of FIG. 11A, the state s j2 The value of is simply written as "Q".

[0076] In step St209, the state s of the vector Q j2 The value of is the maximum value V max If the value is smaller than the j2-th state (the value of the counter j2) in the vector Q (Step St209: Yes), the management device 10 proceeds to Step St210. j2 (shown as "Q" in FIG. 11A) to the maximum value V max In other words, the management device 10 substitutes the value of the state s j2 Update the value of V max In this way, the state s j2 Furthermore, the management device 10 updates the value of the updated state s j2 The operation amount that gives the value of this state s j2 The value is stored in association with the value.

[0077] In step St211, the management device 10 adds the state s to the state list Znxt. j2 As mentioned above, the state list Znxt is a list for storing the target state whose value is to be updated in the next loop accompanying the increment of the value of i (St220). j2 With the update of the value (St210), the maximum value V based on Equation (1) max Since the value of may change, the processing of step St211 is performed. After performing the processing of step St211, the processing of the management device 10 proceeds to step St212.

[0078] In step St209, the state s of the vector Q j2 The value of is the maximum value V max If the value is equal to or greater than the maximum value calculated in step St208 (No in step St209), the management device 10 proceeds to step St212. In this case, the maximum value of the value calculated up to that point is the maximum value calculated as the calculation result in step St208. Vmax Since this is the case, state s j2 There is no particular need to update the values ​​of other states that may transition from.

[0079] In step St212, the management device 10 determines whether the value of the counter j2 is equal to or less than the size Nz2 (number of elements) of the status list Zlist2. If the value of the counter j2 is equal to or less than the size Nz2 of the status list Zlist2 (St212: Yes), the processing of the management device 10 proceeds to step St213.

[0080] In step St213, the management device 10 increments the value of the counter j2 and returns to the processing of step St207. max Among the states to be estimated, the maximum value V max If there is any state for which the estimation of the maximum value V has not yet been performed (Step 212: Yes), the management device 10 max is set as an estimation target (Step 213).

[0081] To give a specific example, the transition destination state of the state list Zlist2 is state s in FIG. 2 If so, the state s 2 The state s as the current state can transition to 1 , s 5 When the value of counter j2 is 1, the state s 1 The maximum value of the manipulated variable V when the current state is max When the value of the counter j2 is 2, the state s 5 The maximum value of the manipulated variable V when the current state is max In this way, the maximum value V of each state included in the state list Zlist2 is calculated. max are calculated sequentially (Step 208).

[0082] Furthermore, if the value of counter j2 is greater than the size Nz2 (number of elements) of the state list Zlist2 in step St212 (St212: No), the processing of the management device 10 proceeds to step St214. In step St214, the management device 10 determines whether the value of counter j1 is less than or equal to the size Nz1 (number of elements) of the state list Zlist1. If the value of counter j1 is less than or equal to the size Nz1 of the state list Zlist1 (St214: Yes), the processing of the management device 10 proceeds to step St215. In other words, if there is any state in the state list Zlist1 (a list of states whose values ​​are to be recalculated) for which values ​​have not yet been recalculated, the processing of the management device 10 proceeds to step St215.

[0083] In step St215, the management device 10 increments the value of counter j1 and returns to the processing of step St205. In other words, if there are any target states among the elements of the state list Zlist1 whose values ​​have not yet been updated (St214: Yes), the management device 10 sequentially updates the values ​​of the other states included in the state list Zlist1. Also, if in step St214 the value of counter j1 is greater than the size Nz1 (number of elements) of the state list Zlist1 (St214: No), the processing of the management device 10 proceeds to step St216.

[0084] Next, in step St216 of FIG. 11B, the management device 10 substitutes the state list Znxt into the state list Zlist1. In this way, by rewriting the elements of the state list Zlist1 to be updated, the elements included in the new state list Zlist1 are updated to the maximum value V max The calculation is set to be subject to recalculation.

[0085] In step St217, the management device 10 determines whether the size Nz1 (number of elements) of the state list Zlist1 is greater than 0. That is, the management device 10 determines whether any goal states that require recalculation of value remain as elements of the state list Zlist1. If the size Nz1 of the state list Zlist1 is greater than 0 (St217: Yes), the processing of the management device 10 returns to step St204 in FIG. 11A.

[0086] If the size Nz1 of the state list Zlist1 is equal to or less than 0 in step St217 (St217: No), the management device 10 proceeds to step St218. In this case, all updates of the elements included in the state list Zlist1 are completed, and the value of each element of the vector Q (i.e., the evaluation value of the i-th column of the control characteristics matrix Mx2) is determined.

[0087] In step St218, the management device 10 assigns the elements included in the vector Q to the i-th column (value of counter i) of the evaluation values ​​of the control characteristics matrix Mx2 (see FIG. 7). Also, although omitted in FIG. 11B, the management device 10 assigns the manipulated variable a corresponding to the element (target state) included in the vector Q to the i-th column of the manipulated variables of the control characteristics matrix Mx2 (see FIG. 7).

[0088] In step St219, the management device 10 checks whether the value of the counter i is equal to the number of states N k Here, it is determined whether the number of states N k is the total number of states that the plant 20 can take. In step St219, the value of the counter i is set to the number of states N k If the evaluation value is equal to or less than the target state s2 (St219: No), the processing of the management device 10 proceeds to step St220 in FIG. 11A. In this case, in the control characteristic matrix Mx2 (see FIG. 7), the evaluation value and the optimal manipulated variable are not calculated. i (i.e., the column in FIG. 7) exists.

[0089] In step St220 of FIG. 11A, the management device 10 increments the value of the counter i and returns to the processing of step St202. In short, the management device 10 increments the value of the counter i in the control characteristic matrix Mx2 (see FIG. 7) in the column (target state s i ) in the next column.

[0090] In step St219, the value of the counter i is set to the number of states N k If the result of Step 219 is Yes, the management device 10 ends the process of estimating the control characteristics (Step 2 in FIG. 5) (END). In this case, the evaluation values ​​and optimal manipulated variables have been calculated for all columns (target states) of the control characteristics matrix Mx2 (see FIG. 7).

[0091] In this way, the control characteristic estimator 14d calculates an evaluation value of each manipulated variable when the plant 20 reaches a target candidate state, which is a candidate for the target state, from a plurality of states that the plant 20 can take, based on the state transition probability, which is the probability of a state transition when the controlled plant 20 is operated with a predetermined manipulated variable. The "target candidate state" mentioned above refers to a plurality of assumed candidate target states (see FIG. 7).

[0092] <Calculation of Control Law> Fig. 12 is a flowchart related to the calculation of the control law (see also Fig. 1 as appropriate). Fig. 12 shows details of the process of step St3 in Fig. 5 (calculation of the control law). That is, the flowchart in Fig. 12 shows a series of processes when the management device 10 calculates a predetermined control law from the reward function and the control characteristic matrix Mx2 (see Fig. 7) using the control law determiner 14e.

[0093] In step St301, the management device 10 calculates the product of each row of the matrix in the first depth layer of the control characteristic matrix Mx2 (see FIG. 7) and the reward function vector. More specifically, the management device 10 first reads the control characteristic matrix Mx2 (see FIG. 7) stored in the control characteristic storage unit 12b and also reads the reward function stored in the reward storage unit 12c. Here, the "reward function" is a function for providing a predetermined reward that serves as an incentive when a predetermined operation amount is selected.

[0094] 13 is an explanatory diagram of the reward function. In the example of FIG. 13, states s1 to s n The reward function is preset as a matrix in which the reward value is associated with the actual goal state s 5 , s 7 For , the value of the element of the reward function is set to a positive real number such as 1 or 1.1. 1 and status 6 and status n For non-goal states, such as

[0095] In step St301 in Fig. 12, the product of each row of the matrix in the first depth layer of the control characteristic matrix Mx2 (see Fig. 7) and the reward function vector (the column vector of the "reward" part in Fig. 13) is calculated. This results in a matrix as shown in Fig. 14.

[0096] FIG. 14 is an explanatory diagram of matrix X, which is the calculation result of the product of the control characteristic matrix and the vector of the reward function. Note that the vertical axis of matrix X shown in FIG. 14 represents the "current state" before a predetermined operation is performed on the plant 20. The horizontal axis of matrix X represents a predetermined "target state" of the plant 20. By performing the processing of step St301 (see FIG. 12) described above, the element values ​​of each column that does not correspond to the actual target state among the multiple target states become 0, while the element values ​​of each column that corresponds to the target state become values ​​that are multiplied by the reward (for example, 1 or 1.1 times) from the element values ​​of the control characteristic matrix Mx2 (see FIG. 7). In the example of FIG. 14, among the evaluation values ​​of the first layer of the control characteristic matrix Mx2 (see FIG. 7), the state s where the reward value (see FIG. 13) is greater than 0 is 5 , s 7 The evaluation value of the remaining state s 1 ~s 4 , s 6 , s 8 ~s n The values ​​of all elements are set to 0.

[0097] 12, the management device 10 acquires the maximum value of the element value (evaluation value) for each row of the matrix X. In the example of FIG. 14, in the matrix X, when the current state is the state s1 In this case, the maximum value of the element value (evaluation value) is 0.9.

[0098] Next, in step St303 of Fig. 12, the management device 10 acquires the operation amount corresponding to the maximum element value (evaluation value). A vector having the acquired operation amount as an element is set as the optimal operation amount of the second depth layer in the control characteristics matrix Mx2 (see Fig. 7).

[0099] FIG. 15 is an explanatory diagram showing the correspondence relationship between the maximum evaluation value of each state and the optimal operation amount. The first column in FIG. 15 shows the maximum evaluation value of each state, and the second column shows the optimal operation amount of each state. For example, in the case of state s 1 The element value of is 0.9, and the optimal operation amount is a 1 In this way, the control law determiner 14e determines, as a control law, an optimal manipulated variable in each state of the plant 20, which is the control target, based on the evaluation value of the manipulated variable. That is, the control law determiner 14e determines the control law based on a reward function including a reward when a predetermined manipulated variable is selected and the evaluation value of the manipulated variable.

[0100] During operation of the plant 20 (control target), the equipment control unit 14f may identify the actual state of the plant 20 based on information including the detection values ​​of the sensors 30 of the plant 20, identify the optimal operation amount for moving from this state to the target state from the control law, and output the optimal operation amount to the plant 20.

[0101] Furthermore, when the reward function is changed, the control law determiner 14e may perform processing to update the control law based on the changed reward function while the plant 20 (control target) is in operation. Because the amount of calculation required to update the control law (processing similar to that in FIG. 12 ) is relatively small, the control law can be updated quickly even while the plant 20 is in operation.

[0102] <Display Screen Example> Fig. 16 is an example of a display screen related to processing by the management device (see also Fig. 1 as appropriate). Fig. 16 illustrates an example in which the control characteristic matrix Mx2 is displayed on the display device 50. As shown in Fig. 16, the output control unit 14b causes the display device 50 to display evaluation values ​​of manipulated variables in association with combinations of multiple states (current states) that the plant 20 (control object) can take and multiple target states. The output control unit 14b also causes the display device 50 to display optimal manipulated variables in association with combinations of multiple states (current states) that the plant 20 (control object) can take and multiple target states. This allows the operator to grasp the evaluation values ​​and optimal manipulated variables at a glance.

[0103] In the example of Fig. 16, a third layer is added to the depth of the control characteristic matrix Mx2. This third layer is configured to show the state with the highest transition probability when the control terminal of the plant 20 is operated with the optimal control amount in the second layer. In this way, the output control unit 14b associates the state with the highest state transition probability when the plant 20 (controlled object) is operated with the optimal control amount with the optimal control amount and displays it on the display device 50. This makes it possible to present the relationship between the control amount and the state transition to the operator in an easy-to-understand manner.

[0104] In FIG. 16, the evaluation values ​​of the first layer in the control characteristics matrix Mx2 are displayed, but it is also possible for the user to operate the input device 40 to display the values ​​of the optimal operation variables of the second layer or the state of the third layer (the state with the highest transition probability).

[0105] Furthermore, display examples related to the processing of the management device 10 are not limited to the example in FIG. 16 . For example, the transition of the state of the plant 20 as shown in FIG. 4 may be displayed on the display device 50. Furthermore, a state transition probability matrix Mx1 (see FIG. 6 ) may be displayed on the display device 50. That is, the output control unit 14b may display on the display device 50 the state transition probability matrix Mx1 (see FIG. 6 ), which is a matrix associating state transition probabilities with combinations of multiple states (current states) that the plant 20 (controlled object) can take, destination states from each state, and manipulated variables of the plant 20. This allows the user to grasp at a glance the state transition probabilities corresponding to combinations of the current state, destination states, and manipulated variables.

[0106] Furthermore, the control law (see FIG. 8) for controlling the plant 20 may be displayed on the display device 50. That is, the output control unit 14b may display, as the control law, a correspondence between a plurality of states that the plant 20 (control target) can take and the optimal manipulated variable in each state on the display device 50. This allows the operator to understand what control law is currently set.

[0107] In addition, the display device 50 may display the state definition of the plant 20 (see FIG. 9 ) and the correspondence between the manipulated variable and the valve opening adjustment variable (see FIG. 10 ). Furthermore, the display device 50 may display the matrix X (see FIG. 14 ), which is the calculation result of the product of the control characteristic matrix Mx2 (see FIG. 7 ) and the reward function (see FIG. 13 ), as well as an explanatory diagram of the reward function (see FIG. 13 ). By viewing such data, the operator can confirm the entire process from information about the behavior of the controlled object (operating data and simulation data) to the calculation of the control law. Therefore, the operator can confirm the data that serves as the basis for setting the manipulated variable at each moment in the plant 20. Safety standards are particularly high in the field of plant control. As described above, the management device 10 uniquely derives the control law from information about the behavior of the controlled object, and by presenting the calculation process, the basis for the manipulated variable can be presented to a third party.

[0108] <Effects> According to this embodiment, the dimensions of the state transition probability matrix Mx1 (see FIG. 6) include not only the current state and the transition destination state but also the manipulated variable of the plant 20. This makes it possible to accurately simulate the relationship between the manipulated variable of the plant 20 and the state transition, thereby improving the prediction accuracy of the model.

[0109] The management device 10 also generates a control characteristic matrix Mx2 (see FIG. 7 ) that includes evaluation values ​​of the manipulated variables in each state relative to the control target and optimal manipulated variables that maximize the evaluation values. Furthermore, a control law for bringing the plant 20 to the target state is derived based on the control characteristic matrix Mx2 and the reward function. Because the amount of calculation required to calculate such a control law (flowchart in FIG. 12 ) is relatively small, the management device 10 can quickly calculate the control law. Therefore, even while the plant 20 is operating, the management device 10 can update the control law as needed and quickly calculate the optimal manipulated variables.

[0110] Furthermore, even if the target state (control target) or predetermined parameters are changed during operation of the plant 20, there is no particular need to recreate the control characteristic matrix Mx2 (see FIG. 7), so the control law can be updated quickly in real time. As described above, the amount of calculation required to update the control law is relatively small, so the control law can be updated quickly even when the plant 20 is large in scale or has a large amount of data.

[0111] Furthermore, conventional model predictive control techniques tend to optimize control that avoids risks, making it difficult to maximize the performance of the controlled object. Conventional model predictive control techniques refer to, for example, techniques that have the function of quickly updating conventional control laws in real time. In contrast, in the present embodiment, the manipulated variable that maximizes the evaluation value is identified in the state transition of the plant 20, assuming all possible combinations. This makes it possible to perform control that maximizes the performance of the plant 20.

[0112] Furthermore, even if the control characteristics of the plant 20 change due to aging of the plant 20 or changes in the surrounding environment, the management device 10 can respond appropriately by appropriately updating the state transition probability matrix Mx1 (see Figure 6) and the control characteristic matrix Mx2 (see Figure 7).

[0113] <<Modifications>> While the management device 10 and the like according to the present disclosure have been described above in the embodiments, the present disclosure is not limited to these descriptions and various modifications can be made. For example, in the embodiments, a case has been described in which one type of sensor 30 is used to acquire the state of the plant 20, but this is not limiting. That is, multiple types of sensors (or multiple sensors of the same type) may be provided. In this case, the state of the control object is identified in association with a combination of ranges of detected values ​​from the multiple sensors. Furthermore, the number of control objects is not limited to one, and the embodiments can also be applied to cases in which there are multiple control objects.

[0114] In the embodiment, the management device 10 (see FIG. 1) includes the communication unit 13 (see FIG. 1), but the present invention is not limited to this. That is, the communication unit 13 may be omitted as appropriate from the configuration in FIG. 1, and the manipulated variable, which is the calculation result of the equipment control unit 14f, may be transmitted to the plant 20 via the output control unit 14b.

[0115] In the embodiment, the device control unit 14f (see FIG. 1) determines the manipulated variable of the plant 20, but the present invention is not limited to this. For example, the calculation unit 14 may select an optimal control law from among the control laws stored in the control law storage unit 12d, and then transmit the optimal control law to the plant 20.

[0116] In the embodiment, the state transition probability matrix Mx1 (see FIG. 6) is used to estimate the operation of the controlled object, but the present invention is not limited to this. For example, a neural network may be used, or information corresponding to the state transition probability matrix Mx1 may be expressed by a predetermined mathematical formula.

[0117] In the embodiment, the case where the action value function (Equation (1)) is calculated using a processing flow based on dynamic programming has been described, but the present invention is not limited to this and a predetermined action value function may be calculated using another calculation formula or method. Furthermore, in the embodiment, the "evaluation value" is defined as an expected value when the controlled object moves from a predetermined state to a target candidate state, but the present invention is not limited to this. In other words, the "evaluation value" may be defined based on a predetermined index different from the expected value.

[0118] In the embodiment, a case has been described in which a signal (control signal) indicating the manipulated variable of the control terminal of the plant 20 is transmitted from the management device 10 to the plant 20, but this is not limiting. For example, a signal indicating the manipulated variable of the control terminal of the plant 20 may be transmitted to a predetermined server or a user's terminal device (i.e., provided as data indicating the manipulated variable). Furthermore, the processing (management method) executed by the management device 10 may be executed as a predetermined program on a computer. The program can be provided via a communication line or can be written to a recording medium such as a CD-ROM and distributed.

[0119] Furthermore, the present disclosure is not limited to the embodiments and includes various modifications. For example, the embodiments have been described in detail to clearly explain the present disclosure, and the present disclosure is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the embodiments with other configurations.

[0120] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card, SD card, or DVD.

[0121] In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected.

[0122] REFERENCE SIGNS LIST 10 Management device 11 Data acquisition unit 12 Memory unit 12a Operation memory unit 12b Control characteristic memory unit 12c Reward memory unit 12d Control law memory unit 13 Communication unit 14 Calculation unit 14a Input control unit 14b Output control unit 14c Model estimation unit 14d Control characteristic estimation unit 14e Control law determination unit 14f Equipment control unit 20 Plant (controlled object) 30 Sensor 40 Input device 50 Display device 60 Cloud 70 External recording medium Mk1 State transition probability matrix Mk2 Control characteristic matrix St2 Step (control characteristic estimation step) St3 Step (control law determination step)

Claims

1. A management device comprising: a control characteristic estimation unit that calculates an evaluation value of each operation amount when a controlled object reaches a target candidate state, which is a candidate for a target state, from multiple states that the controlled object can take, based on a state transition probability, which is the probability of a state transition when the controlled object is operated with a predetermined operation amount; and a control law determination unit that determines, based on the evaluation value, the optimal operation amount for each state of the controlled object as a control law.

2. The management device according to claim 1, wherein the evaluation value is an expected value when the controlled object moves from a predetermined state to the target candidate state.

3. The management device according to claim 1, wherein the control law determination unit determines the control law based on a reward function including a reward when a predetermined manipulated variable is selected and the evaluation value.

4. The management device according to claim 3, characterized in that when the reward function is changed, the control law determination unit performs a process of updating the control law based on the changed reward function while the controlled object is operating.

5. The management device according to claim 1, further comprising an equipment control unit that, while the controlled object is in operation, identifies the actual state of the controlled object based on information including the detected values ​​of the sensors of the controlled object, identifies the optimal manipulated variable from the control law when the controlled object reaches the target state from that state, and outputs the optimal manipulated variable to the controlled object.

6. The management device according to claim 1, further comprising an output control unit that displays on a display device a state transition probability matrix, which is a matrix that associates the state transition probabilities with combinations of multiple states that the controlled object can take, destination states from each of the states, and manipulated variables of the controlled object.

7. The management device according to claim 1, further comprising an output control unit that displays the evaluation values ​​on a display device in association with combinations of a plurality of states that the controlled object can take and a plurality of target states.

8. The management device according to claim 1, further comprising an output control unit that displays the optimal manipulated variable on a display device in association with a combination of a plurality of states that the controlled object can take and a plurality of target states.

9. The management device according to claim 8, wherein the output control unit causes the display device to display the state with the highest state transition probability when the controlled object is operated with the optimal operation amount, in association with the optimal operation amount.

10. The management device according to claim 1, further comprising an output control unit that displays on a display device, as the control law, the correspondence between the multiple states that the controlled object can take and the optimal operation amount for each of the states.

11. A management method including: a control characteristic estimation step of calculating an evaluation value of each operation amount when a controlled object reaches a target candidate state, which is a candidate for a target state, from a plurality of states the controlled object can take, based on a state transition probability, which is the probability of a state transition when the controlled object is operated with a predetermined operation amount; and a control law determination step of determining, as a control law, an optimal operation amount for each state of the controlled object, based on the evaluation value.

12. A control system comprising: a control characteristic estimation unit that calculates an evaluation value of each operation amount when a plant, which is the object of control, reaches a target candidate state, which is a candidate for a target state, from a plurality of states that the plant can take, based on a state transition probability, which is the probability of state transition when the plant is operated with a predetermined operation amount; and a control law determination unit that determines, as a control law, an optimal operation amount for each state of the plant based on the evaluation value; and an equipment control unit that, while the plant is operating, identifies the actual state of the plant based on information including detection values ​​of sensors of the plant, identifies the optimal operation amount from the control law when reaching the target state from that state, and outputs the optimal operation amount to the plant.

Citation Information

Patent Citations

  • Method for improving performance of method for computationally predicting future state of target object, driver assistance system, vehicle including such driver assistance system and corresponding program storage medium and program

    JP2016212872A

  • Future state estimation device and future state estimation method

    JP2019159876A

  • Future state estimating device

    JP2023074434A