Apparatus, method and program
The system addresses instability in control systems by shifting control parameters based on target value changes and environmental adjustments, ensuring stable and accurate control through learning processes.
Patent Information
- Application Number
- JP2022177854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-11-07
Smart Images

Figure 0007735980000001 
Figure 0007735980000002 
Figure 0007735980000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus, a method, and a program. [Background technology]
[0002] Patent Documents 1 to 4 state that "a manipulated variable map is selected based on the target value SV, and the manipulated variable MV is calculated using the selected manipulated variable map" (paragraph 0031 of Patent Document 1). [Prior art document] [Patent documents] [Patent Document 1] JP 2022-156797 A [Patent Document 2] JP 2020-95352 A [Patent Document 3] JP 2021-117699 A [Patent Document 4] JP 2022-014099 A Summary of the Invention
[0003] In a first aspect of the present invention, there is provided an apparatus comprising: a deviation acquisition unit that acquires the deviation between a measured value of a state related to a controlled object and a target value; a control parameter acquisition unit that acquires shifted control parameters obtained by shifting control parameters supplied to the controlled object; a first supply unit that supplies the deviation acquired by the deviation acquisition unit and the shifted control parameters acquired by the control parameter acquisition unit to a control model that outputs recommended control parameters to be supplied to the controlled object in response to the deviation and control parameters being input; and an output unit that outputs the recommended control parameters output from the control model in response to the supply from the first supply unit to the control model.
[0004] The above device may further include a target value acquisition unit that acquires a target value of the state, and the control parameter acquisition unit may shift the control parameter supplied to the controlled object in response to the target value being changed from a reference target value, and acquire the shifted control parameter.
[0005] In the above-mentioned device, the control parameter acquisition unit may, in response to the target value being changed from a reference target value, shift the control parameter supplied to the controlled object by a shift amount corresponding to the target value, and acquire the shifted control parameter.
[0006] In the above-mentioned device that shifts the control parameter by a shift amount corresponding to the target value, the control parameter acquisition unit may determine the shift amount by multiplying the difference between the target value and the reference target value by a predetermined coefficient.
[0007] In the above-mentioned device that shifts the control parameter by a shift amount corresponding to the target value, the control parameter acquisition unit may determine the shift amount using a predetermined relational expression that indicates the relationship between a value set as the target value and the value of the control parameter when the measurement value stabilizes at that value.
[0008] Any of the above devices may further include a first detection unit that detects that the deviation acquired by the deviation acquisition unit after the control object is controlled by the recommended control parameters is not stable within a reference range, and the control parameter acquisition unit may shift the control parameters supplied to the control object and acquire the shifted control parameters in response to the first detection unit detecting that the deviation acquired by the deviation acquisition unit is not stable within the reference range.
[0009] The above device further includes a second supply unit that supplies the deviation acquired by the deviation acquisition unit and the rate of change of the deviation to a shift amount output model that outputs a recommended shift amount recommending a shift of a control parameter by the control parameter acquisition unit in response to input of the deviation and the rate of change of the deviation, and the control parameter acquisition unit may shift the control parameter supplied to the controlled object by the recommended shift amount output from the shift amount output model in response to the first detection unit detecting that the deviation acquired by the deviation acquisition unit is not stable within the reference range and the second supply unit supplying the control parameter to the shift amount output model, thereby acquiring the shifted control parameter.
[0010] In the above device that supplies the deviation acquired by the deviation acquisition unit and the rate of change of the deviation to a shift amount output model, the second supply unit may supply the deviation and the rate of change of the deviation to the shift amount output model at each reference interval.
[0011] The above device that supplies the deviation acquired by the deviation acquisition unit and the rate of change of the deviation to a shift amount output model may further include a first learning processing unit that performs learning processing of the shift amount output model using learning data including the deviation acquired by the deviation acquisition unit, the rate of change of the deviation, and the shift amount of the control parameter acquired by the control parameter acquisition unit, so as to output the recommended shift amount recommended for increasing a reward value determined by a preset reward function in response to input of the deviation and the rate of change of the deviation.
[0012] Any of the above devices may further include a second detection unit that detects that the control object has been changed, and the control parameter acquisition unit may shift the control parameters supplied to the control object in response to the second detection unit detecting that the control object has been changed, and acquire the shifted control parameters.
[0013] In any of the above devices, the control model may include a change amount output model that outputs a recommended change amount that recommends changing the control parameter in response to input of a deviation and a control parameter, and an adder that calculates the recommended control parameter by adding the control parameter supplied to the controlled object and the recommended change amount output from the change amount output model.
[0014] The above device may further include a second learning processing unit that performs learning processing of the change amount output model using learning data including the deviation acquired by the deviation acquisition unit and the control parameter acquired by the control parameter acquisition unit, and that outputs the recommended change amount recommended for increasing a reward value determined by a preset reward function in response to input of the deviation and the control parameter.
[0015] In a second aspect of the present invention, there is provided a method comprising: a deviation acquisition step of acquiring a deviation between a measured value of a state related to a controlled object and a target value; a control parameter acquisition step of acquiring shifted control parameters obtained by shifting control parameters supplied to the controlled object; a first supply step of supplying the deviation acquired in the deviation acquisition step and the shifted control parameters acquired in the control parameter acquisition step to a control model that outputs recommended control parameters to be supplied to the controlled object in response to the deviation and control parameters being input; and an output step of outputting the recommended control parameters output from the control model in response to the supply to the control model by the first supply step.
[0016] In a third aspect of the present invention, there is provided a program that causes a computer to function as a deviation acquisition unit that acquires the deviation between a measured value and a target value of a state related to a controlled object, a control parameter acquisition unit that acquires shifted control parameters obtained by shifting control parameters supplied to the controlled object, a first supply unit that supplies the deviation acquired by the deviation acquisition unit and the shifted control parameters acquired by the control parameter acquisition unit to a control model that outputs recommended control parameters to be supplied to the controlled object in response to the deviation and control parameters being input, and an output unit that outputs the recommended control parameters output from the control model in response to the supply from the first supply unit to the control model.
[0017] The above summary of the invention does not list all of the necessary features of the present invention, and subcombinations of these features may also constitute inventions. [Brief explanation of the drawings]
[0018] [Figure 1] 1 shows a system 1 according to a first embodiment. [Figure 2] An example of the change amount output model 2061 is shown. [Figure 3] 2 shows another example of the change amount output model 2061. [Figure 4] The effect of shifting the control parameters input to the control model 206 is shown. [Figure 5] The fitted curve of the sample data is shown. [Figure 6] 2 illustrates the operation of the device 200. [Figure 7] This shows the transition of the measured value PV and the control parameters when the target value SP is changed. [Figure 8] 1 shows a system 1A according to a second embodiment. [Figure 9] 2 shows an example of a shift amount output model 213A. [Figure 10] 10 shows another example of the shift amount output model 213A. [Figure 11] The operation of the device 200A is shown. [Figure 12] The graph shows the transition of the measured value PV and the control parameters when a disturbance occurs. [Figure 13] 10 shows a system 1B according to a third embodiment. [Figure 14] 22 illustrates an example computer 2200 in which aspects of the present invention may be embodied, in whole or in part. DETAILED DESCRIPTION OF THE INVENTION
[0019] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0020] 1. First Embodiment <1.1. System 1> 1 shows a system 1 according to the first embodiment. The system 1 includes a facility 100 and a device 200.
[0021] <1.1.1.Equipment 100> The facility 100 is a facility or device equipped with a control target 101. For example, the facility 100 may be a plant, or a composite device that combines multiple devices. Examples of plants include industrial plants such as chemical and bio plants, plants that manage and control wellheads and surrounding areas of gas fields and oil fields, plants that manage and control power generation such as hydroelectric, thermal, and nuclear power, plants that manage and control environmental power generation such as solar and wind power, and plants that manage and control water supply and sewage systems, dams, etc.
[0022] The facility 100 is provided with one or more control targets 101. The control target 101 may be a tool, machine, device, or the like to be controlled, and may be a so-called field device. For example, the control target 101 may be a sensor device such as a pressure gauge, flow meter, or temperature sensor, a valve device such as a flow control valve or an on-off valve, or an actuator device such as a fan or motor. The control target 101 is controlled externally via a wired or wireless connection, or may be controlled manually. The control target 101 may be controlled by a control unit 207 in the device 200. In the present embodiment, as an example, the control target 101 may be controlled by receiving an instructed value IV (instructed value) for a manipulated variable MV (manipulated variable) from the control unit 207.
[0023] The facility 100 may also be provided with one or more sensors 102. Each sensor 102 may measure a measurement value of the internal or external state of the facility 100, i.e., a measurement value of a physical quantity indicating the internal or external state. At least one sensor 102 may measure a measurement value PV (Process Variable) of the state of the controlled object 101. The measurement value PV may be operating data indicating the operating state as a result of controlling the controlled object 101, and may indicate a controlled variable to be controlled. As an example, the measurement value PV may indicate the output of the controlled object 101 itself, or may indicate various values that change depending on the output of the controlled object 101. As an example, the measurement value PV may indicate pressure, temperature, pH, speed, flow rate, etc. Each sensor 102 may supply the measured measurement value PV to the device 200.
[0024] <1.1.2.Device 200> The device 200 controls the controlled object 101 and may be, for example, a controller for the controlled object 101. The device 200 may output an indication value IV for a manipulated variable MV of the controlled object 101 to perform process control such as adjusting the temperature, adjusting the liquid level, or adjusting the flow rate.
[0025] The device 200 may be a computer such as a personal computer (PC), a tablet computer, a smartphone, a workstation, a server computer, or a general-purpose computer, or may be a computer system in which multiple computers are connected. Such a computer system is also a computer in a broad sense. The device 200 may also be implemented by one or more virtual computer environments executable within a computer. Alternatively, the device 200 may be a dedicated computer designed for AI control, or may be dedicated hardware realized by dedicated circuits. Furthermore, if the device 200 can connect to the Internet, the device 200 may be realized by cloud computing.
[0026] The apparatus 200 may include a measurement value acquiring unit 201, a target value acquiring unit 202, a deviation acquiring unit 203, a control parameter acquiring unit 204, a first supplying unit 205, a control model 206, a control unit 207, and a learning processing unit 208. Note that these blocks are functionally separated functional blocks and may not necessarily correspond to the actual device configuration. That is, even if shown as a single block in this diagram, it is not limited to being configured by a single device. Furthermore, even if shown as separate blocks in this diagram, it is not limited to being configured by separate devices.
[0027] <1.1.2-1. Measurement value acquisition unit 201> The measurement value acquiring unit 201 acquires a measurement value PV of a state related to the control target 101. In the present embodiment, as an example, the measurement value acquiring unit 201 acquires a measurement value PV for one physical quantity from one sensor 102. However, the measurement value acquiring unit 201 may acquire measurement values PV for each of a plurality of physical quantities from a plurality of sensors 102. The measurement value acquiring unit 201 may supply the acquired measurement value PV to the deviation acquiring unit 203.
[0028] <1.1.2-2. Target value acquisition unit 202> The target value acquiring unit 202 acquires a target value SP (Set Point) of a state related to the controlled object 101. The target value acquiring unit 202 may acquire the target value SP of the measurement value PV acquired by the measurement value acquiring unit 201. The target value acquiring unit 202 may acquire the target value SP from an operator via an input unit (not shown). When the target value SP is not set by the operator, the target value acquiring unit 202 may acquire a preset reference target value as the target value SP. The target value acquiring unit 202 may supply the acquired target value SP to the deviation acquiring unit 203. When the target value SP is set to a value different from the reference target value, the target value acquiring unit 202 may supply a signal indicating the set target value SP (also referred to as a target value change signal) to the control parameter acquiring unit 204.
[0029] <1.1.2-3. Deviation acquisition unit 203> The deviation acquiring unit 203 acquires the deviation between the measurement value PV and the target value SP of the state related to the control object 101. The deviation acquiring unit 203 may acquire the measurement value PV from the measurement value acquiring unit 201 and the target value SP from the target value acquiring unit 202, and calculate the deviation by subtracting the measurement value PV from the target value SP. Alternatively, the deviation acquiring unit 203 may calculate the deviation by subtracting the target value SP from the measurement value PV. The deviation acquiring unit 203 may supply the acquired deviation to the first supplying unit 205. The deviation acquiring unit 203 may store the acquired deviation in a storage unit (not shown).
[0030] <1.1.2-4. Control parameter acquisition unit 204> The control parameter acquiring unit 204 acquires a shifted control parameter P+Δp by shifting the control parameter P supplied to the controlled object 101. The control parameter acquiring unit 204 may acquire the control parameter P from the control unit 207 described below, and in the present embodiment, as an example, the control parameter acquiring unit 204 may acquire the control parameter P each time the control unit 207 supplies the control parameter P to the controlled object 101. The control parameter P may indicate an instruction value IV for an operation amount MV of the controlled object 101. If the controlled object 101 is a valve, the control parameter P may indicate, as an example, a valve opening. The control parameter acquiring unit 204 may shift the control parameter P acquired from the control unit 207 to generate the shifted control parameter P+Δp.
[0031] In response to a change in the target value SP from the reference target value described above, the control parameter acquisition unit 204 may shift the control parameter P supplied to the controlled object 101 to acquire the shifted control parameter P+Δp. In response to receiving a target value change signal from the target value acquisition unit 202, the control parameter acquisition unit 204 may shift the most recent control parameter P acquired from the control unit 207 to generate the shifted control parameter P+Δp.
[0032] In response to a change in the target value SP from the reference target value, the control parameter acquisition unit 204 may shift the control parameter P supplied to the controlled object 101 by a shift amount Δp corresponding to the target value SP to acquire a shifted control parameter P+Δp. The control parameter acquisition unit 204 may determine the shift amount Δp based on the target value SP indicated by the target value change signal. The shift amount Δp will be described in detail later.
[0033] The control parameter acquiring unit 204 may supply the acquired shifted control parameter P+Δp to the first supplying unit 205. Note that, as an example in the present embodiment, when the target value SP is a reference target value, the control parameter acquiring unit 204 supplies the control parameter P acquired from the control unit 207 to the first supplying unit 205 without shifting it, but the control parameter P may be shifted by 0 and supplied to the first supplying unit 205 as a shifted control parameter P+Δp.
[0034] <1.1.2-5. 1st supply section 205> The first supply unit 205 supplies the control model 206 with the deviation acquired by the deviation acquisition unit 203 and the shifted control parameter P+Δp acquired by the control parameter acquisition unit 204. The first supply unit 205 may supply the control model 206 with the shifted control parameter P+Δp obtained by shifting the control parameter P supplied from the control unit 207 to the controlled object 101, and the deviation indicating the operating state resulting from controlling the controlled object 101 using the control parameter P. When the unshifted control parameter P is supplied from the control parameter acquisition unit 204, the first supply unit 205 may supply the control parameter P and the deviation to the control model 206.
[0035] <1.1.2-6. Control Model 206> In response to input of the deviation and the control parameter P, the control model 206 outputs a recommended control parameter Pr that is recommended to be supplied to the controlled object 101. In response to input of the deviation and the control parameter P supplied to one controlled object 101, the control model 206 may output a recommended control parameter Pr that is recommended to be supplied to the one controlled object 101. The recommended control parameter Pr may indicate a recommended instruction value IV for the manipulated variable MV of the controlled object 101.
[0036] The control model 206 may be generated with respect to the control parameter P based on a reference value V (for example, 0), and may output the value of the recommended control parameter Pr based on the reference value V in response to inputting the relative value of the control parameter P with respect to the reference value V together with the deviation. Inputting the deviation and the shifted control parameter P+Δp, instead of inputting the deviation and the control parameter P to the control model 206 based on the reference value V, may be equivalent to inputting the deviation and the control parameter P to another control model based on the reference value V−Δp. Such another control model may be generated in the prior art as disclosed in Patent Document 1 to be used in an environment different from that of the control model 206.
[0037] In response to input of the deviation and the control parameter P, the control model 206 may output recommended control parameters Pr according to a state in which the target value SP is a reference target value. In other words, the control model 206 may output recommended control parameters Pr that are recommended in a state in which the target value is a reference target value. Inputting the deviation and the shifted control parameter P+Δp to this control model 206 instead of inputting the deviation and the control parameter P may be equivalent to inputting the deviation and the control parameter P to another control model that outputs recommended control parameters Pr that are suitable when the target value SP is a value different from the reference target value.
[0038] The control model 206 may output a recommended control parameter Pr to the control unit 207 in response to receiving the deviation and the shifted control parameter P+Δp or the control parameter P from the first supply unit 205. In the present embodiment, as an example, the deviation between the measurement value PV and the target value SP for one physical quantity is described as being input to the control model 206, but the deviations between the measurement value PV and the target value SP for a plurality of physical quantities may also be input. The control model 206 may include a change amount output model 2061 and an adder 2062.
[0039] <1.1.2-6(1). Change amount output model 2061> In response to input of the deviation and the control parameter P, the change amount output model 2061 outputs a recommended change amount that recommends a change to the control parameter P. The change amount output model 2061 may supply the recommended change amount to the adder 2062. The recommended change amount may indicate a recommended change amount from the most recent control parameter P supplied to the controlled object 101. In the present embodiment, as an example, the recommended change amount may indicate a recommended change amount for the most recent indicated value IV for the manipulated variable MV. The change amount output model 2061 may be generated by a learning process by the learning processor 208, and may be stored in a storage unit (not shown).
[0040] <1.1.2-6(2). Addition section 2062> The adder 2062 calculates the recommended control parameter Pr by adding the control parameter P supplied to the controlled object 101 and the recommended change amount output from the change amount output model 2061. The adder 2062 may calculate the recommended control parameter Pr by adding the most recent control parameter P supplied from the control unit 207 and the recommended change amount supplied from the change amount output model 2061.
[0041] The adder 2062 calculates the control parameter P at time t-1 as shown in the following equation (1): (t-1) and the recommended change amount Δu at time t (t) The recommended control parameter Pr at time t is calculated by adding (t) may be calculated. Pr (t) =P (t-1) +Δu (t) (1)
[0042] The adding unit 2062 may store the control parameters P supplied from the control unit 207 and use them to calculate the recommended control parameters Pr. The adding unit 2062 may supply the calculated recommended control parameters Pr to the control unit 207.
[0043] <1.1.2-7. Control unit 207> The control unit 207 is an example of an output unit, and outputs recommended control parameters Pr output from the control model 206 in response to supply from the first supply unit 205 to the control model 206. As an example in the present embodiment, the control unit 207 may output the recommended control parameters Pr as control parameters P to the control object 101 to control the control object 101. The control unit 207 may output the control parameters P input by an operator to the control object 101 to control the control object 101. The control unit 207 may output the control parameters P to the control object 101 in accordance with the control period of the control object 101.
[0044] The control unit 207 may store the control parameter P supplied to the control target 101 in a storage unit (not shown). The control unit 207 may store the control parameter P supplied to the control target 101 in the storage unit in association with the deviation acquired by the deviation acquisition unit 203. The control unit 207 may store the control parameter P supplied to the control target 101 in the storage unit in association with the deviation indicating the operating state as a result of controlling the control target 101 with the control parameter P.
[0045] <1.1.2-8. Learning processing unit 208> The learning processing unit 208 is an example of a second learning processing unit, and performs learning processing of the change amount output model 2061 using learning data including the deviation acquired by the deviation acquiring unit 203 and the control parameter P acquired by the control parameter acquiring unit 204. The deviation and control parameter P included in the learning data may be the deviation and control parameter P acquired when the target value SP is the above-mentioned reference target value, and may be stored in association with each other in a storage unit (not shown). Note that the learning data may be acquired from a simulator (not shown) of the system 1 instead of from the actual system 1. The simulator may be created using actual measurement data of the equipment 100 using any system identification technology.
[0046] The learning processing unit 208 may perform learning of the change amount output model 2061 so as to output a recommended change amount recommended for increasing the reward value in response to input of the deviation and the control parameter P. When a reward value corresponding to the state of the controlled object 101 at a predetermined time point (for example, the time point at which the deviation and the control parameter P are acquired) (for example, a reward value obtained by inputting a value corresponding to the measurement value PV at that time into a reward function) is set as a base reward value, the recommended change amount may be a change amount recommended for increasing the reward value above the base reward value. As an example, the learning processing unit 208 may perform learning using a Kernel Dynamic Policy Programming (KDPP) algorithm. The reward value may be a value determined by a preset reward function. The reward function may be a function based on the deviation, for example, a function in which the smaller the deviation, the larger the reward value. Note that when the deviation acquisition unit 203 acquires deviations for each of multiple physical quantities, the reward function may be a function based on the sum of the multiple deviations or a function based on the result of weighted addition of the multiple deviations.
[0047] According to the above-described device 200, a shifted control parameter P+Δp obtained by shifting the control parameter P is supplied to a control model 206 that outputs a recommended control parameter Pr in response to input of the deviation between the measured value PV and the target value SP and the control parameter P. Therefore, the recommended control parameter Pr output by inputting the control parameter P to another control model whose reference value of the control parameter P is shifted by the shift amount −Δp from that of the control model 206 can be acquired from the control model 206. Therefore, by shifting the control parameter P in accordance with a change in the environment (in this embodiment, a change in the target value is used as an example) and inputting it to the control model 206, a recommended control parameter Pr adapted to the changed environment can be acquired. Furthermore, by shifting the control parameter P, the recommended control parameter Pr output from another control model having a different reference value can be acquired from the control model 206. Therefore, the device 200 can be made smaller than when multiple control models are built into the device 200.
[0048] Furthermore, in the control model 206, a recommended change amount for the control parameter P is output from the change amount output model 2061 according to the deviation and the control parameter P that has already been supplied to the controlled object 101, and the supplied control parameter P is added to the recommended change amount to calculate the recommended control parameter Pr. Therefore, by inputting the control parameter P to the control model 206 as the shifted control parameter P+ΔP, it is possible to obtain the recommended control parameter Pr based on the supplied control parameter P even when the control model 206 is used as another control model. Therefore, it is possible to obtain the appropriate recommended control parameter Pr that is suited to the controlled object 101.
[0049] Furthermore, using learning data including the deviation acquired by the deviation acquisition unit 203 and the control parameter P acquired by the control parameter acquisition unit 204, a learning process is performed on the change amount output model 2061 so as to output a recommended change amount recommended for increasing the reward value determined by the reward function in response to the input of the deviation and the control parameter P. Therefore, it is possible to reliably acquire an appropriate recommended control parameter Pr from the control model 206.
[0050] Furthermore, in response to the deviation and control parameter P being input to the control model 206, recommended control parameters Pr corresponding to the state in which the target value SP is a reference target value are output from the control model 206, and in response to the target value SP being changed from the reference target value, shifted control parameters P+Δp obtained by shifting the control parameters P are supplied to the control model 206. Therefore, using a single control model 206, recommended control parameters Pr corresponding to changes in the target value SP can be obtained.
[0051] Furthermore, in response to a change in the target value SP from the above-described reference target value, the control parameters P supplied to the controlled object 101 are shifted by a shift amount Δp corresponding to the target value SP to obtain shifted control parameters P+Δp. Therefore, using a single control model 206, it is possible to obtain highly accurate recommended control parameters Pr corresponding to changes in the target value SP.
[0052] <1.2. Change Output Model 2061> Fig. 2 shows an example of the change amount output model 2061. In Fig. 2 and Fig. 4 described later, the vertical axis indicates the control parameter P (for example, the command value IV of the valve opening), and the horizontal axis indicates the deviation.
[0053] The change amount output model 2061 may indicate the correspondence relationship between combinations of deviation and control parameter P and recommended change amounts. In this example, the change amount output model 2061 may be a manipulated variable map that maps the correspondence relationship between combinations of deviation and control parameter P and recommended change amounts. The manipulated variable map may be divided into multiple regions, each corresponding to a different recommended change amount, depending on the combination of control parameter P and deviation, and may output recommended change amounts that correspond to the coordinate position of the input combination of control parameter P and deviation. When such a change amount output model 2061 is used, the process becomes stable at a coordinate point where the deviation is 0 and the change amount is 0 (also referred to as an equilibrium point O; in FIG. 2, as an example, the point where the deviation = 0 and the control parameter P = approximately 50).
[0054] The change amount output model 2061 may include information about the entire region of the operation amount map. Alternatively, the change amount output model 2061 may include only information indicating the boundaries of each region (for example, coordinates or function formulas indicating the boundaries) and the recommended change amounts corresponding to each region. In this case, the storage area for storing the change amount output model 2061 can be reduced.
[0055] 3 shows another example of the change amount output model 2061. As shown in this figure, the change amount output model 2061 may be a table that associates combinations of deviations and control parameters P with recommended change amounts.
[0056] <1.3. Shift of control parameter P> FIG. 4 shows the effect of shifting the control parameter P input to the control model 206.
[0057] The manipulated variable map of the change amount output model 2061 according to this embodiment has a coordinate axis for the deviation, and even if the environment of the system 1 changes because the target value SP is changed from the reference target value, the process remains stable on the line of deviation = 0. Therefore, a deviation of the equilibrium point O caused by a change in the environment can be eliminated by shifting the manipulated variable map in the direction of the coordinate axis of the control parameter P (in this figure, the vertical axis direction is used as an example).
[0058] In the device 200 according to this embodiment, when a point Os different from the original equilibrium point O becomes the actual equilibrium point due to a change in the environment, the control parameter P is shifted by a shift amount Δp and input to the change amount output model 2061 of the control model 206. As a result, the manipulated variable map of the change amount output model 2061 shown in the upper part of the figure is shifted by −Δp along the coordinate axis of the control parameter P to obtain the manipulated variable map shown in the lower part of the figure, and a recommended change amount similar to that when the deviation and the control parameter P are input is output. Therefore, by shifting the control parameter P and inputting it to the control model 206, it is possible to obtain an output from another control model whose manipulated variable map has been shifted.
[0059] <1.4. Shift amount> The control parameter acquiring unit 204 may shift the control parameter P supplied to the controlled object 101 by a shift amount Δp corresponding to the set target value SP to generate a shifted control parameter P+Δp. For example, the control parameter acquiring unit 204 may determine the shift amount Δp by multiplying the difference between the target value SP and the reference target value by a preset coefficient. Alternatively, or in addition, the control parameter acquiring unit 204 may determine the shift amount Δp using a preset relational expression that indicates the relationship between the value set for the target value SP and the value of the control parameter P when the measurement value PV stabilizes at that value (also referred to as the value of the control parameter P at the equilibrium point O).
[0060] The measurement value PV being stabilized at the target value SP may mean that the measurement value PV falls within a reference range over a second reference time after a first reference time has elapsed since the control unit 207 outputted the control parameter P to the controlled object 101. The reference range may be a range that includes the target value SP at its center. The first reference time and the second reference time may be set in advance according to the time constant of the controlled object 101. The first reference time and the second reference time may be different from each other. The second reference time may be longer than the first reference time and may be longer than the control period of the controlled object 101 by the control unit 207.
[0061] The above-mentioned coefficients and relational expressions used to determine the shift amount may be set from an approximation curve for a plurality of sample data including the measurement value PV at the equilibrium point O and the control parameter P at the equilibrium point O. If the measurement value PV at the equilibrium point O is considered to be a target value SP for the measurement value PV, the approximation curve may indicate the relationship between the target value SP and the value of the control parameter P at the equilibrium point O. The function of the reference time curve may be set in advance as the above-mentioned relational expression used to determine the shift amount. If the approximation curve is a linear function, the ratio of the change in the control parameter P to the change in the target value SP, i.e., a value indicating the slope of the approximation curve, may be set in advance as the above-mentioned coefficient used to determine the shift amount.
[0062] Each piece of sample data may include the value of the control parameter P when the step response of the measurement value PV stabilizes at a single value while the control parameter P supplied to the controlled object 101 is maintained at a constant value. Alternatively, each piece of sample data may include the value of the control parameter P when the measurement value PV stabilizes at the target value SP while the controlled object 101 is controlled by PID control with a target value SP set. Each piece of sample data may be obtained by actually operating the controlled object 101, or may be obtained by a simulator. An approximation curve of the sample data may be calculated by linear regression such as the least squares method, or may be calculated by nonlinear regression such as polynomial approximation or Gaussian process regression.
[0063] 5 shows an approximation curve of sample data. The horizontal axis (x-axis) in the figure represents the target value SP (or the measurement value PV at the equilibrium point O), and the vertical axis (y-axis) represents the value of the control parameter P at the equilibrium point. In the example shown in this figure, the approximation curve may be y=0.8165x-0.0002, and the target value SP may be changed from 50, which serves as the reference target value, to 80. In this case, the control parameter acquisition unit 204 may multiply 30, which is the difference between the target value SP and the reference target value, by 0.8165, which is a preset coefficient, to determine the shift amount Δp as 24.495. Alternatively, the control parameter acquisition unit 204 may use a preset function formula y=0.8165x-0.0002 to calculate the value of the control parameter P as 40.8248 (=0.8165×50) when the target value SP is the reference target value of 50, and the value of the control parameter P as 65.3198 (=0.8165×80) when the target value SP is 80, and determine the shift amount Δp to be 24.495 from the difference between the two. Note that the value of the control parameter P when the target value SP is the reference target value (40.8248 in this example) may be stored in advance in the control parameter acquisition unit 204.
[0064] According to the device 200 having the control parameter acquisition unit 204 described above, the shift amount is determined by multiplying the difference between the target value SP and the reference target value by a preset coefficient, so that a highly accurate recommended control parameter Pr that matches the changed target value SP can be acquired.
[0065] Furthermore, since the shift amount is determined using a predetermined relational expression that indicates the relationship between the value set for the target value SP and the value of the control parameter P when the measurement value PV stabilizes at that value, the shift amount can be determined by inputting the target value SP into the function expression. Therefore, it is possible to easily obtain highly accurate recommended control parameters Pr that match the changed target value SP.
[0066] <1.5. Operation> 6 shows the operation of the device 200. The device 200 may control the control target 101 by performing the processes of steps S11 to S29. Note that this operation may be started in response to activation of the device 200. Furthermore, the learning process of the change amount output model 2061 may be completed at the start of the operation.
[0067] In step S11, the target value acquisition unit 202 acquires a target value SP of a state related to the control object 101. The target value acquisition unit 202 may acquire a preset reference target value as the target value SP, or may acquire a value different from the reference target value as the target value SP. When the process of step S11 is executed for the first time, the target value acquisition unit 202 may acquire the reference target value as the target value SP. In response to acquiring a value different from the reference target value as the target value SP, the target value acquisition unit 202 may output a target value change signal indicating the set target value SP.
[0068] In step S13, the measurement value acquiring unit 201 acquires the measurement value PV of the state related to the control target 101. The target value acquiring unit 202 may acquire the measurement value PV from the sensor 102 of the facility 100.
[0069] In step S15, the deviation acquisition unit 203 acquires the deviation between the target value SP acquired in step S11 and the measurement value PV acquired in step S13.
[0070] In step S17, the control parameter acquisition unit 204 acquires the control parameter P supplied to the control target 101. The control parameter acquisition unit 204 may acquire the control parameter P supplied to the control target 101 in the most recent control cycle from the control unit 207. As an example, the control parameter acquisition unit 204 may acquire and temporarily store the control parameter P output from the control unit 207 to the control target 101 in the processing of step S27 described below, and read out the control parameter P in step S17. When step S17 is executed for the first time, that is, when the processing of step S27 has not been executed, the control parameter acquisition unit 204 may acquire an initial value of the control parameter P that has been set in advance.
[0071] In step S19, the control parameter acquisition unit 204 determines whether the target value SP acquired in step S11 is a reference target value. The control parameter acquisition unit 204 may make this determination based on whether a target value change signal has been acquired from the target value acquisition unit 202. If it is determined in step S19 that the target value SP is a reference target value (step S19; Yes), the process may proceed to step S25. If it is determined in step S19 that the target value SP is not a reference target value (step S19; No), the process may proceed to step S21.
[0072] In step S21, the control parameter acquisition unit 204 determines the shift amount Δp of the control parameter P. The control parameter acquisition unit 204 may determine the shift amount Δp in accordance with the target value SP acquired in step S11.
[0073] In step S23, the control parameter acquisition unit 204 shifts the control parameter P acquired in step S17 by the determined shift amount Δp to acquire a shifted control parameter P+Δp.
[0074] In step S25, the first supply unit 205 supplies the control parameter P supplied from the control parameter acquisition unit 204 and the deviation supplied from the deviation acquisition unit 203 to the control model 206. If it is determined in step S19 that the target value SP is the reference target value, the first supply unit 205 may supply the control parameter P acquired by the control parameter acquisition unit 204 in step S17 as is to the control model 206. If it is determined in step S19 that the target value SP is not the reference target value, the first supply unit 205 may supply the control parameter P shifted in step S23, i.e., the shifted control parameter P+Δp, to the control model 206.
[0075] As a result, a recommended control parameter Pr corresponding to the input control parameter P and deviation is output from the control model 206. As an example, in the present embodiment, a recommended change amount corresponding to the input control parameter P and deviation may be output from the change amount output model 2061, and the recommended change amount and the control parameter P acquired in step S17 may be added by the adder 2062 to generate the recommended control parameter Pr.
[0076] In step S27, the control unit 207 outputs the recommended control parameters Pr from the control model 206. The control unit 207 may supply the recommended control parameters Pr as control parameters P to the controlled object 101 to control the controlled object 101.
[0077] In step S29, the target value acquisition unit 202 determines whether the target value SP is changed by the operator. If it is determined that the target value SP is not changed (step S29; No), the process may proceed to step S13. If it is determined that the target value SP is changed (step S29; Yes), the process may proceed to step S11.
[0078] <1.6. Example of operation> 7 shows the transition of the measurement value PV and the control parameter P when the target value SP is changed. In the figure, the horizontal axis represents time (seconds), and the vertical axis represents the values of the measurement value PV and the control parameter P. In this figure, as an example, the control parameter P represents the instruction value IV of the valve opening, and the measurement value PV and the control parameter P are normalized within the range of 0 to 100. As shown in this figure, in the device 200 according to this embodiment, even when the target value SP is changed, the controlled object 101 is controlled so that the measurement value PV matches the changed target value SP.
[0079] 2. Second Embodiment <2.1. System 1A> Fig. 8 shows a system 1A according to the second embodiment. Components that are substantially the same as those in the system 1 shown in Fig. 1 are given the same reference numerals, and descriptions thereof will be omitted. The system 1A includes an apparatus 200A. The apparatus 200A may include a deviation acquisition unit 203A, an unstable state detection unit 211A, a second supply unit 212A, a shift amount output model 213A, a control parameter acquisition unit 204A, and a learning processing unit 214A.
[0080] <2.1.1. Deviation acquisition unit 203A> The deviation acquiring unit 203A acquires the deviation between the measured value PV of the state of the control object 101 and the target value SP in the same manner as the deviation acquiring unit 203 in the first embodiment described above. The deviation acquiring unit 203A may further acquire the rate of change of the acquired deviation. For example, each time the deviation acquiring unit 203A acquires a deviation, the deviation acquiring unit 203A may calculate the rate of change of the deviation by dividing the amount of change in the most recent two deviations by the interval between the acquisition timings. The deviation acquiring unit 203A may supply the acquired deviation to the first supplying unit 205 and the unstable state detecting unit 211A. The deviation acquiring unit 203A may supply the acquired deviation and its rate of change to the second supplying unit 212A. The deviation acquiring unit 203A may store the acquired deviation and its rate of change in a storage unit (not shown).
[0081] <2.1.2. Unstable State Detector 211A> The unstable state detection unit 211A is an example of a first detection unit, and detects that the deviation acquired by the deviation acquisition unit 203A after the control target 101 is controlled by the recommended control parameter Pr does not stabilize within a reference range. The reference range may be a range of any size including 0 at the center. The deviation not stabilizing within the reference range may mean that the deviation falls outside the reference range at least at one point within a second reference time after the control parameter P is output from the control unit 207 to the control target 101 and a first reference time has elapsed. The situation in which the deviation does not stabilize within the reference range may occur due to a change (also referred to as a disturbance) in the external environment of the equipment 100. Upon detecting that the deviation does not stabilize within the reference range, the unstable state detection unit 211A may supply a signal (also referred to as an unstable state detection signal) indicating this to the control parameter acquisition unit 204A.
[0082] <2.1.3.Second supply section 212A> The second supply unit 212A supplies the deviation acquired by the deviation acquisition unit 203A and the rate of change of the deviation to the shift amount output model 213A. The second supply unit 212A may perform the supply to the shift amount output model 213A at each reference interval. The reference interval may have a length equal to or longer than the control period of the controlled object 101 by the control unit 207. The reference interval may have a length corresponding to the time constant of the controlled object 101 (for example, a length obtained by multiplying the time constant by an integer).
[0083] <2.1.4. Shift Amount Output Model 213A> In response to input of the deviation and the rate of change of the deviation, the shift amount output model 213A outputs a recommended shift amount Δpr that recommends a shift for the control parameter P by the control parameter acquisition unit 204. In response to input of the deviation and the rate of change from the second supply unit 212A, the shift amount output model 213A may output the recommended shift amount Δpr to the control parameter acquisition unit 204A.
[0084] The shift amount output model 213A may indicate a correspondence relationship between a combination of a deviation and a rate of change of the deviation, and a recommended shift amount Δpr. As an example, the shift amount output model 213A may be a shift amount map that maps a correspondence relationship between a combination of a deviation and a rate of change, and a recommended shift amount Δpr. Alternatively, the shift amount output model 213A may be a table that associates combinations of a deviation and a rate of change with the recommended shift amount Δpr. The shift amount output model 213A may be generated by a learning process by the learning processing unit 214A, and may be stored in a storage unit (not shown).
[0085] <2.1.5. Control parameter acquisition unit 204A> The control parameter acquisition unit 204A acquires the shifted control parameter P+Δp in the same manner as the control parameter acquisition unit 204 in the first embodiment.
[0086] In addition, when the unstable state detection unit 211A detects that the deviation acquired by the deviation acquisition unit 203A is not stable within the reference range, the control parameter acquisition unit 204A may shift the control parameter P supplied to the control object 101 (in this embodiment, as an example, the control parameter P supplied from the control unit 207) and acquire the shifted control parameter P+Δp.
[0087] When the unstable state detection unit 211A detects that the deviation acquired by the deviation acquisition unit 203A is not stable within the reference range and when a supply is made from the second supply unit 212A to the shift amount output model 213A, the control parameter acquisition unit 204A may shift the control parameter P supplied to the control object 101 by the recommended shift amount Δpr output from the shift amount output model 213A.
[0088] The control parameter acquiring unit 204A may supply the acquired shifted control parameter P+Δp to the first supplying unit 205. Note that, as an example in the present embodiment, when the target value SP is a reference target value and the deviation is stable at the target value SP, the control parameter acquiring unit 204A supplies the control parameter P acquired from the control unit 207 to the first supplying unit 205 without shifting it, but the control parameter P may be shifted by 0 and supplied to the first supplying unit 205 as a shifted control parameter P+Δp.
[0089] <2.1.6. Learning Processing Unit 214A> The learning processing unit 214A is an example of a first learning processing unit, and performs learning processing of the shift amount output model 213A using learning data including the deviation acquired by the deviation acquiring unit 203A, the rate of change of the deviation, and the shift amount Δp of the control parameter P acquired by the control parameter acquiring unit 204A. The deviation and rate of change included in the learning data may be the deviation and rate of change acquired when the target value SP is a reference target value.
[0090] The learning data may be generated by a simulator (not shown) of the system 1, and may include a combination of the deviation acquired by the deviation acquiring unit 203A in a simulation in which an arbitrary disturbance is artificially applied to the equipment 100, the rate of change of the deviation, and the shift amount Δp of the control parameter P shifted by the control parameter acquiring unit 204A. The simulator may be created using actual measurement data of the equipment 100 using any system identification technology. After the shift amount output model 213A is generated, the learning data may include the deviation when the measurement value PV does not stabilize at the target value SP when the equipment 100 is actually operated, the rate of change of the deviation, and the shift amount Δp of the control parameter P acquired by the control parameter acquiring unit 204A.
[0091] The learning processing unit 214A may perform learning of the shift amount output model 213A so as to output a recommended shift amount Δpr recommended for increasing the reward value in response to input of the deviation and the rate of change of the deviation. When a reward value corresponding to the state of the controlled object 101 at a predetermined time point (for example, the time point at which the deviation and rate of change are acquired) (for example, a reward value obtained by inputting a value corresponding to the measurement value PV at that time into a reward function) is set as a base reward value, the recommended shift amount Δpr may be a shift amount recommended for increasing the reward value above the base reward value. As an example, the learning processing unit 214A may perform learning using an algorithm based on a kernel dynamic policy programming method. The reward value may be a value determined by a preset reward function. The reward function may be a function based on the deviation, for example, a function in which the smaller the deviation, the greater the reward value. Note that when the deviation acquisition unit 203A acquires deviations for each of multiple physical quantities, the reward function may be a function based on the sum of the multiple deviations or a function based on the result of weighted addition of the multiple deviations.
[0092] According to the above-described device 200A, in response to detection that the acquired deviation is not stable at the target value SP, the control parameter P supplied to the controlled object 101 is shifted to acquire the shifted control parameter P+Δp. Therefore, it is possible to acquire the recommended control parameter Pr in response to the occurrence of a disturbance using a single control model 206. Furthermore, since the recommended control parameter Pr can be acquired even when a disturbance occurs, it is possible to improve the robustness of the device 200A.
[0093] Furthermore, the control parameter P is shifted by the recommended shift amount Δpr output from the shift amount output model 213A in accordance with the deviation and the rate of change of the deviation. Therefore, by generating the shift amount output model 213A in advance, the recommended control parameter Pr when a disturbance occurs can be easily obtained.
[0094] Furthermore, since the deviation and the rate of change of the deviation are supplied to the shift amount output model 213A for each reference interval, it is possible to change the shift amount Δp in accordance with the state of the disturbance and obtain an appropriate recommended control parameter Pr.
[0095] Furthermore, using learning data including the deviation, the rate of change of the deviation, and the shift amount Δp of the control parameter P, the shift amount output model 213A performs a learning process to output a recommended shift amount Δpr recommended for increasing a reward value determined by a predetermined reward function in response to input of the deviation and the rate of change of the deviation. Therefore, an appropriate recommended shift amount Δpr can be reliably acquired from the shift amount output model 213A.
[0096] <2.2. Shift amount output model 213A> Fig. 9 shows an example of the shift amount output model 213A. In Fig. 9, the vertical axis represents the rate of change of the deviation, and the horizontal axis represents the deviation.
[0097] The shift amount output model 213A may indicate a correspondence relationship between a combination of a deviation and a rate of change and a recommended shift amount. The shift amount output model 213A of this example may be a shift amount map that maps a correspondence relationship between a combination of a deviation and a control parameter P and a recommended shift amount Δpr. The shift amount output model 213A may be divided into a plurality of regions, each region corresponding to a different recommended shift amount Δpr, according to the combination of the deviation and the rate of change, and may output a recommended shift amount Δpr corresponding to the coordinate position of the input combination of the deviation and the rate of change.
[0098] The shift amount output model 213A may include information about the entire region of the shift amount map. Alternatively, the shift amount output model 213A may include only information indicating the boundary of each region (for example, coordinates or a function formula indicating the boundary) and the recommended shift amount Δpr corresponding to the region. In this case, the storage area for storing the shift amount output model 213A can be reduced.
[0099] 10 shows another example of the shift amount output model 213A. As shown in this figure, the shift amount output model 213A may be a table in which combinations of deviation and rate of change are associated with recommended shift amounts Δpr.
[0100] <2.3. Operation> 11 shows the operation of the device 200A. The device 200A may control the control target 101 by performing the processes of steps S11 to S45. Note that this operation may start in response to the device 200A being started. Furthermore, the learning processes of the change amount output model 2061 and the shift amount output model 213A may be completed at the start of the operation. The operation of the device 200A according to the second embodiment differs from the operation of the device 200 according to the first embodiment in that the processes of steps S31 to S45 are performed between steps S27 and S29.
[0101] In step S31, the measurement value acquiring unit 201 acquires the measurement value PV of the state of the control target 101. The target value acquiring unit 202 may acquire the measurement value PV in the same manner as in step S13.
[0102] In step S33, the deviation acquisition unit 203A acquires the deviation between the target value SP acquired in step S11 and the measurement value PV acquired in step S31. The deviation acquisition unit 203A also acquires the rate of change of the acquired deviation.
[0103] In step S35, the unstable state detection unit 211A determines whether the deviation is stabilized within the reference range. After the control parameter P is output from the control unit 207 to the controlled object 101 in the process of step S27 and the first reference time has elapsed, the unstable state detection unit 211A may detect that the deviation is not stabilized within the reference range in response to the deviation being outside the reference range at least at one point within the second reference time, and make a determination to that effect.
[0104] If it is determined in step S35 that the deviation is stable within the reference range (step S35; Yes), the process may proceed to step S29. If it is determined in step S35 that the deviation is not stable within the reference range (step S35; No), the process may proceed to step S37.
[0105] In step S37, the control parameter acquisition unit 204A determines a shift amount Δp of the control parameter P. In response to the deviation and rate of change acquired in step S33 being supplied to the shift amount output model 213A, the control parameter acquisition unit 204A may determine, as the shift amount Δp, a recommended shift amount Δpr output from the shift amount output model 213A.
[0106] In step S39, the control parameter acquisition unit 204A acquires the control parameters P supplied to the control target 101. The control parameter acquisition unit 204A may acquire, from the control unit 207A, the control parameters P supplied to the control target 101 in the most recent control cycle. As an example, the control parameter acquisition unit 204A may acquire and temporarily store the control parameters P output from the control unit 207 to the control target 101 in the most recently executed process of step S27 or step S45 described below, and read them out in step S39.
[0107] In step S41, the control parameter acquisition unit 204A shifts the control parameter P acquired in step S39 by the shift amount Δp determined in step S37 to acquire the shifted control parameter P+Δp.
[0108] In step S43, the first supply unit 205 supplies the shifted control parameter P+Δp acquired in step S41 and the deviation acquired in step S33 to the control model 206.
[0109] In step S45, the control unit 207 outputs the recommended control parameters Pr from the control model 206. The control unit 207 may supply the recommended control parameters Pr as control parameters P to the controlled object 101 to control the controlled object 101. When the processing of step S45 is completed, the processing may proceed to step S31.
[0110] <2.4. Example of operation> 12 shows the transition of the measurement value PV and the control parameter P when a disturbance occurs. The horizontal axis in the figure indicates time (seconds), and the vertical axis indicates the values of the measurement value PV and the control parameter P. In this figure, as an example, the control parameter P indicates the instruction value IV of the valve opening, and the measurement value PV and the control parameter P are normalized within the range of 0 to 100. As shown in this figure, in the device 200A according to this embodiment, even when a disturbance occurs, the controlled object 101 is controlled so that the measurement value PV matches the target value SP.
[0111] 3. Third Embodiment <3.1. System 1B> FIG. 13 shows a system 1B according to a third embodiment. Components that are substantially the same as those in the systems 1 and 1A shown in FIGS. 1 and 2 are given the same reference numerals, and descriptions thereof will be omitted. The system 1B includes an apparatus 200B. The apparatus 200B according to this embodiment is capable of changing the controlled object 101. The controlled object 101 before and after the change may have similar process dynamic characteristics, and may be either a first-order lag system or a second-order lag system. In other words, the controlled object 101 before and after the change may have the same order of the denominator of the transfer function indicating the relationship between input and output. The apparatus 200B may include an object information acquisition unit 221B, a control unit 207B, and a control parameter acquisition unit 204B.
[0112] <3.1.1. Target information acquisition unit 221B> The target information acquisition unit 221B is an example of a second detection unit, and detects that the control target 101 has been changed. The target information acquisition unit 221B may detect that the control target 101 has been changed by acquiring identification information of the changed control target 101 from the operator each time the control target 101 is changed. In response to the change in the control target 101, the target information acquisition unit 221B may supply a signal indicating the identification information of the changed control target 101 (also referred to as a control target change signal) to the control unit 207B and the control parameter acquisition unit 204B.
[0113] <3.1.2. Control unit 207B> Similar to the control unit 207B in the first embodiment, in response to supply from the first supply unit 205 to the control model 206, the control unit 207B outputs the recommended control parameters Pr output from the control model 206 as control parameters P. In response to receiving a control target change signal from the target information acquisition unit 221B, the control unit 207B may output the control parameters P to the changed control target 101, and control the changed control target 101.
[0114] <3.1.3. Control parameter acquisition unit 204B> The control parameter acquisition unit 204B acquires the shifted control parameter P+Δp in the same manner as the control parameter acquisition unit 204A in the second embodiment.
[0115] In addition, in response to the object information acquisition unit 221B detecting that the control object 101 has been changed, the control parameter acquisition unit 204B may shift the control parameter P supplied to the control object 101 (the control parameter P supplied from the control unit 207, as an example, in this embodiment) to acquire the shifted control parameter P+Δp. The control parameter P supplied to the control object 101 may be the control parameter P supplied to the control object 101 before the change, or may be the control parameter P supplied to the control object 101 after the change.
[0116] When the controlled object 101 is changed, the control parameter acquisition unit 204B may determine the shift amount Δp based on the difference between the values of the control parameter P at the equilibrium points of the controlled objects 101 before and after the change.
[0117] For example, the control parameter acquisition unit 204B may determine, as the shift amount Δp, the difference between the value of the control parameter P at the equilibrium point O acquired when the target value SP is set as a reference target value in advance and the controlled object 101 before and after the change is used. In this case, the control parameter acquisition unit 204B may store in advance, for each controlled object 101, the value of the control parameter P at the equilibrium point O when the target value SP is set as a reference target value.
[0118] Alternatively, the control parameter acquisition unit 204B may set the target value SP to the same value as the current setting value in advance, and determine as the shift amount Δp the difference between the value of the control parameter P at the equilibrium point O acquired when the controlled object 101 is used before and after the change. In this case, the control parameter acquisition unit 204B may store in advance, for each controlled object 101, a relational expression indicating the relationship between the target value SP and the control parameter P at the equilibrium point O, and may calculate the value of the control parameter P at the equilibrium point corresponding to the current target value SP from the function expression.
[0119] In response to the fact that the target value SP is the reference target value in step S19 described above and the control object 101 has been changed, the control parameter acquisition unit 204B may determine the shift amount Δp based on the difference between the values of the control parameter P at the equilibrium point of each control object 101 before and after the change, and supply the shifted control parameter P+Δp to the first supply unit 205. When the target value SP is changed from the reference target value after the control object 101 is changed, the control parameter acquisition unit 204B may determine the shift amount Δp using the above-described relational expression that indicates the relationship between the target value SP and the value of the control parameter P at the equilibrium point O for the control object 101 after the change.
[0120] As an example in the present embodiment, when the target value SP is a reference target value, the deviation is stable at the target value SP, and the controlled object 101 is not changed, the control parameter acquiring unit 204B may supply the control parameter P acquired from the control unit 207 to the first supplying unit 205 without shifting it. Alternatively, the control parameter acquiring unit 204 may supply the control parameter P to the first supplying unit 205 as a shifted control parameter P+Δp, which is obtained by shifting the control parameter P by 0.
[0121] According to the above-described device 200B, in response to a change in the control object 101, the control parameter P supplied to the control object 101 is shifted to obtain the shifted control parameter P+Δp. Therefore, a single control model 206 can be used to obtain the recommended control parameter Pr in response to a change in the control object 101, and therefore the versatility of the device 200B can be improved and the device 200B can be made smaller than when a separate control model is used for each control object 101.
[0122] Furthermore, in the control model 206, a recommended change amount Δpr of the control parameter P is output from the change amount output model 2061 in accordance with the deviation and the control parameter P that has already been supplied to the controlled object 101, and the supplied control parameter P is added to the recommended change amount Δpr to calculate the recommended control parameter Pr. Therefore, even when the control model 206 is used for another controlled object 101 by inputting the control parameter P to the control model 206 as the shifted control parameter P+ΔP, it is possible to obtain the recommended control parameter Pr based on the supplied control parameter P. Therefore, it is possible to obtain an appropriate recommended control parameter Pr that is suited to the changed controlled object 101, regardless of differences in process gain for each controlled object 101.
[0123] <3. Modifications> In the first to third embodiments, the control model 206 has been described as including the change amount output model 2061 and the adder 2062. However, the control model 206 does not have to include these components as long as it outputs the recommended control parameter Pr in response to input of the deviation and the control parameter P. In this case, the control model 206 may be a learning model generated by an algorithm such as a kernel dynamic policy programming method, deep reinforcement learning, a support vector machine, a logistic regression, a decision tree, or a neural network. The learning processing unit 208 may perform learning processing of the control model 206 using learning data including the deviation acquired by the deviation acquiring unit 203 and the control parameter P acquired by the control parameter acquiring unit 204.
[0124] Furthermore, the change amount output model 2061 and the shift amount output model 213A have been described as maps and tables generated by the learning algorithm of the kernel dynamic policy programming method, but they may be generated by other algorithms such as deep reinforcement learning, support vector machines, logistic regression, decision trees, and neural networks, or may be models in other forms different from maps and tables.
[0125] Furthermore, although the deviation and the control parameter P are described as being input to the change amount output model 2061, other values may also be input. Similarly, although the deviation and the rate of change are described as being input to the shift amount output model 213A, other values may also be input. The other values may be, for example, the differential value or integral value of the measurement value by the sensor 102.
[0126] Furthermore, although the devices 200, 200A, and 200B have been described as having the measurement value acquiring unit 201, the target value acquiring unit 202, and the learning processing unit 208, any of these may be omitted. If the devices 200, 200A, and 200B do not have the measurement value acquiring unit 201 and the target value acquiring unit 202, the deviation acquiring units 203 and 203A may acquire a deviation calculated by an external device. If the devices 200, 200A, and 200B do not have the learning processing unit 208, they may have a change amount output model 2061 that has been trained in advance by an external device.
[0127] Furthermore, although the control parameter acquisition units 204, 204A, and 204B have been described as calculating and acquiring the shifted control parameters P+Δp, the shifted control parameters P+Δp calculated outside the devices 200, 200A, and 200B may be acquired.
[0128] Furthermore, in the second and third embodiments, the devices 200A and 200B have been described as having the learning processing unit 214A, but they may not have the learning processing unit 214A. In this case, the devices 200A and 200B may have a shift amount output model 213A that has been trained in advance by an external device.
[0129] Furthermore, in the above second embodiment, the control parameter acquisition unit 204 was described as shifting the control parameter P when the target value SP is changed from the reference target value and when the deviation is not stabilized within the reference range, but it is not necessary to shift the control parameter P when the target value SP is changed from the reference target value.
[0130] Similarly, in the above third embodiment, the control parameter acquisition unit 204 was described as shifting the control parameter P when the target value SP is changed from the reference target value, when the deviation is not stabilized within the reference range, and when the control object 101 is changed, but it is not necessary to shift the control parameter P when at least one of the cases where the target value SP is changed from the reference target value and the deviation is not stabilized within the reference range.
[0131] Various embodiments of the present invention may also be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which operations are performed or (2) sections of an apparatus responsible for performing the operations. Particular stages and sections may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable medium, and / or a processor provided with computer-readable instructions stored on a computer-readable medium. Dedicated circuitry may include digital and / or analog hardware circuitry, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuitry may include reconfigurable hardware circuitry, including logical AND, OR, XOR, NAND, NOR, and other logic operations, flip-flops, registers, memory elements such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0132] A computer-readable medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that the computer-readable medium having instructions stored thereon comprises an article of manufacture containing instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable media may include electronic, magnetic, optical, electromagnetic, and semiconductor storage media. More specific examples of computer-readable media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray (RTM) disc, memory stick, integrated circuit card, and the like.
[0133] The computer readable instructions may include either assembler instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages such as the “C” programming language or similar programming languages.
[0134] The computer-readable instructions may be provided to a processor or programmable circuitry of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, either locally or over a wide-area network (WAN) such as a local area network (LAN), the Internet, etc., which executes the computer-readable instructions to create means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.
[0135] 14 illustrates an example of a computer 2200 in which aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 may cause the computer 2200 to function as or perform operations associated with an apparatus or one or more sections of the apparatus according to embodiments of the present invention, and / or to perform a process or steps of a process according to embodiments of the present invention. Such programs may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.
[0136] A computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0137] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the graphics controller 2216 itself, and causes the image data to be displayed on the display device 2218.
[0138] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.
[0139] The ROM 2230 stores therein a boot program or the like that is executed by the computer 2200 upon activation, and / or programs that depend on the hardware of the computer 2200. The input / output chip 2240 may also connect various input / output units to the input / output controller 2220 via a parallel port, a serial port, a keyboard port, a mouse port, etc.
[0140] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by realizing information manipulation or processing in accordance with the use of the computer 2200.
[0141] For example, when communication is performed between the computer 2200 and an external device, the CPU 2212 may execute a communication program loaded into the RAM 2214 and instruct the communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of the CPU 2212, the communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in the RAM 2214, the hard disk drive 2224, the DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes reception data received from the network to a reception buffer processing area or the like provided on the recording medium.
[0142] The CPU 2212 may also cause all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and perform various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.
[0143] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 2214. The CPU 2212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored on the recording medium, the CPU 2212 may search for an entry that matches a condition specified by the attribute value of the first attribute from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0144] The above-described programs or software modules may be stored in a computer-readable medium on or near the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable medium, thereby providing the programs to the computer 2200 via the network.
[0145] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.
[0146] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order. [Explanation of symbols]
[0147] 1 System 100 equipment 101 Control Object 102 Sensors 200 equipment 201 Measurement acquisition unit 202 Target value acquisition unit 203 Deviation acquisition unit 204 Control parameter acquisition unit 205 1st supply section 206 Control Model 207 Control Unit 208 Learning processing unit 211 Unstable state detection unit 212 2nd supply section 213 Shift Amount Output Model 214 Learning processing unit 221 Target Information Acquisition Unit 2061 Change Output Model 2062 Addition section 2200 Computer 2201 DVD-ROM 2210 host controller 2212 CPU 2214 RAM 2216 Graphics Controller 2218 Display Device 2220 Input / Output Controller 2222 communication interface 2224 hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 I / O chip 2242 keyboard
Claims
1. a deviation acquisition unit that acquires a deviation between a measured value of a state of a controlled object and a target value; a control parameter acquisition unit that acquires shifted control parameters obtained by shifting the control parameters supplied to the control object; a first supply unit that supplies the deviation acquired by the deviation acquisition unit and the shifted control parameter acquired by the control parameter acquisition unit to a control model that outputs recommended control parameters to be supplied to the controlled object in response to input of the deviation and the control parameter; an output unit that outputs the recommended control parameters output from the control model in response to the supply from the first supply unit to the control model; An apparatus comprising:
2. a target value acquisition unit that acquires a target value of the state, The device according to claim 1 , wherein the control parameter acquisition unit shifts the control parameters supplied to the controlled object in response to the target value being changed from a reference target value, and acquires the shifted control parameters.
3. 3. The device according to claim 2, wherein the control parameter acquisition unit, in response to a change in the target value from a reference target value, shifts the control parameter supplied to the controlled object by a shift amount corresponding to the target value, and acquires the shifted control parameter.
4. The device according to claim 3 , wherein the control parameter acquisition unit determines the shift amount by multiplying a difference between the target value and the reference target value by a preset coefficient.
5. 4. The device according to claim 3, wherein the control parameter acquisition unit determines the shift amount using a predetermined relational expression that indicates a relationship between a value set as the target value and a value of the control parameter when the measurement value stabilizes at that value.
6. a first detection unit that detects that the deviation acquired by the deviation acquisition unit is not stable within a reference range after the control target is controlled using the recommended control parameters; The device according to any one of claims 1 to 3, wherein the control parameter acquisition unit shifts the control parameter supplied to the controlled object in response to detection by the first detection unit that the deviation acquired by the deviation acquisition unit is not stable within the reference range, and acquires the shifted control parameter.
7. a second supply unit that supplies the deviation acquired by the deviation acquisition unit and the rate of change of the deviation to a shift amount output model that outputs a recommended shift amount that recommends a shift of a control parameter by the control parameter acquisition unit in response to input of the deviation and the rate of change of the deviation; 7. The device according to claim 6, wherein the control parameter acquisition unit shifts the control parameters supplied to the controlled object by the recommended shift amount output from the shift amount output model in response to the first detection unit detecting that the deviation acquired by the deviation acquisition unit is not stable within the reference range and the second supply unit supplying the shift amount to the shift amount output model, thereby acquiring the shifted control parameters.
8. The apparatus according to claim 7 , wherein the second supplying unit supplies the shift amount output model for each reference interval.
9. 8. The device according to claim 7, further comprising: a first learning processing unit that performs learning processing of the shift amount output model using learning data including the deviation acquired by the deviation acquiring unit, a rate of change of the deviation, and a shift amount of the control parameter acquired by the control parameter acquiring unit, so as to output the recommended shift amount recommended for increasing a reward value determined by a preset reward function in response to input of the deviation and the rate of change of the deviation.
10. a second detection unit that detects that the control target has been changed; 2. The device according to claim 1, wherein the control parameter acquisition unit shifts the control parameters supplied to the control object in response to the second detection unit detecting that the control object has been changed, and acquires the shifted control parameters.
11. The control model is a change amount output model that outputs a recommended change amount that recommends a change to the control parameter in response to input of the deviation and the control parameter; an adder that calculates the recommended control parameter by adding the control parameter supplied to the controlled object and the recommended change amount output from the change amount output model; 10. The apparatus of claim 1, comprising:
12. 12. The device according to claim 11, further comprising: a second learning processing unit that performs learning processing of the change amount output model using learning data including the deviation acquired by the deviation acquiring unit and the control parameter acquired by the control parameter acquiring unit, and that outputs the recommended change amount recommended for increasing a reward value determined by a preset reward function in response to input of the deviation and the control parameter.
13. a deviation acquisition stage for acquiring a deviation between a measured value of a state of a controlled object and a target value; a control parameter acquisition step of acquiring shifted control parameters by shifting the control parameters supplied to the controlled object; a first supply step of supplying the deviation acquired in the deviation acquisition step and the shifted control parameter acquired in the control parameter acquisition step to a control model that outputs recommended control parameters to be supplied to the controlled object in response to input of the deviation and the control parameter; an output step of outputting the recommended control parameters output from the control model in response to the supply to the control model in the first supply step; A method for providing
14. Computer, a deviation acquisition unit that acquires a deviation between a measured value of a state of a controlled object and a target value; a control parameter acquisition unit that acquires shifted control parameters obtained by shifting the control parameters supplied to the control object; a first supply unit that supplies the deviation acquired by the deviation acquisition unit and the shifted control parameter acquired by the control parameter acquisition unit to a control model that outputs recommended control parameters to be supplied to the controlled object in response to input of the deviation and the control parameter; an output unit that outputs the recommended control parameters output from the control model in response to the supply from the first supply unit to the control model; A program that functions as a
Citation Information
Patent Citations
Process control method and process control device using it
JP2002207502A
Control device and control method
JP2017146665A
Controller, control method, and control program
JP2022143341A
Learning processor, controller, learning processing method, control method, learning program and control program
JP2022156797A
Control device, program therefor, and plant control method
WO2016092872A1