Apparatus, method and program
The control model with sub-control models addresses the inefficiency in existing systems by dynamically adjusting control parameters based on the control object's phase, ensuring rapid and accurate convergence to the target state.
Patent Information
- Application Number
- JP2022180701
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-11-11
AI Technical Summary
Existing control systems lack the ability to efficiently adjust control parameters based on the phase of the control object's state, leading to suboptimal performance in reducing deviations between measured and target values.
A control model with multiple sub-control models, each associated with different phases of the control object's state, outputs recommended control parameters to minimize deviations, using an identification and selection mechanism to determine the appropriate sub-control model based on the measured value and its deviation range.
This approach allows for precise and efficient control adjustments, prioritizing either speed or accuracy based on the deviation, thereby quickly and accurately maintaining the control object's state at the target value.
Smart Images

Figure 0007800387000001 
Figure 0007800387000002 
Figure 0007800387000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus, a method, and a program. [Background technology]
[0002] Patent Documents 1 to 4 state that "a manipulated variable map is selected based on the target value SV, and the manipulated variable MV is calculated using the selected manipulated variable map" (paragraph 0031 of Patent Document 1). [Prior art document] [Patent Documents] [Patent Document 1] JP 2022-156797 A [Patent Document 2] JP 2020-95352 A [Patent Document 3] JP 2021-117699 A [Patent Document 4] JP 2022-014099 A Summary of the Invention
[0003] In a first aspect of the present invention, there is provided an apparatus comprising: a first acquisition unit that acquires a deviation between a measured value of a state related to a control object and a target value; a second acquisition unit that acquires a control parameter supplied to the control object; a control model having a plurality of sub-control models each associated with a plurality of predetermined phases of the state, the control model outputting a recommended control parameter recommended to be supplied to the control object in response to input of the deviation and the control parameter using a sub-control model among the plurality of sub-control models that is associated with a phase corresponding to the measured value; a supply unit that supplies the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit to the control model; and an output unit that outputs the recommended control parameter output from the control model in response to the supply to the control model from the supply unit.
[0004] The above device may further include an identification unit that identifies a phase from the plurality of phases that corresponds to the measurement value, and a selection unit that selects a sub-control model from the plurality of sub-control models that corresponds to the phase identified by the identification unit. The supply unit may supply the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit to a sub-control model selected by the selection unit from among the plurality of sub-control models.
[0005] In the above device having an identification unit and a selection unit, the identification unit may identify the situation depending on which of a plurality of preset numerical ranges the deviation falls within.
[0006] In any of the above devices having an identification unit and a selection unit, at least two of the plurality of sub-control models may output recommended control parameters that are recommended to reduce the deviation between the measured value and a common target value, and in response to input of control parameters.
[0007] In the above device having an identifying unit and a selecting unit, the identifying unit may identify the aspect depending on which of a plurality of preset numerical ranges the measurement value falls within.
[0008] In any of the above devices having an identification unit and a selection unit, at least two of the plurality of sub-control models may be associated with different characteristic target values, and may output recommended control parameters recommended for reducing the deviation between a measured value and the characteristic target value in response to input of control parameters. The device may further include a setting unit that sets one of the characteristic target values of the at least two sub-control models as the target value in response to the situation identified by the identification unit.
[0009] In any of the above devices having an identification unit and a selection unit, each sub-control model may have a change amount output model that outputs a recommended change amount that recommends a change to the control parameter in response to input of a deviation and a control parameter, and an adder that calculates the recommended control parameter by adding the control parameter supplied to the controlled object and the recommended change amount output from the change amount output model. The multiple sub-control models may share the adder.
[0010] In any of the above devices having a specifying unit and a selecting unit, the change amount output models in at least two of the plurality of sub-control models may output the recommended change amounts in different ranges.
[0011] Any of the above devices having an identification unit and a selection unit may further include a learning processing unit that performs learning processing for each change amount output model using learning data including the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit, and that outputs the recommended change amount recommended to increase a reward value determined by a preset reward function in response to input of the deviation and the control parameter.
[0012] In the above device, the learning processing unit may perform learning processing for each change amount output model using separate learning data.
[0013] In a second aspect of the present invention, there is provided a method comprising: a first acquisition step of acquiring a deviation between a measured value of a state related to a controlled object and a target value; a second acquisition step of acquiring a control parameter supplied to the controlled object; a first supply step of supplying the deviation acquired in the first acquisition step and the control parameter acquired in the second acquisition step to a control model having a plurality of sub-control models respectively associated with a plurality of predetermined phases of the state, the control model outputting recommended control parameters recommended to be supplied to the controlled object in response to input of the deviation and control parameters using a sub-control model among the plurality of sub-control models that is associated with a phase corresponding to the measured value; and an output step of outputting the recommended control parameters output from the control model in response to supplying the control model to the control model in the first supply step.
[0014] In a third aspect of the present invention, there is provided a program that causes a computer to function as a first acquisition unit that acquires the deviation between a measured value of a state related to a control object and a target value, a second acquisition unit that acquires a control parameter supplied to the control object, a control model having a plurality of sub-control models each associated with a plurality of predetermined phases of the state, wherein the control model outputs a recommended control parameter recommended to be supplied to the control object in response to input of the deviation and control parameter using a sub-control model among the plurality of sub-control models that is associated with a phase corresponding to the measured value, and an output unit that outputs the recommended control parameter output from the control model in response to supply to the control model from the supply unit.
[0015] The above summary of the invention does not list all of the necessary features of the present invention, and subcombinations of these features may also constitute inventions. [Brief explanation of the drawings]
[0016] [Figure 1] 1 shows a system 1 according to a first embodiment. [Figure 2] The change amount output model 2061 is shown. [Figure 3] The change amount output model 2061 is shown. [Figure 4] 2 shows another example of the change amount output model 2061. [Figure 5] 2 shows another example of the change amount output model 2061. [Figure 6] 2 illustrates the operation of the device 200. [Figure 7] 1 shows the transition of the measured value PV and the control parameter P when the controlled object 101 is controlled. [Figure 8] 1 shows a system 1A according to a modified example. [Figure 9] The table shows the correspondence between the phase ID, the range of the measured value PV, the sub-model ID, and the specific target value. [Figure 10] The operation of the device 200A is shown. [Figure 11] 1 shows the transition of the measured value PV when the controlled object 101 is controlled. [Figure 12] 22 illustrates an example computer 2200 in which aspects of the present invention may be embodied, in whole or in part. DETAILED DESCRIPTION OF THE INVENTION
[0017] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention according to the claims. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0018] <1. System 1> 1 shows a system 1 according to the first embodiment. The system 1 includes a facility 100 and a device 200.
[0019] <1.1.Equipment 100> The facility 100 is a facility or device equipped with a control target 101. For example, the facility 100 may be a plant, or a composite device that combines multiple devices. Examples of plants include industrial plants such as chemical and bio plants, plants that manage and control wellheads and surrounding areas of gas fields and oil fields, plants that manage and control power generation such as hydroelectric, thermal, and nuclear power, plants that manage and control environmental power generation such as solar and wind power, and plants that manage and control water supply and sewage systems, dams, etc.
[0020] The facility 100 is provided with one or more control targets 101. The control target 101 may be a tool, machine, device, or the like to be controlled, and may be a so-called field device. For example, the control target 101 may be a sensor device such as a pressure gauge, flow meter, or temperature sensor, a valve device such as a flow control valve or an on-off valve, or an actuator device such as a fan or a motor. The control target 101 is controlled externally via a wired or wireless connection, or may be controlled manually. The control target 101 may be controlled by a control unit 210 in the device 200. In the present embodiment, as an example, the control target 101 may be controlled by receiving an instructed value IV (instructed value) for a manipulated variable MV (manipulated variable) from the control unit 210.
[0021] The facility 100 may also be provided with one or more sensors 102. Each sensor 102 may measure a measurement value of the internal or external state of the facility 100, i.e., a measurement value of a physical quantity indicating the internal or external state. At least one sensor 102 may measure a measurement value PV (Process Variable) of the state of the controlled object 101. The measurement value PV may be operating data indicating the operating state as a result of controlling the controlled object 101, and may indicate a controlled variable to be controlled. As an example, the measurement value PV may indicate the output of the controlled object 101 itself, or may indicate various values that change depending on the output of the controlled object 101. As an example, the measurement value PV may indicate pressure, temperature, pH, speed, flow rate, etc. Each sensor 102 may supply the measured measurement value PV to the device 200.
[0022] <1.2.Device 200> The device 200 controls the controlled object 101 and may be, for example, a controller for the controlled object 101. The device 200 may output an indication value IV for a manipulated variable MV of the controlled object 101 to perform process control such as adjusting the temperature, adjusting the liquid level, or adjusting the flow rate.
[0023] The device 200 may be a computer such as a personal computer (PC), a tablet computer, a smartphone, a workstation, a server computer, or a general-purpose computer, or may be a computer system in which multiple computers are connected. Such a computer system is also a computer in a broad sense. The device 200 may also be implemented by one or more virtual computer environments executable within a computer. Alternatively, the device 200 may be a dedicated computer designed for AI control, or may be dedicated hardware realized by dedicated circuits. Furthermore, if the device 200 can connect to the Internet, the device 200 may be realized by cloud computing.
[0024] The apparatus 200 may include a measurement value acquiring unit 201, a target value acquiring unit 202, a deviation acquiring unit 203, a control parameter acquiring unit 204, a control model 205, an identifying unit 207, a selecting unit 208, a supplying unit 209, a control unit 210, and a learning processing unit 211. Note that these blocks are functionally separated functional blocks and may not necessarily correspond to the actual device configuration. That is, even if shown as a single block in this diagram, it is not limited to being configured by a single device. Furthermore, even if shown as separate blocks in this diagram, it is not limited to being configured by separate devices.
[0025] <1.2-1. Measurement value acquisition unit 201> The measurement value acquiring unit 201 acquires a measurement value PV of a state related to the control target 101. In the present embodiment, as an example, the measurement value acquiring unit 201 acquires a measurement value PV for one physical quantity from one sensor 102. However, the measurement value acquiring unit 201 may acquire measurement values PV for each of a plurality of physical quantities from a plurality of sensors 102. The measurement value acquiring unit 201 may supply the acquired measurement value PV to the deviation acquiring unit 203.
[0026] <1.2-2. Target value acquisition unit 202> The target value acquiring unit 202 acquires a target value SP (Set Point) of a state related to the control object 101. The target value acquiring unit 202 may acquire the target value SP of the measurement value PV acquired by the measurement value acquiring unit 201. The target value acquiring unit 202 may acquire the target value SP from an operator via an input unit (not shown). In the present embodiment, as an example, the target value acquiring unit 202 may acquire a preset reference target value as the target value SP. The target value acquiring unit 202 may supply the acquired target value SP to the deviation acquiring unit 203.
[0027] <1.2-3. Deviation acquisition unit 203> The deviation acquisition unit 203 is an example of a first acquisition unit, and acquires the deviation between the measurement value PV and the target value SP of the state related to the control target 101. The deviation acquisition unit 203 may acquire the measurement value PV from the measurement value acquisition unit 201 and the target value SP from the target value acquisition unit 202, and calculate the deviation by subtracting the measurement value PV from the target value SP. Alternatively, the deviation acquisition unit 203 may calculate the deviation by subtracting the target value SP from the measurement value PV. The deviation acquisition unit 203 may supply the acquired deviation to the identification unit 207 and the supply unit 209. The deviation acquisition unit 203 may store the acquired deviation in a storage unit (not shown).
[0028] <1.2-4. Control parameter acquisition unit 204> The control parameter acquiring unit 204 is an example of a second acquiring unit, and acquires the control parameter P supplied to the control target 101. The control parameter acquiring unit 204 may acquire the control parameter P from the control unit 210, which will be described later, and in the present embodiment, as an example, may acquire the control parameter P each time the control unit 210 supplies the control parameter P to the control target 101. The control parameter P may indicate an instruction value IV for an operation amount MV of the control target 101. If the control target 101 is a valve, the control parameter P may indicate, for example, a valve opening. The control parameter acquiring unit 204 may supply the acquired control parameter P to the supplying unit 209.
[0029] <1.2-5. Control Model 205> In response to input of the deviation and the control parameter P, the control model 205 outputs a recommended control parameter Pr that is recommended to be supplied to the controlled object 101. In response to input of the deviation and the control parameter P supplied to one controlled object 101, the control model 205 may output a recommended control parameter Pr that is recommended to be supplied to the one controlled object 101. The recommended control parameter Pr may indicate a recommended instruction value IV for the manipulated variable MV of the controlled object 101.
[0030] The control model 205 may output a recommended control parameter Pr to a control unit 210, which will be described later, in response to input of the deviation and the control parameter P from a supply unit 209, which will be described later. In the present embodiment, as an example, the deviation between a measurement value PV and a target value SP for one physical quantity is described as being input to the control model 205, but the deviations between the measurement value PV and the target value SP for a plurality of physical quantities may also be input.
[0031] The control model 205 may have a plurality of sub-control models 206 (two sub-control models 206a and 206b, for example, in this embodiment) respectively associated with a plurality of phases set in advance for the state of the control object 101, and may output the recommended control parameters Pr using the sub-control models 206 associated with a phase according to the measurement value PV. A phase may be a state of the control object 101 at a certain point in time. For example, the plurality of phases may include a first phase in which the measurement value PV is close to the target value SP, and a second phase in which the measurement value PV is far from the target value SP.
[0032] <1.2-5(1). Sub-control model 206> Each sub-control model 206 may output recommended control parameters that are recommended to reduce the deviation in response to the input of the deviation and control parameters. The multiple sub-control models 206 may be provided independently of each other and may be able to acquire the deviation and control parameters independently of each other from a supply unit 209 described below. The two sub-control models 206a, 206b may be associated with a common target value SP, and may output recommended control parameters that are recommended to reduce the deviation in response to the input of the deviation between the common target value SP and the measurement value PV and the control parameters.
[0033] In this embodiment, as an example, the sub-control model 206a may output recommended control parameters at fine intervals or granularity (also referred to as fineness or precision) in the first phase when the measured value PV is close to the target value SP (i.e., the deviation is small). The sub-control model 206a may be designed to control the control target 101 by prioritizing precision over speed, and is also referred to as an precision-oriented sub-control model 206a.
[0034] The sub-control model 206b may output recommended control parameters with larger intervals and granularity than the sub-control model 206a in a second phase in which the measured value PV is far from the target value SP (i.e., the deviation is large). The sub-control model 206b may be designed to control the control target 101 by prioritizing speed over accuracy, and is also referred to as a speed-oriented sub-control model 206b.
[0035] The sub-control models 206a and 206b may be set with different numerical ranges for the input deviation. For example, the sub-control model 206a may be set with a numerical range of the input deviation that includes 0 and has a small absolute value (for example, a range from -1 to 1; also referred to as a first numerical range), and the sub-control model 206b may be set with a range that does not include 0 and has a larger absolute value than the first numerical range (for example, a range of -1 or less and 1 or more; also referred to as a second numerical range). The first numerical range and the second numerical range may not overlap, and the first numerical range may be a range inside the second numerical range. Each sub-control model 206 may include a change amount output model 2061 and an adder 2062.
[0036] <1.2-5(1-1). Change amount output model 2061> Each change amount output model 2061 of each sub-control model 206 outputs a recommended change amount that recommends a change to the control parameter P in response to input of the deviation and the control parameter P. The change amount output model 2061 of each sub-control model 206 may output recommended change amounts in different ranges. As an example, in this embodiment, the recommended change amount output from the change amount output model 2061 (also referred to as change amount output model 2061a) of the sub-control model 206a may have a smaller order (also referred to as number of digits), interval, or granularity than the recommended change amount output from the change amount output model 2061 (also referred to as change amount output model 2061b) of the sub-control model 206b. The recommended change amount of the change amount output model 2061b may be limited to three types: a maximum value, a minimum value, and an intermediate value (for example, 0), and the measured value PV may be brought closer to the target value SP by control approximating full accelerator / full brake control. The recommended change amounts of the change amount output model 2061a may have more values than those of the change amount output model 2061b.
[0037] The change amount output model 2061 may supply the recommended change amount to the adder 2062. The recommended change amount may indicate a change amount recommended to be made to the control parameter P most recently supplied to the controlled object 101. In the present embodiment, as an example, the recommended change amount may indicate a recommended change amount for the most recent indicated value IV for the manipulated variable MV. The change amount output model 2061 may be generated by a learning process by the learning processor 211, and may be stored in a storage unit (not shown).
[0038] <1.2-5(1-2). Addition section 2062> The adder 2062 calculates a recommended control parameter Pr by adding the control parameter P supplied to the controlled object 101 and the recommended change amount output from the change amount output model 2061. The adder 2062 may be shared by the sub-control models 206a and 206b. The adder 2062 may calculate the recommended control parameter Pr by adding the control parameter P supplied to the controlled object 101 and the recommended change amount output from the change amount output model 2061a of the sub-control model 206a, and may also calculate the recommended control parameter Pr by adding the control parameter P supplied to the controlled object 101 and the recommended change amount output from the change amount output model 2061b of the sub-control model 206b.
[0039] The adder 2062 may calculate the recommended control parameter Pr by adding the most recent control parameter P supplied from the control unit 210 and the recommended change amount supplied from the change amount output model 2061. The adder 2062 calculates the recommended control parameter Pr by adding the most recent control parameter P supplied from the control unit 210 and the recommended change amount supplied from the change amount output model 2061. As shown in the following equation (1), the adder 2062 calculates the recommended control parameter Pr by adding the most recent control parameter P (t-1) and the recommended change amount Δu at time t (t) The recommended control parameter Pr at time t is calculated by adding (t) may be calculated. Pr (t) =P (t-1) +Δu (t) (1)
[0040] The adder 2062 may store the control parameters P supplied from the control unit 210 and use them to calculate the recommended control parameters Pr. The adder 2062 may supply the calculated recommended control parameters Pr to the control unit 210.
[0041] <1.2-6. Specific part 207> The identifying unit 207 identifies a phase from among the multiple phases according to the measured value PV. The identifying unit 207 may identify the phase according to the deviation supplied from the deviation acquiring unit 203. The identifying unit 207 may identify the phase according to which of multiple preset numerical ranges the deviation falls within. The identifying unit 207 may identify the phase according to which of a first numerical range and a second numerical range for the deviation to be input, which are preset in the sub-control models 206a and 206b, the deviation from the deviation acquiring unit 203 falls within.
[0042] The identifying unit 207 may store numerical ranges and identification information (also referred to as phase IDs) of phases in association with each other, and may identify a phase associated with a numerical range that includes the deviation from the deviation acquiring unit 203. As an example in the present embodiment, the identifying unit 207 may identify a first phase as a phase corresponding to the measurement value PV when the deviation from the deviation acquiring unit 203 falls within a first numerical range. The identifying unit 207 may identify a second phase as a phase corresponding to the measurement value PV when the deviation from the deviation acquiring unit 203 falls within a second numerical range. The identifying unit 207 may supply the phase ID of the identified phase to the selecting unit 208.
[0043] <1.2-7.Selection section 208> The selection unit 208 selects, from the plurality of sub-control models 206, the sub-control model 206 that corresponds to the phase identified by the identification unit 207. The selection unit 208 may store the phase ID of each phase and identification information (also referred to as sub-model ID) of each sub-control model 206 in association with each other, and may select the sub-control model 206 having the sub-model ID that corresponds to the phase ID supplied from the identification unit 207. The selection unit 208 may supply the sub-model ID of the selected sub-control model 206 to the supply unit 209.
[0044] <1.2-8. Supply section 209> The supply unit 209 supplies the control model 205 with the deviation acquired by the deviation acquisition unit 203 and the control parameter P acquired by the control parameter acquisition unit 204. The supply unit 209 may supply the control model 205 with the control parameter P supplied from the control unit 210 to the controlled object 101 and a deviation indicating the operating state as a result of controlling the controlled object 101 using the control parameter P.
[0045] The supply unit 209 may supply the deviation and control parameters to the sub-control model 206 selected by the selection unit 208 from among the multiple sub-control models 206 in the control model 205. In the present embodiment, as an example, the supply unit 209 may supply the deviation and control parameters to the sub-control model 206 indicated by the sub-model ID supplied from the selection unit 208.
[0046] <1.2-9. Control unit 210> The control unit 210 is an example of an output unit, and outputs recommended control parameters Pr output from the control model 205 in response to supply from the supply unit 209 to the control model 205. As an example in the present embodiment, the control unit 210 may output the recommended control parameters Pr as control parameters P to the control object 101 to control the control object 101. The control unit 210 may output the control parameters P input by an operator to the control object 101 to control the control object 101. The control unit 210 may output the control parameters P to the control object 101 in accordance with the control period of the control object 101.
[0047] The control unit 210 may store the control parameter P supplied to the control target 101 in a storage unit (not shown). The control unit 210 may store the control parameter P supplied to the control target 101 in the storage unit in association with the deviation acquired by the deviation acquisition unit 203. The control unit 210 may store the control parameter P supplied to the control target 101 in the storage unit in association with the deviation indicating the operating state as a result of controlling the control target 101 using the control parameter P.
[0048] <1.2-10. Learning processing unit 211> The learning processing unit 211 performs learning processing for each change amount output model 2061 using learning data including the deviation acquired by the deviation acquisition unit 203 and the control parameter P acquired by the control parameter acquisition unit 204 .
[0049] The learning processing unit 211 may perform learning of the sub-control model 2061 so as to output a recommended change amount recommended for increasing the reward value in response to input of the deviation and the control parameter P. The recommended change amount may be a change amount recommended for increasing the reward value above a base reward value (e.g., a reward value obtained by inputting a value corresponding to the measurement value PV at that time into a reward function) corresponding to the state of the controlled object 101 at a predetermined time point (e.g., the time point at which the deviation and the control parameter P are acquired) when the base reward value is the reward value. The reward value may be a value determined by a preset reward function. The reward function may be a function based on the deviation, e.g., a function that increases the reward value as the deviation decreases. Note that when the deviation acquisition unit 203 acquires deviations for each of multiple physical quantities, the reward function may be a function based on the sum of the multiple deviations or a function based on the result of weighted addition of the multiple deviations. For example, the learning processing unit 211 may perform learning using a Kernel Dynamic Policy Programming (KDPP) algorithm.
[0050] The learning processing unit 211 may perform learning processing for each change amount output model 2061 using different learning data. For example, when performing learning processing for the change amount output model 2061a, the learning processing unit 211 may perform learning processing using learning data whose deviation falls within a first numerical range. As an example, the learning processing unit 211 may perform learning processing using learning data acquired when the control object 101 is successively controlled in a state where the measurement value PV is close to the target value SP. When performing learning processing for the change amount output model 2061b, the learning processing unit 211 may perform learning processing using learning data whose deviation falls within a second numerical range, or may perform learning processing further using learning data whose deviation falls within the first numerical range. As an example, the learning processing unit 211 may perform learning processing using learning data acquired when the control object 101 is successively controlled in a state where the measurement value PV is far from the target value SP.
[0051] The numerical range of the absolute value of the deviation may differ between the learning data of the change amount output model 2061a and the learning data of the change amount output model 2061b. For example, the numerical range of the absolute value of the deviation in the learning data of the change amount output model 2061a may be closer to 0 than the numerical range of the absolute value of the deviation in the learning data of the change amount output model 2061b. As an example, the deviation in the learning data of the change amount output model 2061a may be closer to 0 than the numerical range of the absolute value of the deviation in the learning data of the change amount output model 2061b. 0 , that is, it may be on the order of one digit, and the deviation in the training data of the change amount output model 2061b is 10 1 It may be of the order of , i.e., a two-digit value.
[0052] Furthermore, the numerical range of the control parameter P may be different between the learning data of the change amount output model 2061a and the learning data of the change amount output model 2061b. For example, the numerical range of the control parameter P in the learning data of the change amount output model 2061a may be a value within a third numerical range that includes the value of the control parameter P when the measurement value PV stabilizes at the target value SP (also referred to as the control parameter P at the equilibrium point). The numerical range of the control parameter P in the learning data of the change amount output model 2061b may be a value within a fourth numerical range that is outside the third numerical range, or may be a value within both the third numerical range and the fourth numerical range. The control parameter P in the learning data of the change amount output model 2061a may have smaller intervals and granularity than the control parameter P in the learning data of the change amount output model 2061b.
[0053] The learning processing unit 211 may perform learning processing for each change amount output model 2061 using learning data including the deviation and control parameter P acquired when the target value SP is the same value. Note that the learning data may be acquired from a simulator (not shown) of the system 1 instead of from the actual system 1. The simulator may be created using actual measurement data of the facility 100 using any system identification technology. Each piece of learning data may be stored in a storage unit (not shown).
[0054] According to the above-described device 200, the control model 205 uses the sub-control model 206 associated with the phase corresponding to the measured value, out of a plurality of sub-control models 206 respectively associated with a plurality of phases set in advance for the state, and outputs a recommended control parameter Pr in response to input of the deviation acquired by the deviation acquisition unit 203 and the control parameter P acquired by the control parameter acquisition unit 204. Therefore, by inputting the deviation and the control parameter P to the control model 205, it is possible to acquire the recommended control parameter Pr corresponding to the phase.
[0055] In addition, a phase is identified according to the measurement value PV, and a sub-control model 206 corresponding to the identified phase is selected from among the multiple sub-control models 206, and the deviation and control parameters are supplied to the selected sub-control model 206. Therefore, the recommended control parameters Pr can be obtained by appropriately using the sub-control model 206 according to the situation.
[0056] Furthermore, since the situation is identified depending on which of multiple pre-set numerical ranges the deviation falls within, the situation can be identified depending on the magnitude of the deviation, i.e., the situation depending on the degree of deviation between the target value SP and the measured value PV, and the recommended control parameter Pr depending on the situation can be obtained.
[0057] Furthermore, sub-control models 206a and 206b each output recommended control parameters recommended for reducing the deviation between the measurement value PV and the common target value SP, in response to input of the control parameter P. Therefore, different recommended control parameters Pr for reducing the deviation between the common target value SP and the measurement value PV can be obtained depending on the situation. Therefore, recommended control parameters Pr that rapidly reduce the deviation by prioritizing the speed at which the equilibrium point is reached, and recommended control parameters Pr that gradually reduce the deviation by prioritizing the accuracy at which the equilibrium point is reached, can be obtained depending on the situation.
[0058] Furthermore, in each sub-control model 206, a recommended change amount for the control parameter P is output from the change amount output model 2061 according to the deviation and the control parameter P that has already been supplied to the control target 101, and the supplied control parameter P and the recommended change amount are added together by a common adder 2062 to calculate the recommended control parameter Pr. Therefore, unlike when an adder 2062 is provided for each sub-control model 206, the configuration of the device 200 can be simplified.
[0059] Furthermore, the change amount output models 2061 of the sub-control models 206a and 206b output recommended change amounts in different ranges, so that it is possible to reliably obtain recommended control parameters Pr that rapidly reduce the deviation and recommended control parameters Pr that gradually reduce the deviation depending on the situation.
[0060] Furthermore, using learning data including the deviation acquired by the deviation acquisition unit 203 and the control parameter P acquired by the control parameter acquisition unit 204, a learning process is performed for each change amount output model 2061 so as to output a recommended change amount recommended for increasing a reward value determined by a preset reward function in response to the input of the deviation and the control parameter P. Therefore, an appropriate recommended control parameter Pr can be acquired from each sub-control model 206.
[0061] Furthermore, since the learning process is performed for each sub-control model 206 using different learning data, it is possible to acquire recommended control parameters Pr suited to the situation from each sub-control model 206.
[0062] <2. Change amount output model 2061> 2 and 3 show the change amount output model 2061. In Fig. 2, Fig. 3, etc., the vertical axis indicates the control parameter P (for example, the command value IV of the valve opening), and the horizontal axis indicates the deviation.
[0063] The change amount output model 2061 may indicate the correspondence relationship between combinations of deviations and control parameters P and recommended change amounts. In this example, the change amount output model 2061 may be a manipulated variable map that maps the correspondence relationship between combinations of deviations and control parameters P and recommended change amounts. The manipulated variable map may be divided into multiple regions, each corresponding to a different recommended change amount, depending on the combination of control parameters P and deviations, and may output recommended change amounts that correspond to the coordinate positions of the input combinations of control parameters P and deviations. When such a change amount output model 2061 is used, the process becomes stable at the coordinate point where the deviation is 0 and the recommended change amount is 0 (in this figure, as an example, a point where deviation = 0 and control parameter P = approximately 50), that is, the equilibrium point.
[0064] Here, the change amount output model 2061 in Fig. 2 may be a change amount output model 2061a that outputs recommended control parameters with fine intervals and granularity when the deviation is small, and the change amount output model 2061 in Fig. 3 may be a change amount output model 2061b that outputs recommended control parameters with large intervals and granularity when the deviation is large. A first numerical range of -1.00 to 1.00 may be set for the deviation to be input to the change amount output model 2061a, and a second numerical range of -50 to 50 may be set for the deviation to be input to the change amount output model 2061b. The recommended change amount output from the change amount output model 2061a is 10 -2 ~10 -1 The recommended change amount output from the change amount output model 2061b may be on the order of 10 0 The order of magnitude may be:
[0065] The change amount output model 2061 may include information about the entire region of the operation amount map. Alternatively, the change amount output model 2061 may include only information indicating the boundaries of each region (for example, coordinates or function formulas indicating the boundaries) and the recommended change amounts corresponding to each region. In this case, the storage area for storing the change amount output model 2061 can be reduced.
[0066] 4 and 5 show other examples of the change amount output model 2061. Fig. 4 may show a change amount output model 2061a having the same content as Fig. 2, and Fig. 5 may show a change amount output model 2061b having the same content as Fig. 3. As shown in these figures, the change amount output model 2061 may be a table that associates combinations of deviations and control parameters P with recommended change amounts.
[0067] <3.Operation> 6 shows the operation of the device 200. The device 200 may control the control target 101 by performing the processes of steps S11 to S23. This operation may be started in response to the activation of the device 200. At the start of the operation, the learning process of the change amount output model 2061 may have been completed, and the target value SP may have been set to the reference target value.
[0068] In step S11, the measurement value acquiring unit 201 acquires the measurement value PV of the state of the control target 101. The target value acquiring unit 202 may acquire the measurement value PV from the sensor 102 of the facility 100.
[0069] In step S13, the deviation acquisition unit 203 acquires the deviation between the target value SP (in the present embodiment, a reference target value as an example) and the measurement value PV acquired in step S13.
[0070] In step S15, the identifying unit 207 identifies a phase from among the multiple phases according to the measurement value PV. In the present embodiment, as an example, the identifying unit 207 may identify either the first phase or the second phase depending on whether the deviation from the deviation acquiring unit 203 is included in the first numerical range or the second numerical range.
[0071] In step S17, the selection unit 208 selects, from the plurality of sub-control models 206, the sub-control model 206 that corresponds to the situation identified by the identification unit 207. In the present embodiment, as an example, the selection unit 208 may select the sub-control model 206a in response to the identification of the first situation, and may select the sub-control model 206b in response to the identification of the second situation.
[0072] In step S19, the control parameter acquisition unit 204 acquires the control parameter P supplied to the control target 101. The control parameter acquisition unit 204 may acquire the control parameter P supplied to the control target 101 in the most recent control cycle from the control unit 210. As an example, the control parameter acquisition unit 204 may acquire and temporarily store the control parameter P output from the control unit 210 to the control target 101 in the processing of step S23 described below, and read out the control parameter P in step S19. When step S19 is executed for the first time, that is, when the processing of step S23 has not been executed, the control parameter acquisition unit 204 may acquire a preset initial value of the control parameter P.
[0073] In step S21, the supply unit 209 supplies the control parameter P supplied from the control parameter acquisition unit 204 and the deviation supplied from the deviation acquisition unit 203 to the control model 205. In the present embodiment, as an example, the supply unit 209 supplies the deviation and the control parameter P to a selected sub-control model 206 from the multiple sub-control models 206 in the control model 205. As a result, a recommended control parameter Pr corresponding to the input control parameter P and the deviation is output from the sub-control model 206 corresponding to the situation. In the present embodiment, as an example, a recommended change amount corresponding to the input control parameter P and the deviation may be output from the change amount output model 2061, and the recommended change amount and the control parameter P acquired in step S17 may be added by the adder 2062 to generate the recommended control parameter Pr.
[0074] In step S23, the control unit 210 outputs the recommended control parameters Pr from the control model 205. The control unit 210 may supply the recommended control parameters Pr as control parameters P to the controlled object 101 to control the controlled object 101. When the processing of step S23 is completed, the processing may proceed to step S11.
[0075] <4. Example of operation> 7 shows the transition of the measurement value PV and the control parameter P when the controlled object 101 is controlled. The horizontal axis in the figure represents time (seconds), and the vertical axis represents the measurement value PV and the control parameter P. In this figure, as an example, the control parameter P may represent the indicated value IV of the valve opening.
[0076] As shown in this figure, in the device 200 according to this embodiment, the valve of the controlled object 101 is controlled using the recommended control parameters Pr output from the speed-oriented sub-control model 206b when the deviation falls within the second numerical range. In this figure, as an example, the valve opening is roughly controlled with a change amount of ±10%. Then, the controlled object 101 is controlled using the recommended control parameters Pr output from the accuracy-oriented sub-control model 206a when the deviation falls within the first numerical range. In this figure, as an example, the valve opening is finely controlled with a change amount of ±0.1%. As a result, the controlled object 101 is controlled using the recommended control parameters Pr according to the situation, and the measured value PV can be maintained at the target value SP quickly and accurately.
[0077] <5. Variations> <5.1.System 1A> Fig. 8 shows a system 1A according to a modified example. Components that are substantially the same as those in the system 1 shown in Fig. 1 are given the same reference numerals, and descriptions thereof will be omitted. The system 1A includes an apparatus 200A. The apparatus 200A may include an identification unit 207A, a target value setting unit 212A, a control model 205A, and a learning processing unit 211A.
[0078] <5.1.1. Specific part 207A> The identifying unit 207A identifies a phase from among a plurality of phases according to the measurement value PV. The identifying unit 207A according to this modification may identify a phase according to the measurement value PV supplied from the measurement value acquiring unit 201. The identifying unit 207A may identify a phase according to which of a plurality of preset numerical ranges the measurement value falls within. The identifying unit 207A may identify a phase according to which of third to sixth numerical ranges for the measurement value PV, which are preset for sub-control models 206c to 206f (described later) in the control model 205A, the measurement value PV from the measurement value acquiring unit 201 falls within.
[0079] The identification unit 207A may store a numerical range and a phase ID of each phase in association with each other, and may identify a phase associated with a numerical range that includes the measurement value from the measurement value acquiring unit 201. As an example in the present embodiment, the identification unit 207A may identify a third phase as a phase corresponding to the measurement value PV when the measurement value PV falls within a third numerical range. The identification unit 207A may identify a fourth phase as a phase corresponding to the measurement value PV when the measurement value PV falls within a fourth numerical range. The identification unit 207A may identify a fifth phase as a phase corresponding to the measurement value PV when the measurement value PV falls within a fifth numerical range. The identification unit 207A may identify a sixth phase as a phase corresponding to the measurement value PV when the measurement value PV falls within a sixth numerical range.
[0080] The identification unit 207A may supply the phase ID of the identified phase to the selection unit 208 and the target value setting unit 212A. By supplying the phase ID from the identification unit 207A to the selection unit 208, the selection unit 208 may select the sub-control model 206 corresponding to the identified phase from among the multiple sub-control models 206 in the control model 205A.
[0081] 5.1.2. Target Value Setting Unit 212A The target value setting unit 212A is an example of a setting unit, and sets a target value SP according to the phase identified by the identifying unit 207A. The target value setting unit 212A may set any of the specific target values of each of the sub-control models 206c to 206f (described later) as the target value SP. The target value setting unit 212A may store the specific target value of each of the sub-control models 206c to 206f in association with the phase ID of each phase, and may set the specific target value corresponding to the phase ID supplied from the identifying unit 207A as a new target value SP. The target value setting unit 212A may supply the set target value SP to the target value acquiring unit 202. As a result, the new target value SP may be supplied from the target value acquiring unit 202 to the deviation acquiring unit 203, and the deviation acquiring unit 203 may acquire the deviation between the new target value SP and the measurement value PV.
[0082] <5.1.3. Control Model 205A> Similar to the control model 205 in the above-described embodiment, the control model 205A outputs recommended control parameters Pr that are recommended to be supplied to the controlled object 101 in response to input of the deviation and the control parameters P. The control model 205A according to this modification may have four sub-control models 206 (also referred to as sub-control models 206c to 206f) respectively associated with a plurality of phases that are set in advance for the state of the controlled object 101, and may output the recommended control parameters Pr using the sub-control model 206 associated with the phase according to the measurement value PV.
[0083] The sub-control models 206c to 206f may be provided for each target value and may be associated with different intrinsic target values. Each intrinsic target value may be used as a target value SP by the target value setting unit 212A. The sub-control models 206c to 206f may output recommended control parameters Pr that are recommended to reduce the deviation between the intrinsic target value and the measurement value PV and in response to the input of the control parameter P. In this modified example, as an example, the sub-control models 206c to 206f are selected by the selection unit 208 depending on the situation, and the intrinsic target value of the selected sub-control model 206 is set as the target value SP by the target value setting unit 212A. Therefore, when each of the sub-control models 206c to 206f is selected, it outputs the recommended control parameters Pr in response to the deviation between the intrinsic target value as the target value SP and the measurement value PV and the input of the control parameter P.
[0084] The sub-control models 206c to 206f may have change amount output models 2061c to 2061f, respectively. The change amount output models 2061c to 2061f output recommended change amounts that recommend changes to be made to the control parameter P in response to input of the deviation and the control parameter P. The change amount output models 2061c to 2061f may output recommended change amounts in different ranges, or may output recommended change amounts in the same range. The recommended change amounts output from the change amount output models 2061c to 2061f may have approximately the same intervals and granularity.
[0085] <5.1-4. Learning processing unit 211A> The learning processing unit 211 performs learning processing for each of the change amount output models 2061c to 2061f in the same manner as the learning processing unit 211 in the above embodiment. The learning processing unit 211A may perform learning processing for each of the change amount output models 2061c to 2061f using different learning data.
[0086] For example, when performing learning processing on the change amount output model 2061c of the sub-control model 206c, the learning processing unit 211A may perform the learning processing using learning data in which the deviation falls within a third numerical range. The learning data of the change amount output model 2061c may include the deviation and the control parameter P acquired when the target value SP is set in advance as the inherent target value of the sub-control model 206c.
[0087] When performing learning processing on the change amount output model 2061d of the sub-control model 206d, the learning processing unit 211A may perform the learning processing using learning data in which the deviation falls within a fourth numerical range. The learning data of the change amount output model 2061d may include the deviation and the control parameter P acquired when the target value SP is set in advance as the inherent target value of the sub-control model 206d.
[0088] When performing learning processing on the change amount output model 2061e of the sub-control model 206e, the learning processing unit 211A may perform the learning processing using learning data in which the deviation falls within a fifth numerical range. The learning data of the change amount output model 2061e may include the deviation and the control parameter P acquired when the target value SP is set in advance as the inherent target value of the sub-control model 206e.
[0089] When performing learning processing of the change amount output model 2061f of the sub-control model 206f, the learning processing unit 211A may perform the learning processing using learning data in which the deviation falls within a sixth numerical range. The learning data of the change amount output model 2061f may include the deviation and the control parameter P acquired when the target value SP is set in advance to the inherent target value of the sub-control model 206f.
[0090] The absolute values of the deviations may be approximately the same among the learning data of the change amount output models 2061c to 2061f, and as an example, the orders of the deviations may be the same. The learning data may be acquired from a simulator (not shown) of the system 1A instead of being acquired from the actual system 1A.
[0091] According to the above-described device 200A, the situation is identified depending on which of a plurality of pre-set numerical ranges the measurement value PV falls within, so that the situation can be identified according to the measurement value PV and the recommended control parameters Pr according to the situation can be obtained.
[0092] Furthermore, the inherent target value of one of the sub-control models 206c to 206f is set as the target value SP depending on the situation, and the sub-control model 206 out of the sub-control models 206c to 206f depending on the situation outputs the deviation between the inherent target value as the target value SP and the measurement value PV, and the recommended control parameter Pr depending on the control parameter P. Therefore, it is possible to switch the target value SP depending on the progress of the process, and obtain the recommended control parameter Pr for reducing the deviation between the target value SP and the measurement value PV after switching.
[0093] <5.2. Correspondence Table> FIG. 9 shows the correspondence between the phase ID, the range of the measured value PV, the sub-model ID, and the specific target value. "K3" to "K6" in the figure may be the phase IDs of the third to sixth phases. min ~PVc max ", "PVd min ~PVd max ", "PVe min ~PVe max ", "PVf min ~PVf max " may indicate the numerical range of the measured value PV. "206c" to "206f" may be the sub-model IDs of the sub-control models 206c to 206f. "SPc" to "SPf" may be the specific target values of the sub-control models 206c to 206f.
[0094] The identifying unit 207A may identify a phase ID depending on which of the numerical ranges in the drawing the measurement value PV falls within. The selecting unit 208 may select the sub-control model 206 of the sub-model ID corresponding to the identified phase ID from among the phase IDs in the drawing. The target value setting unit 212A may set the specific target value corresponding to the identified phase ID from among the phase IDs in the drawing as the target value SP.
[0095] <5.3. Operation> 10 shows the operation of the device 200A. The device 200A may control the control target 101 by performing the processes of steps S11 to S23. Note that this operation may start in response to the activation of the device 200. Furthermore, the learning process of the change amount output model 2061 may be completed at the start of the operation. The operation of the device 200A according to the second embodiment differs from the operation of the device 200 according to the first embodiment in that the processes of steps S31 to S35 are performed between steps S17 and S17.
[0096] In step S31, the identifying unit 207A identifies a phase from among a plurality of phases according to the measurement value PV. In this modified example, as an example, the identifying unit 207A may identify any one of the third to sixth phases as the phase according to the measurement value PV.
[0097] In step S33, the target value setting unit 212A sets the target value SP according to the situation identified by the identification unit 207A. The target value setting unit 212A may set, as the target value SP, the specific target value corresponding to the situation identified in step S31, from among the specific target values of the sub-control models 206c to 206f.
[0098] In step S35, the deviation acquisition unit 203 acquires the deviation between the target value SP and the measurement value PV. The deviation acquisition unit 203 may acquire the deviation between the target value SP set in step S33 and the measurement value PV acquired in step S11. When the processing of step S35 ends, the processing may proceed to step S17. As a result, the sub-control model 206 corresponding to the situation identified in step S31 is selected from the multiple sub-control models 206c to 206f.
[0099] <5.4. Example of operation> 11 shows the transition of the measured value PV when the controlled object 101 is controlled. The horizontal axis in the figure represents time (seconds), and the vertical axis represents the measured value PV. In this figure, as an example, the control parameter P may represent the indicated value IV of the temperature inside the furnace.
[0100] As shown in this figure, in the apparatus 200A according to this modification, a third phase is identified as the measured value PV falls within a third numerical range as the process progresses. Then, the control target 101 is controlled using the deviation between the target value SP and the measured value PV according to the third phase and the control parameter P, which is input to and output from the sub-control model 206c.
[0101] Similarly, a fourth phase is identified in response to the measurement value PV being included in a fourth numerical range. Then, the control object 101 is controlled using the deviation between the target value SP and the measurement value PV in response to the fourth phase and the control parameter P, which is input to the sub-control model 206d and output as a recommended control parameter Pr.
[0102] Similarly, a fifth phase is identified in response to the measurement value PV being included in the fifth numerical range. Then, the control target 101 is controlled using the deviation between the target value SP and the measurement value PV in response to the fifth phase and the control parameter P, which is input to the sub-control model 206e and output as a recommended control parameter Pr.
[0103] A sixth phase is identified in response to the measurement value PV being included in a sixth numerical range. The control target 101 is controlled using the deviation between the target value SP and the measurement value PV in response to the sixth phase and the control parameter P, which is input to the sub-control model 206f and output as a recommended control parameter Pr.
[0104] <6. Other Modifications> In the above embodiment and modified examples, the control models 205 and 205A have been described as including the change amount output model 2061 and the adder 2062. However, these may not be included as long as the control models 205 and 205A output recommended control parameters Pr in response to input of the deviation and the control parameter P. In this case, the control models 205 and 205A may be learning models generated by an algorithm such as kernel dynamic policy programming, deep reinforcement learning, a support vector machine, a logistic regression, a decision tree, or a neural network. The learning processing units 211 and 211A may perform learning processing of the control models 205 and 205A using learning data including the deviation acquired by the deviation acquiring unit 203 and the control parameter P acquired by the control parameter acquiring unit 204.
[0105] Furthermore, although the change amount output model 2061 has been described as a map or table generated by a learning algorithm of the kernel dynamic policy programming method, it may be generated by other algorithms such as deep reinforcement learning, support vector machines, logistic regression, decision trees, or neural networks, or may be a model in another form different from a map or table.
[0106] Furthermore, although the deviation and the control parameter P are described as being input to the change amount output model 2061, other values may also be input. The other values may be, for example, the differential value or integral value of the measurement value by the sensor 102.
[0107] Furthermore, although the devices 200 and 200A have been described as including the measurement value acquiring unit 201, the target value acquiring unit 202, and the learning processing unit 211, any of these may be omitted. If the devices 200 and 200A do not include the measurement value acquiring unit 201 and the target value acquiring unit 202, the deviation acquiring unit 203 may acquire a deviation calculated by an external device. If the devices 200 and 200A do not include the learning processing units 211 and 211A, they may include a change amount output model 2061 that has been trained in advance by an external device.
[0108] In addition, although the sub-control models 206 have been described as being provided independently and receiving the deviations and control parameters independently from the supply unit 209, they may be provided in an integrated manner. In this case, each sub-control model 206 may constitute a part of the control models 205, 205A. As an example, the control models 205, 205A may be an operation amount map that maps the correspondence between combinations of deviations and control parameters P and recommended control parameters Pr, and each sub-control model 206 may be a central or peripheral part of the operation amount map. When multiple sub-control models 206 are integrated to constitute the control model 205, the devices 200, 200A do not need to include the identification unit 207 and the selection unit 208. Instead, when the supply unit 209 inputs the deviations and control parameters P to the control models 205, 205A, the recommended control parameters Pr may be output from the corresponding sub-control model 206 in the control models 205, 205A. Such control models 205, 205A may be generated by providing a common input section for the separate sub-control models 206 generated by the learning processing sections 211, 211A, and setting the deviation and control parameter P supplied from the supply section 209 to be input to one of the sub-control models 206 depending on the numerical range.
[0109] Furthermore, in the above embodiment, the control model 205 has been described as having two sub-control models 206a and 206b associated with a common target value SP, but it may have three or more sub-control models 206 associated with the common target value SP. Furthermore, the control model 205 may further have other sub-control models 206 with different target values.
[0110] Various embodiments of the present invention may also be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which operations are performed or (2) sections of an apparatus responsible for performing the operations. Particular stages and sections may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable medium, and / or a processor provided with computer-readable instructions stored on a computer-readable medium. Dedicated circuitry may include digital and / or analog hardware circuitry, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuitry may include reconfigurable hardware circuitry, including logical AND, OR, XOR, NAND, NOR, and other logic operations, flip-flops, registers, memory elements such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0111] A computer-readable medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that the computer-readable medium having instructions stored thereon comprises an article of manufacture containing instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable media may include electronic, magnetic, optical, electromagnetic, and semiconductor storage media. More specific examples of computer-readable media may include floppy disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray (RTM) disc, memory stick, integrated circuit card, and the like.
[0112] The computer readable instructions may include either assembler instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages such as the “C” programming language or similar programming languages.
[0113] The computer-readable instructions may be provided to a processor or programmable circuitry of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, either locally or over a wide-area network (WAN) such as a local area network (LAN), the Internet, etc., which executes the computer-readable instructions to create means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.
[0114] 12 illustrates an example of a computer 2200 in which aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 may cause the computer 2200 to function as or perform operations associated with an apparatus or one or more sections of the apparatus according to embodiments of the present invention, and / or to perform a process or steps of a process according to embodiments of the present invention. Such programs may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.
[0115] A computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0116] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the graphics controller 2216 itself, and causes the image data to be displayed on the display device 2218.
[0117] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.
[0118] The ROM 2230 stores therein a boot program or the like that is executed by the computer 2200 upon activation, and / or programs that depend on the hardware of the computer 2200. The input / output chip 2240 may also connect various input / output units to the input / output controller 2220 via a parallel port, a serial port, a keyboard port, a mouse port, etc.
[0119] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by realizing information manipulation or processing in accordance with the use of the computer 2200.
[0120] For example, when communication is performed between the computer 2200 and an external device, the CPU 2212 may execute a communication program loaded into the RAM 2214 and instruct the communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of the CPU 2212, the communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in the RAM 2214, the hard disk drive 2224, the DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes reception data received from the network to a reception buffer processing area or the like provided on the recording medium.
[0121] The CPU 2212 may also cause all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and perform various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.
[0122] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 2214. The CPU 2212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored on the recording medium, the CPU 2212 may search for an entry that matches a condition specified by the attribute value of the first attribute from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0123] The above-described programs or software modules may be stored in a computer-readable medium on or near the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable medium, thereby providing the programs to the computer 2200 via the network.
[0124] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.
[0125] It should be noted that the execution order of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order. [Explanation of symbols]
[0126] 1 System 100 equipment 101 Control Object 102 Sensors 200 equipment 201 Measurement acquisition unit 202 Target value acquisition unit 203 Deviation acquisition unit 204 Control parameter acquisition unit 205 Control Model 206 Sub-control model 207 Specific section 208 Selection Section 209 Supply Department 210 Control Unit 211 Learning processing unit 212 Target value setting unit 206 Sub-control model 2061 Change Output Model 2062 Addition section 2200 Computer 2201 DVD-ROM 2210 host controller 2212 CPU 2214 RAM 2216 Graphics Controller 2218 Display Device 2220 Input / Output Controller 2222 communication interface 2224 hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 I / O chip 2242 keyboard
Claims
1. a first acquisition unit that acquires a deviation between a measured value of a state of a control object and a target value; a second acquisition unit that acquires control parameters supplied to the control target; a supply unit that supplies the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit to a control model having a plurality of sub-control models respectively associated with a plurality of predetermined phases of the state, the control model using a sub-control model among the plurality of sub-control models that is associated with a phase corresponding to the measurement value, and that outputs a recommended control parameter that is recommended to be supplied to the control target in response to input of the deviation and the control parameter; an output unit that outputs the recommended control parameters output from the control model in response to the supply from the supply unit to the control model; An apparatus comprising:
2. an identifying unit that identifies a phase corresponding to the measurement value from among the plurality of phases; a selection unit that selects, from the plurality of sub-control models, a sub-control model that corresponds to the situation identified by the identification unit; Furthermore, The device according to claim 1, wherein the supply unit supplies the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit to a sub-control model selected by the selection unit from among the plurality of sub-control models.
3. The device according to claim 2 , wherein the specifying unit specifies the phase depending on which of a plurality of preset numerical ranges the deviation falls within.
4. The device according to claim 2, wherein at least two of the plurality of sub-control models output recommended control parameters recommended to reduce the deviation between a measured value and a common target value in response to input of control parameters.
5. The device according to claim 2 , wherein the determination unit determines the aspect depending on which of a plurality of preset numerical ranges the measurement value falls within.
6. At least two of the plurality of sub-control models are associated with different specific target values, and output recommended control parameters that are recommended to reduce the deviation between a measurement value and the specific target value in response to input of control parameters; The device comprises: The device according to claim 2 , further comprising a setting unit that sets one of the specific target values of the at least two sub-control models as the target value in accordance with the situation identified by the identification unit.
7. Each sub-control model is a change amount output model that outputs a recommended change amount that recommends a change to the control parameter in response to input of the deviation and the control parameter; an adder that calculates the recommended control parameter by adding the control parameter supplied to the controlled object and the recommended change amount output from the change amount output model; and The apparatus of claim 1 , wherein the plurality of sub-control models share the adder.
8. The device according to claim 7 , wherein the change amount output models in at least two of the plurality of sub-control models output the recommended change amounts in ranges different from each other.
9. 9. The device according to claim 8, further comprising: a learning processing unit that performs learning processing for each change amount output model using learning data including the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit, and that outputs the recommended change amount recommended for increasing a reward value determined by a preset reward function in response to input of the deviation and the control parameter.
10. The device according to claim 9 , wherein the learning processing unit performs the learning process for each change amount output model using different learning data.
11. a first acquisition step of acquiring a deviation between a measured value of a state of a controlled object and a target value; a second acquisition step of acquiring control parameters supplied to the control object; a first supply step of supplying the deviation acquired in the first acquisition step and the control parameter acquired in the second acquisition step to a control model having a plurality of sub-control models respectively associated with a plurality of predetermined phases of the state, the control model outputting a recommended control parameter to be supplied to the controlled object in response to input of the deviation and the control parameter using a sub-control model among the plurality of sub-control models associated with a phase corresponding to the measured value; an output step of outputting the recommended control parameters output from the control model in response to the supply to the control model in the first supply step; A method for providing
12. Computer, a first acquisition unit that acquires a deviation between a measured value of a state of a control object and a target value; a second acquisition unit that acquires control parameters supplied to the control target; a supply unit that supplies the deviation acquired by the first acquisition unit and the control parameter acquired by the second acquisition unit to a control model having a plurality of sub-control models respectively associated with a plurality of predetermined phases of the state, the control model using a sub-control model among the plurality of sub-control models that is associated with a phase corresponding to the measurement value, and that outputs a recommended control parameter that is recommended to be supplied to the control target in response to input of the deviation and the control parameter; an output unit that outputs the recommended control parameters output from the control model in response to the supply from the supply unit to the control model; A program that functions as a
Citation Information
Patent Citations
Control parameter calculation method, control parameter calculation program, and control parameter calculation device
JP2019204178A
Using predictions to control target systems
JP2019519871A
Device, method and program
JP2021086283A