Plant control system and plant control method
Patent Information
- Application Number
- JP2022167657
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-10-19
AI Technical Summary
Reinforcement learning methods in plant control fail to suppress oscillations in manipulated variables, leading to equipment failures due to vibration, as they lack constraints on output behavior.
A plant control system and method that includes a learning processing device to determine optimal plant behavior, using a state information control unit, action value updating, and an optimal action selection unit to suppress vibrations in manipulated variables.
The system outputs manipulated variables that quickly converge to target values while minimizing vibrations, reducing the risk of equipment failure.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a plant control system and a plant control method. [Background technology]
[0002] In the plant field, reinforcement learning, an AI technology, is increasingly being used as a control to stabilize processes. Reinforcement learning learns optimal control laws by searching through a trial-and-error method using a simulator that mimics the controlled object. The optimal control law here is a control model that can output a manipulated variable that quickly stabilizes the process, in other words, that converges the plant signal value, which is the controlled quantity, to the target value.
[0003] In reinforcement learning, a value is defined for each manipulated variable, and the optimal control law can be learned by updating it using a formula called the value update formula. The value here is a numerical value that indicates how effective a certain manipulated variable is for the purpose of converging the controlled variable to the target value. Reinforcement learning is expected to realize highly accurate control because it is possible to find the optimal operation based on the aforementioned value from the information searched for with the goal of converging the controlled variable to the target value.
[0004] However, while reinforcement learning can quickly converge the controlled variable to the target value, it does not have the function to evaluate the behavior of the output value in the value and its update formula, and it is not possible to set constraints on the behavior itself. Therefore, issues such as the oscillation of the operation variable output by the control law obtained by learning arise. When applied to actual equipment, the oscillation of the operation variable can cause equipment failure, so it needs to be resolved.
[0005] In this context, a method for imposing constraints on the output behavior of control laws acquired through reinforcement learning is desired.
[0006] The method disclosed in Patent Document 1 is able to attenuate the value of updating when a certain output occurs frequently. This makes it possible to learn the optimal control law while evaluating the behavior of the output. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Patent Publication No. 2021-77286 Summary of the Invention [Problem to be solved by the invention]
[0008] However, the reinforcement learning method disclosed in Patent Document 1 does not solve the problem of oscillation of an operation value that may cause a failure during plant control. In the method in Patent Document 1, a restriction is imposed on the number of times a certain output value occurs in the entire control process. Therefore, this method does not take into account the oscillation frequency of the output value, and it is difficult to suppress the oscillation of the output value included in the control process.
[0009] An object of the present invention is to provide a control system and a plant control method capable of outputting a manipulated variable that quickly converges a controlled variable to a target value while suppressing the vibration of the manipulated variable. [Means for solving the problem]
[0010] In view of the above, the present invention provides a plant control system comprising: a learning processing device which determines the optimal action of the plant by learning; and a control processing device which controls the plant in accordance with the optimal action determined by the learning processing device, the learning processing device comprising: a state information control unit which converts a plurality of plant signals into the state of the plant and defines a target state; an action value updating unit which determines an action value, which is the value of the state and action between the previous operation and the current operation, using the state, action, and target state of the plant; and an optimal action selection unit which determines the optimal action for achieving the target state using the action value, and which determines an action that suppresses oscillations in the operating amount of the plant as the optimal action.
[0011] Furthermore, the present invention provides a "plant control method for determining an optimal action for a plant by learning, and controlling the plant in accordance with the optimal action determined by the learning process, the learning process comprising: converting a plurality of plant signals into a plant state to define a target state; determining an action value, which is a value of the state and action between a previous operation and a current operation, using the plant state, action, and target state; determining an optimal action for achieving the target state using the action value; and determining, as the optimal action, an action that suppresses oscillations in the plant's operating amount. Effect of the Invention
[0012] According to the present invention, it is possible to provide a control system capable of outputting a manipulated variable that stabilizes a controlled variable while suppressing the vibration of the manipulated variable. [Brief description of the drawings]
[0013] [Figure 1] FIG. 1 is a diagram showing an example of the configuration of a plant control system according to a first embodiment of the present invention. [Diagram 2] FIG. 4 is a diagram showing an example of the configuration of data input by a user and stored in an input information storage unit; [Figure 3a] FIG. 4 is a diagram showing the relationship between the plant signal value and the state number stored in the signal information storage unit. [Figure 3b]FIG. 13 is a diagram showing the relationship between the operation amount and the action number of the plant stored in the signal information storage unit. [Figure 4] FIG. 4 is a diagram showing an example of a flow of processing performed by the learning processing device. [Diagram 5] FIG. 5 is a diagram showing an example of a processing flow for episode processing performed by the learning processing device in S2 of FIG. 4. [Figure 6a] FIG. 13 is a diagram showing an example of the configuration of data representing the state number and action number of the previous step stored in an action value storage unit. [Figure 6b] FIG. 13 is a diagram showing an example of the configuration of data representing values according to state numbers and action numbers stored in an action value storage unit. [Figure 7a] A schematic diagram showing the shape of the decay function (linear function) in the value update equation. [Figure 7b] A schematic diagram showing the shape of the decay function (quadratic function) in the value update equation. [Figure 7c] A schematic diagram showing the shape of the decay function (step function) in the value update equation. [Figure 8] FIG. 13 is a diagram showing an example of the configuration of data indicating optimal action numbers for each state number and corresponding operation amounts. [Figure 9] FIG. 4 is a diagram showing an example of a flow of processing performed by a control processing device. [Figure 10a] FIG. 13 is a diagram showing an example of the configuration of some data representing control results stored in a control result storage unit; [Figure 10b] 10B is a diagram showing an example of the configuration of some data (a compilation of the data shown in FIG. 10B) showing the control results stored in a control result storage unit; [Figure 11] FIG. 4 is a diagram showing an example of a screen into which a user inputs information required in the processing flow of the present invention. [Figure 12] FIG. 13 is a diagram showing an example of a screen displaying the relationship between the convergence time to a target state and the vibration frequency. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] The following describes examples of the present invention. EXAMPLES
[0015] A plant control system according to a first embodiment of the present invention will be described with reference to Fig. 1. The plant control system in Fig. 1 is composed of a user input / output device 3, a signal information storage unit 2, an input information memory unit 4, a learning processing device 1, and a control processing device 5, and the learning processing device 1 learns information stored in the input information memory unit 4 to obtain an optimal target, which is then provided to the control processing device 5, which then controls a plant 6 to be controlled to the optimal target based on the optimal target. With this configuration, the plant control system targets equipment such as plants and industrial machinery, and is capable of outputting an optimal control manipulated variable with reduced vibration while converging a signal value representing the state of the target to a target state.
[0016] Of these, the input information storage unit 4 obtains conversion information D1 indicating the relationship between the plant signal value and the state number from the signal information storage unit 2, and also receives information D2 input by the user from the user input / output device 3. In addition, various process quantities from the controlled object 6 are input to the input information storage unit 4.
[0017] FIG. 2 shows an example of information D2 input by the user and stored in the input information storage unit 4. All of the information is used by the learning processing device 1. These are selection signals D20 selected by the user from among information candidates (e.g., plant information A, B, C). They are also set values for the discount rate γ (D21), the number of episodes D22, the attenuation coefficient η (D23), and the target value D25, which are arbitrarily set by the user. Or, as the selection function D24, they are functions arbitrarily set by the user from, for example, linear functions, quadratic functions, and step functions. The types of signals to be specified are stored in the selection function D24 and target value D25 columns.
[0018] The present invention is characterized in that the information D2 input by the user, particularly the set values for the damping coefficient η (D23) and the target value D25, are provided in advance from the user input / output device 3. The user input / output device 3 also includes an input section where the user inputs the input information D1 and D2, and a display device that displays line graphs and scatter diagrams for the user to refer to when setting parameters. The display screen will be described in detail later.
[0019] Meanwhile, the signal information storage unit 2 stores conversion information D1 that indicates the relationship between the plant signal value and the state number, etc. Fig. 3a is a table showing the relationship between the range of the plant signal value D1a and the state number D1b stored in the signal information storage unit 2. The rows list the types of state numbers D1b, and the columns list all the types of signal values D1a that the plant can output. The numbers in the table indicate the range of the signal values.
[0020] According to this notation example, the operating state D1b of the plant (states S1, S2, S3, ...) is defined in advance by the magnitude of the pre-specified plant signal value D1a (here, signal A and signal B). State S1 is defined when signal A is in the range of 1 to 2 and signal B is in the range of -5 to -4.5, state S2 is defined when signal A is in the range of 1 to 2 and signal B is in the range of -4.5 to -4, state S3 is defined when signal A is in the range of 1 to 2 and signal B is in the range of -4 to -3.5, state S4 is defined when signal A is in the range of 1 to 2 and signal B is in the range of -3.5 to -3, and state S5 is defined when signal A is in the range of 2 to 3 and signal B is in the range of -5 to -4.5.
[0021] The input plant signal value D1a (signals A and B) is extracted as state S information D1b by referring to the table in FIG. 3a, where it is converted from the plant signal value D1a to state S information D1b.
[0022] 3a, the information transferred from the input information storage unit 4 in FIG. 1 to the learning processing device 1 and the control processing device 5 is converted to state information D1b, not numerical information D1a. The processing in the learning processing device 1 makes it possible to execute learning based on pattern processing of the plant state D1b, which is an advantage of the learning function, rather than numerical learning based on the magnitude of the plant signal value D1a. Specifically, for example, the learning processing device 1 learns that the plant transitions from state S1 to state S5 when it starts up, and the learning result is reflected in the control processing device 5.
[0023] 3b is a conversion table showing the relationship between the range of the operation amount D1c of the plant stored in the signal information storage unit 2 and the action number D1d. According to this notation example, the plant's actions D1b (actions a1, a2, a3, ...) are defined in advance according to the range of the magnitude of the operation amount D1c of the plant specified in advance. The action is defined as follows: action a1 when the operation amount D1c is in the range of 1 to 1.5, action a2 when the operation amount D1c is in the range of 1.5 to 2, action a3 when the operation amount D1c is in the range of 2 to 2.5, action a4 when the operation amount D1c is in the range of 2.5 to 3, and action a5 when the operation amount D1c is in the range of 3 to 3.5.
[0024] The input plant operation amount D1c is extracted as information on the action a by referring to the table in FIG. 3b, and here the plant operation amount D1c is converted into information D1d on the action a.
[0025] By the conversion process shown in FIG. 3b, the plant operation amount D1c is also converted from numerical information into action information D1d shown as a pattern, which is in a form suitable for learning processing, and is provided.
[0026] Here, the state number D1b (state information) will be explained. The state number D1b allows multiple types of signal values to be handled one-dimensionally. First, the plant dynamics simulator 11 outputs multiple types of signal values. As an example, consider a case where two types of signals, signal A and signal B, are used in episode processing in the learning process. Suppose that the two types of signal values are 1 and -5, respectively. Since the signal values are within the range of values indicated by the state S1 in the first row of the table shown in FIG. 3a, the signals A and B are defined as the state S1. This allows the signals A and B to be compressed into one dimension, making it possible to execute the processing described later. A similar compression is also performed by converting the manipulated variable D1c into an action D1d.
[0027] The learning processing device 1 shown in Fig. 1 includes a simulator that simulates a target plant, a processing unit such as a CPU, and a storage unit such as a memory, and acquires a control law that enables output of an optimal control operation amount through repeated interactive processing with the simulator. In the example of Fig. 1, the learning processing device 1 includes a plant dynamics simulator 11, a state information control unit 12, an action value update unit 13, an action value storage unit 14, an optimal action selection unit 15, and an episode number storage unit 16 as functional components.
[0028] 1 is connected to the learning processing device 1 and a control object 6, and is a device that performs optimal control of the control object 6, which is an actual plant, based on the control law acquired by the learning processing device 1. The control processing device 5 includes, as its functional components, a learning information control unit 51, a state information conversion unit 52, an input / output device 53, and a control result storage unit 54. A detailed description of the control processing device 5 will be given later.
[0029] The flow of processing of the learning processing device 1 in Fig. 1 will be described with reference to Fig. 4 and Fig. 5. Fig. 4 is a flow diagram showing the overall processing of the learning processing device 1. In the processing of Fig. 4, in the first processing step S1, information D2 input by the user and signal information are acquired via the input information storage unit 4 to the learning processing device 1. The signal information here specifically refers to the converted information (state number D1b) of Fig. 3a showing the relationship between the plant signal value D1a and the state number D1b, and the converted information (action information D1d) of Fig. 3b showing the relationship between the manipulated variable D1c and the action number D1d.
[0030] Next, in processing step S2, episode processing is performed. An episode is a term used in reinforcement learning algorithms, and learning progresses by repeating episodes. An episode in this device refers to one control simulation using the plant dynamics simulator 11.
[0031] In processing step S3, the number of times episode processing has been performed is updated each time an episode is processed, and in processing step S4, the number of episodes D22 set by the user is compared with the number of times episode processing has been performed, and if the number of times episode processing has been performed is less than or equal to the number of episodes D22 set by the user, the process returns to processing step S2 and the episode processing is performed again. If the number of times episode processing has been performed is greater than or equal to the number of times episode processing has been performed, the process of the learning processing device 1 is terminated. This executes learning and simulation a predetermined number of times.
[0032] Fig. 5 is a flow diagram showing the details of the episode processing. This corresponds to processing step S2 in Fig. 4. The episode processing will be described in detail below with reference to Fig. 5. In the following description, each process constituting one episode will be expressed as a unit called a step. The first process in the repeated processing will be expressed as an initial step.
[0033] In the first processing step S21, a plant signal value D1a representing the state of the plant 6 and a manipulated variable D1c input to the plant 6 are generated by a plant dynamics simulator 11 that simulates the behavior of the target plant 6.
[0034] In processing step S22, the output values (D1a and D1c) of the plant dynamics simulator 11 are input, and in processing step S23, the state information control unit 12 converts the plant signal value D1a into a state number D1b and the operation amount D1c into an action number D1d. When converting the plant signal value D1a into the state number D1b, information on the selection signal (D20 in FIG. 2) input by the user and acquired in processing step S1 in FIG. 4 is used. The selection signal D20 indicates the type of signal value to be used in episode processing selected by the user from among multiple types of signal values output by the plant dynamics simulator 11. In processing step S23, the plant signal value D1a output by the plant dynamics simulator is converted into the state number D1b based on the relationship between the plant signal value D1a and the state number D1b shown in FIG. 3 described above.
[0035] In processing step S24, the target state is defined by the state information control unit 12. Here, the target value (D25 in FIG. 2) acquired in processing step S1 in FIG. 4 and the information of the signal specified as the target value are used. As an example, assume that the target value of signal A is specified as 1.5. In this case, since 1.5 is a numerical value in the range of 1 or more and less than 2, states S1, S2, S3, and S4 in the table shown in FIG. 3a are specified as the target states.
[0036] When the state within a plant is defined by grouping it into several states according to the magnitude of several signals, only the state that meets the condition of the signal magnitude defined as the target value is extracted from the several states defined by the magnitude of several signals, and this is set as the target state.For example, when starting up a plant, if the state of the fluid within the equipment is defined by grouping it into the signals of temperature, pressure, and flow rate, and the main factor is pressure, and you want to raise this value to 1.0, then only the state that satisfies the pressure of 1.0 is extracted from the several states, and this state is set as the target state, which means, for example, the completion of startup.
[0037] In the present invention, a learning process is performed in the subsequent processing. Since the learning process is more suited to pattern processing than numerical processing, the signal is expressed as a state rather than a magnitude, and the target value in the case of a signal is used as the target state in the state processing.
[0038] In the processing step S25, the action value update section 13 acquires information on the state number D1b, the action number D1d, and the goal state in the current step from the state information control section 12.
[0039] In processing step S26, the action value update section 13 obtains the state number D1b and action number D1d of the previous step, the value corresponding to the state number D1b and action number D1d obtained in processing step S25 of the previous step, and the maximum value in the state number D1b obtained in the current processing step S25 from the action value storage section 14. The value is a value stored according to the state number D1b and action number D1d.
[0040] Fig. 6a is a table showing state number D1b and action number D1d stored in the action value storage unit 14. For example, if state number D1b obtained in the previous processing step S25 is state S10 and action number D1d is action a9, state S10 and action a9 are stored in the table of Fig. 6a. The relationship between state S10 and action a9 in the table of Fig. 6a is regarded as value Q.
[0041] FIG. 6b is a table showing the value Q stored in the action value storage unit 14 according to the state number D1b and the action number D1d. In the process step S26, the corresponding value is obtained from the table in FIG. 6b. If the state number D1b and the action number D1d obtained in the process step S25 one step before are the previous example, 941 corresponding to Q(S10, a9) is obtained as the value according to the state number D1b and the action number D1d. If the state number D1b obtained in the current process step S25 is the state S1 and the action number D1d is the action a2, 1990 corresponding to Q(S1, a1) is obtained as the maximum value in the state S1. Here, if this process is the initial step, the state number D1b of the previous step does not exist, so the state number D1b to be obtained is determined randomly.
[0042] In processing step S27, the action value update unit 13 updates the value according to the state number D1b and the action number D1d. When updating, formula (1) is used. In contrast to the value update formula generally used in reinforcement learning, formula (1) used in the device of the present invention has a function f(Δa) added to it in order to suppress the vibration of the manipulated variable. This is the key point of the present invention.
[0043]
number
[0044] The value update process according to formula (1) will be explained in detail below. Formula (1) represents the update calculation for updating the value shown on the left side with the value calculated on the right side. In formula (1), s represents the state number D1b of the previous step, s' represents the state number D1b of the current step, and a represents the action number D1d performed in state number D1b of the previous step.
[0045] Here, it is assumed that the state number D1b of the previous step is state S10, the action number D1d is action a9, and the state number D1b of the current step is state S1, as in the values acquired in processing step S26. γ is the discount rate D21 described in Figure 2, and is a value that can be set arbitrarily by the user.
[0046] In the following explanation, it is assumed that the user set the discount rate γ to 0.99. r(s') is called the reward, and is a function that is 1000 if s' is the goal state, and 0 otherwise. Here, it is assumed that the state S1 of the current step is the goal state, and r(s') is 1000. Q(s, a) represents the value according to the state number D1b and action number D1d of the previous step, and maxQ(s') represents the maximum value according to the current state number D1b. Here, Q(s, a) is 941, and maxQ(s') is 1990.
[0047] f(|Δa|) is the absolute value of the difference between the operation amount in the previous step and the operation amount in the current step, that is, a function of |Δa|. Δa is calculated using the operation amount in the previous step, the operation amount in the current step, and information indicating the relationship between the operation amount and the action number obtained in processing step S1. The difference Δa between the operation amount corresponding to the action number in the previous step and the operation amount corresponding to the action number in the current step is calculated by setting the operation amount to the lower limit of the operation amount range corresponding to the action number in question in the table indicating the relationship between the operation amounts shown in Figure 3b.
[0048] f(|Δa|) is a function whose value changes in the range from 0 to 1 according to |Δa|. This function reduces the value as the fluctuation of the manipulated variable becomes larger, that is, as the difference Δa between the manipulated variables becomes larger, and as a result, the fluctuation of the manipulated variable is suppressed.
[0049] Figures 7a, 7b, and 7c show examples of the function f(|Δa|). This device has three patterns of functions prepared for the user to select from. Furthermore, the degree of damping can be changed by the damping coefficient η (D23). This allows the user to arbitrarily set the degree to which vibrations in the control operation value are suppressed by adjusting the damping coefficient η.
[0050] Figure 7a shows the case where the user specifies a linear function as the function f(|Δa|). In the function shown in Figure 7a, the value of f(|Δa|) decreases linearly as the difference between the manipulated variable one step before and the manipulated variable in the current step increases. In this function, the negative slope becomes larger as the damping coefficient η increases. The function shown in Figure 7a is expressed by equation (2).
[0051]
number
[0052] Figure 7b shows the case where the user specifies a quadratic function as the function f(|Δa|). In the function shown in Figure 7b, the value of f(|Δa|) decreases quadratically as the difference between the manipulated variable one step before and the manipulated variable in the current step increases. In this function, the amount of decrease in f(|Δa|) associated with Δa also increases as the damping coefficient η increases. The function shown in Figure 7b is expressed by equation (3).
[0053]
number
[0054] Figure 7c shows the case where the user specifies a step function as the function f(|Δa|). In the function shown in Figure 7c, the value of f(|Δa|) decreases in a step-like manner depending on the difference between the amount of operation one step before and the amount of operation in the current step. In this function, the damping coefficient η represents the value of the change point Δa of the step function. The function shown in Figure 7c is expressed by equation (4).
[0055]
number
[0056] Returning to the explanation of the value update calculation using equation (1), let us assume that the value of the function f(|Δa|) is 0.5. By substituting the specific number into the right-hand side of equation (1), the calculation result of the right-hand side becomes 1485.1 (=0.5×(1000+0.99×1990)). Therefore, the value of Q(s,a) on the left-hand side is updated from 941 to 1485.1.
[0057] Returning to the explanation of the process flow of FIG. 5, the concrete values used in the above updating process are also used in the following explanation of the process flow. In process step S28, the action value storage unit 14 acquires the value updated by the action value update unit 13, the state number of the current step, and the action number. Considering the above assumption, Q(S10, a9), the state S1 of the current step, and the action a2 are acquired as the updated value, and the acquired information is stored in the action value storage unit 14. The table showing the state number and action number shown in FIG. 6a is updated to state S1 and action a2. The value corresponding to state S10 and action a9 shown in FIG. 6b is updated. That is, the value 941 stored in Q(S10, a9) is updated to 1485.1.
[0058] Fig. 8 is a table storing the optimal action for each state number stored in the action value storage unit 14 and the corresponding operation amount. The optimal action refers to the action number corresponding to the maximum value among the values of each action D1d stored for a certain state number D1b. Specifically, since the action number corresponding to the maximum value among the values for each action number stored in the state S1 shown in Fig. 6b is a1, action a1 (D1d) is stored as the optimal action in state S1 (D1b) in Fig. 8. Then, 1, which is the lower limit value of the operation amount range corresponding to the action a1 in the table showing the relationship with the operation amount D1c shown in Fig. 3b, is stored as the operation amount in the table of Fig. 8.
[0059] In processing step S29, the optimal action selection unit 15 obtains from the action value storage unit 14 a table that stores the optimal action and its operation amount for each state number, as well as the state number of the current step and information on the target state, and outputs the operation amount corresponding to the current step to the simulator. At this time, a random operation amount is output with a probability of 40%. By outputting a random operation amount, the search space can be expanded and control precision can be improved.
[0060] In processing step S30, the end of episode processing is determined based on whether the goal state has been reached. If the goal state has been reached, the episode processing is terminated, and if not, the process returns to processing step S21. In other words, if the state number of the current step acquired by the optimal action selection unit 15 from the action value storage unit 14 is the goal state, the process is terminated. This concludes the explanation of episode processing.
[0061] Returning to the explanation of the processing flow diagram of the learning processing device 1 in Fig. 4, in processing step S3, after the episode processing is completed, the episode number storage unit 16 acquires information indicating that the episode processing is completed from the optimal action selection unit 15, and increments the number of times the episode processing has been performed, which is stored in the episode number storage unit, by +1.
[0062] In processing step S4, if the number of times the episode processing has been performed and stored in the episode number storage unit 16 exceeds the number of episodes set by the user, the episode storage unit 16 sends a stop command to the plant dynamics simulator 11 and also sends information to the learning information control unit that the processing of the learning processing device 1 has ended. This concludes the processing flow of the learning processing device 1.
[0063] Next, a description will be given of the flow of processing by the control processing device 5. Fig. 9 is a flow diagram of processing by the control processing device 5. In Fig. 9, in the first processing step S51, the input / output device 53 acquires a plant signal value from the control target 6. The control target 6 refers to the plant that is the control target.
[0064] In processing step S52, the state information conversion unit 52 obtains, via the learning information control unit 51, from the action value storage unit 14, a table showing the relationship between the plant signal value and the state number, and information on the target state which is information input by the user.
[0065] In process step S53, the state information conversion unit 52 converts the plant signal value into a state number. In process step S54, the learning information control unit 51 obtains the state number from the state information conversion unit 52, and obtains the operation amount corresponding to the state number from the action value storage unit 14 by referring to the table shown in FIG.
[0066] In processing step S55, the input / output device 53 acquires the state number and the manipulated variable from the learning information control unit 51, and outputs the manipulated variable to the controlled object 6. In processing step S56, the input / output device 53 acquires a plant signal value from the controlled object 6. The plant signal value acquired here represents the state of the plant that has changed in response to the manipulated variable. In processing step S57, the control result storage unit 54 acquires the state number and the manipulated variable from the input / output device 53 and stores them.
[0067] 10a and 10b are diagrams showing an example of the configuration of data stored in the control result storage unit 54. The leftmost column in FIG. 10a stores the time when the control result storage unit 54 acquires the state number D1b and the operation amount D1d from the input / output device 53. The second column from the left stores the state number D1b acquired for each time. The third column from the left stores the operation amount D1d acquired for each time. The rightmost column stores the number of times the operation amount has been changed. 0 is stored in the first row of this column of the number of times the operation amount has been changed. In the subsequent rows, when there is a difference between the operation amount stored in the previous row and the operation amount acquired at the current time, 1 is added to the number of times the operation amount has been changed stored in the previous row, and the result is stored in the current row.
[0068] The leftmost column in Figure 10b stores the time when the target state was converged upon. In other words, it represents the time when the acquired state number was the target state. The second column from the left stores the time from the time when the state number was first acquired until the target state was acquired. The rightmost column stores the number of times the manipulated variable was changed before the target state was acquired.
[0069] In processing step S58, the input / output device 53 performs a condition determination as to whether or not the state number acquired from the learning information control unit 51 is the target state. If the acquired state number is not the target state, the process returns to processing step S53. If it is the target state, the process proceeds to processing step S59.
[0070] In processing step S59, the information stored in the control result storage unit 54 is output to the user input / output device. Here, the information to be output has the data configuration example shown in Fig. 10b. The above is the processing flow of the control processing device 5.
[0071] Next, there will be described a display screen output by the user input / output device 3. Fig. 11 shows an example of a screen on which the user inputs information required for executing the processes in the learning processing device 1 and the control processing device 5 described above.
[0072] Item 31 displays multiple types of signals output by the plant 6, which is the object of control. The user selects a signal to be used for processing from the multiple signals in this item 31. Item 32 displays an input field for the user to input the discount rate γ in equation (1) used for the update calculation in the learning processing device 1. The user arbitrarily specifies and inputs the discount rate γ within the range of 0 to 1.
[0073] Item 33 displays an input field for the user to input the number of times episode processing is to be performed in the learning processing device 1. Item 34 displays an input field for the user to input the attenuation coefficient η in equations (2), (3), and (4) used in the update calculations in the learning processing device 1.
[0074] Item 35 displays a selection field that enables the user to select the function to be used for the update calculation in the learning processing device 1 from among equations (2), (3), and (4). The user selects the function to be used for processing from among linear functions, quadratic functions, and step functions. Item 36 selects the signal value to be converged among the types of plant signal values. By selecting the signal value, learning is performed with the aim of converging this signal in the learning processing device 1. Item 37 displays an input field that enables the user to input the value of the signal value to be converged.
[0075] 12 is an example of a screen displaying the relationship between the convergence time to the target state and the vibration frequency. The vertical axis represents the convergence time to the target, and the horizontal axis represents the vibration frequency, and the control results in the control processing device 5 are plotted on a scatter diagram.
[0076] Item 38 indicates supplementary information that is displayed when the mouse cursor is placed over a plot point. The supplementary information includes the number of plant signals selected as the state number, the number of times episode processing input by the user is performed, the convergence time to the target state, the number of vibrations per minute, the value of the damping coefficient η input by the user, and the type of damping coefficient and its associated function (damping function) selected by the user. The user determines the optimal input information while checking this display screen. The optimal input information here refers to a combination of information input by the user that results in the shortest convergence time to the target state and the lowest vibration frequency. This allows the user to determine the optimal damping coefficient value, etc., through a process of error.
[0077] As described above, according to this embodiment, it is possible to acquire a control law that outputs a manipulated variable with suppressed vibration by adding a theory that imposes a penalty on the vibration of the manipulated variable in the value update process in the learning processing device 1. Since the vibration of the manipulated variable leads to failure of plant equipment, controlling the plant using this control law leads to a significant reduction in the risk of failure. [Explanation of symbols]
[0078] 1: Learning processing device 2: Signal information storage section 3: User I / O device 4: Input information storage section 5: Control processing unit 6: Control target 11: Plant dynamics simulator 12: Status information control section 13: Behavioral Value Update Department 14: Action value storage section 15: Optimal action selection section 16: Episode number memory section 51: Learning information control unit 52: Status information conversion unit 53: Input / output device 54: Control result storage section
Claims
1. A learning processing device that determines an optimal behavior of a plant by learning, and a control processing device that controls the plant in accordance with the optimal behavior determined by the learning processing device, The learning processing device includes a state information control unit that converts a plurality of plant signals into plant states and defines a target state, an action value updating unit that uses the plant state, actions, and target state to determine an action value, which is the value of a state and an action between a previous operation and a current operation, and an optimal action selection unit that uses the action value to determine an optimal action for achieving the target state, and determines an action that suppresses oscillations in the plant's operating amount as the optimal action.
2. 2. The plant control system according to claim 1, The plant control system is characterized in that the learning processing device performs learning such that the deviation of the operation amount between the previous operation and the current operation is reduced and its value is increased in the direction approaching the target state for a given attenuation coefficient of the operation amount.
3. 2. The plant control system according to claim 1, The plant control system is characterized in that the learning processing device learns a control law that stabilizes the system by keeping the oscillation frequency of the manipulated variable and the convergence time to the target state within a specified range.
4. 2. The plant control system according to claim 1, The plant control system is characterized in that it is equipped with an input section and a display section, and is capable of inputting and displaying an arbitrary damping coefficient indicating the degree of vibration of a manipulated variable.
5. 5. The plant control system according to claim 4, A plant control system comprising: a display unit that displays a relationship between a convergence time to a target state and a vibration frequency; a damping coefficient that is input from the input unit; and a degree of vibration suppression that is adjusted.
6. 3. The plant control system according to claim 2, A plant control system characterized by suppressing vibrations in manipulated variables by learning to prevent the deviation from becoming extremely large using a function in which the value of updating attenuates linearly as the deviation becomes larger.
7. 3. The plant control system according to claim 2, The plant control system includes a means for suppressing vibration of the manipulated variable by learning so as to prevent the deviation from becoming extremely large by using a function in which the value of updating attenuates quadratically as the deviation increases.
8. 3. The plant control system according to claim 2, A plant control system, characterized in that learning is performed by using a function in which the value of updating attenuates quadratically as the deviation increases, so that the deviation does not become extremely large.
9. A plant control method for determining an optimal behavior of a plant by learning and controlling the plant in accordance with the optimal behavior determined by the learning process, comprising: The learning process includes converting a plurality of plant signals into plant states to define a target state, determining an action value, which is a value of a state and an action between a previous operation and a current operation, using the plant state, actions, and target state, determining an optimal action for achieving the target state using the action value, and determining an action that suppresses oscillations in the plant operation amount as the optimal action, the plant control method being characterized in that: