Control System, Control Method, and Control Program

The control system addresses the challenge of applying multi-agent reinforcement learning to real control systems by using virtual agents for action prediction and adapting to changes in operating states, achieving flexible and optimized control.

JP7696190B1Active Publication Date: 2025-06-20AISING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025031069
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning technologies face challenges in applying to real-world control systems, particularly in coping with changes in the operating state of control devices.

Method used

A control system configuration that includes a plurality of control devices associated with virtual agents performing action prediction, with a state variable acquisition unit, an ordering unit, an action prediction unit, a reward-related value prediction unit, and a cumulative value calculation unit, allowing for flexible control even with changes in operating states.

Benefits of technology

The system achieves overall optimization of control by considering the prediction results of other agents and can adapt to changes in the operating states of control devices, effectively addressing the challenges of applying multi-agent reinforcement learning to real control systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696190000001_ABST
    Figure 0007696190000001_ABST
Patent Text Reader

Abstract

Provided are a control system, a method, and a program using machine learning techniques. 【Solution means】The control system acquires state variables including the operating state of the control device corresponding to each agent from the environment, assigns an order to a plurality of agents, makes an action prediction in the nth agent based on the corresponding state variable and the reward-related cumulative value provided from the (n - 1)th agent, predicts a corresponding reward-related value based on the predicted action, generates, as the reward-related cumulative value provided to the (n + 1)th agent, a value obtained by subtracting the reward-related value corresponding to the (n + 1)th agent from the sum of the reward-related value corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent, operates in order for all agents targeted for action prediction, reward-related value prediction, and cumulative value generation, and controls each control device based on the predicted action of each agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a control system, particularly a control system using machine learning technology and the like.

Background Art

[0002] In recent years, research has been conducted on techniques for controlling a target using multi-agent reinforcement learning technology. For example, Non-Patent Document 1 describes a method of controlling a target using a technique in multi-agent machine learning in which agents arranged in a series sequentially select actions and share the selected actions with the next agent.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, this type of multi-agent machine learning technology as described in Non-Patent Document 1 has still been difficult to apply to actual control scenarios.

[0005] For example, in a real environment, it may happen that some of the plurality of control devices stop operating and the operating state is dynamically changed. The conventional system cannot cope with such a situation where the number of agents is changed.

[0006] The present invention has been made under the above-described technical background, and an object thereof is to solve the problems in applying multi-agent reinforcement learning technology to an actual control system.

Means for Solving the Problems

[0007] The above-described technical problems can be solved by a control system, method, program, etc. having the following configuration.

[0008] That is, the control device according to the present invention includes a plurality of control devices that control objects in the environment, and each of the control devices is associated with a virtual agent that respectively performs action prediction regarding the control of each of the control devices. The control system includes: a state variable acquisition unit that acquires, from the environment, state variables corresponding to each of the agents, including the operating state of the control device corresponding to each of the agents; an ordering unit that assigns an order to the plurality of agents; an action prediction unit that, in the nth agent, predicts an action based on the state variable corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent; a reward-related value prediction unit that predicts a reward-related value corresponding to the nth agent based on the predicted action; a cumulative value calculation unit that generates, as a reward-related cumulative value provided to the (n + 1)th agent, a value obtained by subtracting the reward-related value corresponding to the (n + 1)th agent from the sum of the reward-related value corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent; a repetitive operation unit that operates the action prediction unit, the reward-related value prediction unit, and the cumulative value calculation unit in order for all the agents targeted, thereby causing all the agents to predict actions; and a control unit that controls each of the control devices based on the actions predicted by each of the agents.

[0009] According to such a configuration, since each agent performs action prediction in consideration of the prediction results of other agents, overall optimization of control can be achieved. In addition, since the state variables include the operating states of the control devices corresponding to other agents, control can be performed flexibly even when the operating states of the control devices are changed. That is, problems in applying multi-agent reinforcement learning technology to a real control system, such as coping with changes in the operating state of a control device, can be solved.

[0010] Each of the agents is assigned a normalized coefficient. In the action prediction unit described above, for the nth agent, in addition to the state variable corresponding to the nth agent and the reward-related cumulative value provided by the (n - 1)th agent, the prediction of the action is further performed based on the sum of the coefficients for each of the agents up to the (n - 1)th agent. It may be such.

[0011] According to such a configuration, the influence of the reference order of the agents can be reduced by using coefficients, so that the action prediction accuracy can be improved.

[0012] The reward-related value prediction unit may predict the reward-related value based on a predetermined input value assigned to each agent in addition to the predicted action.

[0013] According to such a configuration, by adding a predetermined input such as a state variable, the reward-related value can be predicted more accurately.

[0014] It may further include a second repetition operation unit that repeats the operation of the repetition operation unit a predetermined number of times.

[0015] According to such a configuration, the accuracy of action prediction can be further improved.

[0016] The action prediction unit may predict the action using a regression neural network.

[0017] According to such a configuration, it is possible to predict an action taking into account the context.

[0018] The regression neural network may be an LSTM.

[0019] According to such a configuration, the problem of vanishing gradients can be solved, short-term memory and long-term memory can be integrated, and long-term dependency relationships can be learned. This is suitable for dealing with real environments that do not necessarily have the Markov property (or are partial Markov processes).

[0020] The reward-related value prediction unit may be one that predicts the reward-related value using a decision tree.

[0021] According to such a configuration, the reward-related value can be predicted in a manner that is relatively easy to interpret and has high generality.

[0022] The decision tree may be a gradient boosting decision tree.

[0023] According to such a configuration, the reward-related value can be predicted with a model that has a good balance between the memory capacity used and the learning and prediction accuracy.

[0024] The agent may function as an actor in actor-critic type reinforcement learning.

[0025] According to such a configuration, both the actor and the critic can be learned to increase the future reward using the result of action generation from the agent related to the actor.

[0026] In the critic of the actor-critic type reinforcement learning, learning of the action value function may be performed.

[0027] According to such a configuration, more desirable actions with higher rewards can be induced through the learning of the critic.

[0028] Viewed from another perspective, the present invention is a control method for a control system. That is, the control method according to the present invention includes a plurality of control devices for controlling objects in an environment, and each of the control devices is associated with a virtual agent that respectively performs action prediction regarding the control of each of the control devices. The control method of the control system is as follows: a state variable acquisition step of acquiring, from the environment, state variables corresponding to each of the agents, including the operating states of the control devices corresponding to each of the agents; an ordering step of assigning an order to the plurality of agents; an action prediction step of, in the nth agent, predicting an action based on the state variable corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent; a reward-related value prediction step of predicting a reward-related value corresponding to the nth agent based on the predicted action; a cumulative value calculation step of generating, as the reward-related cumulative value provided to the (n + 1)th agent, a value obtained by subtracting the reward-related value corresponding to the (n + 1)th agent from the sum of the reward-related value corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent; a repetitive operation step of sequentially applying the action prediction step, the reward-related value prediction step, and the cumulative value calculation step to all of the agents targeted, thereby causing actions to be predicted in all of the agents; and a control step of controlling each of the control devices based on the predicted actions of each of the agents.

[0029] Viewed from another perspective, the present invention is a control program for a control system. That is, the control program according to the present invention includes a plurality of control devices that control objects in the environment, and each of the control devices is associated with a virtual agent that respectively performs action prediction regarding the control of each of the control devices. It is a control program for a control system, and in the state variable acquisition step, state variables corresponding to each of the agents are acquired from the environment, including the operating state of the control device corresponding to each of the agents; in the ordering step, an order is given to the plurality of agents; in the action prediction step, in the nth agent, an action prediction is performed based on the state variable corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent; in the reward-related value prediction step, a reward-related value corresponding to the nth agent is predicted based on the predicted action; in the cumulative value calculation step, a value obtained by subtracting the reward-related value corresponding to the (n + 1)th agent from the sum of the reward-related value corresponding to the nth agent and the reward-related cumulative value provided from the (n - 1)th agent is generated as the reward-related cumulative value provided to the (n + 1)th agent; and in the repetitive operation step, the action prediction step, the reward-related value prediction step, and the cumulative value calculation step are sequentially applied to all the target agents, thereby causing actions to be predicted in all the agents, and in the control step, each of the control devices is controlled based on the predicted action of each of the agents.

Advantages of the Invention

[0030] According to the present invention, it is possible to solve the problems in the application of multi-agent reinforcement learning technology to a real control system.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0032] Hereinafter, an embodiment of the present invention will be described in detail with reference to the accompanying drawings.

[0033] (1. First Embodiment) The first embodiment will be described with reference to the drawings. In this embodiment, an example in which the control system according to the present invention is applied to the control of the input of raw materials and stirring in a chemical plant, particularly a reactor, will be described.

[0034] (1.1 Configuration) Figure 1 is a detailed configuration diagram of the information processing apparatus 1 included in the control system 40 according to this embodiment. As is clear from the figure, the information processing apparatus 1 includes a processor 10, a storage unit 11, a communication unit 12, a display control unit 15, an I / O processing unit 16, and an audio control unit 17.

[0035] The processor 10 is an arithmetic unit such as a CPU, and executes a program to realize various operations described later. The storage unit 11 is a storage medium such as a ROM / RAM, a flash memory, or a hard disk (including a non-temporary computer-readable storage medium), and stores a program and data for realizing various operations described later. The communication unit 12 is a communication unit that enables the exchange of information with an external device. The display control unit 15 performs processing to control image information and the like to be displayed on the display device 19. The I / O processing unit 16 processes input / output signals with an external device. The audio control unit 17 performs processing to control audio information and the like output to a speaker.

[0036] The processor 10 provides various functions together with the storage unit 11 and the like.

[0037] FIG. 2 is a functional block diagram of the information processing apparatus 1. As is clear from the figure, the processor 10 includes a data acquisition unit 101, a preprocessing unit 102, a prediction processing unit 103, a storage processing unit 104, and a learning processing unit 105.

[0038] The data acquisition unit 101 acquires data from the storage unit 11 and performs processing to provide the data to the preprocessing unit 102, the prediction processing unit 103, and the learning processing unit 105. The preprocessing unit 102 performs preprocessing on the data acquired from the data acquisition unit 101 and performs processing to provide the processed data to the prediction processing unit 103.

[0039] In this embodiment, the term "prediction" represents the output of a learned model and can be paraphrased using other terms. For example, terms such as "inference" may be used.

[0040] The prediction processing unit 103 performs prediction processing based on the data provided from the data acquisition unit 101 or the preprocessing unit 102. The prediction processing result is provided to the storage processing unit 104. The storage processing unit 104 performs processing to provide the provided data to the storage unit 11 for storage.

[0041] The learning processing unit 105 performs machine learning processing based on the data provided by the data acquisition unit 101. The learning processing result is provided to the memory processing unit 104. The memory processing unit 104 performs a process of providing the provided data to the storage unit 11 for storage.

[0042] Note that the hardware configuration is not limited to the configuration according to this embodiment. Therefore, other configurations may be adopted. For example, it may be configured as a server-client system, or a server for storing and providing data may be provided separately.

[0043] Also, in this embodiment, the computer program may be provided as a computer program product or a recording medium for recording the computer program. Note that the function corresponding to the processor 10 may be realized circuitously by an IC such as an FPGA.

[0044] FIG. 3 is an overall configuration diagram of the control system 40 of the chemical plant according to this embodiment. In the example of this figure, the control system 40 controls the input amount of various raw materials provided from the tank 30 to the reactor 45 and the stirring amount of the input raw materials in the reactor 45.

[0045] In the example of this figure, for raw material A, (N - 3) tanks 31 are provided, and the input amount of raw material A from each tank 31 is controlled by controlling the output (in this embodiment, the rotational frequency) of the corresponding pump 310. For raw material B, one tank 32 is provided, and the input amount of raw material B from the tank 32 is controlled by controlling the output (in this embodiment, the rotational frequency) of the corresponding pump 320. For raw material C, one tank 33 is provided, and the input amount of raw material C from the tank 33 is controlled by controlling the output (in this embodiment, the rotational frequency) of the corresponding pump 330. Also, a stirrer is arranged in the reactor 45, and the stirring amount by the stirrer is controlled by controlling the output (in this embodiment, the rotational speed) of the stirrer motor 43 connected to the stirrer.

[0046] Note that the N control devices each consisting of a pump (310, 320, 330) and a stirrer motor 43 are each connected to an information processing device 1 (not shown) and are controlled by N virtual agents.

[0047] FIG. 4 is a conceptual diagram of the reinforcement learning adopted by the control system 40 according to the present embodiment. As is clear from the figure, in the present embodiment, a so-called soft actor-critic type reinforcement learning technique that includes an actor 50 and a critic 60 and maximizes the entropy of the policy is also adopted, which is suitable for learning in the real world.

[0048] The actor 50 predicts an action a based on information such as a state s obtained from the environment 70. The actor 50 includes a policy network 51 including a plurality of learned models corresponding to virtual agents that predict an action a based on a given state s or the like. The action a predicted by each agent is the output to each control device, that is, the rotation frequency of each pump (310, 320, 330) and the rotation speed of the stirrer motor 43.

[0049] In the present embodiment, as this learned model, a regression type neural network (RNN), particularly, LSTM (Long Short Term Memory) is adopted.

[0050] According to such a configuration, it is possible to predict an action taking into account the context. In particular, by adopting LSTM, the problem of gradient disappearance can be solved, short-term memory and long-term memory can be integrated, and long-term dependency relationships can be learned. This is suitable for dealing with a real environment that does not necessarily have Markov property (or becomes a partial Markov process).

[0051] In this embodiment, the actor 50 further includes a reward-related value prediction model 52 that corresponds to each trained model (LSTM) and predicts a reward-related value based on the predicted action a or the like. The reward-related value is a value of the same dimension as the reward that facilitates reward calculation. In this embodiment, the reward is the total consumption energy of all agents or an amount corresponding thereto, and the reward-related value is the consumption energy e related to each agent. Further, the trained model related to the reward-related value prediction model 52 is a decision tree, particularly, a gradient boosting decision tree (GBDT).

[0052] According to such a configuration, using a decision tree, the reward-related value can be predicted in a manner that is relatively easy to interpret and has high versatility. In particular, by using GBDT, the reward-related value can be predicted with a model that has a good balance between the memory capacity used and the learning and prediction accuracy.

[0053] Note that in this embodiment, although the reward is set to a value that increases as the total consumption energy of the entire system decreases, the present invention is not limited to such a configuration. Therefore, for example, it may be set from the viewpoints of production speed, target characteristics of products, etc.

[0054] The policy network 51 is updated based on the loss function (P-Loss) 53 related to the actor 50. Although the update method is not particularly limited, in this embodiment, it is performed by the mini-batch stochastic gradient descent method.

[0055] All information regarding the environment 70 is aggregated and buffered in the storage unit 11 while being stored. The information regarding the environment 70 includes all variables described later, that is, the state s, the action a, a predetermined coefficient c, the reward-related value e, the reward r, and the like. Note that in this embodiment, the reward r is set as a variable that increases as the future total energy consumption of the entire system decreases.

[0056] Critic 60 includes a Q-network 61 that includes an action-value function (Q-function) and a target Q-network 62 that includes a target Q-function. Both the Q-function and the target Q-function are pre-trained models such as neural networks. The Q-network 61 and the target Q-network 62 output Q-values and target Q-values based on the values read from the memory unit 11.

[0057] The loss function (P-Loss) 53 of the actor 50 and the loss function (Q-Loss) 63 of the critic are calculated based on these Q-values and target Q-values, etc.

[0058] The Q-network 61 is updated based on the loss function (Q-Loss) 63 of the critic 60. Although the update method is not particularly limited, in this embodiment, it is performed by the mini-batch stochastic gradient descent method.

[0059] In this embodiment, the target Q-network 62 is updated by the soft update method. That is, in this embodiment, the target Q-network 62 is updated by taking the weighted average of the parameters of the models related to the Q-network 61 and the target Q-network 62.

[0060] (1.2.1 Prediction operation) Next, the prediction operation using the control system 40, that is, the actual control operation, will be described.

[0061] FIG. 5 is a general flowchart regarding the prediction operation of the control system 40. As is clear from the figure, when the process starts, the data acquisition unit 101 reads and refers to the data buffered in the memory unit 11 (S10). Note that the buffer amount of this data is the amount of data for the time series length required for prediction.

[0062] As a result of this reference process, if it is determined that the data has not accumulated more than a predetermined amount (S11: NO), random action generation is performed, and the generated action is written to the storage unit 11 by the storage processing unit 104. This written action is output to each pump (310, 320, 330) and the agitator motor 43 (S12). Each pump (310, 320, 330) and the agitator motor 43 perform control with the value related to the predicted action as the target value. After that, the process returns to the beginning, and a series of processes are performed again.

[0063] On the other hand, if it is determined that the data has accumulated more than a predetermined amount (S11: YES), the prediction processing unit 103 executes a prediction process and stores the prediction process result in the storage unit 11 by the storage processing unit 104 (S13). After this prediction process, this written action is output to each pump (310, 320, 330) and the agitator motor 43 (S15). Each pump (310, 320, 330) and the agitator motor 43 perform control with the value related to the predicted action as the target value. After this output, the process returns to the beginning, and a series of processes are performed again.

[0064] With reference to FIGS. 6 and 7, the details of the prediction process will be described (S13). FIG. 6 is a detailed flowchart of the prediction process (S13), and FIG. 7 is a conceptual diagram of the prediction process.

[0065] When the process starts, the prediction processing unit 103 performs a process of selecting a target agent based on the information about the control device acquired by the data acquisition unit 101, and performs a process of arranging the selected agents in an arbitrary order (S131). In the present embodiment, an agent is provided corresponding to each of the N control devices including each pump (310, 320, 330) and the agitator motor 43, and an order is assigned to those agents (that is, agent 1 to agent N in FIG. 7). Note that each agent selects an action a when a state s is given based on the policy π.

[0066] After the process of setting the order, the data acquisition unit 101 obtains an observation value о from the environment 70 for the N agents. nPerform a process of obtaining, and the preprocessing unit 102 performs preprocessing on the observed value о n and performs a process of converting it into the state variable s n . This converted state variable s n is provided to each agent.

[0067] Note that this state variable s n includes various variables related to the environment 70. In this embodiment, the state variable s n includes various parameters related to the environment 70, that is, temperature, humidity, monitoring values to be optimized (in-tank temperature and pressure in the reactor 45), monitoring values of the control device (material quality values (viscosity, particle size of raw materials), etc.). Further, the state variable s n includes information on whether each control device is currently in an operating state. For example, when only raw material A is charged into the reactor and stirred, only the agents corresponding to the pump 310 related to the tank 31 for raw material A and the stirrer motor 43 are set to the operating state (flag is 1), while the agents corresponding to the pumps (320, 330) related to the tank 32 for raw material B and the tank 33 for raw material C are set to the non-operating state (flag is 0).

[0068] According to such a configuration, since the operating state of the control device corresponding to other agents is included in the state variable s n , even if the operating state of the control device is changed, control can be performed flexibly.

[0069] When the conversion process of the state variable s n is completed, the prediction processing unit 103 performs a process of normalizing a predetermined coefficient corresponding to each agent (S133). By this normalization, the sum of all N coefficients of each agent becomes 1. In this embodiment, the energy consumption coefficient c n is adopted as the coefficient. The energy consumption coefficient c n is a value indicating the tendency of energy consumption in the control device corresponding to each agent.

[0070] According to such a configuration, by using the normalized coefficients, it is possible to emphasize agents that receive more information and are arranged in the latter half in terms of order, so that the action prediction accuracy can be improved.

[0071] After the normalization process, the prediction processing unit 103 performs a process of initializing the values of the cumulative total of consumed energy e_cum and the cumulative total of coefficients c_cum (S135). In this embodiment, the values of e_cum0 and c_cum0 are each set to 0.

[0072] After this initialization process, the prediction processing unit 103 performs a process of initializing a variable k indicating the number of times of the repeated process described later (S136). For example, the variable k is set to 1.

[0073] Also, after the initialization process of the variable k, the prediction processing unit 103 performs a process of initializing a variable n indicating the agent to be referred to (S137). For example, the variable n is set to 1.

[0074] After the initialization process of the variable n, the prediction processing unit 103 inputs the state variable s n to the nth agent, the cumulative total of consumed energy e_cum n-1 of the (n - 1)th agent, and the cumulative total of coefficients c_cum n-1 of the (n - 1)th agent to the corresponding learned model (LSTM) of the policy network 51, and predicts the action a n in the agent (S138).

[0075] After the action prediction process, the prediction processing unit 103 inputs the predicted action a n and a predetermined input x n to the learned model (gradient boosting decision tree) of the reward-related value prediction model 52 corresponding to the Nth agent, and performs a process of calculating the predicted consumed energy e_hat n as the reward-related value in the agent (S140). In this embodiment, the input x n is the state variable s n corresponding to the agent.

[0076] According to such a configuration, in addition to the predicted action a n the state variable s n at that time is also input as the input x n Therefore, the reward-related value can be predicted with higher accuracy.

[0077] After the arithmetic processing, a process of updating various cumulative values is performed. That is, the prediction processing unit 103 adds the coefficient cumulative value c_cum n-1 up to the (n - 1)-th one and the (n - 1)-th coefficient c n-1 and substitutes the sum into the n-th coefficient cumulative value c_cum n-1 (S141).

[0078] Also, the prediction processing unit 103 subtracts the (n + 1)-th predicted energy consumption e_hat n-1 from the sum of the cumulative total e_cum n of the energy consumption up to the (n - 1)-th one and the n-th predicted energy consumption e_hat n+1 and substitutes the resulting value into the n-th cumulative total e_cum n of the energy consumption (S142). Note that the (n + 1)-th predicted energy consumption e_hat n+1 refers to the value one step before in the iterative process using the variable k.

[0079] After these update processes, the prediction processing unit 103 determines whether all agents have been referred to, that is, whether N agents have been referred to (S143). If not all agents have been referred to yet (S143 NO), the variable n is incremented by 1 (S145), and the same processes (S138 to S142) are executed for the next (n + 1)-th agent. On the other hand, if it is determined that all agents have been referred to (S143 YES), the process proceeds to the next stage assuming that the action prediction for all agents has been completed.

[0080] When the references to all agents are completed, the prediction processing unit 103 determines whether a predetermined repetitive process has been completed, that is, whether a series of processes has been repeated m times (S146). If the repetition has not been performed the predetermined number of times (S146 NO), the variable k is incremented by 1 (S147), and the same processes (S137 to S143) are executed in the next (k + 1) round.

[0081] According to such a configuration, the accuracy of action prediction can be further improved by the repetitive process.

[0082] On the other hand, when the repetitive process has been performed m times (S146 YES), the prediction processing unit 103 provides the generated action a n and the predicted energy consumption value e_hat n to the storage processing unit 104 and performs the process of storing them in the storage unit 11 (S148), and the process ends. The stored prediction result, that is, the generated action a n is used to execute the action output process (S15).

[0083] According to the above configuration, since each agent performs action prediction in consideration of the prediction results of other agents, overall optimization of control can be achieved. Further, since the state variable includes the operating state of the control device corresponding to the agent, control can be flexibly performed even when the operating state of the control device is changed. That is, problems in applying the multi-agent reinforcement learning technology to a real control system, such as flexible response to changes in the operating state of the control device, can be solved.

[0084] (1.2.2 Reinforcement learning operation) Next, machine learning for generating each learned model and the like used in the prediction operation, particularly the flow of reinforcement learning, will be described.

[0085] Figure 8 is a detailed flowchart regarding the reinforcement learning process executed by the learning processing unit 105. As is clear from the figure, when the process starts, the learning processing unit 105 performs a process of initializing the agent to be referred to, that is, a process of initializing the variable n corresponding to the agent (S21). For example, the variable n is set to 1.

[0086] After the initialization process, the learning processing unit 105 performs a process of reading the data stored in the storage unit 11 (S22). More specifically, for the batch size, the state variable s n , action a n , reward r, action a n and the next state s' n after that, etc., are read.

[0087] After the reading process, the learning processing unit 105 uses the corresponding learned model of the policy network 51 to perform a prediction process based on the state s' n and generates a predicted action a' n (S23).

[0088] After this process, the learning processing unit 105 inputs the reward r, state variable s n , action a n , state s' n , and predicted action a' n to the target Q network 62, thereby performing a process of predicting the target Q value (S25). This target Q value is for the Q network 61 to learn with that value as the target and contributes to the stabilization of the learning of the Q network 61.

[0089] After the process of predicting the target Q value, the learning processing unit 105 inputs the reward r, state variable s n , action a n to the Q network 61, thereby performing a process of predicting the Q value (S26).

[0090] After the target Q-value and the Q-value prediction process, the learning processing unit 105 performs a process of calculating a loss function (Q-Loss) 63 of the critic 60 based on these target Q-values and Q-values (S27). After that, the parameters of the Q-function related to the Q-network 62 are updated using this loss function (Q-Loss) (S28). This update is performed, for example, by the mini-batch stochastic gradient descent method.

[0091] When the learning process in the Q-function is completed, the learning processing unit 105 inputs the state variable s n to the policy network 51 to execute a process of predicting the action a'' n (S30). After this prediction process, the learning processing unit 105 calculates a loss function (P-Loss) 53 of the policy network 51 (or actor 50) based on these state variables s n , action a'' n (S31). After that, the parameters of the learned model (LSTM) related to the policy network 51 are updated using this loss function (P-Loss) (S32). This update is performed, for example, by the mini-batch stochastic gradient descent method.

[0092] After the learning process of the policy function, the learning processing unit 105 executes a learning process for the target Q-network 62 to update the parameters related to its learned model (S33). This learning process is executed by the soft update method. That is, in this embodiment, the target Q-network 62 is updated by taking a weighted average of the parameters of the models related to the Q-network 61 and the target Q-network 62 using the weight coefficient τ.

[0093] When the learning process of the target Q function is completed, the learning processing unit 105 determines whether it has referred to all agents (S35). If it has not yet referred to all agents (S35: NO), it increments the variable n by 1 and repeats the series of processes again (S22 to S33). On the other hand, if it is determined that all agents have been referred to (S35: YES), the learning process ends, and all learning results are stored in the storage unit 11 by the storage processing unit 104.

[0094] According to such a configuration, both the actor 50 and the critic 60 can be learned to increase the future rewards by using the result of action generation from the agent related to the actor 50.

[0095] Although the detailed flowchart of FIG. 8 does not mention the learning or updating of the reward-related value prediction model 52, this model may also be learned sequentially or batchwise. That is, when an action a n is taken in a certain state variable s n , the energy consumption e n of the target control device corresponding to each agent is observed, and supervised learning may be performed in the model (GBDT) based on these data.

[0096] Also, each step in the above flow may be reordered, except for those that require the calculation result of the previous step.

[0097] (2. Modification Example) The present invention can be implemented in various modified forms.

[0098] In the above embodiment, instead of the action a n predicted by the agent, it is converted into the predicted value (e_hat n ) of the reward-related value and provided to the next agent, but the present invention is not limited to such a configuration. Therefore, the reward-related value prediction model 52 or the like may not be provided, and simply the action a n may be provided to the next agent.

[0099] In the above-described embodiment, an example of application to a chemical plant has been described. However, the present invention can be applied to various objects. Therefore, it can be applied to a system that performs control using a plurality of devices (agents). For example, it may be applied to an energy management system (EMS) related to HEMS, district heating and cooling, data centers, solar and wind power generation, etc. Further, it may be applied to an automatic driving system, robot control in factories and homes, stock trading, smart cities, customer support, etc.

[0100] FIG. 9 shows a modified example in which the control system according to the present invention is applied to a temperature control system 90 (air conditioning system) of a building 95 as an example of an energy management system (EMS). The building 95 is, for example, an office building, a commercial facility, a hospital, or the like.

[0101] An air handling unit (AHU) 89 is provided in the building 95, and cold air is provided from the AHU 89. At this time, the AHU 89 provides cold air using chilled water supplied from an air-cooled chiller 81, a turbo chiller 82, and an absorption chiller 83.

[0102] Chilled water pumps 883 and 884 are provided between the turbo chiller 82 and the AHU 89, and between the absorption chiller 83 and the AHU 89, respectively.

[0103] Further, the turbo chiller 82 is operated by being supplied with power from a commercial power cogeneration 86, and performs heat exchange by circulating cooling water between the cooling tower 85. Similarly, the absorption chiller 83 is operated by being supplied with power from a boiler cogeneration 88, and performs heat exchange by circulating cooling water between the cooling tower 87. For these circulations, cooling water pumps 881 and 882 are provided between the cooling towers (88, 87) and the chillers (82, 83), respectively.

[0104] In this system, the control devices are the chilled water pumps 883, 884, and the cooling water pumps 881, 882. That is, various state variables s provided from the environment n, for example, based on temperature, humidity, wind speed, monitoring values to be optimized (such as room temperature, occupancy status, etc.), monitoring values of the control device (such as chilled water temperature, cooling water temperature, operating status, etc.), each agent (881 to 884) performs action a n , that is, in this example, the pump frequency is predicted. At this time, the reward r is, for example, a value that becomes larger as the total sum of the overall energy consumption becomes smaller. Note that the reward is not limited to such a value and can be arbitrarily set. For example, it may be an index value for stable control (such as frequency fluctuation, etc.).

[0105] Although the embodiments of the present invention have been described above, the above embodiments merely show a part of the application examples of the present invention, and are not intended to limit the technical scope of the present invention to the specific configurations of the above embodiments. Also, the above embodiments can be appropriately combined within a range where no contradiction occurs.

Industrial Applicability

[0106] The present invention can be used in various industries that utilize machine learning technology.

Explanation of Signs

[0107] 1 Information processing device 10 Processor 101 Data acquisition unit 102 Preprocessing unit 103 Prediction processing unit 104 Memory processing unit 105 Learning processing unit 11 Memory unit 12 Communication unit 15 Display control unit 16 I / O processing unit 17 Audio control unit 30 Tank 31 Tank (raw material A) 310 Feed rate control pump 32 Tank (raw material B) 320 Feed rate control pump 33 Tank (raw material C) 330 Feed rate control pump 40 Control System 43 Stirrer Motor 45 Reactor 50 Actor 51 Policy Network 52 Reward-Related Value Prediction Model 53 Loss Function (Policy Network / Actor) 60 Critic 61 Q-Network 62 Target Q-Network 63 Loss Function (Critic) 70 Environment 90 Temperature Control System 81 Air-Cooled Chiller 82 Turbo Refrigerator 83 Absorption Refrigerator 85 Cooling Tower 86 Power Supply (Commercial Power / Cogeneration) 87 Cooling Tower 88 Heat Source (Boiler / Cogeneration) 881 Cooling Water Pump 882 Cooling Water Pump 883 Chilled Water Pump 884 Chilled Water Pump 89 AHU (Air Handling Unit)

Claims

1. A control system including a plurality of control devices for controlling objects in an environment, each of the control devices being associated with a virtual agent that performs behavioral prediction regarding the control of each of the control devices, a state variable acquisition unit that acquires state variables corresponding to each of the agents, including an operating state of the control device corresponding to each of the agents, from the environment; an order assigning unit that assigns an order to the plurality of agents; a behavior prediction unit for predicting a behavior of the n-th agent based on the state variable corresponding to the n-th agent and a reward-related cumulative value provided from the n-1-th agent; a reward-related value prediction unit that predicts a reward-related value corresponding to the n-th agent based on the predicted behavior; an accumulation value calculation unit that generates a value obtained by subtracting a reward-related value corresponding to the (n+1)th agent from the sum of the reward-related value corresponding to the (n)th agent and the reward-related accumulation value provided from the (n-1)th agent, as a reward-related accumulation value to be provided to the (n+1)th agent; a repeating operation unit which operates the action prediction unit, the reward-related value prediction unit, and the cumulative value calculation unit in sequence for all of the agents to be targeted, thereby predicting actions for all of the agents; A control system comprising: a control unit that controls each of the control devices based on the predicted behavior of each of the agents.

2. Each agent is assigned a respective normalized coefficient; 2. The control system of claim 1, wherein in the behavior prediction unit, the nth agent predicts the behavior based on the state variable corresponding to the nth agent, the reward-related cumulative value provided from the n-1th agent, and further based on the sum of the coefficients related to each of the agents up to the n-1th agent.

3. The control system according to claim 1 , wherein the reward-related value prediction unit predicts the reward-related value based on a predetermined input value assigned to each of the agents in addition to the predicted behavior.

4. The control system according to claim 1 , further comprising a second repeating operation unit that causes the repeating operation unit to repeatedly operate a predetermined number of times.

5. The control system according to claim 1 , wherein the behavior prediction unit predicts the behavior using a recurrent neural network.

6. The control system of claim 5 , wherein the recurrent neural network is a LSTM.

7. The control system according to claim 1 , wherein the reward-related value prediction unit predicts the reward-related value using a decision tree.

8. The control system of claim 7 , wherein the decision tree is a gradient boosting decision tree.

9. The control system according to claim 1 , wherein the agent functions as an actor in actor-critic type reinforcement learning.

10. The control system according to claim 9 , wherein a critic in the actor-critic type reinforcement learning learns an action-value function.

11. A control method for a control system including a plurality of control devices that control objects in an environment, each of the control devices being associated with a virtual agent that performs behavior prediction regarding control of each of the control devices, the method comprising: a state variable acquisition step of acquiring state variables corresponding to each of the agents, including an operating state of the control device corresponding to each of the agents, from the environment; an ordering step of ordering the agents; a behavior prediction step of predicting a behavior of the n-th agent based on the state variables corresponding to the n-th agent and a reward-related cumulative value provided by the n-1-th agent; a reward-related value prediction step of predicting a reward-related value corresponding to the n-th agent based on the predicted behavior; an accumulation value calculation step of subtracting a reward-related value corresponding to the n+1th agent from the sum of the reward-related value corresponding to the nth agent and the reward-related accumulation value provided from the n-1th agent, as a reward-related accumulation value to be provided to the n+1th agent; a repeating operation step of applying the action prediction step, the reward-related value prediction step, and the cumulative value calculation step to all of the agents in turn, thereby predicting actions for all of the agents; and a control step of controlling each of the control devices based on the predicted behavior of each of the agents.

12. A control program for a control system including a plurality of control devices for controlling objects in an environment, each of the control devices being associated with a virtual agent that performs behavior prediction regarding control of each of the control devices, the control program comprising: a state variable acquisition step of acquiring state variables corresponding to each of the agents, including an operating state of the control device corresponding to each of the agents, from the environment; an ordering step of ordering the agents; a behavior prediction step of predicting a behavior of the n-th agent based on the state variables corresponding to the n-th agent and a reward-related cumulative value provided by the n-1-th agent; a reward-related value prediction step of predicting a reward-related value corresponding to the n-th agent based on the predicted behavior; an accumulation value calculation step of subtracting a reward-related value corresponding to the n+1th agent from the sum of the reward-related value corresponding to the nth agent and the reward-related accumulation value provided from the n-1th agent, as a reward-related accumulation value to be provided to the n+1th agent; a repeating operation step of applying the action prediction step, the reward-related value prediction step, and the cumulative value calculation step to all of the agents in turn, thereby predicting actions for all of the agents; and a control step of controlling each of the control devices based on the predicted behavior of each of the agents.

Citation Information

Patent Citations

  • Parallel management and control system and method for intelligent workshop scheduling

    CN116700160A

  • Multi-microgrid collaborative optimization operation method, device, equipment and medium

    CN117374937A

  • Dynamic scheduling method and system for unstable hybrid job shop with randomly arrived orders

    CN118780540A