Control actions of an automatically dimension-reduced and optimized control system via a mathematical representation of the control system

By combining CRL and CMDP technologies, the subset of system state variables are automatically identified and the CMDP model is simplified, and the problems of computing complexity and resource requirements in the existing technology are solved, achieving rapid and efficient operational optimization.

CN115605817BActive Publication Date: 2025-06-24INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180034423.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-11
Filing Date
2021-04-21
Publication Date
2025-06-24
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to calculate the large number of transfer probability values ​​of controlled application systems within a reasonable time, and the constrained mathematical model is complex and difficult to calculate, resulting in the control system being unable to effectively optimize operations.

Method used

By combining constrained reinforcement learning (CRL) and constrained Markov decision-making process (CMDP) technology, we automatically identify a subset of system state variables, reduce the state space dimension, simplify the transfer probability matrix of the CMDP model, and formulate it into a linear programming (LP) problem.

Benefits of technology

The operation of controlled application systems is achieved quickly and efficiently optimized while taking into account the constraints imposed on the system, reducing computational complexity and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115605817B_ABST
    Figure CN115605817B_ABST
Patent Text Reader

Abstract

A method for automatically reducing the dimension of a mathematical representation of a controlled application system is provided. The method includes receiving, at a control system, data corresponding to a control action and system state variables related to the controlled application system, fitting a Constrained Reinforcement Learning (CRL) model to the controlled application system based on the data, and automatically identifying a subset of the system state variables by selecting control action variables of interest and identifying the system state variables that drive the CRL model to recommend each control action variable of interest. The method further includes automatically performing state space reduction of the CRL model using the subset of the system state variables, estimating a transition probability matrix of a Constrained Markov Decision Process (CMDP) model of the controlled application system, and formulating the CMDP model as a Linear Programming (LP) problem using the transition probability matrix and a number of costs.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] The present disclosure relates to the field of machine learning. More specifically, the present disclosure relates to optimizing control actions of a control system that guides the operation of a controlled application system by automatically reducing the dimension of a mathematical representation of the controlled application system. SUMMARY OF THE INVENTION

[0002] According to an embodiment described herein, a method is provided for automatically reducing the dimension of a mathematical representation of a controlled application system that is subject to one or more constraint values. The method includes receiving, at a control system that guides the operation of the controlled application system, data corresponding to control action variables and system state variables related to the controlled application system, and fitting a Constrained Reinforcement Learning (CLR) model to the controlled application system via a processor of the control system based on the data corresponding to the control action variables and the system state variables. The method further includes automatically identifying a subset of the system state variables via the processor by selecting control action variables of interest and, for each control action variable of interest, identifying the system state variables that drive the CRL model to recommend the control action variables of interest. The method further includes automatically performing state space reduction of the CRL model using the subset of the system state variables via the processor, estimating a transition probability matrix of a Constrained Markov Decision Process (CMDP) model of the controlled application system that is subject to one or more constraint values after the state space reduction via the processor, and formulating the CMDP model as a Linear Programming (LP) problem using the transition probability matrix, one or more cost objectives, and one or more constraint-related costs.

[0003] In another embodiment, a control system is provided that guides the operation of a controlled application system that is subject to one or more constraint values. The control system includes an interface for receiving data corresponding to control action variables and system state variables related to the controlled application system. The control system further includes a processor and a computer-readable storage medium. The computer-readable storage medium stores program instructions that direct the processor to fit a CRL model to the controlled application system based on the data corresponding to the control action variables and the system state variables, and to automatically identify a subset of the system state variables by selecting control action variables of interest and, for each control action variable of interest, identifying the system state variables that drive the CRL model to recommend the control action variables of interest. The computer-readable storage medium further stores program instructions that direct the processor to automatically perform state space reduction of the CRL model using the subset of the system state variables, estimate a transition probability matrix of a CMDP model of the controlled application system that is subject to one or more constraint values after the state space reduction, and formulate the CMDP model as an LP problem using the transition probability matrix, one or more cost objectives, and one or more constraint-related costs.

[0004] In yet another embodiment, a computer program product is provided for automatically reducing the dimension of a mathematical representation of a controlled application system that complies with one or more constraint values. The computer program product includes a computer-readable storage medium having program instructions embodied therewith, where the computer-readable storage medium itself is not a transitory signal. The program instructions are executable by a processor to cause the processor to fit a CRL model to the controlled application system, where the CRL model is based on data corresponding to control action variables and system state variables, and automatically identify a subset of the system state variables by selecting control action variables of interest and identifying, for each control action variable of interest, the system state variables that drive the CRL model to recommend the control action variables of interest. The program instructions are further executable by the processor to cause the processor to automatically perform state space reduction of the CRL model using the subset of the system state variables, estimate a transition probability matrix of a CMDP model of the controlled application system that complies with one or more constraint values after the state space reduction, and formulate the CMDP model as an LP problem using the transition probability matrix, one or more cost objectives, and one or more constraint-related costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a simplified block diagram of an exemplary wastewater treatment plant (WWTP) in which the automatic state space reduction techniques described herein may be implemented; and

[0006] Figure 2 is a process flow diagram of a method for automatically reducing the dimension of a mathematical representation of a controlled application system that complies with one or more constraint values. DETAILED DESCRIPTION

[0007] Machine learning techniques are commonly used for optimizing the operation of control systems associated with various different types of controlled application systems, such as wastewater treatment plants, healthcare systems, agricultural systems, water resource management systems, queuing systems, epidemic handling systems, robotic motion planning systems, and the like. More specifically, machine learning techniques are used to generate models that mathematically describe the operation of such controlled application systems. For example, reinforcement learning (RL) techniques can be used to learn a strategy that maps the system state of a particular controlled application system relative to its environment to control actions taken by an operator of a control system associated with the controlled application system in accordance with the concept of maximizing cumulative reward. An extension of reinforcement learning, known as deep reinforcement learning (DRL), utilizes deep neural networks to model the environmental behavior of a customer's / problem's. Additionally, when constraints are imposed on the operation of a controlled application system, constraint reinforcement learning (CRL) and / or constrained deep reinforcement learning (CDRL) techniques can be used to learn a strategy that takes such constraints into account. Additionally, Markov decision process (MDP) techniques involve a complex framework for making decisions using explicitly computed / estimated probabilities of transitioning between different system states in response to particular control actions. When using constrained Markov decision process (CMDP) techniques, the control model is such that specific constraints associated with the controlled application system are satisfied.

[0008] In operation, such controlled application systems typically involve a very large number of possible system states, as well as a very large number of available control actions that can be taken by the corresponding control system. For example, some controlled application systems involve hundreds of possible system state variables, as well as over a hundred possible control actions. Thus, implementing a constrained mathematical model for the controlled application system may involve computing millions of transition probability values or learning a very large neural network. However, control systems for these applications typically do not have sufficient processing power to compute such a large number of transition probability values in a reasonable amount of time, and many controlled application systems do not include sufficient sensors to collect enough data to perform computations at such a level. Additionally, when one or more constraints are imposed on the controlled application system, the constrained mathematical model becomes even more complex and difficult to compute. Thus, it is generally desirable to reduce the size of such a constrained mathematical model.

[0009] Accordingly, the present disclosure describes techniques for automatically reducing the state space dimension of such a constrained mathematical model by reducing the number of possible system states from, for example, thousands of variables to fewer than 10 or fewer than 20 variables. This is achieved using a combination of CMDP and CRL (or CDRL) techniques. The resulting simplified model allows the control system to quickly and effectively optimize the operation of the controlled application system while still taking into account one or more constraints imposed on the controlled application system.

[0010] The techniques described herein can be applied to any suitable type of controlled application system, such as a wastewater treatment plant, a healthcare system, an agricultural system, a water resource management system, a queuing system, an epidemic handling system, a robotic motion planning system, and the like. However, for ease of discussion, the embodiments described herein relate to applying the techniques to a wastewater treatment plant.

[0011] Figure 1 FIG. 4 is a simplified block diagram of an exemplary wastewater treatment plant (WWTP) 100 in which the automatic state space reduction techniques described herein can be implemented. The WWTP 100 can be any suitable type of wastewater treatment unit, device, or system configured to treat influent wastewater (referred to as influent 102). The influent 102 travels through a liquid line 104 within the WWTP 100, which can include any number of screens, pumps, reactors, sedimentation tanks, separation devices, blowers, etc. for treating the influent 102 to a level set according to local regulations or protocols. Two main outputs of the liquid line 104 are treated liquid (referred to as effluent 106) and treated biosolids (referred to as sludge 108). The sludge 108 then passes through a sludge line 110, which can include any number of screens, pumps, reactors, sedimentation tanks, separation devices, dewatering systems, aeration tanks, blowers, etc. for treating the sludge 108. The resulting treated sludge 112 (which can include filter cake) is then sent, for example, for disposal 114, while the recycle effluent 116 (which can include activated sludge for the reactor) is recycled, for example, back into the liquid line 104. In addition, methane gas 118 separated from the sludge 108 in the sludge line 110 can be sent to a gas line 120, which can be sold or used for various purposes in the gas line 120.

[0012] The WWTP 100 also includes a control system 122 configured to control the functions of various devices and systems within the WWTP 100, such as (multiple) screens, (multiple) pumps, (multiple) reactors, (multiple) sedimentation tanks, (multiple) separation devices, (multiple) blowers, (multiple) dewatering systems, (multiple) aeration tanks, etc. included in the liquid line 104 and / or the sludge line 110. In various embodiments, the control system 122 of the WWTP 100 includes, for example, one or more servers, one or more general-purpose computing devices, one or more special-purpose computing devices, and / or one or more virtual machines. Additionally, the control system 122 includes one or more processors 124, and an interface 126 for receiving readings from a plurality of sensors within the WWTP 100, such as one or more sensors 128A connected to the liquid line 104 and one or more sensors 128B connected to the sludge line 110. The (multiple) sensors 128A and 128B can be used to monitor various conditions related to the WWTP 100. For example, the sensors 128A and 128B can be used to monitor flow variables, such as the total nitrogen flow variable at the effluent and the total phosphorus flow variable at the effluent. Readings can be received directly from the (multiple) sensors, or can be received via a proxy or input device.

[0013] The interface 126 can also be used to obtain data from one or more databases 130 associated with the WWTP 100 via a network 132 (e.g., the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network). The network 132 can include associated copper transmission cables, optical transmission fibers, wireless transmission devices, routers, firewalls, switches, gateway computers, edge servers, etc. The data obtained from the database 130 can include, for example, information related to the operating requirements of the WWTP 100, as well as historical operating data of the WWTP 100 or other similar WWTPs.

[0014] The control system 122 also includes a computer-readable storage medium (or media) 134 that includes program instructions executable by the processor(s) 124 to control the operation of the WWTP 100, as further described herein. The computer-readable storage medium 134 may be integrated with the control system 122 or may be an external device connected to the control system 122 in use. The computer-readable storage medium 134 may include, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium 134 includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any appropriate combination of the foregoing. Additionally, as used herein, the term "computer-readable storage medium" should not be construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire. In some embodiments, the interface 126 includes a network adapter card or network interface that receives program instructions from the network 132 and forwards the program instructions for storage in the computer-readable storage medium 134 within the control system 122.

[0015] The program instructions may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program instructions may be executed entirely on the control system 122, partially on the control system 122, executed as a stand-alone software package, partially on the control system 122 and partially on a remote computer or server connected to the control system 122 via the network 132, or entirely on such remote computer or server. In some embodiments, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the program instructions by utilizing the state information of the program instructions to customize the electronic circuit so as to perform aspects of the techniques described herein.

[0016] In operation, the properties of the effluent 106 and the treated sludge 112 leaving the WWTP 100 must comply with regulatory constraints specific to the location of the WWTP 100. Such constraints can include, for example, limits on the total nitrogen concentration and the total phosphorus concentration of the effluent 106 over a specific time period. An example of such a constraint is an upper limit of 15 milligrams per liter (mg / L) on the average total nitrogen concentration of the effluent 106 per month (or per day). To treat the effluent 106 and the treated sludge 112 to such an extent, multiple steps are carried out in the liquid line 104 and the sludge line 110. For example, the liquid line 104 can include various different types of equipment for treating the influent 102 via biophysical and / or chemical subprocesses, as well as equipment for treating the influent 102 via any number of iterative subprocesses such that the same influent portion visits some of the subprocesses multiple times.

[0017] To comply with such a constraint, the control system 122 can be used to control various control action variables and system state variables related to the WWTP 100. As used herein, the term "control action" refers to a possible action taken by the control system to control the operation of the controlled application system, while the term "system state" refers to a possible state of the controlled application system after the control system has taken one or more control actions. Examples of system state variables include the following: (1) operating at a low, medium, or high influent flow rate (in cubic meters per hour (m 3 / hr)); (2) operating with one or more feedback or recirculation loops set to an on or off setting; (3) operating with an upper limit on the total nitrogen concentration (in mg / L) and / or the total phosphorus concentration (in mg / L) of the effluent 106 set to low, medium, or high; (4) operating over a time period related to an electricity cost set to an expensive, medium, or cheap setting; and (5) operating with a total operating cost set to low, medium, or high. Examples of control action variables include the following: (1) increasing or decreasing the speed of one or more blowers within the WWTP 100; (2) increasing or decreasing the amount of chemicals added to the liquid line 104 and / or the sludge line 110 within the WWTP 100; (3) increasing or decreasing the pumping rate (in m 3 / hr) of one or more pumps within the WWTP 100; and (4) controlling feedback devices within the WWTP 100, such as those for sending the recirculated effluent 116 back to the liquid line 104.

[0018] Due to this complexity, WWTPs typically have high operating costs. For example, the total operating costs can include the cost of electricity required to operate components within the WWTP 100, the cost of various chemicals used in the wastewater treatment process, and the cost of disposing of the resulting treated sludge 112. Today, most WWTPs operate in a risk-averse mode that is conservative and inefficient, without the ability to quantify risk or truly optimize costs. Some subprocesses are locally optimized; however, local optimization of one subprocess can have a negative impact on one or more other subprocesses that are not considered. This local optimization can even have an adverse effect on the overall process.

[0019] Accordingly, the computer-readable storage medium 134 of the control system 122 includes an operation optimization module 136. The operation optimization module 136 is configured to generate a mathematical representation of the WWTP 100 that can be used to select control actions that maximize the operating efficiency of the WWTP 100 and minimize the operating costs, while taking into account one or more imposed constraints.

[0020] In operation, due to the large number of variables associated with the wastewater treatment process, the CMDP model associated with the WWTP 100 involves a high degree of complexity. For example, the WWTP 100 can involve more than 100 possible control actions that can be taken by the control system 122 and hundreds of possible system state variables. As a result, implementing the CMDP model for the WWTP 100 can include calculating or estimating more than 2^200x100 transition probability values. However, the control system 100 may not have sufficient processing power to calculate such a large number of transition probability values within a reasonable time. In addition, the WWTP 100 includes only a limited number of sensors 128A and 128B. For example, the entire WWTP 100 can include only 5-10 sensors. As a result, the control system 122 may not be able to collect enough data to perform the calculations at this level. In addition, because constraints are imposed on the WWTP 100, the CMDP model becomes more complex and difficult to calculate.

[0021] Thus, according to an embodiment described herein, an operation optimization module 136 within a computer-readable storage medium 134 of a control system 122 includes a sub-module referred to as an automatic state space reduction module 138. The automatic state space reduction module 138 includes instructions that direct one or more processors 124 to automatically reduce the state space dimension of a CRL model associated with the WWTP 100 by reducing the number of possible system states used within the CRL model from, for example, thousands of variables to fewer than 10 or fewer than 20 variables. This in turn simplifies the transition probability matrix of the CMDP model of the controlled application system, allowing the CMDP model to be formulated as a linear programming (LP) problem. The control system 122 can then use the LP problem to quickly and efficiently control the operation of the WWTP 100 while fully considering the imposed constraints.

[0022] It will be appreciated that Figure 1 the block diagrams are not intended to indicate that the WWTP 100 (or the control system 122 within the WWTP 100) will include Figure 1 all of the components shown in Figure 1 Rather, the WWTP 100 (and / or the control system 122) may include Figure 1 any number of additional or alternative components not shown in For example, the computer-readable storage medium 134 of the control system 122 may include any number of additional modules for controlling and / or optimizing the operation of the WWTP 100. Additionally, the control system 122 may include and / or be connected to an analog system that provides analog data related to the operation of the WWTP 1000. In some embodiments, the control system 122 may use a combination of analog data and data obtained from sensors 128A and 128B to optimize the operation of the WWTP 100.

[0023] Now refer to the mathematical interpretation of the CMDP technique that may be used according to an embodiment described herein. Define the CMDP as a 5-tuple where S is a finite set of system states (|S| = n), U is a finite set of control actions (|U| = k), P denotes a transition probability function (X 2 × U → [0; 1]), c denotes a cost function (X × U → R) that includes a vector of costs for each state-action pair, and denotes a vector of costs associated with the constraints. Denote the finite set of system states S as <s1, s2,... s s > ∈ S, where the elements from S are vectors, and denote the finite set of control actions U as <u1, u2,... u u > ∈ U, where the elements from U are vectors. Represent the transition probability function P as P(s j |s i , u k) ∈ [0, 1]. Additionally, s t is a random variable representing the state at time t, and u t is a random variable representing the action at time t.

[0024] The transition probability function P(y|x, u) quantifies the probability of transitioning from one system state x to another system state y when action u is chosen. The cost associated with choosing action u when the system state x equals c(x, u). Based on different probabilities and outcomes, a policy π is generated, where each policy induces a probability measure over state-action pairs. The goal of the CMDP model is to find a policy π that minimizes the total cost C(π) while ensuring that the total value of the constraint remains below a specified value.

[0025] For the CMDP discounted cost model, the parameter β ∈ (0, 1) specifies the rate at which future costs are discounted, where p π (x t = x; u t = u) indicates the probability of the event x t = x and u t = u when the initial state equals x0 and the policy is π. According to this model, the discounted occupation measure is defined as shown in Equation (1), where (0 < β < 1 - (discount factor)).

[0026]

[0027] As shown in Equation (2), the expected value of the cost at time t with policy π is defined.

[0028]

[0029] As shown in Equation (3), the cost is defined, where (0 < β < 1 - (discount factor)).

[0030]

[0031] As shown in Equation (4), the constraint is similarly defined

[0032]

[0033] As described above, the goal of the CMDP technique is to find a policy π that minimizes the total cost C(π) while ensuring that the total value of the constraint remains below a specified value. The resulting policy π can be used for optimizing the operation of a controlled application system. However, as Figure 1As described in the illustrated example, many controlled application systems involve a very large number of possible system states and possible control actions. This large number of variables results in the generation of highly complex models that cannot be solved within a reasonable time. Therefore, the embodiments described herein propose an automatic state space reduction technique that can be used to quickly and effectively reduce the size of a CMDP model by automatically selecting a subset of the system state variables that are most important for each CMDP model. In various embodiments, this is achieved using a combination of CRL (or CDRL) and CMDP techniques, as described in more detail with respect to Figure 2 Furthermore, after state space reduction, the CMDP model can be formulated as a linear programming (LP) problem, which can be solved quickly and effectively by a typical solver (such as the IBM Optimizer).

[0034] Figure 2 FIG. 200 is a process flow diagram of a method 200 for automatically reducing the dimension of a mathematical representation of a controlled application system that is subject to one or more constraint values. The method 200 is implemented by a control system that guides the operation of the controlled application system. In some embodiments, the controlled application system is a WWTP, such as the WWTP 100 described with respect to Figure 1 and the control system includes a processor and a computer-readable storage medium that includes program instructions for guiding the processor to perform the steps of the method 200. Additionally, in other embodiments, the controlled application system can be a healthcare system, an agricultural system, a water resource management system, a queuing system, an epidemic handling system, a robotic motion planning system, or any other suitable type of controlled application system.

[0035] The method 200 begins at block 202, where the control system receives data corresponding to control action variables and system state variables related to the controlled application system. In some embodiments, at least a portion of the data is calculated via sensors coupled to the controlled application system. Additionally, in some embodiments, the control system calculates at least a portion of the data based on a simulation of the controlled application system.

[0036] As described herein, the control action variables include a large number of possible actions that the control system can execute to control the operation of the controlled application system, and the system state variables include a large number of possible states of the controlled application system after the control system executes one or more actions. In various embodiments, optimizing the control actions of the control system includes solving a large-scale stochastic problem related to the relationship between the control action variables and the system state variables. Additionally, the large-scale stochastic problem takes into account a cost function related to the controlled application system and one or more constraint values for the controlled application system.

[0037] At block 204, a processor of the control system fits a Constrained Reinforcement Learning (CRL) model to the controlled application system based on data corresponding to control action variables and system state variables. In other words, the processor of the control system considers large-scale stochastic problems within a machine learning framework and outputs a CRL model that best describes the operation of the controlled application system, given one or more constraint values, control action variables, and system state variables that define the problem. In some embodiments, the CRL model is a Constrained Deep Reinforcement Learning (CDRL) model.

[0038] In various embodiments, the model fitting process includes performing standard reinforcement learning techniques known in the art, such as parameter tuning and optimization, to generate a policy π: S × U → [0, 1]. A policy is a mapping that provides the probability P (usually between 0 and 1) of taking a control action u when the environment of the controlled application system indicates a system state s: The policy can be stationary (i.e., time-independent) or non-stationary (i.e., time-dependent), as is known in the art. Additionally, based on the policy, the joint probability of a particular pair of system state variables and control action variables (i.e., a state-action pair) can be calculated as where is calculated based on data received by the control system.

[0039] In various embodiments, the policy of the CRL (or CDRL) model is implicit. Specifically, given a system state that is an input to the CRL (or CDRL) model, a distribution of control actions (or the most likely control action) is provided as an output. Since the total number of system states may be huge, or even infinite, the policy is not typically calculated explicitly. Instead, the CRL (and CDRL) process acts like an "oracle". However, according to the embodiments described herein, probabilities can be estimated using the data received at the input during the fitting process, thus providing an implicit policy.

[0040] One particular process that can be used to generate the policy of a CRL (or CDRL) model is called the Reward Constrained Policy Optimization (RCPO) process. This process is described in a paper by Tessler, Mankowitz, and Mannor titled "Reward Constrained Policy Optimization" (2018). As described herein, for a constrained optimization problem, the task is to find a policy that maximizes an objective function f(x) while satisfying the inequality constraint g(x) ≤ α. The RCPO algorithm achieves this by incorporating the constraint as a penalty signal into the reward function. The penalty signal guides the policy towards a solution that satisfies the constraint.

[0041] More specifically, for the RCPO algorithm, the Lagrangian relaxation technique is first used to formulate large-scale stochastic problems, which involves transforming a constrained optimization problem into an equivalent unconstrained optimization problem. A penalty term is added for infeasibility, making infeasible solutions suboptimal. Thus, given a CMDP-based model as shown in Equation (5), an unconstrained optimization problem is defined, where L is the Lagrangian and λ ≥ 0 is the Lagrange multiplier (penalty coefficient).

[0042]

[0043] The objective of Equation (5) is to find the saddle point (θ * (λ * ), λ * )(i.e., a feasible solution), where the feasible solution is a solution that satisfies .

[0044] Next, the gradients are estimated. Specifically, the data related to the control action variables and the system state variables are used to define the algorithms for the constrained optimization problem as shown in Equations (6) and (7).

[0045]

[0046]

[0047] In Equations (6) and (7), Г θ is the projection operator, which maintains the stability of the iteration θ k by projecting onto a compact convex set. Г λ projects λ onto the range . Derived from Equation (5) and where the log-likelihood trick shown in Equations (8) and (9) is used to derive

[0048]

[0049]

[0050] Furthermore, in Equations (6) and (7), η1(k) and η2(k) are the step sizes that ensure policy updates are performed on a time scale faster than the penalty coefficient λ. This leads to the assumption shown in Equation (10), which states that the iteration (θ n , λ n ) will converge to a fixed point (i.e., a local minimum), which is a feasible solution.

[0051]

[0052] and

[0053] The RCPO algorithm extends this process by using an actor-critic based approach, where the actor learns the policy π and the critic learns the value using temporal difference learning (i.e., via the recursive Bellman equation). More specifically, under the RCPO algorithm, an alternative guided penalty (called a discounted penalty) is used to train the actor and the critic. The discounted penalty is defined as shown in equation (11).

[0054]

[0055] In addition, as shown in equations (12) and (13), a reward function of penalty is defined.

[0056]

[0057]

[0058] The penalty value of equation (13) can be estimated using a TD learning critic. The RCPO algorithm is a three-timescale (constrained actor critic) process, where the actor and critic are updated after solving equation (13), and λ is updated after solving equation (6). The RCPO algorithm converges to a feasible solution under the following conditions: (1) represents the feasible solution set; (2) by Θ γ instruct The local minimum set of ; and (3) Assume Then the RCPO algorithm almost certainly converges to a fixed point (θ * (λ * ),v * ),v * (λ * ),λ * ), which is a feasible solution.

[0059] In various embodiments, another actor-critic process can be used to fit the CRL model to the controlled application system. The steps that the control system can follow to execute this process are summarized as follows: (1) Use the gradient ascent method to learn the Lagrange multipliers that fit the mathematical representation of the controlled application system, where the gradient descent method is used for the main method (e.g., the Actor-Criticar algorithm); and (2) Use the learned Lagrange multipliers to formulate the mathematical representation into a CRL model. This is due to the time separation property, i.e., the slow time scale and fast time scale of the gradient descent method and the gradient ascent method. In various embodiments, this process can be applied to more general constraints using CDRL techniques. In addition, this process is described in more detail in the article by Borkar titled "Actor-Criticar Algorithm for Constrained Markov Decision Processes - Systems & Control Letters" (2005).

[0060] In block 206, the processor automatically identifies a subset of system state variables by selecting the control action variables of interest and, for each control action variable of interest, identifying the system state variables that drive the CRL model to recommend the control action variables of interest. In various embodiments, this step involves reducing, for example, thousands of variables to less than 10 or less than 20 variables. In addition, the subset of the identified system state variables is described by one or more constraint values and cost-related variables associated with the constrained mathematical model.

[0061] In various embodiments, the automatic identification of a subset of system state variables in block 206 involves four steps. First, calculate the occupancy measure of the control action variable and system state variable pairs, i.e., the state-action pairs described with respect to block 204. As shown in equation (14), calculate the occupancy measure.

[0062]

[0063] This calculation is based on the state-action pair probabilities calculated during the model fitting process described with respect to block 204. The occupancy of each state-action pair represents the frequency of accessing that state-action pair, which is between 0 and 1. In various embodiments, the occupancy measure can also be calculated during the model fitting process described with respect to block 204. In addition, as shown in equation (15), formulate the occupancy measure with the discount factor β (for discounting future rewards) and an infinite horizon.

[0064]

[0065] Second, the control system selects multiple control actions of interest, where each control action of interest is represented by u iRepresentation. In some embodiments, the control action of interest is manually selected by an operator of the control system. For example, the user interface of the control system can present a list of control actions to the operator from which the operator can conveniently select, or the operator can manually type in the computer-readable identifier of the control action of interest. In other embodiments, the control action of interest is automatically selected by the control system in response to an automatically generated indication of a trigger event pre-programmed into the control system.

[0066] Third, the control system automatically determines specific state-action pairs that include each control action of interest and have an occupancy metric that meets a predetermined threshold. For example, for each control action of interest, D state-action pairs (D≥1) with the highest occupancy metric can be determined, i.e., These state-action pairs may include the system state variables that drive the CRL model to recommend a specific control action u of interest i Fourth, the state-action pairs are decomposed and analyzed to identify a specific subset of the system state variables that have the most substantial impact on driving the CRL model to recommend each control action of interest.

[0067] At block 208, the processor automatically performs state space reduction of the CRL (or CDRL) model using the subset of system state variables. This involves incorporating the subset of system state variables identified at block 206 into the CRL model. In various embodiments, performing state space reduction of the CRL model in this manner transforms a complex large-scale stochastic problem into a more concise stochastic problem that can be solved quickly and efficiently by the control system.

[0068] At block 210, the processor estimates a transition probability matrix of a Constrained Markov Decision Process (CMDP) model of a controlled application system subject to one or more constraint values after state space dimensionality reduction. In various embodiments, this transition probability estimation process can be broken down into three steps. First, the processor analyzes the state space and the action space with as much coverage as possible. In other words, the processor analyzes as many transitions as possible between different system states in response to different control actions. To this end, the processor can run simulations of possible transitions using a set of control action variables of interest and reduced system state variables, focusing on the least used control actions for each system state. Second, the processor estimates the transition probability by dividing the number of transitions of each change of the system state for a particular control action by the total number of system state variables and control action variables within a subset of variables. Third, the processor performs interpolation to generate the transition probability matrix. This can be achieved by using adjacent system states for action-related system state variables and / or action-unrelated system state variables, and / or by performing irreducibility assurance methods. The resulting transition probability matrix represents a concise model of the operation of the controlled application system (subject to one or more constraint values) based on the subset of variables.

[0069] Now referring to the mathematical explanation of the transition probability estimation process, let N a represent the set of control action variables, where the number of values for each control action variable is and the total number of control action variables is Let N s represent the set of reduced system state variables, where the number of values for each system state variable is and the total number of system state variables in the reduced set is Let represent the number of visits to system state j after control action i, and let represent the number of times of transitioning from system state j to system state k after executing control action i. Using these definitions, Equation (16) shows the basic estimation of the transition probability matrix.

[0070]

[0071] In various embodiments, the system state variables are divided into two groups: controllable and uncontrollable. For embodiments where the controlled application system is a WWTP, the uncontrollable system state variables can include, for example, influent flow rate, influent chemical load, and time period type according to electricity cost. Control actions and controllable system state variables do not affect the transitions between uncontrollable system state variables. The controllable system state variables describe the internal or effluent characteristics of the controlled application system and may be affected by any type of past actions and system state variables. Let I, 0 ≤ I ≤ N sDenote the number of uncontrollable system state variables, where the system state variables with indices 1, …, I are uncontrollable. X N and X C are the state spaces corresponding to the uncontrollable and controllable system state variables respectively. The full state space is the Cartesian product of these two spaces, i.e., X = X N × X C . Their dimensions are respectively and are The set of system states from the state space corresponding to state j from the uncontrollable state space X N is defined as NS(j), j ∈ X N . In other words, NS(j) = {j} ∈ X C . The set of system states from the state space corresponding to state j from the controllable state space X c is defined as CS(j), j ∈ X C . In other words, CS(j) = X N ∈ {j}. Further, if k ∈ X, then H(k) ∈ X N represents the uncontrollable part of state k, and E(k) ∈ X C represents the controllable part of state k. In other words, state k is the concatenation of states H(k) and E(k).

[0072] For action-driven coverage, assume the current system state is j, j ∈ X. Select the control operation that has been selected the least so far, i.e., i * = argmin V j i . If there are several such control actions, randomly select among them according to a uniform distribution. Similarly, for state-driven coverage, assume the current system state is j, j ∈ X. Available denotes the total number of times the system state k has been visited up to the current time. Further,[[]] where,[[]] is the current estimate of the transition probability. Select the control action i * = argmax W i . If there are several such control actions, randomly select among them according to a uniform distribution. The control action selected by state-driven coverage means visiting system states with a high , or in other words, system states that are rarely visited.

[0073] In various embodiments, use action-driven coverage to repeat this process over a relatively long period of time to obtain a reasonable initial estimate of the transition probability. Then, alternately repeat this process between action-driven coverage and state-driven coverage.

[0074] For the interpolation of the transition probability matrix, the first process involves interpolation based on adjacent system states. An adjacent system state j (j ∈ X) of a system state is the following system state such that the values of all system state variables except one are identical to those of the system state of j and the difference from one system state variable is equal to an adjacent interval. The set of adjacent system states of state j and the system state j itself are represented by L j is denoted. A subset of L j is denoted by and includes system states in which control action i is executed at least once. Further, denotes the size of this set.

[0075] Using a parameter M ≥ 0, interpolation is performed for control actions and system states with V i j ≤ M. Usually, a value of M = 10 is used, and the interpolation probability is calculated as shown in Equation (17).

[0076]

[0077] If the number of times of accessing adjacent system states replaces V in the threshold criteria for system state j and control action i j i , this process can be iteratively executed.

[0078] The second process of transition probability matrix interpolation involves ensuring the irreducibility of the transition probability matrix. This is important because if the transition probability matrix is not irreducible for a certain control action i, then the optimal solution may imply that some system states are not visited under that control action, which is incorrect for many controlled application systems. Therefore, an irreducibility guarantee method is used to solve this problem. The irreducibility guarantee method uses a parameter ε, 0 ≤ ε ≤ 1, where usually ε = 0.01 is assumed. Given s0 as the initial system state, the irreducibility guarantee method includes the following steps. First, as shown in Equations (18) and (19), the values for all control actions i ∈ U, system states k and j.

[0079]

[0080]

[0081] Then, a simple algorithm is used to normalize the transition probability matrix such that the sum of all rows is equal to 1.

[0082] In some embodiments, a third process may be used to transform the interpolation of the probability matrix. Specifically, the transition probabilities obtained via the first and second processes may include significant distortions with respect to the uncontrollable system state variables. Thus, interpolation of the uncontrollable system state variables may be performed to correct the distortions. This process is further described in the paper "Optimal Operation of Wastewater Treatment Plants: A CMDP-Based Decomposition Approach" by Zadorojniy, Shwartz, Wasserkrug, and Zeltyn (2016).

[0083] At block 212, the processor formulates the CMDP model as a linear programming (LP) problem using a transition probability matrix, one or more cost objectives, and one or more constraint-related costs. In some embodiments, the LP problem may be formulated using the Python / Optimization Programming Language (OPL) and then solved, for example, using an IBM optimizer. The mathematical interpretation of block 212 is shown in equations (20)-(25), where the non-zero entries correspond to the initial system state.

[0084]

[0085]

[0086]

[0087]

[0088] A |S|×|S||U| =[I - β×P(u0),…,I - β×P(u |U|-1 )] (24)

[0089] b = (1 - β, 0,…, 0) T (25)

[0090] Figure 2The block diagram is not intended to represent that the blocks 202-212 of the method 200 will be executed in any particular order, or that all of the blocks 202-212 of the method 200 will be included in each instance. Additionally, depending on the details of the specific implementation, any number of additional blocks may be included within the method 200. For example, in some embodiments, the method 200 includes analyzing an LP problem to determine one or more control actions that will optimize the operation of a controlled application system that is subject to one or more constraint values, and executing the one or more control actions via a control system. Additionally, in some embodiments, the method 200 is executed multiple times, and the resulting LP problems are used to analyze multiple different prediction scenarios related to the operation of the controlled application system. For example, the method 200 can be used to provide predictions (or "what-if" analyses) regarding the operation of the controlled application system under various different constraint values and / or various different types of constraints.

[0091] The description of the various embodiments of the invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or improvements made to the technology found in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.

Claims

1. A method for automatically reducing the dimension of a mathematical representation of a controlled application system, the controlled application system being subject to one or more constraint values, the method comprising: Receiving, at a control system that guides the operation of the controlled application system, data corresponding to control action variables and system state variables associated with the controlled application system; Via a processor of the control system, fitting a Constrained Reinforcement Learning (CRL) model to the controlled application system based on the data corresponding to the control action variables and the system state variables; Via the processor, automatically identifying a subset of the system state variables by selecting control action variables of interest and, for each control action variable of interest, identifying the system state variables that drive the CRL model to recommend the control action variable of interest; Via the processor, automatically performing state space reduction of the CRL model using the subset of the system state variables; Via the processor, estimating a transition probability matrix of a Constrained Markov Decision Process (CMDP) model of the controlled application system that obeys the one or more constraint values after state space reduction; And Via the processor, formulating the CMDP model as a Linear Programming (LP) problem using the transition probability matrix, one or more cost objectives, and one or more constraint-related costs.

2. The method according to claim 1, wherein, The CRL model includes a Constrained Deep Reinforcement Learning (CDRL) model.

3. The method according to claim 1, comprising: Via the processor, solving the LP problem to determine one or more control actions that will optimize the operation of the controlled application system that obeys the one or more constraint values; And Via the control system, executing the one or more control actions.

4. The method according to claim 1, comprising at least one of the following: Receiving at least a portion of the data from sensors coupled to the controlled application system; and Via the processor, calculating at least a portion of the data based on a simulation of the controlled application system.

5. The method according to claim 1, wherein Fitting the CRL model to the controlled application system includes performing a Reward Constraint Policy Optimization (RCPO) process.

6. The method according to claim 1, wherein Fitting the CRL model to the controlled application system includes: Using the gradient ascent method to learn Lagrange multipliers that fit the mathematical representation of the controlled application system; and Using the learned Lagrange multipliers to formulate the mathematical representation as a CRL model.

7. The method according to claim 1, wherein Automatically identifying the subset of the system state variables includes: Calculating an occupancy measure associated with pairs of the control action variables and the system state variables; Automatically determining a specific pair of the control action variables and the system state variables, the specific pair including each control action variable of interest and having an occupancy measure that meets a predefined threshold; and Analyzing the specific pair of the control action variables and the system state variables to determine the subset of the system state variables that has the most significant impact on driving the CRL model to recommend each control action variable of interest.

8. The method according to claim 1, wherein Estimating the transition probability matrix of the CMDP model includes: Analyze the state space and action space related to a subset of the control action variables of interest and the system state variables; Estimate the transition probability by dividing the number of transitions of each change in the system state variables of a specific control action variable by the total number of control action variables and system state variables within the subset of the control action variables of interest and the system state variables; and Perform interpolation to generate the transition probability matrix by: Using adjacent system states of at least one of the action-related system state variables or action-unrelated system state variables; and / or Performing an irreducibility guarantee method.

9. The method according to claim 1, including solving the LP problem via the processor to analyze a plurality of different prediction scenarios related to the operation of the controlled application system.

10. A control system that guides the operation of a controlled application system, the controlled application system subject to one or more constraint values, the control system comprising: An interface for receiving data corresponding to control action variables and system state variables related to the controlled application system; A processor; And A computer-readable storage medium storing program instructions that direct the processor to: Fit a Constrained Reinforcement Learning (CRL) model to the controlled application system based on the data corresponding to the control action variables and the system state variables; Automatically identify a subset of the system state variables by selecting control action variables of interest and, for each control action variable of interest, identifying the system state variables that drive the CRL model to recommend the control action variable of interest; Automatically perform state space reduction of the CRL model using the subset of the system state variables; Estimate the transition probability matrix of a Constrained Markov Decision Process (CMDP) model of the controlled application system that complies with the one or more constraint values after state space reduction; And Formulate the CMDP model as a Linear Programming (LP) problem using the transition probability matrix, one or more cost objectives, and one or more constraint-related costs.

11. The control system according to claim 10, wherein, The computer-readable storage medium stores program instructions that direct the processor to: Solve the LP problem to determine one or more control actions that will optimize the operation of the controlled application system that complies with the one or more constraint values; And Execute the one or more control actions via the control system.

12. The control system according to claim 10, wherein, The CRL model includes a Constrained Deep Reinforcement Learning (CDRL) model.

13. The control system according to claim 10, wherein, The controlled application system includes a Wastewater Treatment Plant (WWTP), and wherein the one or more constraint values include at least one of a limit on the total nitrogen concentration or a limit on the total phosphorus concentration of the effluent produced by the WWTP.

14. The control system according to claim 10, wherein, The computer-readable storage medium stores program instructions that direct the processor to fit the CRL model to the controlled application system by performing a Reward-Constrained Policy Optimization (RCPO) process.

15. The control system according to claim 10, wherein, The computer-readable storage medium stores program instructions that direct the processor to fit the CRL model to the controlled application system by: learning Lagrange multipliers that fit a mathematical representation of the controlled application system using gradient ascent; and formulating the mathematical representation as a CRL model using the learned Lagrange multipliers.

16. The control system according to claim 10, wherein The computer-readable storage medium stores program instructions that direct the processor to identify the subset of the system state variables by: computing an occupancy measure associated with pairs of the control action variables and the system state variables; automatically determining a specific pair of the control action variables and the system state variables that includes each control action variable of interest and has an occupancy measure that meets a predefined threshold; and analyzing the specific pair of the control action variables and the system state variables to determine the subset of the system state variables that has the most significant impact on driving the CRL model to recommend each control action variable of interest.

17. The control system according to claim 10, wherein, The computer-readable storage medium stores program instructions that direct the processor to estimate the transition probability matrix of the CMDP model by: analyzing the state space and action space associated with the subset of the control action variables of interest and the system state variables; estimating a transition probability by dividing the number of transitions of each change in the system state variables for a specific control action variable by the total number of control action variables and system state variables within the subset of the control action variables of interest and the system state variables; and performing interpolation to generate the transition probability matrix by: using adjacent system states of at least one of the action-related system state variables or action-unrelated system state variables; and / or performing an irreducibility guarantee method.

18. The control system according to claim 10, wherein, The computer-readable storage medium stores program instructions that direct the processor to analyze a plurality of different prediction scenarios related to the operation of the controlled application system by solving the LP problem.

19. The control system according to claim 10, wherein, The computer-readable storage medium stores program instructions that direct the processor to calculate at least a portion of the data based on a simulation of the controlled application system.

20. A computer program product for automatically reducing the dimension of a mathematical representation of a controlled application system that complies with one or more constraint values, the computer program product including a computer-readable storage medium having program instructions embodied therewith, wherein, The computer-readable storage medium itself is not a transient signal, and wherein the program instructions are executable by the processor to cause the processor to: fit a constrained reinforcement learning CRL model to the controlled application system, wherein the CRL model is based on data corresponding to control action variables and system state variables; automatically identify the subset of the system state variables by selecting control action variables of interest and, for each control action variable of interest, identifying the system state variables that drive the CRL model to recommend the control action variable of interest; automatically perform state space reduction of the CRL model using the subset of the system state variables; estimate the transition probability matrix of the constrained Markov decision process CMDP model of the controlled application system that obeys the one or more constraint values after state space reduction; and Formulate the CMDP model as a linear programming (LP) problem using the transition probability matrix, one or more cost objectives, and one or more constraint-related costs.

Citation Information

Patent Citations

  • Method and system to estimate variables in an integrated gasification combined cycle (IGCC) plant

    CN103430121A

  • Demand response capability assessment method of intelligent power grid and calculating equipment

    CN108470233A