Simulation of Industrial Equipment for Control

The framework addresses the inadequacy of existing control policy training by incorporating non-deterministic elements into simulations, enabling robust control policies for industrial equipment that account for real-world imperfections.

JP2025521355AActive Publication Date: 2025-07-08ジーディーエム·ホールディング·エルエルシー
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024575577
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2023-06-23
Publication Date
2025-07-08
Estimated Expiration
2043-06-23

AI Technical Summary

Technical Problem

Existing frameworks for training control policies for industrial equipment are inadequate due to the need for robustness against real-world imperfections such as noisy sensors and changing environmental conditions, making deterministic simulators ineffective.

Method used

A framework that introduces non-deterministic elements into the simulation environment by adding noise and changing configuration parameters, allowing for robust control policy training and evaluation without modifying the simulator or reinforcement learning agent.

Benefits of technology

Enables the training and evaluation of control policies that are resilient to real-world imperfections, ensuring reliable operation of industrial equipment by simulating non-deterministic conditions effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521355000001_ABST
    Figure 2025521355000001_ABST
Patent Text Reader

Abstract

A method, system, and apparatus comprising a computer program encoded on a computer storage medium for simulating industrial equipment for control. One of the methods includes, at each of a plurality of time steps during a task episode, receiving, from a computer simulator of the industrial equipment, measurements representing the current state of the equipment; generating observations from the measurements; providing the observations as input to a control policy for controlling the equipment; receiving, as output, an action for controlling one or more set points of the equipment; generating, from the action, one or more control inputs for one or more set points of the equipment; and providing, as input to the simulator, (i) the control inputs and (ii) the current values of one or more configuration parameters of the simulator to cause the simulator to generate, as output, new measurements representing a new state of the equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Application No. 63 / 354,930, filed on June 23, 2022, which is hereby incorporated by reference in its entirety.

[0002] This specification relates to controlling industrial equipment using a machine learning model.

Background Art

[0003] A neural network is a machine learning model that uses one or more layers of non - linear units to predict an output for a given input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer of the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of each set of parameters.

Summary of the Invention

Means for Solving the Problems

[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations that simulates the operation of industrial equipment to enable a machine learning model to be trained to control the equipment.

[0005] The subject matter described in this specification may be implemented in certain embodiments so as to realize one or more of the following advantages.

[0006] This specification describes techniques for training, evaluating, or both, a control policy for industrial equipment using computer simulation of the industrial equipment. When the control policy is trained and / or evaluated in the simulation, the control policy may be deployed and used to control the (real-world) industrial equipment.

[0007] More specifically, computer simulations of industrial equipment are deterministic, and given an initial configuration, the state of the industrial equipment, and control inputs, the computer simulation will always update the state of the industrial equipment in the same way. This can make existing frameworks for training control policies in simulation an inappropriate choice for training control policies for industrial equipment. This is because the control of industrial equipment requires a control policy that is robust to several numbers of real-world imperfections where a given control input may have different effects on the state of the equipment. For example, the sensors of the equipment may be noisy or may not function well, external situations in the real-world equipment environment may change rapidly, the setpoint may not function well, and so on. This specification describes a framework for training a control policy to be robust to such imperfections, or for evaluating a control policy to determine whether the policy is robust to such imperfections, or for both, without the need to modify the simulator or RL agent being trained. That is, this specification describes a framework that enables a deterministic simulator of industrial equipment to be effectively used to simulate real-world non-determinism. Specifically, by using an environment subsystem to interface between the RL agent and the simulator, the system can incorporate various aspects of non-determinism into the interaction, for example, by introducing noise into the control input, or the measurement, or both, or by changing the configuration parameters of the simulator between task episodes and within task episodes. Further, the same framework can be employed for multiple different simulators of different equipment and for multiple different tasks, to introduce these different degrees of non-determinism. Specifically, this framework allows for extremely extended configurability, where each of the task, simulator, scenario, and noise is an independent axis of configurability, enabling the user to combine them.

[0008] In one example described in this specification, a method performed by one or more computers includes, at each of a plurality of time steps during a task episode, receiving, from a computer simulator of industrial equipment, measurements representing the current state of the industrial equipment; generating observations from the measurements; providing the observations as inputs to a control policy for controlling the industrial equipment; receiving, as outputs from the control policy, actions for controlling one or more setpoints of the industrial equipment; generating, from the actions, one or more control inputs for one or more setpoints of the industrial equipment; and providing, as inputs to the computer simulator, (i) the one or more control inputs and (ii) the current values of one or more configuration parameters of the computer simulator, to cause the computer simulator to generate, as an output, new measurements representing a new state of the industrial equipment for a subsequent time step.

[0009] The configuration parameters may specify additional information (in addition to the control inputs) used by the computer simulator to represent the state of the industrial equipment. Some exemplary configuration parameters are described below.

[0010] Generating observations from the measurements may include adding noise to the measurements. Generating, from the actions, one or more control inputs for one or more setpoints of the industrial equipment may include adding noise to one or more control inputs defined by the observations. The method may further include identifying a scenario for the task episode. The scenario may specify, for each of the plurality of time steps, respective changes applied to one or more of the configuration parameters, or one or more of the control inputs, or one or more of the measurements. The scenario may specify changes applied to one or more of the configuration parameters.

[0011] The method may further include sampling the configuration for each task episode that specifies an initial value for each of the configuration parameters. The method may further include, for each of the time steps and for each of one or more configuration parameters, applying a change specified by the scenario of the time step to the initial value of the configuration parameter to generate a current value of the configuration parameter. The scenario may specify a change to be applied to one or more of the measured values. Generating the observed value from the measured value may further include, for each of the one or more measured values, applying a change specified by the scenario of the time step to the measured value. The scenario may specify a change to be applied to one or more of the control inputs. Generating one or more control inputs from the action may further include, for each of the one or more control inputs, applying a change specified by the scenario of the time step to the control input.

[0012] The computer simulator may be a deterministic simulator of the dynamics of industrial equipment. The method may further include training a control policy based at least on the task episode and, after training, deploying the control policy to control the industrial equipment. The method may further include evaluating a control policy based at least on the task episode and, after evaluation, deploying the control policy to control the industrial equipment.

[0013] The method may further include, after deployment of the control policy, receiving a measurement of the current state of the industrial equipment from the industrial equipment, generating a second observed value from the measurement of the current state of the industrial equipment, providing the second observed value as an input to the control policy for controlling the industrial equipment, receiving, as an output from the control policy, a second action for controlling one or more setpoints of the industrial equipment, generating, from the second action, one or more second control inputs for the one or more setpoints of the industrial equipment, and controlling the one or more setpoints of the industrial equipment based on the one or more second control inputs.

[0014] The method may further include controlling a second industrial facility using a second control policy to generate a dataset, and a computer simulator of the industrial facility is configured to generate measurement values representing a current new state of the industrial facility based on the dataset. The second industrial facility may be the same industrial facility as the industrial facility.

[0015] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0017] Like reference numerals and names in the various drawings indicate like elements.

[0018] This specification describes a system implemented as a computer program on one or more computers in one or more locations that simulates the operation of an industrial facility while the industrial facility is being controlled by a control policy.

[0019] Specifically, the control policy receives, as input, observed values that describe the characteristics of the state of the industrial equipment, and in response, generates an action that specifies each setting for one or more set points of the industrial equipment. Each set point is a different controllable element of the industrial equipment. That is, the control policy controls the equipment by repeatedly updating the settings for one or more set points of the industrial equipment.

[0020] For example, the control policy may be implemented as a neural network or other machine learning model, and the system may be used to train the control policy in simulation before deploying the control policy to control real-world industrial equipment. For example, the control policy may be trained through reinforcement learning to maximize the received reward representing the performance of the policy on a specified task.

[0021] As another example, the system may be controlled using one control policy, such as a pre-trained neural network or a fixed or heuristic-based control policy, to generate a dataset. This dataset can be further used to train another control policy, for example, through offline reinforcement learning, without the need to use other control policies to control the industrial equipment. Alternatively or additionally, the dataset can be used to evaluate the performance of another control policy, for example, to determine whether the control policy is suitable for deployment to control real-world industrial equipment.

[0022] Generally, industrial equipment includes one or more items of electronic devices, mechanical devices, or both that are controllable by a control policy. The control policy operates to control the industrial equipment to perform a specified task.

[0023] In some implementations, the facility is a service facility, or any service facility, that includes a plurality of items of electronic equipment, such as a server farm or data center, e.g., a telecommunications data center, or a computer data center for storing or processing data. The service facility may include an auxiliary control device that controls the operating environment of the items of the device, such as an environmental control device for temperature control, e.g., a cooling device, or air flow control or air conditioning device. This device can include, for example, an air-cooled chiller, a water-cooled chiller, or both. The task may include a task that controls the use of resources, e.g., minimizes, such as the power consumption or water consumption while the facility is operating. In some cases, the optimization may be subject to one or more constraints.

[0024] Generally, the action may be any action that affects the observed state of the environment, e.g., an action configured to adjust any of the detected parameters described below. These may include actions that control the items of the device or auxiliary control device, or impose operating conditions on them, e.g., actions that cause a change in settings to adjust, control, or switch on or off the operation of the items of the device or auxiliary control device. As a specific example, the action may include an action that controls one or more chillers operating within the facility.

[0025] Generally, the observed value of the state of the environment may include any electronic signal that represents the function of the facility or the devices within the facility. For example, the representation of the state of the environment may be an observed value created by a sensor detecting the state of the physical environment of the facility, or derived from an observed value created by a sensor detecting the state of one or more of the items of the device or one or more of the items of the auxiliary control device. These include sensors configured to detect electrical conditions such as current, voltage, power or energy, the temperature of the facility, the flow, temperature, or pressure of a fluid within the facility or the cooling system of the facility, or the physical facility configuration such as whether a vent is open or not.

[0026] Rewards or benefits may be related to metrics of task performance. For example, in the case of a task that controls the use of a resource, such as a task that controls, e.g., minimizes the use of a resource, such as a task that controls the use of power or water, the metric may include any metric of resource use.

[0027] In some implementations, the facility is a power generation facility, e.g., a renewable power generation facility such as a solar power plant or a wind power plant. The task may include a control task that controls the power generated by the facility, e.g., to meet demand, or to reduce the risk of mismatch between elements of the grid, or to maximize the power generated by the facility, e.g., a control task that controls the supply of power to the power distribution network. The action may include an action that controls the electrical or mechanical configuration of one or more power generation elements, such as a wind turbine or the configuration of one or more solar panels or solar mirrors, or the electrical or mechanical configuration of a rotating electric power generator, e.g., for controlling the electrical or mechanical configuration of a power generator. The mechanical control action may include an action that controls the conversion of energy input to electrical energy output, e.g., the efficiency of the conversion or the degree of coupling of the energy input to the electrical energy output. The electrical control action may include an action that controls one or more of the voltage, current, frequency, or phase of the generated power.

[0028] Rewards or benefits may be related to metrics of task performance. For example, in the case of a task that controls the supply of power to the power distribution network, the metric may be related to the amount of power transmitted, or the amount of electrical mismatch between the power generation facility and the grid, such as voltage, current, frequency, or phase mismatch, or the amount of power or energy loss in the power generation facility. In the case of a task that maximizes the supply of power to the power distribution network, the metric may be related to the amount of power or energy transmitted to the grid, or the amount of power or energy loss in the power generation facility.

[0029] Generally, the observed value of the environmental state may include any electronic signal representing the electrical or mechanical function of a power generation device within a power generation facility. For example, the representation of the environmental state may be derived from observed values created by sensors detecting the physical or electrical state of a device within a power generation facility that is generating power, or the physical environment of such a device, or the situation of an auxiliary device supporting the power generation device. Such sensors may include sensors configured to detect observed values of the electrical situation of a device, such as current, voltage, power, or energy, the temperature or cooling of the physical environment, the flow of a fluid, or the physical configuration of the device, and the electrical situation of the grid, such as from local or remote sensors. The observed value of the environmental state may also include one or more predictions regarding the future situation of the operation of the power generation device, such as predictions of future wind levels or solar irradiance, or predictions of the future electrical situation of the grid.

[0030] FIG. 1 is a diagram of an exemplary simulation system 100. The simulation system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations that can implement the systems, components, and techniques described below.

[0031] System 100 is described as being used to control a heating, ventilating, and air conditioning (HVAC) system 120 of an industrial facility 110 to be simulated.

[0032] However, more generally, system 100 can be used to control any aspect of the operation of any type of industrial facility 110, such as one of the aspects described above.

[0033] The simulated industrial facility 110 (also referred to as simulator 110) is a computer simulation of a real-world industrial facility, i.e., it models the state and dynamics of a real-world industrial facility as observed in various contexts using one or more computer programs. That is, simulator 110 maintains the state of a real-world industrial facility, such as the current readings of sensors within the facility and optionally additional information, and takes as inputs (i) the current value of a configuration parameter that specifies the configuration of the simulator, and (ii) control inputs for one or more set points of the industrial facility, and provides as output measured values, i.e., updated readings of sensors within the facility that reflect the updated state of the facility as a result of the control inputs, via one or more software programs.

[0034] System 100 can utilize any suitable computer simulator. For example, a user of the system can provide system 100 with access to a computer simulator of a real-world facility of interest to the user, for example, by enabling system 100 to access the simulator via an API or other interface, or by enabling system 100 to execute the simulator.

[0035] In general, the problem of controlling an industrial facility to perform a specified task can be formulated as a constrained multi-objective optimization.

[0036] In the HVAC example, controller 130 controls several set points that adjust the temperature exchange characteristics of HVAC system 110 in order to perform a task, such as attempting to keep the facility temperature at a constant level. In the HVAC example, the set points can include activating and deactivating selected chillers and optionally setting the chiller leaving temperature.

[0037] The HVAC component draws power from the grid, and thus the next goal of the controller 130 may be to reduce power consumption. Thus, the overall task performed by the controller 130 can be assembled as minimizing the power consumption by the HVAC system 110 while satisfying one or more constraints regarding the facility temperature.

[0038] If the controller fails in its task, the controller risks overheating the facility, which can lead to tragic consequences. For example, a failure of computer components can cause data loss or downtime of electrical or mechanical components essential for the operation of the facility. To prevent this from happening, the manufacturer of the controller 130 introduces a set of fail-safe constraints to prevent such events from occurring. Violating the constraints not only impairs the reliability of the controller but also usually disconnects the controller from the facility and it can no longer optimize power consumption.

[0039] The system 100 can be used to provide a set of simulated scenarios (e.g., a control policy implemented as a machine learning model) that can be used to safely and efficiently train and evaluate the controller. That is, the system 100 can be used to train, evaluate, or both, a control policy that controls one or more of the set points specified by the controller 130 for the simulator 110.

[0040] More specifically, during operation of the system 100, the control policy 150 (e.g., a reinforcement learning agent) uses the simulator 110 as a ground truth model of the facility dynamics to perform tasks in a closed-loop control system.

[0041] The system 100 uses the simulator 110 to evaluate the effect of the action proposed by the policy 150 with respect to the current state of the simulation.

[0042] The simulator 110 returns results in the form of measurement values, which are a subset of the simulation state. The measurement values can include current readings from any of the various sensors of the industrial equipment.

[0043] The system 100 processes the measurement values into observation values 160, which are provided as inputs to the control policy 150.

[0044] However, HVAC simulation is deterministic, i.e., performing a given action in a given simulated state will always result in the same updated state. However, the control of real-world HVAC systems needs to account for any of the various non-deterministic factors that may be encountered during operation and that can change how an action affects the state of the equipment. Examples of these non-deterministic factors (also called "imperfections") will be described in more detail below with reference to FIGS. 2 and 3.

[0045] To introduce imperfections, the system 100 can introduce noise into one or more of various aspects of the control pipeline, such as control inputs, simulation configurations, and observation values. This will be described in more detail below with reference to FIGS. 2 and 3.

[0046] FIG. 2 shows a more detailed view of the simulation system 100.

[0047] As shown in FIG. 2, the simulation system 100 includes a simulator 110. In some implementations, the system 100 can also include a simulator data storage 210 that stores specifications for multiple different simulators so that, for example, an appropriate simulator can be selected for a given task to control a given real-world piece of equipment.

[0048] During operation, the simulation system 100 represents its interaction with the simulator 110 as an interaction with the environment subsystem 220.

[0049] The environment subsystem 220 is implemented as one or more computer programs and controls the interaction with the simulator 110 by the RL agent 230. The RL agent 230 can include a control policy and associated components for training the control policy through reinforcement learning based on the interaction of the control policy with the simulator 110.

[0050] As can be seen from FIG. 2, the RL agent 230 receives the observation value as input and provides, as output, an action 234 for controlling one or more setpoints of the simulated facility. The input observation value includes the environmental observation value 232 and, optionally, a "task observation value" 272 that includes additional information specific to the task being performed (e.g., generated from the environmental observation value 232 according to some task parameters).

[0051] The environment subsystem 220 converts the action 234 into a control input 236 and provides the control input 236 to the simulator 110. For example, converting the action 234 can include converting a high-level action (e.g., an indicator that a cooler should be disabled) into an instruction or other command (e.g., a machine-readable instruction to disable the cooler) that can be executed within the facility to perform the high-level action.

[0052] The environment subsystem 220 also provides the value of the configuration parameter 262 as input to the simulator 110.

[0053] The configuration parameter 262 specifies additional information (in addition to the control input 236) required by the simulator 110 to adequately represent the state of the real-world industrial facility. That is, the configuration parameter is a parameter required to initialize the simulator, i.e., to adequately represent the state of the real-world facility.

[0054] As an example, the configuration parameter 262 can specify the value of a setpoint that is not controlled by the RL agent 250 but is required to be specified by the controller. For example, the setpoint can include activating and deactivating a selected chiller and configuring the chiller outlet temperature. When the RL agent 230 simply controls the activation and deactivation of the chiller, the configuration parameter specifies the chiller outlet temperature of the chiller.

[0055] The configuration parameter 262 also specifies the properties of the external environment of the real-world industrial facility. For example, the configuration parameter 262 can specify the temperature of the external environment, the humidity of the external environment, the precipitation of the external environment, and the like.

[0056] For example, before starting a task episode, the environment subsystem 220 can sample from the configuration storage 264 a configuration that specifies an initial value for each of the configuration parameters, for example, a configuration that models the real-world configuration of an actual industrial facility. In some cases, as described below, the subsystem 220 can change the initial value during the course of the task episode, and in other cases, the subsystem 220 can maintain the initial value throughout the task episode.

[0057] When the simulator 110 is in the configuration specified by the configuration parameter 238, the simulator 110 returns a measurement value 238 that reflects the updated state of the simulator 110 as a result of the application of the control input 236.

[0058] The environment subsystem 220 converts the measured value 238 into the observed value 232, and the observed value 232 is provided as an input to the RL agent 230. For example, the user can provide the system with the specifications of the input received by the RL agent 230, i.e., which sensor measurements are provided as inputs, the expected range of the sensor measurements, the numerical format of the sensor measurements, etc. The subsystem 220 can then standardize the measured value 238 so that it conforms to the specifications of the observed value 232 provided by the user.

[0059] As described above, the RL agent 230 controls the simulator 110 to execute the specified task, e.g., to optimize one or more metrics of performance that are subject to one or more constraints. The constraints can include constraints on measurements, e.g., the temperature not exceeding a threshold, constraints on actions, e.g., a given cooler not being enabled for a time window exceeding a certain duration, or both.

[0060] To determine whether any of the constraints are violated by a given action or measurement, the system 100 includes a constraint evaluator 260 that maintains data specifying the current set of constraints of the task being performed by the RL agent 230. Specifically, each configuration in the storage 264 is associated with a set of constraints for a given task.

[0061] The evaluator 260 receives an input that includes the current action, the current set of measurements, or both, and determines whether any of the current set of constraints specified by the configuration of the current task episode are violated. The evaluator 260 can then provide the environment subsystem 220, which can provide this information as part of the corresponding observed value to the RL agent 230, with data identifying whether any constraint violations have occurred.

[0062] The constraints for a given task can include soft constraints, hard constraints, or both.

[0063] Soft constraints are constraints that can be violated and simply have an adverse effect on the evaluation of the performance of the RL agent 230.

[0064] Hard constraints are constraints that must not be violated, that is, a violation of the constraint will result in the controller being disconnected from the facility. When the evaluator 260 determines that a hard constraint has been violated, the environment subsystem 220 terminates the current episode of control, that is, provides the RL agent 230 with an indication that a hard constraint has been violated and that the RL agent 230 can no longer continue this instance of the simulation.

[0065] To train the RL agent 230 or evaluate the performance of an already trained RL agent 230, the RL agent requests a training signal. In reinforcement learning, this is represented in the form of a set of rewards 272, which are numerical values generated by the task subsystem 270 based on appropriate information such as the results of constraint evaluation, measurements, and control inputs. The mapping from this information to one or more numerical values representing the rewards 272 can be specified by the user of the system 100.

[0066] The RL agent 230 can use the rewards 272, environmental observations 232, and actions 234 to train a control policy using appropriate reinforcement learning techniques, such as on-policy or off-policy reinforcement learning algorithms.

[0067] Alternatively, as described above, the RL agent 230 can store the rewards 272, environmental observations 232, and actions 234 for use when training another policy through offline reinforcement learning or evaluating another policy as described above.

[0068] Once a given policy is trained, the given policy can be used to control real-world facilities simulated by simulator 110.

[0069] Rather than directly providing action 234 as control input 236 and directly providing observation value 238 that results in a deterministic control loop, environment subsystem 220 uses any of a variety of components to introduce imperfection into the control process.

[0070] Specifically, environment subsystem 220 can utilize one or more of scenario 226 or noise generator 290.

[0071] As a specific example, environment subsystem 220 can use the noise generated by noise generator 290 to add noise to the control input as part of converting an action to the control input, to the observation value as part of converting a measurement to the observation value, or to both. The noise is added to simulate sensor / pipeline imperfections and effectively diversify the distributions of the simulated states (control noise) and observation values (observation noise). The parameters of noise generator 290, such as the noise distribution from which the noise is sampled, or the time and location at which the noise is applied, or both parameters, can be specified by the user or sampled by system 100 from a set of possible parameters.

[0072] Scenario 226 models a real-world scenario of interest, and the scenario 226 used for a given task episode can be specified by the user or sampled by system 100 from a set of scenarios. Examples of scenario 226 are those used to model environmental instabilities during the operation of real-world facilities that cannot be effectively captured by the operation of simulator 110. An example of such a real-world scenario is simulating changing weather conditions in the environment of a facility that can affect the effect of an action on the state of the facility.

[0073] More specifically, scenario 226 is implemented as a change to the input to the simulator, i.e., the configuration parameters 262 used for a given task and / or the control input given to the simulator, a change to the output of the simulator 110, i.e., a change to the measured values generated by the simulator 110, or both.

[0074] As described above, the configuration parameters 262 include information about the state of the facility. When scenario 226 that changes the configuration parameters is selected, after the environment subsystem 220 changes the configuration parameters 262 using scenario 226, it provides the configuration parameters 262 as an input to the simulator 110. When scenario 226 that changes the control input is selected, after the environment subsystem 220 changes the control input using scenario 226, it provides the control input as an input to the simulator 110. When scenario 226 that changes the measured values is selected, after the environment subsystem 220 changes the measured values using scenario 226, it converts the measured values into observed values.

[0075] Therefore, at each episode step, in addition to sending the control input, the environment subsystem 220 also changes one or more values of the selected configuration parameters 262, the selected control input itself, or the selected measured values provided in response to the control input according to scenario 226.

[0076] Specifically, scenario 226 can be implemented as a time-dependent function that creates values to be used as modifiers for one or more of (i) one or more configuration parameters, (ii) one or more control inputs, or (iii) one or more measured values. That is, scenario 226 maps the time index during the control episode to the respective modifiers for one or more of (i), (ii), or (iii).

[0077] Some specific examples of scenario 226 follow.

[0078]

[0078] An example of Scenario 226 is a baseline scenario that does not utilize a configuration trajectory. The baseline scenario can be used to test an agent's performance while controlling a simulated facility that is not subject to perturbations, and to develop a fitness baseline for comparison with other tasks. Thus, in this scenario, the initial parameter values specified by the configuration are used throughout the episode.

[0079]

[0079] Another example of Scenario 226 is a sensor drift scenario. This scenario introduces time-correlated noise into a selected set of measurement components. For example, the components can be randomly selected at the beginning of each episode. This scenario can test an agent's resilience to partially incorrect information.

[0080]

[0080] Another example of Scenario 226 is a frozen control scenario. This scenario freezes the values of selected controls for a random amount of time. That is, instead of applying a value for the control specified by an action, the environment subsystem 220 instead samples the length of a random time interval and provides the environment with the value of the selected control that was selected immediately prior to the start of that time interval for the length of that time interval. For example, the control can be randomly selected at the beginning of each episode. This scenario can test an agent's ability to detect when a selected policy is not functioning and adapt by switching to an alternative.

[0081] Another example of scenario 226 is a non - stationary dynamics scenario. This scenario uses a set of configuration trajectories to change selected simulation configuration parameters that represent aspects of the real - world environment of real - world facilities over the course of an episode. The configuration trajectories create changes to the selected parameters, and subsystem 220 adds these to their baseline values and then passes them to simulator 110. Examples of such parameters include external environmental temperature, humidity, wind speed, precipitation, etc. This scenario can test an agent's resilience to changing environmental conditions, building loads, and other variables outside the agent's control domain.

[0082] Another example of a scenario is a scenario of equipment degradation. This scenario uses a set of configuration trajectories to change selected simulation configuration parameters that represent the performance efficiency or other metrics of equipment within the facility, such as pumps, heat exchangers, cooling towers, chillers, etc. This scenario can test an agent's resilience to a decrease in equipment performance, for example, as a result of wear, during the operation of the facility.

[0083] In this way, a given task episode is specified by the selection of simulator 110 from simulator storage 210, the simulator configuration that specifies the initial configuration parameter values 262, scenario 226, and optionally the noise parameters of noise generator 290. Once these are specified, for example, sampled by the system or specified by the user, system 100 can execute the task episode to generate training data for RL agent 230.

[0084] Figure 3 shows an example 300 of the operation of simulation system 100 during a task episode.

[0085] The task "episode" is a series of time steps during which agent 230 controls simulator 110. A "time step" is the time interval during which measurements are received from simulator 110 and control inputs are provided to the simulator according to the measurements. A task episode can end, for example, when a predetermined number of time steps have occurred, when a hard constraint is violated, or when an error occurs in the simulator.

[0086] Before starting a task episode, the system selects configuration 304. For example, the system can select a predetermined or randomly sampled initial configuration for the configuration parameters of simulator 110.

[0087] The system also identifies a scenario, which is represented as a configuration trajectory 302 that assigns respective values or respective changes to one or more of the configuration parameters, one or more of the control inputs, or one or more of the measurements at each time step during the episode. That is, the scenario defines a time-dependent function for updating the initial configuration 304, the measurements, and / or the control inputs.

[0088] At each time step, agent 230 receives an observation 330 and selects an action 340 that specifies the value of one or more setpoints of simulator 110.

[0089] Action converter 350 converts action 340 into control input 354 for simulator 110. As part of the conversion, action converter 350 can add noise 352 to the control input.

[0090] The simulator receives control input 354 and the values of the configuration parameters generated by applying trajectory 302 to configuration 304, and generates a measurement 306 that includes the respective current values of each set of sensors of the facility as a result of applying control input 354 given the configuration parameter values.

[0091] The observed value converter 310 (e.g., part of the environment subsystem 220) converts the measured value 306 into the next observed value 330 of the agent 230. Specifically, as part of the conversion of the observed value, the converter 310 can add the observed noise 312 to one or more of the measured values 306. If the scenario requires that there be sensor drift in one of the sensors, the observed noise 312 can reflect the readings with the specific noise of the selected sensor.

[0092] The constraint evaluator 260 then evaluates the control input generated from the action proposed by the agent 230 and the observed value generated from the measured value 306 to determine whether any of the constraints are violated. If so, information identifying which constraint is violated can be added to the observed value 330 before the next observed value 330 is passed to the agent 230.

[0093] FIG. 4 is a flowchart of an exemplary process 400 for executing a task episode using a simulator. For convenience, process 400 is described as being performed by a system of one or more computers located in one or more locations. For example, a simulation system appropriately programmed in accordance with this specification, such as the simulation system 100 shown in FIG. 1, can perform process 400.

[0094] Prior to the task episode, the system can select a simulator configuration and a scenario. For example, the system can randomly sample a configuration from a set of configurations capable of modeling the real-world operating conditions of the facility and receive a scenario as user input. As another example, the system can randomly sample both a configuration and a scenario.

[0095] The system then performs the following steps at each time step during the task episode.

[0096] The system receives, as output from the simulator, measurement values representing the current state of the industrial equipment modeled by the simulator (step 402).

[0097] The system converts the measurement values into observation values (step 404).

[0098] The system provides the observation values as input to the control policy (step 406). For example, the control policy may be a policy trained by an RL agent.

[0099] The system receives an action as output from the control policy (step 408).

[0100] The system converts the action into a control input for the simulator (step 410).

[0101] The system provides, as input to the simulator, i.e., for use in generating new measurement values representing the next state of the industrial equipment, the control input and the current values of the configuration parameters (412).

[0102] In this specification, the term "configured" is used in relation to systems and computer program components. For a system of one or more computers to be "configured" to perform a particular operation or action means that the system has installed therein software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action during operation. For one or more computer programs to be configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.

[0103] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. A computer storage medium may be, or include, a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to an appropriate receiver apparatus for execution by a data processing apparatus.

[0104] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, all kinds of apparatus, devices, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can be, or further include, a special purpose logic circuit, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), or both. The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations of them.

[0105] A computer program, which may also be referred to as or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including a compiler language or an interpreter language or a declarative language or a procedural language, and may be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in the file system. The program may be stored in another program or data, for example, a part of a file containing one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple cooperating files, such as one or more modules, subprograms, or files storing a part of the code. The computer program may be executed on one computer or deployed to be executed on multiple computers located at one site or distributed across multiple sites and interconnected by a data communication network.

[0106] As used herein, the term "database" is used broadly to refer to any collection of data, which need not be structured in any particular way or at all and may be stored in one or more locations on a storage device. Thus, for example, an indexed database may contain multiple collections of data, each of which may be organized and accessed differently.

[0107] Similarly, as used herein, the term "engine" is widely used to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers located in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines may be installed on the same computer and may operate on the same computer.

[0108] The processes and logical flows described herein may be implemented by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows may also be performed by, for example, a special purpose logic circuit such as an FPGA or ASIC, or by a combination of special purpose logic circuits and one or more programmed computers.

[0109] A computer suitable for executing a computer program can be based on a general-purpose microprocessor or a dedicated microprocessor, or both, or other types of central processing units. Generally, the central processing unit will receive instructions and data from a read-only memory, or a random access memory, or both. Essential elements of a computer are a central processing unit for performing or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory may be supplemented by, or incorporated in, dedicated logic circuitry. Generally, a computer also includes, or is operatively coupled to receive data from, or transfer data to, one or more mass storage devices, such as magnetic disks, magneto-optical disks, or optical disks, for storing data. However, a computer need not have such devices. Further, a computer may be embedded in another device, such as, by way of a few examples, a cellular phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash device.

[0110] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0111] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, as well as a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user can be in any form, including acoustic input, voice input, or tactile input. Additionally, the computer can interact with the user by sending a document to the device used by the user and receiving the document from that device, such as by sending a web page to a web browser on the user's device in response to a request received from the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user as a reply.

[0112] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive portions of machine learning training or production, i.e., inference, workload.

[0113] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework.

[0114] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes, for example, backend components as a data server, or includes middleware components such as an application server, or includes frontend components such as a graphical user interface, a web browser, or a client computer having an app with which a user can interact with an implementation of the subject matter described in this specification, or includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, for example, by a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0115] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers that have a client-server relationship to each other. In some embodiments, the server displays data, such as HTML pages, to a user who interacts with a user device, for example, a device acting as a client, and sends the data for receiving user input from the user. Data generated at the user device, for example, the result of a user interaction, can be received at the server from the device.

[0116] This specification includes many specific implementation details, but these should not be construed as limitations on the scope of the invention or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described in the context of separate embodiments herein may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately in multiple embodiments, or in any suitable partial combination. Further, features may be described above as functioning in a certain combination and may even be initially claimed as such, but one or more features from the claimed combination may in some cases be excised from that combination, and the claimed combination may be directed to a partial combination or a variation of a partial combination.

[0117] Similarly, although operations are shown in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown, or in a sequential order, or that all of the illustrated operations be performed to achieve a desired result. In some environments, multitasking and parallel processing may be advantageous. Further, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0118] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still achieve a desired result. As one example, the processes shown in the accompanying figures do not necessarily require the particular order shown, or a sequential order, to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous.

Explanation of Symbols

[0119] 100 Simulation System 110 Industrial Equipment to be Simulated / Simulator 120 Heating, Ventilation, and Air Conditioning System 130 Controller 150 Control Policy 160 Observed Value 210 Simulator Data Storage 220 Environment Subsystem 226 Scenario 230 RL Agent 232 Environment Observed Value 234 Action 236 Control Input 238 Configuration Parameter / Measured Value 250 RL Agent 260 Constraint Evaluator 262 Configuration Parameter 264 Configuration Storage 270 Task Subsystem 272 Task Observed Value / Reward 290 Noise Generator 302 Trajectory 304 Configuration 306 Measured Value 310 Observed Value Converter 312 Observation Noise 330 Observed Value 340 Action 350 Action Converter 352 Noise 354 Control Input

Claims

1. A method performed by one or more computers, comprising: At each of a plurality of time steps during a task episode, Receiving, from a computer simulator of industrial equipment, measurement values representing the current state of the industrial equipment; Generating observation values from the measurement values; Providing the observation values as an input to a control policy for controlling the industrial equipment; Receiving, as an output from the control policy, an action for controlling one or more setpoints of the industrial equipment; Generating, from the action, one or more control inputs for the one or more setpoints of the industrial equipment; Providing, as an input to the computer simulator, (i) the one or more control inputs and (ii) current values of one or more configuration parameters of the computer simulator, to cause the computer simulator to generate, as an output, new measurement values representing a new state of the industrial equipment for a subsequent time step; A method comprising the above steps.

2. Generating the observation values from the measurement values includes Adding noise to the measurement values The method according to claim 1, comprising the above steps.

3. Generating the one or more control inputs for the one or more setpoints of the industrial equipment from the action includes Adding noise to one or more control inputs defined by the observation values The method according to claim 1 or 2, comprising the above steps.

4. Identifying a scenario for the task episode, the scenario specifying, for each of the plurality of time steps, One or more of the configuration parameters, One or more of the control inputs, or One or more of the measurement values, Each change applied to one or more of them. The method according to any one of claims 1 to 3, further comprising the above step.

5. The scenario specifies changes applied to one or more of the configuration parameters, and the method includes Sampling a configuration for the task episode by specifying respective initial values for each of the configuration parameters, and at each time step, For each of the one or more configuration parameters, applying the change specified by the scenario of the time step to the initial value of the configuration parameter to generate the current value of the configuration parameter The method according to claim 4, further comprising.

6. The scenario specifies changes to be applied to one or more of the measured values, and generating observed values from the measured values For each of the one or more measured values, applying the change specified by the scenario of the time step to the measured value The method according to claim 4 or claim 5, comprising.

7. The scenario specifies changes to be applied to one or more of the control inputs, and generating one or more control inputs from the action For each of the one or more control inputs, applying the change specified by the scenario of the time step to the control input The method according to any one of claims 4 to 6, comprising.

8. The method according to any one of claims 1 to 7, wherein the computer simulator is a deterministic simulator of the dynamics of the industrial facility.

9. Training the control policy based at least on the task episode After the training, deploying the control policy for controlling the industrial facility The method according to any one of claims 1 to 8, further comprising.

10. Evaluating the control policy based at least on the task episode After the evaluation, deploying the control policy for controlling the industrial facility The method according to any one of claims 1 to 9, further comprising.

11. After deploying the control policy, receiving a measured value of the current state of the industrial facility from the industrial facility Generating a second observed value from the measured value of the current state of the industrial facility Providing the second observed value as an input to the control policy for controlling the industrial facility Receiving, as an output from the control policy, a second action for controlling one or more setpoints of the industrial facility Generating a second one or more control inputs for the one or more setpoints of the industrial facility from the second action controlling the one or more set points of the industrial facility based on the second one or more control inputs The method according to claim 9 or 10, further comprising. **Claim 12** controlling a second industrial facility using a second control policy to generate a data set further comprising the computer simulator of the industrial facility is configured to generate the measurement values representing the current new state of the industrial facility based on the data set The method according to any one of claims 1 to 11. **Claim 13** One or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the operations of the respective methods according to any one of claims 1 to 12. A system comprising. **Claim 14** One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the respective methods according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Plant operation condition setting support system, learning device, and operation condition setting support device

    JP2019197315A

  • Robust adjustment device and model creation method

    JP2020003893A

  • Artificial intelligence system for learning robotic control policies

    US10792810B1