Method and system for controlling a fluid delivery system
By automatically selecting a subset of input signals and updating the control strategy through a self-learning control process, the complexity of existing HVAC systems and zone heating network control methods is solved, achieving efficient and fast fluid transport system control that is adaptable to different types of system configurations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GRUNDFOS HLDG
- Filing Date
- 2021-05-21
- Publication Date
- 2026-05-19
AI Technical Summary
Existing control methods for HVAC systems and district heating networks require extensive manual configuration, and existing model predictive control methods are difficult to adapt to differences between different buildings or areas, resulting in complex and inefficient control system installation.
By adopting a self-learning control process, a data-driven control method is achieved by automatically selecting a subset of input signals and updating the control strategy based on a performance index function. This reduces the dependence on the system model and adapts to different types of fluid transport systems.
It enables efficient and rapid control in different types of fluid transport systems, reduces the need for expert knowledge, simplifies the system configuration process, and improves control quality and system adaptability.
Smart Images

Figure CN116157751B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method and system for controlling fluid transport systems such as heating systems, water supply systems, wastewater systems, etc. Background Technology
[0002] Operating heating systems, such as HVAC systems used in buildings and district heating networks, at the lowest possible cost while still providing good comfort to the end user has been a problem that has been addressed for decades. However, the problem is complex, and conventional industrial controllers are used in many existing systems.
[0003] Previous attempts have been made to develop control methods that can be used in ecosystems with abundant data, see, for example, EP2807527 or “Model Predictive Control (MPC) for Improving Energy Efficiency of Buildings and HVAC Systems: Problem Formulation, Applications and Opportunities” by Gianluca Serale, Massimo Fiorentini, Alfonso Capozzoli, Daniele Bernardini, and Alberto Bemporad, Energies, November 2018. These existing control methods are model-based, and in particular, use a model predictive control (MPC) framework. However, this requires an appropriate model structure for the building or for consumers in a district heating network. Such models are not readily available due to differences between buildings or between district heating networks. In particular, these differences may relate to variations in the loads involved, the availability of data points, structural details, etc. Therefore, each control system may have to be manually configured one by one by appropriate experts based on posterior knowledge of the specific system.
[0004] A growing number of devices, such as IoT devices, are connected to data networks like the Internet, thereby gaining increased access to device data points. Therefore, large pools of data points can be used to improve the control of HVAC systems and district heating networks, particularly for providing end-users with systems that operate with good comfort while keeping operating costs as low as possible. Recently, data-driven (i.e., model-free) control methods have been proposed in an attempt to reduce the need for manual configuration of each device individually.
[0005] Overgaard, in the proceedings of the 13th REHVA World Congress CLIMA 2019 JDBendtsen and BKNielsen's "Hybrid Loop Control Using Reinforcement Learning" describes a reinforcement learning method called Q-learning in the context of hybrid loops used to control temperature and pressure in liquid circulating heating systems.
[0006] Although the above studies have shown that the proposed method outperforms some commercial industrial controllers, the practical application of reinforcement learning involves many challenges.
[0007] Augmented agents typically require long training times to achieve sufficiently high or even acceptable quality of control. Furthermore, heating systems can vary significantly depending on the installation. This reduces the applicability of results obtained from a generic model to individual installations.
[0008] Therefore, it remains desirable to provide a method for controlling a fluid transport system that partially or completely solves one or more of the problems mentioned above and / or provides other benefits. Summary of the Invention
[0009] According to one aspect, this document discloses an embodiment of a computer-implemented method for controlling the operation of a fluid transport system by applying a self-learning control process, the method comprising:
[0010] - During the operation of the fluid delivery system in the first time period, the values of multiple input signals are received, wherein the operation of the fluid delivery system during the first time period is controlled by a predetermined control process.
[0011] - Based on the received values of multiple input signals, automatically select a subset of the multiple input signals.
[0012] - During the operation of the fluid delivery system in the second time period, values of at least a subset of selected input signals are received, wherein the operation of the fluid delivery system in the second time period is controlled by applying a self-learning control process, wherein the self-learning control process is configured to control the operation of the fluid delivery system based only on the selected subset of input signals, and wherein applying the self-learning control process includes updating the self-learning control process based on the received values of the selected subset of input signals and at least based on an approximation of a performance index function.
[0013] Therefore, embodiments of the process disclosed herein control a fluid transport system by applying a self-learning process that is updated continuously or intermittently during operation. Updates are based on input signals and on a performance index function, or at least its approximation. Since the input signals to the self-learning control process are automatically selected as a subset from multiple available input signals, a potentially large number of available input signals can be effectively managed. By making the self-learning control process based only on the selected subset of input signals, the time required for the update process to produce the expected value of the performance index function, particularly the time required to produce at least a near-optimal value of the performance index function, can be significantly reduced. The automatic selection of the subset of input signals is based on the values of multiple input signals received while the fluid transport system is being controlled by a predetermined control process. Therefore, the process is able to select a subset of input signals that are highly correlated with the updates of the self-learning process, thereby facilitating rapid convergence of the update process. Furthermore, since the subset of input signals is selected from multiple input signals available during the operation of the fluid transport system, embodiments of the process disclosed herein can be effectively applied to different types of fluid transport systems and provide high-quality control for various types of fluid transport systems.
[0014] The embodiments of the methods described herein are data-driven and do not heavily depend on the model of the fluid delivery system to be controlled. In particular, the automatic selection of input signals can be data-driven. Furthermore, the need for posterior domain knowledge of the configuration control process is greatly reduced, if not completely eliminated.
[0015] Typically, examples of fluid transport systems include systems that utilize fluids as a heat transfer medium, such as heating and / or cooling systems, particularly heating and / or cooling systems for one or more buildings. Therefore, examples of fluid transport systems include liquid circulation systems, heating and / or ventilation and / or air conditioning systems—also known as HVAC systems. Other examples of fluid transport systems include district heating systems. District heating systems can include district heating networks comprising multiple heating systems for individual buildings. Yet another example of a fluid transport system includes fluid supply systems or fluid disposal systems, such as water supply systems or wastewater systems. Examples of fluids transported by fluid transport systems include liquids, such as water or aqueous liquids, for example, wastewater or liquids comprising mostly water and a small portion of other components. Other examples of fluids include other forms of liquid or gaseous heating or cooling agents. Other examples of fluids include air or other gases.
[0016] Typically, each control process (i.e., predetermined control process and self-learning control process) controls the operation of the fluid delivery system by controlling one or more controllable control variables. Examples of controllable control parameters include valve settings (especially valve opening degree), temperature setpoint, pump pressure setpoint or pump speed setpoint, damper opening degree and / or fan speed, etc. Typically, each control process (i.e., predetermined control process and self-learning control process) can control the fluid delivery system in response to one or more input signals, particularly one or more control variables of the fluid delivery system. To this end, each control process can implement its own control strategy, which determines how to control the fluid delivery system in response to input signals, particularly how to control the control variables of the fluid delivery system in response to input signals.
[0017] The predetermined control process can implement predetermined control strategies, such as model-based and / or rule-based control strategies, such as feedforward control, feedback control, fixed setpoint routines, and / or another conventional control process. In particular, the predetermined control process can implement non-adaptive control strategies, meaning the control strategy applied by the predetermined control process can be static during the first time period, i.e., unchanged. Therefore, the initial configuration of the self-learning control process, particularly the selection of input signals, requires neither extensive expertise in the specific fluid transport system to be controlled nor data on the performance of the fluid transport system while it is being controlled by the self-learning control process to be configured. Instead, data for selecting input signals can be collected while the fluid transport system is controlled by another suitable control process, such as a conventional non-adaptive control process. Therefore, embodiments of the methods described herein can even be implemented in new fluid transport systems in a plug-and-play manner. Alternatively or additionally, in some embodiments, the predetermined control process may be a prior self-learning control process, thereby facilitating the reconfiguration of the earlier self-learning control process, such as periodically or by user commands or other triggering events, for example, periodically or by user commands or other triggering events in the event that an existing fluid control system has been modified.
[0018] A self-learning control process refers to an adaptive control process configured to automatically change (i.e., update) the control strategy, for example, by updating a data-driven estimate of the future response (specifically, future performance) of a fluid transport system in response to control actions taken by the self-learning control system. The resulting self-learning control process updates over time based on one or more criteria, specifically to improve (or at least approximate) performance measurements. An example of a performance measurement is the so-called gain, which can be expressed as a weighted sum of performance metrics obtained at corresponding times. Performance metrics, or combinations of performance metrics, can also be referred to as rewards.
[0019] The performance of a fluid transport system at a given time typically depends on its operating state at one or more previous times. Therefore, control actions taken at a given time to change the operating state of the fluid transport system affect the system's future performance. These control actions will also be referred to simply as actions. Examples of actions may include setting and / or adjusting one or more control variables of the fluid transport system, such as setting or adjusting valve settings (especially valve opening degree), temperature setpoints, pump pressure setpoints or pump speed setpoints, damper opening degree, and / or fan speed, etc.
[0020] Since the future performance of the system has not yet been measured when decisions are made regarding potential control actions, a self-learning control process can seek to select actions that improve the expected future performance measurement. The expected future performance measurement is also known as the performance index function. The performance index function can depend on the current operating state and optionally on the actions taken. Depending on the type of performance index function, the self-learning control process can choose to seek actions that increase or decrease the value of the performance index function.
[0021] Since a self-learning control process may not know the exact form of the performance index function, i.e., how the system's expected future performance depends on the current state and / or current action, some embodiments of the self-learning process can maintain an estimate or approximation of the performance index function, particularly a parameterized approximation. This approximation will also be referred to as the performance approximator function, as it can represent an approximation of how the estimated future performance measurement depends on the current operating state of the fluid transport system and / or on the current control action taken by the self-learning control process.
[0022] Examples of self-learning control processes include control processes that implement reward-based learning agents, such as reinforcement learning processes for adapting control policies. Some embodiments of self-learning control processes select control actions to be taken based on one or more selection rules, also referred to as policies. The one or more selection rules may indicate which actions to take based on the current state of the fluid transport system (particularly based on a subset of received input signals). Examples of selection rules include selection rules that select actions that improve the output of an estimated future performance measurement. Some examples of selection rules, such as the so-called ε-greedy policy, are exploratory selection rules that select actions to improve the output of an estimated future performance measurement while allowing exploratory actions with a certain probability. Exploratory actions may include suboptimal actions based on the current performance approximator function.
[0023] Some embodiments of the self-learning control process can update the performance approximator function based on observed performance metrics. Observed performance metrics can be measured values and / or values calculated based on measured parameters, and can indicate the performance of the fluid delivery system. Performance metrics can also be referred to as rewards, as they can represent the reward associated with a previous action taken by the self-learning agent. Therefore, updates to the resulting self-learning control process over time can include updates to the performance approximator function to reduce error measurements, which can indicate the deviation between one or more observed performance metrics and the output of the performance approximator function derived from one or more previous actions.
[0024] Other examples of self-learning control processes include adaptive control and iterative learning. For the purposes of this specification, self-learning control processes will also be referred to as “learning agents.” Implementations of self-learning processes are capable of learning from their experience, particularly based on observed input variables and performance metrics from one or more observed parameters, to adapt the control policy. The self-learning process starts with an initial version of the control policy and is then able to act and adapt autonomously through learning to improve its own performance. The current control policy may include a current performance approximator function and a selection rule (also called a policy), by which the self-learning process selects one or more actions given the output of the current performance approximator function. Updating the self-learning process may include updating the current performance metric function and / or the selection rule based on one or more observed performance metrics.
[0025] Typically, each input signal can indicate a sensed, determined, or otherwise acquired input variable. For example, some input signals can be sensed by and received directly or indirectly from suitable sensors such as temperature sensors, pressure sensors, flow sensors, etc. Some input signals can be determined by the process from one or more sensed signals and / or from other input data. Some input signals can be received from components of the fluid delivery system. Additionally, some input signals can be received from a remote data processing system, from the Internet, or from another source. Sensed or otherwise acquired input signals can indicate characteristics of the fluid at a location within the fluid delivery system, such as fluid temperature, fluid pressure, or fluid flow rate. Other examples of sensed or otherwise acquired input signals can indicate other characteristics of the fluid delivery system, such as the operating temperature or power consumption of a pump or another functional component of the fluid delivery system. Still other examples of sensed or otherwise acquired input signals can indicate characteristics of the environment surrounding the fluid delivery system, such as room temperature, outside temperature, wind speed, wind direction, etc. In some embodiments, one or more of the input signals are received from various IoT-connected devices such as weather stations, calorimeters, etc. Therefore, selecting an input signal may include selecting the type of input variable, such as room temperature measured by a temperature sensor, fluid pressure measured by a pressure sensor, etc.
[0026] Each input signal can be associated with the moment when an input variable is sensed or otherwise acquired (e.g., absolute or relative time), such as relative to the current time or another suitable reference time. In some embodiments, the associated time can simply be the time when the process receives the input signal; alternatively or additionally, one or more input signals can have a timestamp associated with them, such as indicating the time of measurement of the sensed input variable, etc. Thus, each input signal can indicate the type of the input variable, the value of the input variable sensed or otherwise acquired, and the moment when the value of the input variable was sensed or otherwise acquired (e.g., received). In particular, in some embodiments, one or more input variables in the input signals can represent various time series of the values of the sensed or otherwise acquired input variables. Selecting an input signal can include selecting the type of input variable and also includes selecting a time-shift delay, i.e., selecting a specific relative time at which the self-learning control process will consider the observed values of the input variable.
[0027] The time shift delay can be selected relative to a suitable reference time (such as the current time when the control decision is made, i.e., the time when the control process determines what control action to take). For example, selecting the input signal could include selecting a time series of room temperature measured by a temperature sensor shifted by 1 / 2 h, or a time series of fluid pressure measured by a pressure sensor shifted by 2 minutes. In order to select a subset of input signals from multiple input signals, an input variable measured at a first time and the same input variable measured at a second time different from the first time can therefore be considered as different input signals.
[0028] In many fluid transport systems, the impact of numerous input variables on system performance is associated with delays; that is, changes in input variables do not immediately affect system performance but rather at a later point in time. These delays can depend on the fluid flow rate within the system. Specifically, high flow rates can cause changes in input variables to affect system performance more quickly than low flow rates. To better account for this situation, the selection of input signals can be conditional, for example, dependent on the selection of one or more operating conditions. For instance, the time delay associated with the selected input signal can be chosen based on the fluid flow rate, particularly with respect to the time series representing the values of time-indexed input variables. Therefore, depending on the variable indicating the fluid flow rate in the system, the selected time delay can be a constant or variable time delay, particularly dependent on the flow rate.
[0029] Automatic selection preferably results in the selection of a true subset of the input signals or in another form of dimensionality reduction of the input space defined by multiple input signals. Thus, in some embodiments, multiple input signals define an input space with a first dimension, wherein the selected subset of input signals defines an input space with a reduced dimension less than the first dimension. The selected subset may include the selection of all or only a subset of the input variables. As described above, for each input variable represented as a time series value (i.e., a sequence of values indexed by a time variable), selection may include selecting a time-shift delay indicating when the corresponding observed value of the time series should be used as an input signal for the self-learning control process. Therefore, at any given time step, the self-learning control process may not need to evaluate the entire time series, but only the individual values of the time series, i.e., those values of the time series at the selected time-shift delay. Thus, the selection of one or more individual time-shift delays for the time series can also contribute to reducing the dimensionality of the input space of the self-learning control process, for example, by additionally or alternatively selecting a subset of the input variables.
[0030] Automatic selection can be performed based on one or more selection criteria (i.e., represented as a sequence of values indexed by a time variable). While input selection can be based on a model of the system, it is preferred that the input selection be model-free and driven solely by the received input signals, and optionally by suitable performance measurements, particularly observed performance measurements. Data-driven (especially model-free) selection criteria facilitate the application of the processes disclosed herein to a wide variety of fluid control systems, not all of which are known a priori. In some embodiments, selection can apply nonlinear selection processes, such as information theory selection processes, thereby providing more reliable control for a wide range of fluid transport systems. An information theory selection process refers to a selection process that applies information theory selection criteria. The term "information theory selection criterion" refers to a criterion used for information theory-based selection, particularly a criterion based on measurements of entropy and / or mutual information. Entropy-based measurements determine the level of information in the input signal, while mutual information-based measurements determine the level of information shared between two or more variables, for example, the level of information shared between the input signal and observed performance measurements. In some embodiments, the selection is based on mutual information criteria, particularly selection criteria based on the mutual information between one or more input signals and observed performance measurements. Here and below, the term "observed performance measurement" refers to the output of a performance measurement determined (especially calculated) from actually observed data, particularly from input signals obtained during a first time period. The performance measurement used for the purpose of selecting the input signal may be the same performance measurement sought to be improved through a self-learning control process during a second time period. In some embodiments, selection includes selecting an input variable and a time-shift delay associated with the maximum mutual information of the observed performance measurement.
[0031] Automatic selection of a subset of input signals can be performed during a transition period after the first time period and before the second time period. The transition period can be shorter than either the first or second time period. Specifically, the selection of a subset of input signals can be performed after the completion of the first time period. Upon completion of the input signal selection, the control of the fluid delivery system can switch from a predetermined control process to a self-learning control process, thereby initiating the second time period.
[0032] In some embodiments, the method further includes configuring an initial version of the self-learning control process, particularly based on a selected subset of input signals. Thus, the initial version of the self-learning control process can apply an initial control strategy, which is subsequently updated during operation of the self-learning control process. The configuration of the initial version may include pre-training the initial version of the self-learning control process based on values of multiple input signals received during a first time period and performance index values recorded during operation of the fluid delivery system during the first time period. The configuration of the initial version may also be performed during transition periods. Therefore, the pre-trained self-learning control process provides relatively high-quality control from the outset and evolves towards optimization more rapidly.
[0033] Performance measurements define the quality criteria or other success criteria of a control strategy. Performance measurements may depend on one or more variables, such as temperature, humidity, airflow, energy consumption, etc. Performance measurements may depend on the time of day and / or the time of year and / or the workday, or may otherwise depend on time. The value of a performance measurement may be a scalar. Performance measurements may represent a single performance metric or a combination of multiple performance metrics (e.g., a weighted sum). Examples of performance metrics may be indicators of operating costs, such as power consumption. In the context of heating systems, examples of cost metrics include metrics based on one or more temperatures measured in the system or any other variable that can be associated with the cost of operating a building or district heating network. Another example of performance metrics in the context of heating or HVAC systems may include metrics indicating comfort levels, such as the difference between room temperature and target temperature, the rate of room temperature fluctuation, room temperature variation across the building, humidity, airflow, etc. Therefore, this process allows the control process to consider different performance criteria.
[0034] In some embodiments, the performance index function represents the expected future performance measurement of the fluid transport system; that is, the function value of the performance index function can represent the expected future value of the performance measurement. The expected future performance measurement can be determined based on the current state of the fluid transport system, particularly when the current state of the fluid transport system is represented by a subset of the received input signals. The expected future performance measurement can be further determined based on the control actions determined by a self-learning control process.
[0035] In some embodiments, performance measurement relies on performance metrics evaluated at multiple times, optionally with time-dependent weighting of the performance metric values, particularly time-dependent weighting based on the rate of fluid flow in the fluid delivery system. In many fluid delivery systems, the importance of earlier values of the performance metric to the current performance of the fluid delivery system depends on the flow rate. Similarly, the importance of the current value of the performance metric to the expected future performance of the fluid delivery system depends on the flow rate. Therefore, incorporating weighting (particularly flow rate-dependent weighting) in the updates of the self-learning control process, particularly in the updates of the performance approximator function, helps to more accurately compensate for the flow rate variable delivery delay in the fluid delivery system. In some embodiments, the weighting includes time-dependent weighting of how previously observed performance metrics affect the updates of the self-learning control process. To this end, in some embodiments, the self-learning control process implements a multi-step approach using an eligibility trace. Such embodiments have been found to exhibit good performance. In particular, some embodiments employ flow rate-dependent eligibility trace decay. In some embodiments, volumetric flow rate is used in both the selection of the input signal and the self-learning control process to compensate for the flow rate variable delay between signals.
[0036] In some embodiments, the self-learning control process includes at least one stochastic component, thereby facilitating the exploration of new variations of the control strategy through the self-learning control process, and thus promoting the evolution of the self-learning control process toward an improved control strategy.
[0037] Typically, it may be unknown how the performance metrics of a fluid transport system depend on selected input variables, particularly how the expected future performance of the fluid transport system depends on the selected input variables and / or the control actions taken by the control process. Therefore, in some embodiments, a self-learning control process is based on a performance approximator function that approximates the dependence of the performance metric function on a selected subset of input signals, and optionally on the dependence on the current control action. Specifically, the performance approximator function may be parameterized by multiple weight parameters, and updating the self-learning control process may include updating one or more of these weight parameters. Thus, over time, the self-learning control process learns how to approximate the dependence of the performance metric function (i.e., the expected future performance measurement) on a selected subset of input signals, and optionally, on the dependence on the control actions taken by the self-learning control process.
[0038] During the operation of the fluid delivery system in the second time period, operating conditions may change, and components of the fluid delivery system may be replaced, added, removed, or otherwise altered. Therefore, a subset of the initially selected input signals may no longer be the optimal selection. To better account for such changes, some embodiments of the process perform a reselection of the input signals. Therefore, in some implementations, the method further includes:
[0039] - Based on the received values of multiple input signals, a new subset of the multiple input signals is automatically selected, and this new subset of the multiple input signals is received during the second time period.
[0040] - During the operation of the fluid delivery system in the third time period, the values of a new subset of at least selected input signals are received, wherein the operation of the fluid delivery system in the third time period is controlled by applying a new self-learning control process adapted to the new subset of selected input signals, wherein the new self-learning process is configured to control the operation of the fluid delivery system based only on the new subset of selected input signals, and wherein applying the new self-learning control process includes updating the new self-learning control process based on the values of the new subset of selected input signals received and based on a performance index function or at least an approximation of the performance index function.
[0041] It should be noted that the features of the various embodiments of the computer-implemented methods described above and below can be implemented at least in part in the form of software or firmware, and executed on a data processing system or other processing unit caused by the execution of program code means such as computer-executable instructions. Here and below, the term "processing unit" includes any circuitry and / or means suitable for performing the functions described above. In particular, the above terms include general-purpose programmable microprocessors or application-specific programmable microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable logic arrays (PLAs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), special-purpose electronic circuits, and combinations thereof.
[0042] This disclosure relates to various aspects, including the methods, other methods, systems, apparatuses, and product apparatuses described above and below, each of which produces one or more benefits and advantages in conjunction with one or more other aspects, and each of which has one or more embodiments corresponding to the embodiments described herein and / or as disclosed in one or more other aspects as disclosed in the appended claims.
[0043] In particular, another aspect disclosed herein relates to embodiments of a control system for controlling a fluid delivery system. Embodiments of the control system are configured to perform the actions of the methods described herein. For this purpose, the control system may include one or more processing units, particularly control units and remote data processing systems. The one or more processing units may have program code stored thereon configured to cause the control system to perform the actions of the methods described herein when executed by the one or more processing units. It should be understood that the control process may include multiple processing units, each configured to perform a subset of the actions of the methods described herein.
[0044] Specifically, in some embodiments, the control system includes a control unit communicatively coupled to one or more controllable components of the fluid delivery system; wherein the control unit is configured to receive values of at least a selected subset of input signals acquired during operation of the fluid delivery system, and to selectively control the operation of the fluid delivery system by applying a predetermined control procedure or by applying a self-learning control procedure. Specifically, the control unit may be configured to control the operation of the fluid delivery system by applying a predetermined control procedure during a first time period and optionally during a transition time period, and by applying a self-learning control procedure during a second time period.
[0045] In some embodiments, the control system includes a data processing system configured to receive values of a plurality of input signals acquired during operation of the fluid delivery system during a first time period, and to automatically select a subset of the plurality of input signals based on the received values. In some embodiments, the data processing system is a remote data processing system located away from the control unit, particularly a cloud service. In some embodiments, the data processing system is further configured to configure an initial version of a self-learning control process based on the selected subset of input signals; wherein configuring the initial version of the self-learning control process includes training the initial version of the self-learning control process based on the values of the plurality of input signals received during operation of the fluid delivery system during the first time period, and based on performance index values recorded during operation of the fluid delivery system during the first time period.
[0046] According to another aspect, embodiments of fluid delivery systems are disclosed herein, including embodiments of control systems as described herein.
[0047] Another aspect disclosed herein relates to embodiments of a computer program configured to cause a control system to perform the actions of the computer-implemented methods described above and below. The computer program may include program code means adapted, when executed on one or more processing units, to cause the one or more processing units to perform the actions of the computer-implemented methods disclosed above and below. The computer program may be stored on a computer-readable storage medium (particularly a non-transient storage medium) or implemented as a data signal. The non-transient storage medium may include any suitable circuitry or means for storing data, such as RAM, ROM, EPROM, EEPROM, flash memory, magnetic or optical storage devices such as CD-ROM, DVD, and / or hard disks. Attached Figure Description
[0048] The above and other aspects will become apparent and elucidated from the embodiments described below with reference to the accompanying drawings, wherein:
[0049] Figure 1 An example of a fluid transport system is illustrated schematically.
[0050] Figure 2 Another example of a fluid transport system is illustrated schematically.
[0051] Figure 3 A control system for controlling a fluid delivery system is shown schematically.
[0052] Figure 4 The process for controlling a fluid delivery system is illustrated schematically.
[0053] Figure 5 The process for controlling a fluid delivery system by a self-learning control process is illustrated schematically.
[0054] Figure 6 The process for selecting input signals for a self-learning control process is illustrated schematically. Detailed Implementation
[0055] In the following description, embodiments of the aspects disclosed herein will be described in the context of heating systems such as HVAC systems or district heating networks.
[0056] Against this backdrop, embodiments of the methods and systems disclosed herein provide a data-driven approach for self-commissioning and optimal control of building or district heating networks. Here, the term self-commissioning refers to a system that automatically selects which data points from a large pool to be used as inputs for a data-driven control method.
[0057] The embodiments of the processes and systems described herein employ a self-learning control process that, even without minimization, reduces the operating costs of building HVAC systems or district heating networks while maintaining good comfort. In contrast to existing solutions, the embodiments of the systems and methods described herein do not require extensive a posteriori knowledge for configuring the control system. Furthermore, the embodiments of the methods and systems described herein can be applied to new fluid delivery systems where no data logs are yet available for configuration.
[0058] Figure 1 An embodiment of a fluid transport system, particularly a heating system, is illustrated schematically.
[0059] The system includes one or more controllable components 40, also referred to as actuators. Examples of controllable components in a heating system include valves, pumps, and / or dampers. For example, controllable component 40 may include a valve / pump combination constituting a so-called hybrid loop. It should be understood that some examples of fluid delivery systems may include alternative and / or additional types of controllable components, such as fans, etc. In addition to controllable component 40, the system also includes additional components (not explicitly shown), such as pipes, reservoirs, radiators, etc. Some or all of the additional components may be directly or indirectly coupled to controllable component 40 (e.g., in fluid communication with controllable component 40).
[0060] The heating system includes a control system 10, which is operatively coupled to the controllable component 40 and configured to control one or more controllable variables of the fluid delivery system by controlling the controllable component 40. Examples of controllable variables include valve settings (particularly valve opening degree), temperature setpoints, pump pressure setpoints or pump speed setpoints, damper opening degree and / or fan speed, etc.
[0061] The control system 10 can be implemented by a suitably programmed data processing system, such as a suitably programmed computer, or by another data processing device or control unit. In some embodiments, the control system 10 can be implemented as a distributed system including multiple computers, data processing devices, or control units. The control system 10 is communicatively coupled to the controllable component 40, for example, via a wired or wireless connection. Communication between the control system and the controllable component can be via a direct connection or via an indirect connection, for example, via one or more nodes of a communication network. Examples of wired connections include local area networks, serial communication links, control buses, direct control lines, etc. Examples of wireless connections include radio frequency communication links, such as Wi-Fi, Bluetooth, cellular communication, etc.
[0062] The control system 10 includes a suitably programmed processing unit 11, such as a CPU, microprocessor, etc. The control system also includes a memory 12, which may store computer programs and / or data for use by the processing unit 10. It should be understood that the control system 10 may include, for example, additional components such as a graphical user interface displayed on a display of the data processing system (e.g., on a touchscreen), such as one or more communication interfaces and / or user interfaces. Examples of communication interfaces include wired network adapters or wireless network adapters, serial data interfaces, Bluetooth transceivers, etc.
[0063] The heating system includes multiple sensors 30. Examples of sensors include temperature sensors, pressure sensors, sensors for sensing wind speed, sensors for sensing the operational status of windows, doors, etc.
[0064] Sensor 30 is communicatively coupled to control system 10, for example, via a wired or wireless connection. Communication between sensor 30 and control system 10 can be via a direct or indirect connection, such as via one or more nodes of a communication network. Examples of wired connections include local area networks, serial communication links, data buses, direct data lines, etc. Examples of wireless connections include radio frequency communication links, such as Wi-Fi, Bluetooth, cellular communication, etc. Figure 1 In this example, the control system is directly coupled to some sensors and indirectly coupled to others. Indirect coupling can occur via the building management system 20 or by receiving sensor signals from multiple sensors and / or other data sources, and forwarding some or all of the sensor data to another form of data server of the control system 10. In some embodiments, the building management system may implement both a data server and the control system 10.
[0065] Typically, examples of input signals include sensor data, such as sensor data from a load-indicating sensor associated with the heating system (e.g., temperature, pressure, flow rate, etc.), or sensor data from an event indicator such as a window or door switch, or sensor data from other types of sensors.
[0066] In addition to or in lieu of the sensor signals from sensor 30, the control system 10 may also receive data or other input signals from other sources. For example, the control system may receive weather forecast data from a weather service, occupancy data from a reservation system, or information about energy prices from an external system.
[0067] Therefore, during operation, the control system 10 receives sensor input from sensor 30 and optionally receives further input from other sources. Typically, the input from sensor 30 and optionally from other sources can be received in the form of a digital signal or an analog signal, which can then be converted into a digital signal. For the purposes of this specification, the received input will be referred to as the input signal. The control system 10 may receive the input signal intermittently, for example periodically, such that the control system 10 receives one or more time-series sensed values indicating corresponding input variables sensed by the sensor at different points in time.
[0068] The control system 10 is configured to execute a control process that controls the controllable component 40 in response to received input signals or at least a selected subset of input signals, as described herein. Specifically, the control system is configured to execute the process as described herein, for example, with reference to the following... Figures 3 to 6 One or more processes described in the process.
[0069] Figure 2 Another embodiment of the fluid delivery system is shown. Besides the fact that the control system 10 is a distributed system, Figure 2 The system and Figure 1 The system is the same. Specifically, the control system includes a local control unit 10A, which is communicatively coupled to the controllable component 40 and the sensor 30. The control system also includes a remote data processing system 10B, such as a remote computer, a distributed computing environment, etc. The local control unit 10A includes a processing unit 11A and a memory 12A, such as in combination... Figure 1 The remote data processing system 10B also includes one or more processing units 11B (e.g., one or more CPUs) and at least one memory 12B. The local control unit 10A and the remote data processing system 10B are communicatively coupled to each other, for example, via a direct or indirect communication link, which may be wired or wireless. For example, the local control unit 10A and the remote data processing system 10B may be communicatively coupled via the Internet or another suitable computer network. In some embodiments, the remote data processing system may receive input directly from the building management system 20 and / or from sensors 30.
[0070] exist Figure 2In this embodiment, the local control unit 10A and the remote data processing system 10B can cooperate to implement embodiments of the processes described herein. For example, the remote data processing system 10B can implement the selection of input signals and / or configurations, and optionally, implement pre-training of the self-learning control process, while the local control unit 10A can execute the self-learning control process. Alternatively, the remote data processing system 10B can also execute a portion of the self-learning control process, such as determining the action to be taken and / or updating the self-learning control process. In such an embodiment, the local control unit 10A can receive information from the remote data processing system about the action to be taken and translate information in a specific control command to the controllable component 40. The local control unit 10A can also collect input signals and forward at least selected input signals to the remote data processing system 10B.
[0071] Figure 3 A self-learning control process (generally designated 310) for controlling a fluid delivery system, particularly a heating system for controlling a heated building, is schematically illustrated. The self-learning control process 310 may be a reinforcement learning control process or other self-learning control process for controlling the controllable component 40 of the fluid delivery system 60 (e.g., a liquid circulation system such as an HVAC or district heating system).
[0072] The self-learning control process 310 receives multiple input signals, such as signals from sensors or other sources. The input signals may include controllable control variables through which the self-learning control process applies control actions. The input signals may also include other signals, such as signals describing the load on the system; these can be considered disturbances to the control. The self-learning control process uses a subset of the available input signals as a representation of the state of the fluid transport system. The selected variables may include some or all of the control variables and / or signals describing disturbances. In the following text, the total pool of available input variables will be represented by x. The value of x at time t is a vector x. t The selected input signal will be referred to as state s. The value of s at time t is determined by the state vector s. t The state vector represents the state of the fluid transport system at time t. Therefore, the process of automatically selecting which subset of input signals to use by the self-learning control process is also called state selection. The self-learning control process can receive input signals directly from the corresponding signal source or from a data server or another suitable data store that is otherwise aggregated. The self-learning control process and the data server can be implemented as separate modules or integrated with each other. In particular, they can be controlled by a local control unit, a building management system, or by, for example, in a system such as... Figure 1The same data processing system in the illustrated system is implemented in other ways. Alternatively, the self-learning control process and / or data server can be communicatively coupled to a local control unit, for example, in... Figure 2 The remote data processing system is implemented in the system shown.
[0073] The self-learning control process adjusts one or more control variables for a controllable component. The control variable can be a setpoint for a local control loop in the heating system, such as a temperature, pressure, or flow control loop. For example, in a hybrid loop, the control setpoint could be both temperature and pump pressure. For this purpose, the control system executing the self-learning control process 310 has an interface to the controllable component 40 of the heating system, through which the control system applies actions to the heating system. The adjustment of the control variables by the self-learning control process in response to received state information is called a set of actions. This set of actions applied at time t will be designated as vector a. t The controllable component 40 can be, for example, a valve, a pump, or a damper. In one example, the controllable component includes a valve / pump combination that forms a combination called a hybrid loop. The process responds to a selected subset of received input signals, i.e., in response to state s. t Determine action a t .
[0074] The control process is self-learning, meaning that the control actions for a given state change over time to improve performance because new knowledge about the consequences of previous control actions has been learned.
[0075] To this end, the self-learning control process receives one or more performance metrics, which indicate one or more results of the control actions taken by the self-learning control process. This process can make the updates of the self-learning process based on a weighted combination of multiple performance metrics and / or another function of one or more performance metrics. The value of the performance metric at time t will also be referred to as the reward r. t It should be understood that performance metrics can be holistic, such as a combination of multiple different performance metrics. Therefore, taking action a... t And cause the system to enter state s t+1 , generating a reward r t+1 The reward may include one or more performance metrics, such as control error, and optionally include the operating cost of the heating system. The reward may be a function of the selected input signal.
[0076] Typically, reinforcement learning control processes can be configured to, for example, Figure 3 The expected behavior reinforced by rewards updates itself, i.e., learning. The self-learning control process 310 seeks to select actions that optimize the combination of rewards over time, also known as the return G. In particular, the return can be defined as the cumulative reward over n steps:
[0077]
[0078] Here, 0 ≤ γ ≤ 1 is the discount rate that reduces the impact of future rewards on payoff. The increased discount rate ensures that payoff is well defined as time becomes infinite (i.e., n → ∞), while also ensuring the higher importance of rewards occurring quickly. In reinforcement learning, the control policy used to control the system is often referred to as the policy.
[0079] At the moment an action is taken, the reward cannot be measured, and therefore the benefit gained from that action cannot be measured. Therefore, a self-learning control process can consider anticipated future benefits. An action-value function can be defined, which describes the state s. t Take action a t And follow the expected return of strategy π, for example, defined as:
[0080] Q π (s,a)=Ε[G t+∞ |s t =s,a t =a].
[0081] Therefore, the action-value function is an example of a performance metric function. Other examples of performance metric functions include value functions that describe the expected reward of being in state s and following a policy.
[0082] A policy that chooses action a, given a state s, to maximize the expected reward is called a greedy policy. A problem associated with greedy policies is the lack of new exploration of potentially more rewarding actions. The trade-off between exploration and developing current knowledge is a concern in reinforcement learning, as optimality and fitness are desirable. Therefore, stochastic policies capable of maximizing the amount of exploration can be used. An example of such a policy is the so-called ε-greedy policy, which takes a random action with probability ε.
[0083] Self-learning control processes seek to improve system performance by applying a learning process. For example, in Q-learning, the optimal policy that maximizes reward is sought by finding the optimal action-value function, regardless of the policy followed.
[0084]
[0085] Here, α∈[0,1] is the learning rate. While one goal of learning an agent is to approximate the action-value function or other performance metric function, another goal of the agent is to find the optimal policy that maximizes the reward for each state. Since the value function and the optimal policy are interdependent, they are typically optimized in an iterative manner called value-policy iteration.
[0086] Typically, self-learning control processes can use approximations of performance index functions, particularly functional approximations of performance index functions, where a selected subset of input signals and optionally one or more actions are inputs to the function approximation. This approximation is also referred to as a performance approximator function. The performance approximator function can be a parameterized approximation, for example, a parameterized approximation in the form of a neural network function approximation, or it can have a function structure derived from domain knowledge. For example, in some embodiments, the control system can maintain an estimate or approximation of the action value function as a performance approximator function, hereinafter referred to as... Where w represents making Parameterized weight vector. This is achieved by taking action a. t And for state s t and reward r t Sampling and self-learning control process 310 can improve the performance index function. The estimate.
[0087] The update of the performance index function estimate can employ a suitable learning process, such as temporal difference learning. Temporal difference learning is another aspect of concern in reinforcement learning. It can be described as a hybrid of Monte Carlo and dynamic programming. In Monte Carlo, the complete event of action, state transition, and reward is measured, and the estimate of the state-action-value function is then computed purely from the measurements. In dynamic programming, the model of the Markov Decision Process is known, so the estimate from this knowledge can be used for bootstrapping. In temporal difference learning, the bootstrap target is computed from both the sampled reward and the acquired system knowledge. The temporal difference error is the error between the current estimate and the new estimate of the state-action-value function.
[0088] In some implementations, a multi-step method is employed. Multi-step methods generally perform better than single-step methods because they utilize more samples. To achieve this, the gain can be parameterized using trace decay λ, allowing the method to span from λ = 0 (corresponding to a one-step method) up to ... = 1 (corresponding to a Monte Carlo method):
[0089]
[0090] Due to computational advantages, multi-step methods can preferably be implemented as qualification traces. Qualification traces utilize a trace vector z, which varies according to the partial derivative of the estimated performance index function with respect to weights that parameterize the estimated performance index function. The trace vector decays through γλ:
[0091]
[0092] Then adjust the weights according to the following formula:
[0093] w t+1 =w t +αδ t z t
[0094] Among them, the time difference error δ t It is the error between the current estimate and the new estimate of the state-action-value function.
[0095] In some embodiments, the self-learning control process uses volumetric flow rate to schedule a time range on which rewards are considered for the purpose of updating the performance approximator function. Multiple ranges are possible for multiple signals. The time range can be defined explicitly or implicitly, for example, by appropriate weighting or attenuation (such as by trace attenuation as described herein).
[0096] A subset of the input signals used in the self-learning control process is selected from a pool of available input signals. This selection is performed automatically, preferably by a data-driven method, also known as a state selection method.
[0097] A state selection method can be performed before the fluid delivery system is first controlled using a self-learning control process. For this purpose, data can initially be collected from the fluid delivery system when it is controlled by a standard non-adaptive control process. Furthermore, new state selections can subsequently be performed, for example, periodically and / or triggered by a triggering event. In some embodiments, state selection is performed as frequently as possible based on domain knowledge, for example, to keep a subset of the selected input signals updated, such as by combining... Figure 4 The subsequent state selection can be based on data collected during the control of the fluid delivery system by the current self-learning control process.
[0098] Typically, data-driven state selection identifies the input signals containing the most relevant information for a self-learning control process. A faster learning rate for the self-learning control process is achieved using only a subset of the input signals. State selection methods can be computationally expensive and can be advantageously performed by a remote data processing system with access to the available input signals x, for example, via a cloud computing environment, such as in combination with... Figure 2 Described. For example, state selection methods apply mutual information methods, such as, in combination with... Figure 6 Described.
[0099] Figure 4 The process for controlling a fluid delivery system is illustrated schematically.
[0100] In the initial steps S1 and S2, during the first time period, the process controls the fluid delivery system by applying a predetermined control process, particularly a non-adaptive control process. The predetermined control process can be implemented, for example, as a conventional, commercially available controller, such as a PLC known in the art. While the predetermined control process controls the fluid delivery system, it samples and records multiple input signals. The process also samples and records performance indicators during the first cycle, which indicate the performance of the fluid delivery system under the control of the predetermined control process. During data collection, in step S1, training data can be collected for state selection. This can be done within a predetermined time period t. s Training data is collected during this period. The training data used for state selection includes a set of input signals x. s During time period t s Actions taken by the predetermined control process during the period, and during the time period t s The performance metrics recorded during the period, also known as the reward r s Here, the subscript 's' refers to the state selection.
[0101] Additionally, in step S2, further data can be collected for use as verification data, particularly for defining stopping criteria. This can be done within a predetermined time period t. v Verification data is collected during this period. The verification data used for state selection includes a set of input signals x. v During time period t v Actions to be taken during the predetermined control process a v And in time period t v The recorded reward r during the period v Here, the subscript v refers to the validation data.
[0102] During steps S1 and / or S2, the process may record a flow rate value q, which can be used as an input signal and / or for the purpose of performing flow rate correction.
[0103] The first time period t1 = t can be determined in advance.s +t v Duration and time period t s and t v The choice of an appropriate duration may depend on various parameters, such as the complexity of the system to be controlled, the number of input signals, and the frequency of sampling the input signals.
[0104] Some embodiments of the process described herein apply compensation for flow-dependent effects through a self-learning control process, particularly the flow-dependent selection of the input signal and / or the quantity-dependent weighting of the reward. To this end, in optional step S3, the process selects a constant. This is intended for use in compensation of flow-dependent effects, particularly for calculating the trace decay of flow variables. Given a chosen constant... The trace decay λ of the flow variable can be calculated as follows: Where, q η(t) This represents the normalized fluid flow rate in the fluid transport system at time t. (Constant) From the normalized lumped volume v η Calculate: φ=h(v) η It can be determined based on domain knowledge about fluid transport systems. For example, the lumped volume factor v η The decision can be based on the delay between two variables, which results in maximum mutual information between them, or it can be based on parameters that indicate general transmission delays in the system. Specifically, for HVAC or district heating systems, the two variables can be the supply temperature and the return temperature, the latter being a function of the supply temperature at an earlier time point.
[0105] In the subsequent step S4, when flow compensation is applied for state selection purposes, the attenuation rate λ determined in step S3 is used to calculate the gain Gt or other performance measurements.
[0106] In the subsequent step S5, the process selects a subset of the input signals for use by the self-learning control process during subsequent control of the fluid delivery system. This selection process is also referred to as state selection. State selection preferably uses a mutual information criterion or another suitable selection criterion, such as another information theory selection criterion. In particular, a mutual information selection criterion can make the selection based on the mutual information between the respective input signals and the calculated benefits or other suitable performance measurements. For this purpose, the process can apply the flow compensation benefit G calculated in step S4. t The following will be a reference. Figure 6 An example describing the state selection process in more detail.
[0107] In step S6, the process configures an initial version of the self-learning control process based on the selected input signal. Specifically, the process pre-trains the self-learning control process based on training data collected during step S1 and / or data collected during step S2. This pre-training can be performed using a self-learning scheme, which is then applied during actual use of the self-learning control process, for example, as referenced below. Figure 5 The description is as follows. However, pre-training is "off-policy" execution because the actions used during pre-training are actions executed by a pre-defined control process, not by a self-learning control process being pre-trained. Pre-training is also based on the optional flow-compensated gain G. t .
[0108] Finally, in step S7, the process uses a pre-trained self-learning control process to control the fluid delivery system. The control of the fluid delivery system by the self-learning control process includes updating the self-learning control process according to a suitable self-learning scheme (e.g., reinforcement learning). References will follow below. Figure 5 An example of this process is described in more detail. The self-learning control process can control the fluid delivery system during a second time period (e.g., until it is stopped). For example, triggered at predetermined intervals and / or by a suitable triggering event (such as a user command), the process can determine (step S8) whether a re-state selection should be performed. This determination can be based on the time elapsed since the previous state selection and / or on one or more performance metrics and / or on user commands. If the process initiates a re-state selection, it returns to step S5. The re-state selection can then be based on data recorded during step S7 (i.e., during the current self-learning control process controlling the fluid delivery system). If the re-state selection results in the selection of alternative and / or additional input signals, a new self-learning control process can be configured and pre-trained, replacing the current one.
[0109] Since state selection can be computationally expensive, it is preferable to perform it via a cloud computing environment or other remote data processing system. Furthermore, state selection is preferably performed only when the controlled system undergoes significant changes, thereby making other signals more readily available for learning agent usage. Significant changes might be, for example, due to structural changes in the controlled system, the introduction of new sensors, or variations in radial load.
[0110] If the learning agent uses only a few input signals, the self-learning control process can be computationally cheap enough to be implemented on the local computing unit of the heating system (e.g., the local computing unit of an HVAC or district heating system), or even through the control unit of a component of the heating system (such as a smart valve or centrifugal pump). However, in some embodiments, the self-learning control process can be implemented by a cloud computing environment or other remote data processing system.
[0111] Typically, during the first time period, the process controls the fluid delivery system through a predetermined, such as conventional, control procedure, which may sample the time series of multiple input signals. During the transition period, the process selects a subset of the input signals that provide the most information compared to performance measurements. Also during the transition period, the process pre-trains a self-learning control procedure using the selected input signals. In the subsequent second time period, the process enables the self-learning control procedure to operate in real-time on the selected input signals, and during this second time period, the self-learning control procedure controls the fluid delivery system and continues to optimize itself, particularly by adapting a parameterized estimate of the performance index function.
[0112] For example, when Figure 4 When this process is applied to the mixing loop of a heating system, specific examples can be summarized as follows:
[0113] Result: Plug-and-play control scheme
[0114] Initialization: Commercial Controller
[0115] Parameter: m s m v
[0116]
[0117] Sure
[0118] From v * Sure
[0119] use From the recorded data
[0120] Perform as Figure 6 Status selection
[0121] Use the dataset to select the state departure strategy Pre-training RL agents
[0122]
[0123]
[0124] Here and in the following text, when the symbol ' is used, it refers to a causal relationship between the values of the same parameter. An example is that s is the state vector at a certain iteration / step, and s' is the state vector in the next iteration / step.
[0125] Figure 5 The process for controlling a fluid transport system via a self-learning control procedure is illustrated schematically. In the initialization step S71, the process loads initial weights w, specifically weights obtained from a pre-trained model. If pre-training has not yet been performed, or for pre-training purposes, the weights can be initialized to, for example, random values or zero in another suitable manner. The process further loads an initial trace vector z, which can also be obtained from pre-training. The process also observes the current state s of the fluid transport system. t For example, it receives the current values of a subset of input signals selected during state selection. The state vector *s* can describe the state of a controlled system following Markov properties. Once sufficient input signals have been observed (e.g., considering input signals with delays determined during state selection), the process computes action *a* according to the current control strategy. Action vector *a* describes what actions the learning agent takes to control the system; that is, which control variables are modified and how.
[0126] This action can move the fluid transport system to different states. Being in a state generates a reward r. The self-learning control process seeks to maximize the weighted sum of rewards over time. This weighted sum is called the return G. t Learn about agent-maintained state-action-value functions. State-action-value functions describe how well a system performs with respect to expected returns, given a state, action, and control policy. Via... Performance approximator functions of the form (such as, for example, neural networks) are used to approximate state-action values. Here, It is the performance approximator function value Q that approximates a given state s and action a. The performance approximator function depends on the set of weights w, for example, according to:
[0127]
[0128] Among them, b i These are suitable basis functions, such as radial basis functions.
[0129] The performance approximator function continuously improves over time to better match the system. This is accomplished by the learning agent simultaneously taking different actions based on the measured state and reward, and updating the weights according to a suitable backup function. The backup function utilizes the time difference error (δ), which describes the difference between the current knowledge in the state-action-value space and the newly acquired knowledge, which forms the target that the function should move towards.
[0130] To ensure exploration and facilitate updates to the state-action-value function through new information, the agent can take suboptimal exploratory actions based on its current knowledge. The amount of exploration the agent performs, rather than optimizing the system by developing current knowledge, is determined by the control policy.
[0131] A trail vector z with a decay rate λ is used to determine how quickly historical rewards decay. For higher λ, the effect of the trail following the reward decays more rapidly.
[0132] Specifically, once the process is initialized, it enters a control and update loop, which is repeated until the process terminates. Specifically, in step S72, the process observes the reward r. t and the state s obtained from the previous action taken. t+1 .
[0133] In step S73, the process calculates action a based on the selected control strategy (e.g., based on an ε-greedy strategy). t+1 .
[0134] In step S74, the process is performed from the observed state s via basis functions (e.g., radial basis functions). t+1 and the chosen action a t+1 Calculate the basis vector b t+1 .
[0135] In step S75, the process calculates the time difference error δ.
[0136] The choice of time difference error may differ from that of pre-training (off-policy) and online training (on-policy).
[0137] For example, during strategy training, the time difference error can be chosen to be,
[0138]
[0139] In off-policy approaches, the behavioral policy used by the agent differs from the target policy learned. For example, Q-learning is an off-policy approach because the target policy is the optimal policy seen by the facilitator, which maximizes the action taken.
[0140]
[0141] During pre-training, that is... Figure 4 In step S3 of the process, the data acquired during the first time period is used for the initial training of the self-learning control process before it takes over control; that is, pre-training occurs during the transition period before the second time period. During the first time period, the self-learning control process is used to control the fluid delivery system. Therefore, pre-training is a policy-agnostic process. To achieve knowledge sharing, the time difference error can be based on parameterization, which provides a way to switch between the policy-agnostic and policy-agnostic approaches.
[0142]
[0143] Knowledge transfer can be achieved by setting σ=0 and training the reinforcement learning algorithm on the data recorded in the first time period.
[0144] In step S76, the process observes the flow rate q and calculates the result from q. max normalized q n Therefore, the process calculates the flow variable trace decay λ and then updates the trace vector z. Specifically, to implement reinforcement learning as a multi-step method, a Dutch trace can be used. Flow-dependent trace decay alters the range at which an action affects the reward. Some embodiments apply trace decay proportional to flow. In this and other embodiments of this method, a concentrated pipeline volume approximation is used. This means that only the volume with the highest impact on input-output delay is used. At each sample, the trace decay can be calculated as:
[0145]
[0146] Where, q η(t) ∈[q η,min ,1] is the flow rate normalized by the maximum flow rate, and where φ∈[0,1] is a constant that can be empirically determined as a function of the lumped volume, for example, in the following relationship between the forward and return temperatures:
[0147] φ=h(ν η In the relationship.
[0148] The flow rate and lumped volume can be normalized relative to the system's maximum flow rate:
[0149]
[0150] Here, the description of the lumped volume ν is given in the context of a hybrid loop. Consider an example of a system without terminal units and with only piping connections between the supply and return of the hybrid loop, where the return temperature is a function of the forward temperature, which acts with different delays due to the different piping routes:
[0151] T r (t)=h(T s ,q)
[0152] in,
[0153]
[0154] q = [q1,...,q N ] T .
[0155] The individual flow rates in different pipeline routes are not always known; for example, in some applications of mixed loops, only the total flow rate leaving the mixed loop is known. Therefore, a flow ratio β can be introduced, where the sum of the flow ratios of the p pipeline paths is:
[0156]
[0157] q N (T)=β N q(t).
[0158] Since only the main flow rate is known, the flow ratio β can be assumed to be constant. The terminal unit is controlled by a regulating valve, which can change how the flow ratio is distributed. Because it affects all areas, variations in external temperature may keep the ratio almost constant, while solar radiation reaching only one side of the building may significantly alter the ratio, and this dependence on the specific building makes the assumed approximation less accurate. v can now be defined. N for:
[0159]
[0160] Applying this to the example above gives:
[0161]
[0162] Therefore, a minimum flow threshold can be used.
[0163] In step S77, the process updates the weights w of the performance approximator function. This process stores the basis vector b. t+1 The state action value Q is used for the next iteration.
[0164] In step S78, the process performs action a. t+1 Then return to step S72.
[0165] It should be understood that different implementations may use different types of self-learning control processes, such as different types of backup functions and / or different types of time difference measurements.
[0166] For example, when Figure 5The specific implementation methods of this process, when applied to hybrid circuits, can be summarized as follows:
[0167] Result: Online
[0168] Initialization: weight w, trace vector z, according to ε-greedy π(.|s0)
[0169] Take action a′ and calculate the characteristic state.
[0170] b = b(s0′a′), Q old =0
[0171] Parameters: ε, α, γ, φ, σ
[0172]
[0173] The aspects of interest in the embodiments described herein include the approximation of the state-action space by a performance approximator function. To ensure stability, it is preferable, for example, to employ a performance approximator function with linear weights that guarantees convergence. Multi-step methods implemented using qualification traces have shown good performance.
[0174] The above and other embodiments employ flow-dependent qualification trace decay. Flow-variable trace decay is used to compensate for the transmission delay of flow-variables in fluid transport systems. With this compensation, lower flow rates result in slower decay, which in turn increases the benefit of larger delays. A lumped volume parameter is used to determine the decay rate λ. The lumped volume can be found by analyzing the correlation between supply and return temperatures, which provides information about the delay characteristics of the system. An example of how flow-variable qualification traces can be applied is described in Overgaard, A., Nielsen, B.K., Kallesoe, CS, & Bendtsen, and JD's "Reinforcement Learning with Flow-Variable Qualification Traces for Hybrid Loop Control" (2019) (https: / / doi.org / 10.1109 / CCTA.2019.8920398) presented at the 3rd IEEE Conference on Control Techniques and Applications 2019 (CCTA 2019) (https: / / doi.org / 10.1109 / CCTA.2019.8920398).
[0175] The reward function used in the above embodiments is a weighted sum of user comfort and cost. User comfort can be indicated by one or more variables related to comfort within a building. Examples of metrics for user comfort include the average of temperature and / or humidity errors in different areas of a building heated by a heating system. Cost can be a measure of the cost of thermal power and actuator energy consumption. The reward function can be constructed such that during a backoff period, which may be user-defined, comfort measurements are removed from the reward function and replaced by a soft boundary at a specified temperature. This soft boundary temperature determines how low the area temperature is allowed to be during the backoff period. Because the learning agent optimizes the reward (gain) over time, it will learn how to perform the optimal backoff for cost-effective reheating based on the cost structure of the heat source.
[0176] Figure 6 The diagram schematically illustrates the process of selecting input signals for a self-learning control process. Selecting input signals (also known as state selection) reduces the dimensionality of the learning agent's state space to improve training speed. Dimensionality reduction can be achieved in various ways, such as principal component analysis or pruning.
[0177] For an agent or other self-learning control process to learn effectively, it needs to be able to predict future gains from states and actions. This means the states should possess sufficient information for the process to make reasonable predictions about gains. For buildings heated and cooled via, for example, hybrid loops, this depends on the specific building in which the hybrid loop is installed.
[0178] A building might have large windows through which observations of solar radiation provide information about free heating. Another building might be poorly insulated and leaky, where observations of wind speed provide more information. One could argue that if all available inputs were fed into the reinforcement learning agent, it would still converge if the required information were available. However, using input variables that lack information or even possess redundant information will reduce the learning rate of the algorithm due to dimensionality; while the dimensionality of the input set increases linearly, the total volume of the model domain increases exponentially. In this disclosure, the inventors propose data-driven state selection, which allows states to be selected based on the specific building without requiring expertise on that specific building.
[0179] Therefore, the problem of automatic state selection is treated here as a prediction problem, in which the process determines a set of input signals carrying information suitable for predicting future returns.
[0180] The reinforcement learning method used in the above embodiments employs an action-value function, in which the expected reward is predicted based on the system's state and the actions taken. The action space includes controllable variables that can be controlled by the control process. For example, in the context of a hybrid loop, this includes pump speed and forward temperature.
[0181] Since this information provides predictions about returns, it can be removed from the prediction target before selecting the input signal.
[0182] This embodiment and other embodiments described herein employ variable selection via mutual information. Features of interest for this method include its ability to handle nonlinear correlations, its model-free nature, and its status as a filter method in the sense of removing the entire input via a filter. Further details of mutual information can be found in "Input Selection for Estimating Return Temperature with Partial Mutual Information with Flow Variable Delay in Hybrid Loops" by Overgaard, CSKallesoe, JDBendtsen, and BKNielsen, 1st IEEE Conference on Control Techniques and Applications (CCTA) 2017 vol. 2017-January 2017, pp. 1372-1377. The aforementioned article describes the application of mutual information to the estimation of return temperature in a heating system. Embodiments of the methods described in this disclosure apply mutual information criteria to the selection of input signals for a self-learning control process used to control a fluid transport system. In particular, in some embodiments of the methods disclosed herein, mutual information is applied to determine whether the input signal contains information about future gains or other suitable performance measurements.
[0183] This process can be performed iteratively: after finding a first input signal, a second input signal is sought that provides the highest mutual information after the first input signal has been given, the first signal containing the highest mutual information of the gain. Information already given by the first input signal is removed by estimating the gain using only the selected input signal and then subtracting the estimate from the gain. For this purpose, a functional approximator can be used to make the estimate based on observed states and performance measurements. An example of such a functional approximator includes a neural network [f(s t:t+k ,w)≈G t:t+k ], where the weights w are adjusted to make the absolute error between the predictor output and the observed performance measurement. Minimize. Similarly, estimate the input signal using only the selected input and subtract it from the remaining input signal, leaving a set of residuals.
[0184] For details, please refer to the following: Figure 6 In step S51, the process loads the training dataset and the validation dataset, specifically in... Figure 4The training and validation datasets collected in steps S1 and S2. Each dataset contains a complete set of input signals x for the system. One of the input signals is also the volumetric flow rate variable q. t and revenue G t The rewards can be calculated by describing how well the system is controlled.
[0185] In step S52, the process selects the input signal x to be analyzed from the set of input signals x. i A loop iterates through all the input signals in the set.
[0186] In steps S53 to S55, the process calculates the mutual information of the input signals.
[0187] Specifically, in step S53, the process determines the volume constant v for each state, which gives the time offset as: This allows the time offset to maximize mutual information about the gains. The improvement from using a flow-dependent time shift delay is due to the transmission delay in the system. For signals without transmission delay, the minimum offset is typically zero (via the volume constant v). min =0 is given). For a given signal, the mutual information value is in the constant volume interval (v min to v max Maximize on ).
[0188] In step S54, this process calculates the offset signal vector x in the training dataset. i With revenue G t Mutual information I between them.
[0189] Mutual information between two variables can be defined as:
[0190]
[0191] This process can compute an approximation of the mutual information. Specifically, a discrete approximation of the mutual information can be computed using m samples of the input signal and the payoff, based on estimates of the marginal and joint probability density functions of the input signal and the payoff.
[0192] In step S55, the process checks whether v, which generates the highest mutual information, has been found. min With v max If the value between v is specified, the process continues at step S56; otherwise, the process returns to step S53.
[0193] In step S56, the process updates the sorted index vector to track how much mutual information the selected input signal contains compared to other input signals already studied. Specifically, the process stores the i'-th index in a vector sorted according to the maximum mutual information.
[0194] In step S57, the process determines whether all input signals have been analyzed. If so, the process continues at step S58; otherwise, the process returns to step S52, that is, repeats steps S52 to S57, until the signal index vectors have been sorted according to a mutual information level having a corresponding volume constant v for all signals.
[0195] In step S58, the process adds the signal with the highest current mutual information s to the signal vector s.
[0196] In step S59, the process uses signals from the validation dataset s to calculate the gain G. t The estimate.
[0197] In step S510, the process checks whether the stopping criterion is met. If not, the process continues at step S511 and calculates a new training dataset that does not contain information from the added signal. This calculation is performed in steps S511 to S514, and then steps S52 to S57 are repeated until the sorted index vector of the new dataset is obtained.
[0198] This method uses partial mutual information iteratively. Specifically, it selects the input variable with the most information, removes that information from the prediction target, and leaves a new residual prediction target, thus utilizing partial mutual information. It then finds the input that provides the most information about the residual prediction target, and so on, until all input variables have been stopped or sorted.
[0199] Therefore, the process can compare the mutual information of the corresponding input variables at the corresponding time shift delay, especially at the time shift delay that leads to the highest mutual information for the corresponding variables.
[0200] Specifically, in step S511, the process generates an estimated return based on s. And generate the estimated state vector based on s.
[0201] In step S512, the process calculates the new state vector by subtracting the estimated state vector from the new state vector:
[0202] In step S513, the process calculates the new revenue by subtracting the estimated revenue from the previous revenue:
[0203] In step S514, the process will G t =G t,j+1 and x t =x t,j+1 Set it as a new training set and return to step S52.
[0204] If the stopping criterion of step S510 is met, the process continues at step S55 and selects a state vector as s. Then, the state selection process is completed.
[0205] An example of a suitable stopping criterion to be applied in step S510 could be based on a white noise comparison, i.e., determining whether the signal with the highest current mutual information provides a greater description of the gain than white noise. If this is not the case, the signal does not contain information about the gain, and a stopping state can be selected.
[0206] Another possible stopping criterion could be improved based on RMSE (Root Mean Square Error), that is, based on the gain G of the validation dataset when adding a signal with the highest mutual information. t The process can be terminated if the estimated RMSE has improved to a certain level.
[0207] For example, when Figure 6 When the specific implementation method of the process is applied to a hybrid circuit, it can be summarized as follows:
[0208] Result: State Selection
[0209] Initialization: Load all inputs income and
[0210] flow The training data, load n inputs
[0211] income and traffic Valid data.
[0212] Parameter: to /
[0213]
[0214] Although embodiments of the various aspects disclosed herein have been described primarily in the context of heating systems for buildings, it should be understood that embodiments of the methods and systems described herein can also be applied to the control of other types of fluid transport systems.
[0215] For example, embodiments of the methods and systems described herein can also be applied to the control of water supply systems. In the context of a water supply system, examples of suitable input signals may include one or more of the following:
[0216] - Flow and pressure measurements at each pumping station, including controllable flow and / or pressure and / or flow / pressure that cannot be directly controlled, for example, at uncontrolled pumping stations.
[0217] - The water level in each tank or other storage container.
[0218] - Weather forecast data, such as information about precipitation and / or temperature.
[0219] - Data from pressure sensors within the network.
[0220] - Data from meters that measure water consumption at one or more consumers.
[0221] Control variables that can be controlled by the control process may include flow rate and / or pressure at one or more pump stations in the water supply system.
[0222] Similarly, embodiments of the methods and systems described herein can also be applied to the control of wastewater systems. In the context of wastewater systems, examples of suitable input signals may include one or more of the following:
[0223] - Flow and pressure measurements at each pumping station, including controllable flow and / or pressure and / or flow / pressure that cannot be directly controlled, for example, at uncontrolled pumping stations.
[0224] - Water level in the wastewater conduit and / or reservoir based on gravity.
[0225] Weather forecast data, for example, regarding precipitation.
[0226] Data on water consumption.
[0227] - Wastewater production data from wastewater sources (e.g., from one or more large industrial wastewater sources).
[0228] Control variables that can be controlled by the control process may include flow rate and / or pressure and / or setpoints of levels at one or more pump stations in the wastewater system.
[0229] Embodiments of the methods described herein can be implemented by means of hardware comprising several different elements and / or at least in part by means of a suitably programmed microprocessor. In the apparatus claims enumerating several means, several of these means may be embodied by the same element, component, or hardware item. The mere fact that certain measures are recited in mutually different dependent claims or described in different embodiments does not indicate that combinations of these measures cannot be advantageously used.
[0230] It should be emphasized that, when used in this specification, the term "comprising / including" is used to specify the presence of the said feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, components, or groups thereof.
Claims
1. A computer-implemented method for controlling the operation of a fluid transport system by applying a self-learning control process, characterized in that, The method includes: - During the operation of the fluid delivery system during a first time period, the values of a plurality of input signals are received, wherein the operation of the fluid delivery system during the first time period is controlled by a predetermined control process, wherein the plurality of input signals define an input space having a first dimension. -Based on the values of the multiple input signals received during the first time period while the operation of the fluid delivery system is being controlled by the predetermined control process, a subset of the multiple input signals is automatically selected. - During the operation of the fluid delivery system during the second time period, values of at least a selected subset of input signals are received, wherein the operation of the fluid delivery system during the second time period is controlled by applying the self-learning control process, wherein the self-learning control process is configured to control the operation of the fluid delivery system based only on the selected subset of input signals, and wherein applying the self-learning control process includes updating the self-learning control process based on the received values of the selected subset of input signals and at least based on an approximation of a performance index function, wherein the self-learning control process implements a reward-based learning agent, and wherein the selected subset of input signals defines a reduced input space with a reduced dimension, the reduced dimension being less than the first dimension, such that the dimension of the state space of the learning agent is reduced.
2. The computer-implemented method according to claim 1, wherein, The predetermined control process is a non-adaptive control process.
3. The computer-implemented method according to any one of the preceding claims, wherein, Automatic selection involves applying one or more information theory selection criteria.
4. The computer-implemented method according to claim 3, wherein, The one or more information theory selection criteria include a mutual information criterion, which is based on a determined mutual information measurement between each of the plurality of input signals and the observed performance measurement.
5. The computer-implemented method according to claim 4, wherein, The observed performance measurement includes at least one observed performance index evaluated over multiple time periods, wherein the performance index value is subject to time-dependent weighting, and the time-dependent weighting depends on the rate of fluid in the fluid delivery system.
6. The computer-implemented method according to any one of the preceding claims, wherein, Automatic selection includes selecting at least one input signal associated with a time shift delay, wherein the time shift delay includes a variable time shift delay that depends on the flow rate of the fluid in the fluid delivery system.
7. A computer-implemented method according to any one of the preceding claims, comprising configuring an initial version of the self-learning control process based on a subset of the selected input signals; wherein, Configuring the initial version of the self-learning control process includes pre-training the initial version of the self-learning control process based on the values of the multiple input signals received during the first time period and based on the performance index values recorded during the operation of the fluid delivery system during the first time period.
8. The computer-implemented method according to claim 7, wherein, The automatic selection and configuration of the initial version of the self-learning control process are performed during a transition period, which is after the first period and before the second period.
9. The computer-implemented method according to claim 1, wherein, The reward-based learning agent is a reinforcement learning agent.
10. The computer-implemented method according to claim 1 or 9, wherein, The self-learning control process is updated based on one or more observed performance metrics, which are observed over a time period, including a flow-dependent time period.
11. The computer-implemented method according to any one of the preceding claims, wherein, The self-learning control process includes at least one random component.
12. The computer-implemented method according to any one of the preceding claims, wherein, The self-learning control process is updated based on an approximation of the performance index function, wherein the approximation is a performance approximator function that approximates the dependence of the performance index function on a subset of the selected input signals and / or on one or more control actions taken by the self-learning control process to control the fluid delivery system.
13. The computer-implemented method according to claim 12, wherein, The performance approximator function is parameterized by a plurality of weight parameters, wherein updating the self-learning control process includes updating one or more of the plurality of weight parameters.
14. The computer-implemented method according to any one of the preceding claims, wherein, The performance index function includes comfort indexes and / or cost indexes.
15. The computer-implemented method according to any one of the preceding claims further includes: - Based on the received values of the multiple input signals, a new subset of the multiple input signals is automatically selected, which is received during the second time period. - During the operation of the fluid delivery system during the third time period, values of a new subset of at least selected input signals are received, wherein the operation of the fluid delivery system during the third time period is controlled by applying a new self-learning control process adapted to the new subset of the selected input signals, wherein the new self-learning control process is configured to control the operation of the fluid delivery system based solely on the new subset of the selected input signals, and wherein applying the new self-learning control process includes updating the new self-learning control process based on the received values of the new subset of the selected input signals and at least based on an approximation of the performance index function.
16. A control system for controlling a fluid transport system, characterized in that, The control system is configured to perform the steps of a computer-implemented method according to any one of the preceding claims.
17. The control system of claim 16, comprising a control unit communicatively coupled to one or more controllable components of the fluid delivery system; wherein, The control unit is configured to receive values of at least a selected subset of input signals acquired during operation of the fluid delivery system, and to selectively control the operation of the fluid delivery system by applying the predetermined control procedure or by applying the self-learning control procedure.
18. The control system of claim 17, further comprising a data processing system configured to receive values of the acquired plurality of input signals during operation of the fluid delivery system during the first time period, and to automatically select a subset of the plurality of input signals based on the received values of the acquired plurality of input signals.
19. The control system according to claim 18, wherein, The data processing system is a remote data processing system located far from the control unit, and the remote data processing system includes cloud services.
20. The control system according to claim 18 or 19, wherein, The data processing system is further configured to configure an initial version of the self-learning control process based on a subset of the selected input signals; wherein configuring the initial version of the self-learning control process includes training the initial version of the self-learning control process based on the values of the multiple input signals received during the operation of the fluid transport system during the first time period and based on the performance index values recorded during the operation of the fluid transport system during the first time period.