Multi-parameter coupling optimization control method based on deep reinforcement learning

By employing a multi-parameter coupled optimization control method based on deep reinforcement learning, real-time data from the chilled water system is collected, a hydraulic state model is constructed, and a differential pressure compensation signal is introduced. This solves the problem of hydraulic imbalance in high-rise buildings and achieves hydraulic balance and energy consumption optimization.

CN121364640BActive Publication Date: 2026-03-17NANJING DEEPCTRLS TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Chilled water systems in high-rise buildings suffer from hydraulic imbalance, leading to uneven heating and cooling. Existing technologies are insufficient to effectively address the hydraulic imbalance problem in multi-story high-rise buildings.

Method used

A multi-parameter coupled optimization control method based on deep reinforcement learning is adopted. By collecting the opening degree and differential pressure data of the terminal valve in real time, a hydraulic state model is constructed, the differential pressure compensation signal of the partition is extended, and hydraulic constraints are introduced into the deep reinforcement learning strategy to generate control commands to adjust the settings of the chiller, pump and fan to achieve hydraulic balance.

Benefits of technology

It effectively solves the problem of hydraulic imbalance in high-rise buildings, achieves dynamic balance of water distribution on different floors, and ensures comfort and energy consumption optimization within the building.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121364640B_ABST
    Figure CN121364640B_ABST
Patent Text Reader

Abstract

A multi-parameter coupling optimization control method based on deep reinforcement learning, comprising: collecting the end valve opening degree and pressure difference data of the chilled water pipe network in different height partitions of the building in real time, combining the pump house outlet pressure and the building height difference to construct a hydraulic state model. Based on the hydraulic state model, the key compensation parameters are extracted, and the partition differential pressure compensation signal is expanded in the initial action space of the chilled water pipe network to construct a multi-parameter action space. Based on the multi-parameter action space, the deep reinforcement learning strategy is trained in the simulation environment, and the hydraulic constraint is introduced into the deep reinforcement learning strategy. The deep reinforcement learning strategy is called to generate multiple control instructions according to the real-time partition hydraulic state to control the setting of the chiller, pump and fan in the chilled water pipe network and the partition differential pressure compensation signal. The water supply temperature and flow rate of the high-end terminal after the execution of the control instruction are monitored in real time to verify whether the hydraulic state of the building in different height partitions meets the preset target threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chilled water system regulation technology, and in particular to a multi-parameter coupled optimization control method based on deep reinforcement learning. Background Technology

[0002] Chilled water systems are complex piping networks with interconnected loops. Changes in flow rate in one loop often lead to mutual influence across different loops, requiring careful consideration of numerous factors for regulation. Furthermore, the abundance of unknowns on-site and the correlation between the characteristics of pumps, valves, and other components and the hydraulic characteristics of the piping network make implementation extremely difficult. Currently, the chilled water systems of central air conditioning systems in large public buildings are typically massive in scale and complex in piping. Due to a series of problems arising during the design, construction, and later operation phases, hydraulic imbalance is a common issue. This hydraulic imbalance causes the flow distribution in each loop and terminal equipment to deviate from design values, directly resulting in uneven heating and cooling across different air-conditioned areas.

[0003] Currently, in high-rise office buildings or mixed-use commercial and residential buildings, the static pressure difference of chilled water pipe networks varies significantly with the floor height. Even if reinforcement learning optimizes the overall energy consumption, it is still easy for the valves at the end of the upper floors to not receive enough water, making it difficult to guarantee the comfort of the building's interior. Therefore, it is necessary to provide a multi-parameter coupled optimization control method based on deep reinforcement learning to set up special compensation control for hydraulic imbalance in three-dimensional high-rise buildings. Summary of the Invention

[0004] Therefore, it is necessary to propose a multi-parameter coupled optimization control method based on deep reinforcement learning to address the above-mentioned technical problems.

[0005] The present invention adopts the following technical solution.

[0006] The first aspect of this invention discloses a multi-parameter coupled optimization control method based on deep reinforcement learning, the method comprising:

[0007] Real-time data on the opening degree and differential pressure of the end valves of the chilled water pipe network in different building height zones are collected. Combined with the pump station outlet pressure and the building height difference, a hydraulic state model of the building height zone is constructed.

[0008] Based on the hydraulic state model, key compensation parameters are extracted, and the zone differential pressure compensation signal is amplified in the initial action space of the chilled water pipe network to construct a multi-parameter action space.

[0009] Based on the multi-parameter action space, a deep reinforcement learning strategy is trained in a simulation environment, and hydraulic constraints are introduced into the deep reinforcement learning strategy.

[0010] The deep reinforcement learning strategy is invoked to generate multiple control commands based on the real-time hydraulic state of each zone, in order to control the settings of chillers, pumps and fans in the chilled water network as well as the differential pressure compensation signal of each zone.

[0011] Real-time monitoring of the water supply temperature and flow rate at the high-level terminal after the execution of the control command is carried out in order to verify whether the hydraulic status of different height zones of the building meets the preset target threshold.

[0012] The building is divided into high, middle and low zones according to different heights; the key compensation parameters are the minimum differential pressure data at the end of the high zone and the valve margin in the middle and low zones; the initial operating space is the chilled water supply temperature, water pump speed, fan speed and valve opening.

[0013] Furthermore, the real-time acquisition of terminal valve opening and differential pressure data of chilled water pipe networks in different building height zones, combined with the pump station outlet pressure and building height difference, constructs a hydraulic state model for each building height zone, including:

[0014] Differential pressure sensors are installed at the ends of the chilled water pipe network in the high, middle and low zones of the building, and the differential pressure data of different height zones of the building are collected in real time through the differential pressure sensors. At the same time, the valve opening degree at the end of each zone and the pump room outlet pressure are obtained.

[0015] The differential pressure data, valve opening, and pump station outlet pressure are filtered and normalized to construct a standardized set of monitoring parameters.

[0016] Furthermore, the real-time acquisition of terminal valve opening and differential pressure data of chilled water pipe networks in different building height zones, combined with the pump station outlet pressure and building height difference, to construct a hydraulic state model for the building height zones, also includes:

[0017] Based on the height values ​​of the high, middle and low zones of the building, the theoretical static pressure of different height zones is calculated, and the effective differential pressure is determined according to the differential pressure data collected by the differential pressure sensor. The effective differential pressure is the sum of the theoretical static pressure and the differential pressure data.

[0018] Based on the effective differential pressure, a valve flow characteristic curve is constructed to estimate the water supply of each zone and obtain the overall water supply vector of the building.

[0019] Based on the water supply vector and the preset target water supply for each zone, combined with the valve flow characteristic curve and valve opening, the hydraulic state model is generated.

[0020] Furthermore, based on the hydraulic state model, key compensation parameters are extracted, and the zone differential pressure compensation signal is amplified in the initial action space of the chilled water pipe network to construct a multi-parameter action space, including:

[0021] Based on the hydraulic state model, the minimum differential pressure data at the end of the high zone and the valve margins in the middle and low zones are extracted according to the time sequence to construct the differential pressure sequence at the end of the high zone and the valve opening sequence in the middle and low zones.

[0022] Within the set monitoring period, the minimum value of the high-zone end pressure difference is calculated based on the high-zone end pressure difference sequence, and the set of valve margin parameters for the middle and low zones is determined based on the middle and low zone valve opening sequence.

[0023] The initial operating space of the chilled water pipeline network is obtained, and the partition differential pressure compensation signal is determined based on the minimum value of the valve margin parameter set of the low and medium zones and the terminal pressure difference of the high zone. The partition differential pressure compensation signal is then written into the initial operating space to obtain the expanded multi-parameter operating space.

[0024] Furthermore, the step of training a deep reinforcement learning policy in a simulation environment based on the multi-parameter action space, and introducing hydraulic constraints into the deep reinforcement learning policy, includes:

[0025] The indoor temperature and humidity of each height zone of the building are obtained and written into the multi-parameter action space. At the same time, the multi-parameter action space and the indoor temperature and humidity are normalized to a preset range to obtain a normalized state space set.

[0026] Based on the state space set, a reward function is constructed with the total power consumption of the chilled water pipe network system as the optimization objective. Peak penalty, hydraulic constraint and comfort constraint are introduced into the reward function to obtain a composite reward function.

[0027] The hydraulic constraints are that the pressure difference at the end of the high zone is not allowed to be lower than a set threshold and the valve opening in the middle and low zones is not allowed to exceed a set upper limit; the comfort constraints are that the indoor temperature and humidity meet the set temperature and humidity range.

[0028] Furthermore, the step of training a deep reinforcement learning policy in a simulation environment based on the multi-parameter action space, and introducing hydraulic constraints into the deep reinforcement learning policy, further includes:

[0029] The multi-parameter action space is used as the candidate actions of the deep reinforcement learning policy, and the state space set is used as the input of the deep learning framework. Deep reinforcement learning is trained in a simulation environment to obtain the deep reinforcement learning policy.

[0030] The composite reward function is used as the optimization objective and is introduced into the deep reinforcement learning strategy.

[0031] Furthermore, the invocation of the deep reinforcement learning strategy to generate multiple control commands based on the real-time partition hydraulic state to control the settings of chillers, pumps, and fans in the chilled water network, as well as the partition differential pressure compensation signal, includes:

[0032] During the execution of the deep reinforcement learning strategy in the chilled water pipe network, the indoor temperature and humidity, water supply, effective differential pressure, valve opening and pump room outlet pressure of each height zone of the building are collected in real time to construct a real-time hydraulic state set.

[0033] The real-time hydraulic state set is input into the deep reinforcement learning strategy to output the optimal action command, and the partition differential pressure compensation signal is parsed in response to the optimal action command to obtain the partition differential pressure adjustment parameters.

[0034] Based on the partition differential pressure regulation parameters, control instructions to be executed are generated, resulting in a set of control instructions for controlling the settings of chillers, pumps, and fans in the chilled water network, as well as the partition differential pressure compensation signal.

[0035] Furthermore, the real-time monitoring of the water supply temperature and flow rate at the high-level terminal after the execution of the control command, to verify whether the hydraulic status of different height zones of the building meets the preset target threshold, includes:

[0036] In response to the control command, the control command is executed according to the differential pressure adjustment parameters of the zones to adjust the chilled water supply temperature, water pump speed, fan speed, valve opening and differential pressure target of each height zone, while collecting hydraulic status feedback data.

[0037] Based on the hydraulic state feedback data, a judgment function is constructed based on high-zone water supply constraints, mid- and low-zone valve margin constraints, comfort constraints, and energy efficiency constraints. The judgment function is then called to output a comprehensive judgment result, and the comprehensive judgment result is verified to see if it meets the preset conditions.

[0038] The second aspect of this invention discloses a multi-parameter coupled optimization control device based on deep reinforcement learning, used to implement the multi-parameter coupled optimization control method based on deep reinforcement learning as described in any one of the first aspects, the device comprising:

[0039] The building zoning modeling module is used to collect real-time data on the opening degree and pressure difference of the end valves of the chilled water pipe network in different building height zones. Combined with the pump room outlet pressure and the building height difference, it constructs a hydraulic state model of the building height zone.

[0040] The action space construction module is used to extract key compensation parameters based on the hydraulic state model and amplify the partition differential pressure compensation signal in the initial action space of the chilled water pipe network to construct a multi-parameter action space.

[0041] The strategy reinforcement training module is used to train a deep reinforcement learning strategy in a simulation environment based on the multi-parameter action space, and to introduce hydraulic constraints into the deep reinforcement learning strategy.

[0042] The instruction generation module is used to call the deep reinforcement learning strategy to generate multiple control instructions based on the real-time partition hydraulic state, so as to control the settings of chillers, pumps and fans in the chilled water pipeline network and the partition differential pressure compensation signal.

[0043] The monitoring and verification module is used to monitor the water supply temperature and flow rate at the high-level terminal after the control command is executed in real time, so as to verify whether the hydraulic status of different height zones of the building meets the preset target threshold.

[0044] The building is divided into high, middle and low zones according to different heights; the key compensation parameters are the minimum differential pressure data at the end of the high zone and the valve margin in the middle and low zones; the initial operating space is the chilled water supply temperature, water pump speed, fan speed and valve opening.

[0045] A third aspect of the present invention discloses a terminal, including a processor and a storage medium;

[0046] The storage medium is used to store instructions;

[0047] The processor is configured to operate according to the instructions to perform the steps of the method described in the first aspect.

[0048] A fourth aspect of the present invention discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0049] The present invention has the following advantages:

[0050] (1) This invention uses differential pressure sensors arranged in the chilled water network to collect the opening degree and pressure difference of typical terminal valves in high, medium and low zones. Combined with the pump room outlet pressure and the floor height difference, a hydraulic state model of each floor zone is formed. By comparing the opening degree of valves in each zone with the actual water supply, it is determined whether there is a significant imbalance in hydraulic distribution, and the influence of static pressure difference in high-rise buildings is clarified, providing a basis for subsequent compensation control. At the same time, based on the action space of initial chilled water supply temperature, pump speed, fan speed and valve opening, the zone differential pressure compensation signal is added as a new control variable, so that the control object can explicitly include the hydraulic balance target, solving the problem of three-dimensional hydraulic balance and introducing an operable compensation method for water cooling optimization.

[0051] (2) This invention utilizes the parameter action space to train a deep reinforcement learning strategy in a simulation environment. The training objectives include not only minimizing energy consumption and maintaining comfort, but also introducing additional hydraulic constraints, namely, the pressure difference at the end of the high zone must not be lower than the threshold and the valve opening degree in the middle and low zones must not exceed the upper limit. This enables the reinforcement learning strategy to simultaneously consider energy consumption optimization and hydraulic balance in a multi-story cooling scenario, incorporating the hydraulic balance objective into the optimization process, and effectively avoiding the imbalance between water shortage in the high zone and oversupply in the low zone of the building.

[0052] (3) In actual system operation, the present invention calls reinforcement learning strategy to generate control commands in combination with real-time partition hydraulic status. These commands not only include the conventional settings of chiller, pump and fan, but also the regulation of partition differential pressure compensation signal. For example, dynamically adjust the start and stop logic of high zone water supply pump and set differential pressure maintenance point. Through online reasoning and real-time feedback, the water distribution of different floor areas is kept dynamically balanced, which solves the water supply imbalance caused by the floor height difference in the operation of water cooling system and realizes the hydraulic regulation. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0054] Figure 1 This is a flowchart illustrating the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention.

[0055] Figure 2 This is a schematic diagram of the structure of the multi-parameter coupled optimization control device based on deep reinforcement learning provided by the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] like Figure 1 As shown, in one embodiment, a multi-parameter coupled optimization control method based on deep reinforcement learning includes the following steps:

[0058] Step S110: Real-time data on the opening degree and pressure difference of the end valves of the chilled water pipe network in different building height zones are collected. Combined with the pump room outlet pressure and the building height difference, a hydraulic state model of the building height zone is constructed.

[0059] The building is divided into high zone, middle zone and low zone according to different heights.

[0060] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention includes the following steps in step S110:

[0061] Step S111: Differential pressure sensors are installed at the ends of the chilled water pipe networks in the high, middle and low zones of the building. The differential pressure data of different height zones of the building are collected in real time through the differential pressure sensors. At the same time, the valve opening degree at the end of each zone and the pump room outlet pressure are obtained.

[0062] Step S112 involves filtering and normalizing the differential pressure data, valve opening, and pump station outlet pressure to construct a standardized set of monitoring parameters.

[0063] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention further includes the following steps in step S110:

[0064] Step S113: Based on the height values ​​of the high, middle and low zones of the building, calculate the theoretical static pressure of different height zones, and determine the effective differential pressure according to the differential pressure data collected by the differential pressure sensor. The effective differential pressure is the sum of the theoretical static pressure and the differential pressure data.

[0065] Step S114: Construct valve flow characteristic curves based on effective differential pressure to estimate the water supply of each zone and obtain the overall water supply vector of the building.

[0066] Step S115: Based on the water supply vector and the preset target water supply for each zone, and combined with the valve flow characteristic curve and valve opening, generate a hydraulic state model.

[0067] In a specific embodiment, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention includes steps 1 to 5:

[0068] Step 1: Full-floor hydraulic condition monitoring and zonal modeling.

[0069] Differential pressure sensors are deployed in the chilled water pipe network to collect the opening degrees and pressure differentials of typical terminal valves in high, medium, and low-pressure zones. Combined with the pump station outlet pressure and floor height differences, a hydraulic state model for each floor zone is formed. By comparing the valve opening degrees and actual water supply in each zone, it is determined whether there is a significant imbalance in hydraulic distribution, thus clarifying the impact of static pressure difference in high-rise buildings and providing a basis for subsequent compensation control. This includes the following sub-steps:

[0070] Sub-step 1.1: Acquisition and preprocessing of raw hydraulic parameters.

[0071] Specifically, the building is divided into high, middle and low zones according to its height, and differential pressure sensors are installed at typical ends of the chilled water pipe network in the high, middle and low zones to obtain the differential pressure value of each zone in real time. At the same time, the valve opening degree at the end of each zone and the pump room outlet pressure (200~600kPa) are collected. Then, the collected differential pressure value, valve opening degree and pump room outlet pressure of each zone are filtered and normalized to obtain a standardized set of monitoring parameters.

[0072] Sub-step 1.2: Calculation of floor height difference and static pressure compensation.

[0073] Specifically, since the static pressure difference in high-rise buildings increases linearly with height, assuming the floor height of a certain zone is known, its theoretical static pressure is the product of the floor height, water density, and gravitational acceleration. Then, the actual pressure difference value measured by the differential pressure sensor at the corresponding floor height is added to the theoretical static pressure at that floor height to obtain the effective differential pressure for that floor height. The effective differential pressures of all zones in the building are calculated in this way, ultimately constructing the effective differential pressure set for the entire building.

[0074] Sub-step 1.3: Fitting the relationship between valve opening and water supply.

[0075] Specifically, a valve flow characteristic curve is constructed based on the effective differential pressure set, and the water supply of each zone is estimated based on this curve. The expression is as follows:

[0076]

[0077] In the formula, This represents the water supply volume of region i at time t; The valve flow coefficient ranges from 0.5 to 2.0. Let be the effective differential pressure of the i-th region at time t; Let be the valve opening degree of region i at time t.

[0078] Finally, the water supply volume of each zone on the entire floor is calculated according to the above formula, and then the water supply volume vector corresponding to the entire floor is output. This vector is composed of the water supply volume of each zone.

[0079] Sub-step 1.4: Construction and imbalance determination of the zoned hydraulic state model.

[0080] Specifically, a target water supply volume is defined for each zone. This target volume is determined by design load calculations or operational experience. If the water supply volume of the corresponding zone is less than the target volume and the valve opening is close to 1, it indicates that the zone is under-supplied. If the water supply volume of the corresponding zone is greater than the target volume and the valve opening is close to 0.2, it indicates that the zone is over-supplied. Therefore, by combining the target water supply volume with the aforementioned water supply vector, effective differential pressure, and valve opening, the required zone hydraulic state model can be constructed.

[0081] Step S120: Based on the hydraulic state model, key compensation parameters are extracted, and the partition differential pressure compensation signal is amplified in the initial action space of the chilled water network to construct a multi-parameter action space.

[0082] The key compensation parameters are the minimum differential pressure data at the end of the high zone and the valve margins in the middle and low zones; the initial operating space is the chilled water supply temperature, pump speed, fan speed, and valve opening.

[0083] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention includes the following steps in step S120:

[0084] Step S121: Based on the hydraulic state model, extract the minimum differential pressure data at the end of the high zone and the valve margins in the middle and low zones according to the time sequence to construct the differential pressure sequence at the end of the high zone and the valve opening sequence in the middle and low zones.

[0085] Step S122: Within the set monitoring period, calculate the minimum value of the high-zone end pressure difference based on the high-zone end pressure difference sequence, and determine the set of valve margin parameters for the middle and low zones based on the valve opening sequence for the middle and low zones.

[0086] Step S123: Obtain the initial action space of the chilled water pipe network, and determine the partition differential pressure compensation signal based on the set of valve margin parameters in the low and medium zones and the minimum value of the pressure difference at the end of the high zone, so as to write the partition differential pressure compensation signal into the initial action space and obtain the expanded multi-parameter action space.

[0087] In a specific embodiment, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by this invention includes step 2: hydraulic compensation parameter extraction and control variable expansion. Based on the partitioned hydraulic state model output in step 1, key compensation parameters are extracted, including the minimum differential pressure at the high-zone end and the valve margin in the middle and low-zone areas. Based on the original action space of chilled water supply temperature, pump speed, fan speed, and valve opening, partitioned differential pressure compensation signals are added as new control variables, enabling the controlled object to explicitly include the hydraulic balance objective, thus solving the problem of three-dimensional hydraulic balance and introducing an operable compensation method for optimization. This includes the following sub-steps:

[0088] Sub-step 2.1: Analysis and index extraction of the zonal hydraulic state model.

[0089] Specifically, based on the aforementioned zonal hydraulic state model, the minimum differential pressure at the end of the building is extracted as a key indicator for measuring the water supply capacity of high-rise buildings in the high-rise area; and the valve opening degree in the middle and low-rise areas is extracted as a key indicator for judging whether there is excessive valve throttling.

[0090] Sub-step 2.2: Calculation of minimum differential pressure at the end of the high zone.

[0091] Specifically, within the set monitoring period T, the minimum value of the differential pressure at the end of the high zone is calculated. In this example, T is taken as an optimized control period of 15~60 minutes. If the minimum value of the differential pressure in the high zone is lower than the set threshold (30-50 kPa), it indicates that there is a risk of insufficient water supply in the high zone.

[0092] Sub-step 2.3: Calculation of valve margin in the low and medium zones.

[0093] Specifically, valve margin is defined as the difference between the adjustable range of the valve and its current opening. The valve margin in the low and medium zones within the monitoring period T is determined as the valve margin parameter for the low and medium zones. In this example, the valve margin range is 0 to 1. Simultaneously, the valve margin parameter for the low and medium zones is correlated with the minimum differential pressure in the high zone to determine whether the conditions for improving the water supply in the high zone through zone differential pressure regulation are met.

[0094] Sub-step 2.4: Control variable expansion and action space reconstruction.

[0095] Specifically, first, the original operating space is determined, including the chilled water supply temperature, pump speed, fan speed, and the opening degree of the valves at the end of each zone. Then, the zone differential pressure compensation signal is added to the original operating space. This zone differential pressure compensation signal is determined by the minimum differential pressure in the high zone and the valve margin, and is used to dynamically adjust the start / stop of the water supply pump or the differential pressure setpoint, ultimately obtaining the extended operating space.

[0096] Step S130: Based on the multi-parameter action space, a deep reinforcement learning policy is trained in the simulation environment, and hydraulic constraints are introduced into the deep reinforcement learning policy.

[0097] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention includes the following steps in step S130:

[0098] Step S131: Obtain the indoor temperature and humidity of each height zone of the building, and write the indoor temperature and humidity into the multi-parameter action space. At the same time, normalize the multi-parameter action space and the indoor temperature and humidity to a preset range to obtain a normalized state space set.

[0099] Step S132: Based on the state space set, construct a reward function with the total power consumption of the chilled water pipe network system as the optimization objective, and introduce peak penalty, hydraulic constraint and comfort constraint into the reward function to obtain a composite reward function.

[0100] Among them, the hydraulic constraint is that the pressure difference at the end of the high zone must not be lower than the set threshold and the valve opening in the middle and low zones must not exceed the set upper limit; the comfort constraint is that the indoor temperature and humidity meet the set temperature and humidity range.

[0101] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention further includes the following steps in step S130:

[0102] Step S133: Using the multi-parameter action space as the candidate actions for the deep reinforcement learning policy, and using the state space set as the input to the deep learning framework, deep reinforcement learning is trained in a simulation environment to obtain the deep reinforcement learning policy.

[0103] Step S134: Using the composite reward function as the optimization objective, the composite reward function is introduced into the deep reinforcement learning strategy.

[0104] In a specific embodiment, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by this invention includes step 3: training the deep reinforcement learning strategy and introducing hydraulic constraints. Using the multi-parameter action space output in step 2, a deep reinforcement learning strategy is trained in a simulation environment. The training objectives include not only minimizing energy consumption and maintaining comfort, but also additionally introducing hydraulic constraints: the pressure difference at the high-floor end must not fall below a threshold, and the valve opening degree in the middle and low-floor areas must not exceed the upper limit. In this way, the reinforcement learning strategy can simultaneously consider energy consumption optimization and hydraulic balance in multi-floor cooling scenarios, incorporating the hydraulic balance objective into the optimization process and avoiding water shortages in high-floor areas and oversupply in low-floor areas. This includes the following sub-steps:

[0105] Sub-step 3.1, State space construction and normalization.

[0106] Specifically, in multi-story cooling scenarios, the state space not only includes the aforementioned extended action space, but also needs to include the temperature and humidity inside the building in the extended action space in order to balance cooling energy consumption and indoor comfort. The temperature, humidity and hydraulic parameters in the extended action space are then normalized to the interval [0,1] to obtain the normalized state space.

[0107] Sub-step 3.2: Construction of reward function and introduction of energy consumption target.

[0108] Specifically, the reward function takes the total power consumption of the chilled water pipe network system as the main optimization objective, while also considering peak penalty (i.e., peak penalty coefficient). (In this example, we take 5-10), combined with the chiller power. Water pump power Fan power and power demand limit (Range range: 200-2000kW), constructed reward function The expression is:

[0109]

[0110] In the formula, This represents the total power of the system.

[0111] Sub-step 3.3, introduction of hydraulic constraints and comfort constraints.

[0112] Specifically, based on the total power consumption as the optimization objective in sub-step 3.2, hydraulic constraint costs and comfort constraint costs are introduced. Both hydraulic and comfort constraint costs have been normalized to baseline values ​​and, after dimensionless processing, can be directly weighted. The reward function expression with hydraulic and comfort constraints is updated as follows:

[0113]

[0114] In the formula, The hydraulic constraint cost at time t. The weighting coefficient for hydraulic constraint costs; The comfort constraint cost at time t, The weighting coefficients for the cost of comfort constraints; This is the composite reward function after introducing hydraulic constraint costs and comfort constraint costs.

[0115] Among them, hydraulic constraint cost The expression is:

[0116]

[0117] In the formula, For the target differential pressure in the high-area region, The effective differential pressure in the high zone at time t; Let t be the valve opening degree in the middle zone. This represents the maximum valve opening value in the middle zone. Let t be the valve opening in the lower zone. This represents the maximum valve opening in the low-zone area.

[0118] Comfort Constraints The expression is:

[0119]

[0120] In the formula, Total number of indoor units; Let be the indoor temperature of zone i at time t. To set a comfortable temperature threshold, in this example, it is 24°C; This is the tolerance, which in this example is equal to 1℃; Let be the indoor humidity of region i at time t. This represents the upper limit of humidity, which is 60% in this example.

[0121] Sub-step 3.4: Strategy training and generation of water conservancy compensation mechanism.

[0122] Specifically, the aforementioned extended action space is used as the policy's selectable actions, the normalized state space is used as the input, and the composite reward function is used as the optimization objective. Deep reinforcement learning is then performed in the simulation space to obtain the optimal policy as the deep reinforcement learning policy.

[0123] The extended action space includes zone differential pressure compensation signals, so the final deep reinforcement learning strategy can automatically generate compensation actions for insufficient water supply in high-rise buildings, such as adjusting the differential pressure setpoint or starting and stopping the water supply pump, based on the conventional control of chilled water supply and return water temperature, pump and fan speed.

[0124] Step S140: Invoke the deep reinforcement learning strategy to generate multiple control commands based on the real-time zonal hydraulic status, so as to control the settings of chillers, pumps and fans in the chilled water pipeline network as well as the zonal differential pressure compensation signal.

[0125] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention includes the following steps in step S140:

[0126] Step S141: During the execution of the deep reinforcement learning strategy in the chilled water pipe network, the indoor temperature and humidity, water supply, effective differential pressure, valve opening and pump room outlet pressure of each height zone of the building are collected in real time to construct a real-time hydraulic state set.

[0127] Step S142: Input the real-time hydraulic state set into the deep reinforcement learning strategy to output the optimal action command, and parse the partition differential pressure compensation signal in response to the optimal action command to obtain the partition differential pressure regulation parameters.

[0128] Step S143: Generate control instructions to be executed based on the zone differential pressure adjustment parameters, and obtain a set of control instructions for controlling the settings of chillers, pumps and fans in the chilled water network and the zone differential pressure compensation signal.

[0129] In a specific embodiment, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by this invention includes step 4: online inference and adaptive adjustment of zone differential pressure. During actual system operation, the deep reinforcement learning strategy output in step 3 is invoked, combined with real-time zone hydraulic status, to generate control commands. These commands include not only the conventional settings for chillers, pumps, and fans, but also zone differential pressure compensation signals, such as dynamically adjusting the start-stop logic of the high-zone water supply pump and setting the differential pressure maintenance point. Through online inference and real-time feedback, the water distribution in different floor zones is kept dynamically balanced, solving the water supply imbalance caused by floor height differences during operation and achieving real-time hydraulic regulation. This includes the following sub-steps:

[0130] Sub-step 4.1: Strategy invocation and real-time status input.

[0131] Specifically, during system operation, real-time monitoring data of each zone on each floor is collected, including indoor temperature and humidity, water supply, effective differential pressure, valve opening, and pump room outlet pressure, to construct a real-time status set.

[0132] Sub-step 4.2, Action reasoning and compensation signal generation.

[0133] Specifically, the deep reinforcement learning strategy of the aforementioned output is invoked, and the real-time state set is used as input to output the optimal action, and finally the optimal action command set for the entire floor is obtained, which is used to control the chilled water supply temperature setting, water pump speed, fan speed, valve opening, and adjust the high zone differential pressure set point or the operation of the water supply pump according to the zone differential pressure compensation signal.

[0134] Sub-step 4.3: Compensation logic analysis and zone differential pressure regulation.

[0135] Specifically, the differential pressure compensation signal for each zone is analyzed, defining the target differential pressure for the high zone as: current measured differential pressure in the high zone + compensation ratio coefficient (range 0.5~1.5) × differential pressure compensation signal (-20~+20kPa). If the target differential pressure in the high zone is lower than the set threshold (30-50kPa), the high zone water supply pump is triggered. If the average valve opening in the middle and low zones is less than 0.4, the differential pressure setpoint in the middle and low zones is lowered to release margin for use in the high zone. Finally, a set of zone adjustment parameters is output, including the target differential pressures for the high, middle, and low zones.

[0136] Sub-step 4.4: Issuance of control commands and real-time feedback.

[0137] Specifically, based on the aforementioned set of zoned adjustment parameters, executable control commands are generated. These commands are used to: adjust the water supply temperature setting on the chiller side to ensure that the cooling output matches the demand in the high zone; adjust the start / stop or frequency conversion setting of the high zone makeup water pump on the pump side according to the target pressure difference in the high zone, and adjust the pressure setting of the main pump in the medium and low zones according to the target differential pressure in the medium and low zones; adjust the fan speed on the fan side according to the set of action commands to coordinate with the water distribution to stabilize the terminal air handling capacity; and issue valve opening commands on the valve side to appropriately close valves in over-supply zones and appropriately open valves in under-supply zones.

[0138] Step S150: Real-time monitoring of the water supply temperature and flow rate at the high-level terminal after the execution of the control command, in order to verify whether the hydraulic status of different height zones of the building meets the preset target threshold.

[0139] In some embodiments, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by the present invention includes the following steps in step S150:

[0140] Step S151: In response to the control command, execute the control command according to the zone differential pressure adjustment parameters to adjust the chilled water supply temperature, water pump speed, fan speed, valve opening and pressure difference target of each height zone, and collect hydraulic status feedback data at the same time.

[0141] Step S152: Construct a judgment function based on high-zone water supply constraints, mid- and low-zone valve margin constraints, comfort constraints, and energy efficiency constraints according to the hydraulic state feedback data. Call the judgment function to output the comprehensive judgment result and verify whether the comprehensive judgment result meets the preset conditions.

[0142] In a specific embodiment, the multi-parameter coupled optimization control method based on deep reinforcement learning provided by this invention includes step 5: closed-loop feedback verification and end-point comfort assurance. After executing the zone differential pressure compensation control command output in step 4, the water supply temperature and flow rate of the high-zone end are monitored in real time and compared with the target threshold. When the temperature and humidity of the high-zone end remain within the comfort range, while the valve opening in the middle and low zones does not exceed the available margin, and the total system energy consumption is not higher than the benchmark value, it can be determined that the hydraulic balance and energy efficiency optimization goals are achieved simultaneously. The feedback result is input again into the deep reinforcement learning strategy for continuous strategy improvement, ultimately achieving the unification of comfort and energy efficiency across all floors, further solving the problem of hydraulic imbalance at the end of high-rise complexes. This includes the following sub-steps:

[0143] Sub-step 5.1: Execute control commands and monitor data collection.

[0144] Specifically, the aforementioned zone differential pressure compensation control command is executed at the control layer, and feedback data is collected in real time after execution, including indoor temperature and humidity, water supply, effective differential pressure, valve opening degree and total system power.

[0145] Sub-step 5.2, comfort and energy efficiency assessment.

[0146] Specifically, based on the collected feedback data, a judgment function for comfort and energy efficiency is constructed to determine whether the water supply in the high zone is guaranteed, whether the valve margin in the middle and low zones meets the opening requirements, whether the indoor comfort meets the set standards, and whether the system energy efficiency exceeds the benchmark, based on the judgment results, and to impose conditional constraints on the chilled water pipeline system.

[0147] Sub-step 5.3: Feedback input and continuous strategy improvement.

[0148] Specifically, if the judgment result obtained in sub-step 5.2 meets the set conditions, then the high-zone water supply, the valve margin in the middle and low zones, indoor comfort, and system energy efficiency are all determined to meet the constraints, and the system operating status that meets the standards is output. Furthermore, if any constraint is not met, the judgment result and its monitoring data are fed back to the simulation space as incremental training samples to continuously optimize the deep reinforcement learning strategy.

[0149] The following describes the multi-parameter coupled optimization control device based on deep reinforcement learning provided by the present invention. The multi-parameter coupled optimization control device based on deep reinforcement learning described below can be referred to in correspondence with the multi-parameter coupled optimization control method based on deep reinforcement learning described above.

[0150] like Figure 2 As shown, in one embodiment, a multi-parameter coupled optimization control device based on deep reinforcement learning includes a building zoning modeling module, an action space construction module, a strategy reinforcement training module, an instruction generation module, and a monitoring and verification module.

[0151] The building zoning modeling module is used to collect real-time data on the opening degree and pressure difference of the end valves of the chilled water pipe network in different building height zones. Combined with the pump room outlet pressure and the building height difference, it constructs a hydraulic state model of the building height zone.

[0152] The action space construction module is used to extract key compensation parameters based on the hydraulic state model and amplify the partition differential pressure compensation signal in the initial action space of the chilled water network to construct a multi-parameter action space.

[0153] The policy reinforcement training module is used to train deep reinforcement learning policies in a simulation environment based on a multi-parameter action space, and to introduce hydraulic constraints into the deep reinforcement learning policies.

[0154] The instruction generation module is used to call a deep reinforcement learning strategy to generate multiple control instructions based on the real-time hydraulic status of each zone, in order to control the settings of chillers, pumps and fans in the chilled water network as well as the zone differential pressure compensation signal.

[0155] The monitoring and verification module is used to monitor the water supply temperature and flow rate at the high-level terminal after the control command is executed in real time, so as to verify whether the hydraulic status of different height zones of the building meets the preset target threshold.

[0156] The building is divided into high, middle and low zones according to different heights; the key compensation parameters are the minimum differential pressure data at the end of the high zone and the valve margin in the middle and low zones; the initial action space is the chilled water supply temperature, water pump speed, fan speed and valve opening.

[0157] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0158] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A multi-parameter coupling optimal control method based on deep reinforcement learning, characterized in that, The method comprises: Real-time acquisition of end valve opening and differential pressure data of the chilled water pipe network in different height partitions of the building, combination of pump house outlet pressure and building height difference, construction of a hydraulic state model of the building height partitions; Based on the hydraulic state model, key compensation parameters are extracted, and a partition differential pressure compensation signal is expanded in the initial action space of the chilled water pipe network to construct a multi-parameter action space; Based on the multi-parameter action space, a deep reinforcement learning strategy is trained in a simulation environment, and a hydraulic constraint is introduced into the deep reinforcement learning strategy; The deep reinforcement learning strategy is called to generate multiple control instructions according to real-time partition hydraulic states to control the settings of chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal; Real-time monitoring of the water supply temperature and flow rate of the high-end terminal after the execution of the control instructions to verify whether the hydraulic states of the different height partitions of the building meet the preset target threshold; Wherein, the different height partitions of the building are high, middle and low zones; the key compensation parameters are the minimum differential pressure data of the high-end terminal and the valve allowance of the middle and low zones; the initial action space is the chilled water supply temperature, the water pump speed, the fan speed and the valve opening; The deep reinforcement learning strategy is trained in a simulation environment based on the multi-parameter action space, and a hydraulic constraint is introduced into the deep reinforcement learning strategy, which comprises: Obtain the indoor temperature and humidity of each height partition of the building, and write the indoor temperature and humidity into the multi-parameter action space, and normalize the multi-parameter action space and the indoor temperature and humidity to a preset interval to obtain a normalized state space set; Based on the state space set, a reward function is constructed with the total power consumption of the chilled water pipe network system as the optimization target, and a peak penalty, the hydraulic constraint and a comfort constraint are introduced into the reward function to obtain a composite reward function; Wherein, the hydraulic constraint is that the high-end terminal differential pressure is not allowed to be lower than the set threshold, and the valve opening of the middle and low zones is not allowed to exceed the set upper limit; the comfort constraint is that the indoor temperature and humidity meet the set temperature and humidity range; The deep reinforcement learning strategy is trained in a simulation environment based on the multi-parameter action space, and a hydraulic constraint is introduced into the deep reinforcement learning strategy, which further comprises: Correlate the multi-parameter action space with the deep reinforcement learning strategy, and take the state space set as the input of the deep learning framework to perform deep reinforcement learning training in the simulation environment to obtain the deep reinforcement learning strategy; Take the composite reward function as the optimization target, and introduce the composite reward function into the deep reinforcement learning strategy.

2. The deep-reinforcement learning based multi-parameter coupled optimal control method according to claim 1, wherein, The real-time acquisition of end valve opening and differential pressure data of the chilled water pipe network in different height partitions of the building, combination of pump house outlet pressure and building height difference, construction of a hydraulic state model of the building height partitions, comprises: Differential pressure sensors are arranged at the ends of the chilled water pipe network in the high, middle and low zones of the building, and the differential pressure data of the different height partitions of the building are acquired in real time through the differential pressure sensors, and the valve openings of the end terminals of each partition and the pump house outlet pressure are obtained; Filter and normalize the differential pressure data, valve opening and pump house outlet pressure.

3. The deep-reinforcement learning based multi-parameter coupled optimal control method according to claim 2, wherein, The real-time acquisition of the chilled water pipe network terminal valve opening and differential pressure data in different height partitions of the building, combined with the pump house outlet pressure and the building height difference, constructs a hydraulic state model of the building height partition, further comprising: Based on the height values of the high, middle and low zones of the building, the theoretical static pressure of different height partitions is calculated, and the effective differential pressure is determined according to the differential pressure data collected by the differential pressure sensor, which is the sum of the theoretical static pressure and the differential pressure data; Based on the effective differential pressure, a valve flow characteristic curve is constructed to estimate the water supply of each partition, and the water supply vector of the whole building is obtained; According to the water supply vector and the preset target water supply of each partition, combined with the valve flow characteristic curve and the valve opening, the hydraulic state model is generated.

4. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 1, characterized in that, Based on the hydraulic state model, key compensation parameters are extracted, and partition differential pressure compensation signals are amplified in the initial action space of the chilled water pipe network to construct a multi-parameter action space, including: Based on the hydraulic state model, the minimum terminal differential pressure data of the high zone and the valve margin of the middle and low zones are extracted in time sequence to construct the high zone terminal differential pressure sequence and the middle and low zone valve opening sequence; In a set monitoring period, the minimum value of the high zone terminal differential pressure is calculated according to the high zone terminal differential pressure sequence, and the middle and low zone valve margin parameter set is determined based on the middle and low zone valve opening sequence; The initial action space of the chilled water pipe network is obtained, and the partition differential pressure compensation signal is determined according to the middle and low zone valve margin parameter set and the minimum value of the high zone terminal differential pressure, and the partition differential pressure compensation signal is written into the initial action space to obtain the expanded multi-parameter action space.

5. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 1, wherein, The calling of the deep reinforcement learning strategy generates multiple control instructions according to the real-time partition hydraulic state to control the settings of chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal, including: In the process of executing the deep reinforcement learning strategy in the chilled water pipe network, the indoor temperature and humidity, water supply, effective differential pressure, valve opening and pump house outlet pressure of each height partition of the building are collected in real time to construct a real-time hydraulic state set; The real-time hydraulic state set is input into the deep reinforcement learning strategy to output the optimal action instruction, and the partition differential pressure compensation signal is analyzed in response to the optimal action instruction to obtain the partition differential pressure adjustment parameter; According to the partition differential pressure adjustment parameter, the control instruction to be executed is generated to obtain a control instruction set, which is used to control the settings of chillers, pumps and fans in the chilled water pipe network and the partition differential pressure compensation signal.

6. The deep reinforcement learning based multi-parameter coupled optimal control method according to claim 5, characterized in that, The real-time monitoring of the water supply temperature and flow at the high zone terminal after the execution of the control instruction verifies whether the hydraulic state of different height partitions of the building meets the preset target threshold, including: In response to the control instruction, the control instruction is executed according to the partition differential pressure adjustment parameter to adjust the chilled water supply temperature, water pump speed, fan speed, valve opening and differential pressure target of each height partition, while collecting hydraulic state feedback data; A decision function is constructed based on high area water supply constraints, medium and low area valve surplus constraints, comfort constraints and energy efficiency constraints according to the hydraulic state feedback data, a comprehensive decision result is output by calling the decision function, and it is verified whether the comprehensive decision result meets a preset condition. 7.A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is configured to store instructions; The processor is configured to operate according to the instructions to perform the steps of the multi-parameter coupled optimization control method based on deep reinforcement learning according to any one of claims 1-6.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the steps of the multi-parameter coupled optimization control method based on deep reinforcement learning according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-cold-source annular cold supply system pressure difference control method based on reinforcement learning

    CN116819945A

  • Intelligent control method of evaporative cooling air conditioning system based on deep reinforcement learning

    CN120194407A